HyperAI
Command Palette
Search for a command to run...
HalluQA 中国大型モデル幻覚評価データセット

このリポジトリには、HalluQA (中国の幻覚質問応答) ベンチマークのデータと評価スクリプトが含まれています。 HalluQA の完全なデータは HalluQA.json にあります。 HalluQA を紹介する論文と複数の中国語大規模言語モデルの詳細な実験結果は、ここ。 HalluQA には、複数の分野にまたがり、中国の歴史、文化、習慣、社会現象を考慮して、慎重に設計された 450 の敵対的な質問が含まれています。
引用
@article{DBLP:journals/corr/abs-2310-03368,
author = {Qinyuan Cheng and
Tianxiang Sun and
Wenwei Zhang and
Siyin Wang and
Xiangyang Liu and
Mozhi Zhang and
Junliang He and
Mianqiu Huang and
Zhangyue Yin and
Kai Chen and
Xipeng Qiu},
title = {Evaluating Hallucinations in Chinese Large Language Models},
journal = {CoRR},
volume = {abs/2310.03368},
year = {2023},
url = {https://doi.org/10.48550/arXiv.2310.03368},
doi = {10.48550/arXiv.2310.03368},
eprinttype = {arXiv},
eprint = {2310.03368},
timestamp = {Thu, 19 Oct 2023 13:12:52 +0200},
biburl = {https://dblp.org/rec/journals/corr/abs-2310-03368.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}
このデータセットはコミュニティユーザーによって提供されており、教育および情報提供のみを目的としています。著作権侵害に関わるコンテンツがある場合は、[email protected]までご連絡ください。速やかに確認し、削除いたします。