HyperAI
Command Palette
Search for a command to run...
MMLU-CF 无污染多任务语言理解基准数据集
MMLU-CF 是由 Microsoft 于 2024 年发布的一个无污染且更具挑战性的多项选择题基准数据集,相关论文成果为「MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark」,旨在解决开源基准因大语言模型训练数据泄露导致的评估不可靠问题。
该数据集包含 20,000 道选择题,覆盖生物、数学、化学、物理、法律、工程等多个学科领域。 通过严格的去污染规则及闭源测试集设计,MMLU-CF 有效避免了模型记忆效应,为评估大语言模型的真实性推理能力提供了更纯净的基准。
数据集组成
-
验证集(val):包含 10,000 道选择题,按照学科划分为 Biology、Math、Chemistry、Physics、Law、Engineering、Other、Economics、Health、Psychology、Business、Philosophy、Computer Science、History 等子集。
-
开发集(dev):包含对应的开发数据,按照相同的学科类别划分为多个子集,用于模型评估与实验。
Citation
@misc{zhao2024mmlucfcontaminationfreemultitasklanguage,
title={MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark},
author={Qihao Zhao and Yangyu Huang and Tengchao Lv and Lei Cui and Qinzheng Sun and Shaoguang Mao and Xin Zhang and Ying Xin and Qiufeng Yin and Scarlett Li and Furu Wei},
year={2024},
eprint={2412.15194},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.15194},
}
此数据集由社区用户贡献,仅用于教育和信息目的。如有任何内容涉及版权侵权,请通过 [email protected] 联系我们,我们将及时审核并删除。