HyperAI
Command Palette
Search for a command to run...
Nemotron-3-Nano-RL-Training-Blend 强化学习训练数据集
Nemotron-3-Nano-RL-Training-Blend 是由 NVIDIA 于 2025 年发布的一个用于训练 Nemotron-3-Nano-30B-A3B 模型的精选数据集,包含 93,244 个样本,总存储量为 6.92 GB,数据规模适中且经过精心混合与预处理。由多个子数据集按比例混合而成,包括指令遵循、知识问答、职场助手、结构化输出、竞技编程及数学推理等任务。旨在通过强化学习从可验证奖励(RLVR)框架提升大语言模型性能。
数据样本按照从易到难的课程学习顺序排列,以确保平衡的学习进度,专为 NeMo Gym 框架的后训练阶段设计,支持商业使用。
数据集组成
该数据集是一个混合数据集,由以下组件数据集按指定比例混合而成:
- nvidia/Nemotron-RL-instruction_following (0.17)
- nvidia/Nemotron-RL-knowledge-mcqa (0.20)
- nvidia/Nemotron-RL-agent-workplace_assistant (0.10)
- nvidia/Nemotron-RL-instruction_following-structured_outputs (0.05)
- nvidia/Nemotron-RL-coding-competitive_coding (0.25)
- BytedTsinghua-SIA/DAPO-Math-17k (0.10)
- Skywork/Skywork-OR1-RL-Data (excluding OmniMath) (0.12)
此数据集由社区用户贡献,仅用于教育和信息目的。如有任何内容涉及版权侵权,请通过 [email protected] 联系我们,我们将及时审核并删除。