Command Palette
Search for a command to run...
Nemotron-3-Nano-RL-Training-Blend Reinforcement Learning Training Dataset
Nemotron-3-Nano-RL-Training-Blend is a curated dataset released by NVIDIA in 2025 for training the Nemotron-3-Nano-30B-A3B model, containing 93,244 samples with a total storage size of 6.92 GB. The data scale is moderate and has been carefully blended and preprocessed. It is composed of multiple sub-datasets mixed in specific proportions, covering tasks such as instruction following, knowledge question answering, workplace assistance, structured outputs, competitive programming, and mathematical reasoning. It is designed to improve large language model performance through reinforcement learning from verifiable rewards (RLVR).
The data samples are arranged in a curriculum learning order from easy to hard to ensure a balanced learning progression. It is specifically designed for the post-training phase of the NeMo Gym framework and supports commercial use.
Dataset Composition
This dataset is a blended dataset composed of the following component datasets mixed in the specified proportions:
- nvidia/Nemotron-RL-instruction_following (0.17)
- nvidia/Nemotron-RL-knowledge-mcqa (0.20)
- nvidia/Nemotron-RL-agent-workplace_assistant (0.10)
- nvidia/Nemotron-RL-instruction_following-structured_outputs (0.05)
- nvidia/Nemotron-RL-coding-competitive_coding (0.25)
- BytedTsinghua-SIA/DAPO-Math-17k (0.10)
- Skywork/Skywork-OR1-RL-Data (excluding OmniMath) (0.12)
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.