HyperAIHyperAI

Command Palette

Search for a command to run...

Nemotron-3-Nano-RL-Training-Blend Reinforcement Learning Training Dataset

Date

Organization

NVIDIA

License

Other

Nemotron-3-Nano-RL-Training-Blend is a curated dataset released by NVIDIA in 2025 for training the Nemotron-3-Nano-30B-A3B model, containing 93,244 samples with a total storage size of 6.92 GB. The data scale is moderate and has been carefully blended and preprocessed. It is composed of multiple sub-datasets mixed in specific proportions, covering tasks such as instruction following, knowledge question answering, workplace assistance, structured outputs, competitive programming, and mathematical reasoning. It is designed to improve large language model performance through reinforcement learning from verifiable rewards (RLVR).

The data samples are arranged in a curriculum learning order from easy to hard to ensure a balanced learning progression. It is specifically designed for the post-training phase of the NeMo Gym framework and supports commercial use.

Dataset Composition

This dataset is a blended dataset composed of the following component datasets mixed in the specified proportions:

  • nvidia/Nemotron-RL-instruction_following (0.17)
  • nvidia/Nemotron-RL-knowledge-mcqa (0.20)
  • nvidia/Nemotron-RL-agent-workplace_assistant (0.10)
  • nvidia/Nemotron-RL-instruction_following-structured_outputs (0.05)
  • nvidia/Nemotron-RL-coding-competitive_coding (0.25)
  • BytedTsinghua-SIA/DAPO-Math-17k (0.10)
  • Skywork/Skywork-OR1-RL-Data (excluding OmniMath) (0.12)

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp