Command Palette
Search for a command to run...
Nemotron-3-Super-RL-Training-Blends Super Reinforcement Learning Training Blends Dataset
Nemotron-3-Super-RL-Training-Blends is a reinforcement learning training data mixture released by NVIDIA in 2026, used to train the Nemotron-3-Super-120B-A12B model. It contains 479,303 samples with a total size of approximately 26.3 GB, and the data is in JSONL format, designed to provide proportioned and preprocessed data for reinforcement learning from verifiable rewards (RLVR) and post-training of large language models.
Dataset Composition
-
RLVR 1: Contains data for various tasks including mathematical reasoning, tool use, code competitions, instruction following, and safety, totaling 138,712 samples.
-
RLVR 2: Adjusts the mixing ratios of different data sources based on RLVR 1, totaling 156,278 samples.
-
RLVR 3: Further adjusts data proportions, adding more data related to agent capabilities, function calling, and reinforcement learning, totaling 107,037 samples.
-
SWE 1: Mainly composed of R2E-Gym-Subset (79.70%) and SWE-Gym (20.30%), totaling 50,661 samples.
-
SWE 2: Mainly composed of R2E-Gym-Subset (81.18%) and SWE-Gym (18.82%), totaling 1,444 samples.
-
RLHF: Mainly composed of Nemotron-RLHF-GenRM-v1 (77.00%), Nemotron-RL-Agentic-Conversational-Tool-Use-v1 (20.00%), and Nemotron-RL-Identity-Following-v1 (3.00%), totaling 25,171 samples.
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.