HyperAIHyperAI

Command Palette

Search for a command to run...

Nemotron-3-Super-RL-Training-Blends Super Reinforcement Learning Training Blends Dataset

Date

Organization

NVIDIA

License

CC BY 4.0

Nemotron-3-Super-RL-Training-Blends is a reinforcement learning training data mixture released by NVIDIA in 2026, used to train the Nemotron-3-Super-120B-A12B model. It contains 479,303 samples with a total size of approximately 26.3 GB, and the data is in JSONL format, designed to provide proportioned and preprocessed data for reinforcement learning from verifiable rewards (RLVR) and post-training of large language models.

Dataset Composition

  • RLVR 1: Contains data for various tasks including mathematical reasoning, tool use, code competitions, instruction following, and safety, totaling 138,712 samples.

  • RLVR 2: Adjusts the mixing ratios of different data sources based on RLVR 1, totaling 156,278 samples.

  • RLVR 3: Further adjusts data proportions, adding more data related to agent capabilities, function calling, and reinforcement learning, totaling 107,037 samples.

  • SWE 1: Mainly composed of R2E-Gym-Subset (79.70%) and SWE-Gym (20.30%), totaling 50,661 samples.

  • SWE 2: Mainly composed of R2E-Gym-Subset (81.18%) and SWE-Gym (18.82%), totaling 1,444 samples.

  • RLHF: Mainly composed of Nemotron-RLHF-GenRM-v1 (77.00%), Nemotron-RL-Agentic-Conversational-Tool-Use-v1 (20.00%), and Nemotron-RL-Identity-Following-v1 (3.00%), totaling 25,171 samples.

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp