Command Palette
Search for a command to run...
Reasoning Corpus Large Model Mind Chain Dataset
Reasoning Corpus is a large model thinking dataset released by SupraLabs, primarily used for supervised fine-tuning (SFT), model distillation, and instruction training. This dataset contains approximately 5 million samples, with each sequence length limited to 5,000 tokens. Each sample includes the original user prompt, inference trajectory, final assistant response, source information, estimated token length, and a pre-formatted ChatML representation. The data originates from over 60 public inference data warehouses, covering multiple fields such as science, mathematics, coding, logical reasoning, finance and economics, medicine, and multilingual STEM. It integrates thought chain data generated by mainstream large models such as DeepSeek-v4, DeepSeek-r1, Qwen3, and Gemma4.
Data fields:
- repo_id: Identifier of the upstream data source repository
- tok_len: Pre-computed estimate of the sample token length
- user: Original user prompt or task command
- thought_trace: The intermediate reasoning thought chain generated by the model
- assistant: The final answer derived from the reasoning process.
- ChatML: Preformatted standard ChatML formatted text
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.