Command Palette
Search for a command to run...
Grug-think Agent Reasoning Simplified Dataset
Grug-think is a simplified dataset for agent inference, designed to optimize agent inference efficiency before tool invocation. Its aim is to reduce agent inference token consumption while maintaining the accuracy of external interactions and tool invocation formats. This dataset contains 100,891 real agent execution trajectories from five large open datasets, covering 45% of code-based intelligent agent tasks and 55% of general API tool calls. Some samples intentionally omit tool calls to train the model's autonomous judgment capabilities. The dataset retains the complete original message structure, tool definitions, and tool call content, rewriting only the inference part. A very concise inference text is inserted before each response statement, resulting in a median of only 11 words for the optimized inference text.
Dataset composition:
- glaive-function-calling-v2: Approximately 40,800 basic API tool call dialogues.
- hermes-function-calling-v1: Approximately 6,200 multi-turn tool usage dialogues.
- ToolACE: Approximately 8,900 API calls for multiple tools.
- SWE-smith trajectories (tool & ticks): Approximately 30,000 real-world code debugging agent trajectories.
- Nebius SWE-agent trajectories: Approximately 15,000, code fixes agent trajectories.
Data fields:
- id: Unique identifier for the sample
- source: Data source identifier
- tools: List of available tool definitions
- message: A complete sequence of dialogue messages, each message organized according to the structure of role, content, and tool_calls.
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.