HyperAIHyperAI

Command Palette

Search for a command to run...

MiniCPM-RobotManip, With Its 1.5 Billion Parameters, Is Officially Open-sourced; NVIDIA Releases the full-modal Model Cosmos3-Edge, Covering Text, Images, Videos, and Motion sequences.

Featured Image

In the field of embodied intelligence, enabling robots to perceive, understand, and act is key to propelling intelligent agents into the real world. MiniCPM's MiniCPM-Robot series explores two core tasks: robot manipulation and target tracking.MiniCPM-RobotManip is a visual-language action model with only 1.5 billion parameters. Based on the MiniCPM-V 4.6 visual-language backbone and combined with a diffused action head, it achieves end-to-end mapping from visual observation and language commands to robot actions.MiniCPM-RobotTrack focuses on entity target tracking, achieving real-time edge visual tracking capabilities with 0.9B parameters. Through lightweight architecture, streaming context memory, and efficient visual compression technology, the MiniCPM-Robot series demonstrates the potential for efficient inference in real-world robotic scenarios using small-scale models.

The HyperAI website now features "MiniCPM-Robot: Embodied Intelligent Models for the Real World," so come and try it out!

Online use:https://go.hyper.ai/0VsYI

A quick overview of hyper.ai's official website updates from July 25th to July 30th:

* High-quality public datasets: 9

* High-quality tutorial selection: 8

* Community article analysis: 2 articles

* Popular encyclopedia entries: 5

Visit the official website:hyper.ai

Selected public datasets

1.Math-Graph Mathematical Theorem Dependency Graph Dataset

Math-Graph is a dataset of mathematical theorem dependency graphs released in 2026 by the Mathematical Artificial Intelligence Laboratory at the University of Washington. It constructs a mathematical theorem dependency graph at the statement level. The related paper is titled TheoremGraph: Bridging Formal and Informal Mathematics.

This dataset contains 47,952 pairs of informal and formal mathematical statements, integrating the two types of mathematical knowledge into a coherent dependency graph that covers both informal and formal mathematics. It includes paper metadata, mathematical statements (formal/informal), dependency edges, and LLM-generated natural language descriptions.

Run online:https://go.hyper.ai/bwVQf

2.Nemotron-SFT-Math-v4 Mathematical Inference SFT Dataset

Nemotron-SFT-Math-v4 is a mathematical reasoning dataset released by NVIDIA in May 2026. The related paper is titled Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision. It aims to solve the problems of inconsistent quality, non-standard reasoning trajectories, low accuracy, and limited scenarios in traditional mathematical datasets. It effectively improves the model's structured reasoning, multi-trajectory reasoning, and answer verification capabilities. It is widely used for fine-tuning large-scale mathematical reasoning models, reasoning trajectory analysis, answer verification algorithm development, long-context reasoning system construction, and model reasoning robustness evaluation. 

This dataset contains 545,431 training samples, including 285,516 COT reasoning samples and 259,915 TIR tool reasoning samples. It covers mathematical scenarios in competitions and university research in algebra, geometry, number theory, combinatorics, etc. The data is annotated using a hybrid manual and automated method and includes standardized fields such as unique number, question text, multi-turn dialogue, standard answer, source, and protocol.

Run online:https://go.hyper.ai/Kv9E4

3.MathNet Multimodal Mathematical Benchmark Inference Dataset

MathNet is a large-scale, multilingual, multimodal mathematical reasoning dataset released in 2026 by a team from MIT in collaboration with King Abdullah University of Science and Technology and other institutions. The related paper is titled "MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval."It aims to evaluate and improve the capabilities of large models in Olympic-level mathematical reasoning and structured retrieval tasks, and is widely used in mathematical reasoning evaluation, RAG research, and multimodal AI training.

This dataset, version v0, contains 27,817 expert-level math problems and their standard solutions. It covers official math competition problems from 58 countries and regions in 17 languages, including 5,148 illustrated problems with a total of 7,541 geometric and graphical illustrations. The dataset covers algebra, geometry, number theory, combinatorics, calculus, probability and statistics, and other Olympiad math knowledge systems. It supports three benchmark tasks: solving math problems, mathematical semantic retrieval (identifying structurally equivalent and similar problems), and retrieval enhancement problem solving.

Run online:https://go.hyper.ai/AWKaV

Dataset Example

4Nemotron-Math-v2 Mathematical Inference Dataset

Nemotron-Math-v2 is a mathematical reasoning dataset released by NVIDIA Corporation in 2025. Related research is titled "Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision." It is primarily used to train LLMs to perform structured mathematical reasoning, to study the differences between tool-enhanced reasoning and pure language reasoning, and to build long-context or multi-path reasoning systems.

This dataset contains approximately 347,000 high-quality mathematical problems and 7 million model-generated inference trajectories. Each problem is solved in six configurations: high/medium/low inference depth and with or without Python TIR, and the answers are validated via a pipeline using an LLM as the arbiter.

Run online:https://go.hyper.ai/rYdfW

5. CHIMERA General Inference Synthetic Dataset

CHIMERA is a synthetic reasoning dataset designed specifically for reasoning training. The related paper is titled CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning.

This dataset covers a wide range of STEM subjects and provides Long Chain Thinking (CoT) trajectories, containing 9,225 questions across 8 subjects (mathematics, computer science, chemistry, physics, literature, history, biology, and phonetics). All examples are generated by a large language model (LLM) and are automatically validated without manual annotation.

Run online:https://go.hyper.ai/JskxG

6. Open-RL Inference Problem Dataset

Open-RL is a multi-domain reasoning problem dataset released by Turing in 2026, containing independent, verifiable, and explicit STEM reasoning problems in physics, mathematics, biology, and chemistry. Each problem requires multi-step reasoning, involves symbolic operations and/or numerical computation, and has an objectively verifiable final answer.

This dataset is suitable for fine-tuning reinforcement learning, reward modeling, outcome-supervised training, and verifiable inference benchmarking.Each problem requires multiple steps of reasoning and involves symbolic operations and numerical computation, with a verifiable final answer.

Run online:https://go.hyper.ai/WYDck

7. GPT-5.4 Stepwise Inference Dataset

The GPT-5.4 step-by-step reasoning dataset is a high-density synthetic reasoning dataset designed for long-chain reasoning (CoT) modeling and complex problem-solving tasks. The dataset is built on a "Master-Architect" workflow, using gemini 3 flash to generate challenging hints, and combining this with the GPT-5.4 Reasoning Core to complete multi-step reasoning and result generation.

This dataset contains approximately 1,500 elite-level samples, covering highly complex fields such as mathematics, programming, and medicine. The task difficulty is uniformly set at the "Grandmaster" and "Beyond-PhD" levels. The data samples include complete multi-step reasoning processes and verified final solutions, making them suitable for scenarios involving long-chain logical derivation and assessment of extreme reasoning abilities.

Run online:https://go.hyper.ai/CeBEp

8.Sutra 10B Pretraining Teaching and Training Dataset

Sutra 10B Pretraining is a high-quality teaching dataset for pretraining large language models. Generated by the Sutra framework, it creates structured educational content and optimizes the pretraining of language models. This is the largest dataset in the Sutra series, designed to demonstrate how dense, well-curated datasets can provide optimal pretraining performance for small language models.

This dataset contains 10,193,029 teaching records, totaling over 10 billion tokens, covering nine major areas: interdisciplinary, technology, science, social studies, mathematics, life skills, arts and creativity, language arts, and philosophy and ethics. The data follows a well-established teaching paradigm, with 10 levels of difficulty from basic to advanced, demonstrating good hierarchy and systematic organization.

Run online:https://go.hyper.ai/skT3q

9. VisCoR-55K Visual Inference Dataset

VisCoR-55K is a high-quality visual reasoning dataset released in 2026 by Huazhong University of Science and Technology in collaboration with Alibaba Cloud. This dataset contains approximately 55,000 visual reasoning samples, each of which generates a corresponding reasoning process using contrasting samples.A high-quality visual reasoning dataset covering five categories: general, reasoning, mathematical, graphing, and OCR.The aim is to promote research on visual language models in reliable and robust visual reasoning.

Run online:https://go.hyper.ai/kIlWk

Selected Public Tutorials

1. MiniCPM-Robot: An Embodied Intelligence Model for the Real World

MiniCPM-Robot is a family of embodied intelligence models from OpenBMB for real-world perception, decision-making, and action. The initial release includes two models: MiniCPM-RobotManip, a general VLA model for robot manipulation, and MiniCPM-RobotTrack, an edge model for embodied target tracking.

Run online:https://go.hyper.ai/0VsYI

Demo Page

2. Diamond-1.0: Autoregressive Speech Restoration Model

Diamond is a speech restoration model developed by nineninesix.ai in July 2026. It uses an autoregressive seq2seq architecture to restore degraded audio (low bitrate encoding, background noise, clipping, narrowband) to 44.1 kHz near-studio quality speech.

Run online:https://go.hyper.ai/OHZ6o

Demo Page

3. Nanbeige4.2-3B: Compact Intelligent Agent Model

Nanbeige4.2-3B is a compact agent model developed by Nanbeige in July 2026, built on Nanbeige4.2-3B-Base. This model employs a Looped Transformer architecture, reusing Transformer layers to increase model capacity without increasing the number of parameters—using only 3B of non-embedded parameters (out of a total of 4B), it surpasses larger models such as Qwen3.5-9B and Gemma4-12B in agent benchmark tests.

Run online:https://go.hyper.ai/KI1QR

Demo Page

4. Fara 1.5-27B: Multimodal Computers Using Intelligent Agents

Fara1.5-27B is a multimodal computer agent released by Microsoft Research AI Frontiers in May 2026. Its core innovation is a vision-only perception architecture: it observes the browser through screenshots and completes end-to-end tasks using structured tool calls (clicking, typing, scrolling, web page searching, etc.) without accessing the DOM. This model is based on Qwen3.5-27B and fine-tuned under supervision, with training data generated from the FaraGen1.5 multi-agent process.

Run online:https://go.hyper.ai/9AwuJ

Demo Page

5. NVIDIA Cosmos3-Edge: Full-Modal Inference Model

Cosmos3-Edge is a 4B-parameter multimodal world foundation model released by the NVIDIA team in July 2026. Its core innovation and feature is the adoption of a hybrid Transformer (MoT) architecture, which integrates autoregressive Transformer and diffusion Transformer within a unified framework to handle text generation and multimodal generation such as images/videos/actions, respectively.

Run online:https://go.hyper.ai/yB4bz

Demo Page

6. Mage-Flow: A fundamental model for efficient native resolution image generation and editing

Mage-Flow is a compact 4B image generation and editing foundation model released by Microsoft in July 2026. Through a carefully designed tokenizer-backbone-system co-optimization, it achieves quality comparable to larger models (Qwen-Image 20B, FLUX.2 32B) at a 4B parameter scale, while maintaining faster inference speed and lower memory footprint.

Run online:https://go.hyper.ai/3cw36

Demo Page

7. LingBot-World-V2: Image-to-Video Generation with Infinite Worlds and Diverse Interactions

LingBot-World-V2 is an image-to-video generation model released by the Robbyant team in July 2026. This model achieves infinitely long interactions through causal pre-training while maintaining stable output quality; it also supports real-time 720p 60fps video generation and possesses diverse interactive capabilities such as attacking, archery, and spellcasting. Furthermore, LingBot-World-V2 is the first to introduce an intelligent agent framework into the world model, achieving more intelligent scene understanding and interaction generation through behavior planning agents and environment synthesis agents.

Run online:https://go.hyper.ai/wecWC

Demo Page

8. MOSS-Transcribe-Diarize: End-to-end multi-speaker audio transcription and speaker separation

MOSS-Transcribe-Diarize is an end-to-end multi-speaker audio understanding model released by the MOSI.AI (OpenMOSS) team in July 2026. With just one inference, the model can simultaneously complete speech transcription and speaker separation, outputting structured text with timestamps and anonymous speaker labels (such as [S01], [S02]).

Run online:https://go.hyper.ai/TElI9

Demo Page

💡We have also established a Stable Diffusion tutorial exchange group. Welcome friends to scan the QR code and remark [SD tutorial] to join the group to discuss various technical issues and share application results~

Community article interpretation

1. Argonne National Laboratory in the United States proposed ChemGraph, which uses 13 benchmark tests to evaluate the value of agents in the field of computational chemistry.

Argonne National Laboratory has launched ChemGraph, an LLM-powered intelligent agent specifically designed to automate computational chemistry molecular simulation workflows. Benchmark tests show that large models such as GPT-4o perform stably on complex tasks; however, by strategically breaking down tasks, smaller models (such as GPT-4o-mini) can also significantly improve performance, reaching or even surpassing the performance of larger models. This research provides a new approach for low-cost automation of AI in scientific research.

View the full report:https://go.hyper.ai/Cgo56

2. Based on large model inference and the use of the MCP tool, AI X-ray scientists at Stanford University autonomously completed single-crystal diffraction alignment at a synchrotron radiation source.

A research team from Stanford University and SLAC National Accelerator Laboratory has developed an "AI X-ray Scientist" system and deployed it on a real beamline at the Stanford synchrotron radiation source. The team first debugged the system in a virtual beamline before migrating it to the real equipment, using the task of solving a single-crystal orientation matrix to test whether the AI could autonomously complete multiple experimental steps and handle unexpected deviations in a real environment.

View the full report:https://go.hyper.ai/jF9qk

Popular Encyclopedia Articles

1. Human-machine loop (HITL)

2. Automatic Speech Recognition (ASR)

3. Optical Character Recognition (OCR)

4. World Action Model WAM

5. The Thinking-Driven Reinforcement Learning Framework (GTR)

Here are hundreds of AI-related terms compiled to help you understand "artificial intelligence" here:

https://go.hyper.ai/wiki

August deadline for the summit

The above is all the content of this week’s editor’s selection. If you have resources that you want to include on the hyper.ai official website, you are also welcome to leave a message or submit an article to tell us!

See you next week!