HyperAIHyperAI

Command Palette

Search for a command to run...

Text/LaTeX/HTML Tables All in One Step! OvisOCR2 Achieves end-to-end Intelligent Document Parsing; 14,000+ Element Annotations, Tens of Thousands of Language Commands! Voxel51 Releases the SceneFun3D Indoor Scene micro-interaction dataset.

OvisOCR2, launched in July 2026 by ATH-MaaS, Alibaba, and other teams, is a compact end-to-end document parsing model with only 0.8B (approximately 1.7GB) of parameters. Built on Qwen3.5-0.8B, this model can directly convert document images into a structured Markdown format that preserves the natural reading order, and accurately parse text, formulas, tables, and visible areas.With its innovative hybrid data engine and multi-stage (SFT, RL, and OPD) training scheme, OvisOCR2 achieved a score of 96.58 in the OmniDocBench v1.6 benchmark, becoming the first end-to-end model to outperform traditional pipelined methods.It perfectly achieves a balance between extremely low deployment costs and outstanding performance.

The HyperAI website now features the "OvisOCR2: 0.8B End-to-End Document Parsing Model," so give it a try!

Online use:https://go.hyper.ai/1qOnc

Welcome to visit our official website for more information:

https://hyper.ai

A quick overview of hyper.ai's official website updates from July 18th to July 24th:

* High-quality public datasets: 9

* A selection of high-quality tutorials: 9

* Community article analysis: 1 article

* Popular encyclopedia entries: 5

Top conferences with August deadlines: 6

Visit the official website:hyper.ai

Selected public datasets

1. SceneFun3D 3D Scene Function Interaction Dataset

SceneFun3D is a 3D scene functional interaction dataset released by Voxel51 in 2024. It focuses on fine-grained understanding of tiny interactive parts, covering part localization, Gibsonian function inference, motion parameter estimation, and natural language task description. The dataset includes 710 high-resolution real-world indoor scenes, containing 14,867 functional interaction element annotations, 9 function categories, 14,279 motion parameters, and over 10,913 natural language task instructions.

Online use:https://go.hyper.ai/g81mC

2. Vāgdhenu Sanskrit Recitation Corpus Dataset

Vāgdhenu is a corpus of single-person Sanskrit chanting recordings, primarily intended for Sanskrit speech synthesis training, prosody research, and language accessibility applications. The dataset contains 1,467 segments, with a total audio duration of approximately 5.3 hours, in 24 kHz mono WAV format. The dataset includes two independent subsets, style_a and style_b, covering a large number of different verses.

Online use:https://go.hyper.ai/D7Q6H

3. Math-Graph Mathematical Theorem Dependency Graph Dataset

Math-Graph, released in 2026 by the Mathematical Artificial Intelligence Laboratory at the University of Washington, is a dataset of mathematical theorem dependency graphs that constructs a statement-level dependency graph of mathematical theorems. This dataset contains 47,952 matching pairs of informal and formal mathematical statements, integrating these two types of mathematical knowledge into a coherent dependency graph. It covers both informal and formal mathematics, encompassing paper metadata, mathematical statements, dependency edges, and LLM-generated natural language descriptions.

Online use:https://go.hyper.ai/CtIDb

4. GLM-5.2 Agent Interaction Trajectory Dataset

GLM-5.2 Agent is a dataset of raw agent interaction trajectories for GLM-5.2 models generated by TeichAI using Teich. It aims to provide standardized, parsable session records for the training and design of agent models, supporting subsequent fine-tuning and enhanced tool invocation capabilities.

Online use:https://go.hyper.ai/itLRp

5. EVA-Bench end-to-end speech agent benchmark dataset

EVA-Bench, released by ServiceNow, is an end-to-end speech agent evaluation benchmark dataset designed to assess the task completion capabilities and interactive experience of speech agents. The dataset contains 213 scenarios covering three enterprise domains: airline customer service management, healthcare human resources services, and enterprise IT service management. It supports comprehensive evaluation of multi-turn voice interaction, tool invocation, task completion, and robustness of voice scenarios.

Online use:https://go.hyper.ai/51B4l

6. Metacognition-Bench dataset

Metacognition-Bench, released by ginigen-ai, is a benchmark dataset for the metacognitive abilities of large language models. It aims to measure the model's functional metacognitive ability to detect and self-correct its own reasoning errors, rather than simply evaluating the accuracy of the final answer. The dataset contains 300 metacognitive trap questions covering 121 fields, including mathematics, physics, biology, law, medicine, economics, statistics, ethics, and computer science. It encompasses eight metacognitive behavior types: self-correction, trap escape, transformation detection, conflict resolution, incremental discovery, multi-constraint handling, expert panel, and uncertainty decision-making.

Online use:https://go.hyper.ai/8qdca

7. SVG Benchmark: A benchmark dataset for still image generation.

SVG Benchmark is a large-scale model static image generation evaluation dataset released by Rapida in 2025. It aims to compare the ability of 30 state-of-the-art large language models to generate static SVGs based on text prompts through human evaluation. The dataset is collected in a pairwise comparison format, containing 500 English prompts from human-written or publicly available datasets, 188,754 image comparison samples, and 1,355,161 preference responses based on human voting.

Online use:https://go.hyper.ai/13Rn2

8. IFStruct v1.0 Structured Output Compliance Benchmark Dataset

IFStruct v1.0 is a structured output compliance benchmark dataset released by Liquid AI in 2026. It aims to systematically evaluate the ability of large language models to generate structured output conforming to a specified schema. The dataset contains 2,000 prompt words, each requiring the model to generate several instances based on a target pattern. The prompt words are presented in various styles, including natural dialogue, explicit path instructions, raw JSON schema, code block requirements, and no additional text, all accompanied by strict underlying schema validation specifications.

Online use:https://go.hyper.ai/3NUcg

9. EdgeBench Real-World Learning Benchmark Dataset for Intelligent Agents

EdgeBench is a real-world learning benchmark dataset for intelligent agents released by ByteDance Seed in 2026. It aims to evaluate the ability of autonomous AI agents to learn from real-world environments. The dataset contains 134 real-world tasks, 51 of which are open-source, covering six ability categories: scientific computing and machine learning, systems and software engineering, optimization, knowledge reasoning, formal reasoning, and game theory.

Online use:https://go.hyper.ai/FqmX7

Selected Public Tutorials

1. OvisOCR2: 0.8B End-to-End Document Parsing Model

Released by the ATH-MaaS team in July 2026, the OvisOCR2 model is a compact 0.8B end-to-end document parsing model. This model is built by post-training on Qwen3.5-0.8B, employing a carefully designed data engine (combining real and synthetic data) and a multi-stage training strategy integrating SFT, RL, and OPD.

Run online:https://go.hyper.ai/1qOnc

Demo Page

2. RoBERTa Tweet Sentiment Text Extraction

The RoBERTa model was released by the Facebook AI team in July 2019. Its core innovations and features include the introduction of a dynamic masking strategy, larger batch training, and more training data on top of BERT, which significantly improves NLP benchmark performance.

Run online:https://go.hyper.ai/8iZIz

3. handson-ml2 Machine Learning in Practice Tutorial

This is a complete collection of practical machine learning tutorials, covering everything from machine learning basics to deep learning. This tutorial is based on Aurélien Géron's classic book, "Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow" (2nd edition). 

Run online:https://go.hyper.ai/jEyUh

4. Keras Deep Learning Fundamentals Tutorial—Breast Cancer Classification

This tutorial originates from Paras Jindal's introductory deep learning project released on Kaggle in July 2026. It uses the Keras Sequential API to build a neural network model to perform binary classification of benign and malignant breast cancer cell feature data.

Run online:https://go.hyper.ai/s1X9X

Model Summary and Comparison

5. Recommendation Engine (Collective Purchase Recommendation Engine)

The co-purchase recommendation engine is a recommendation system based on Chris Deotte's Kaggle Notebook "Recommend Items Purchased Together," released in February 2022. Its core innovation lies in constructing a co-purchase matrix by analyzing historical transaction data to recommend frequently purchased items together, integrating three strategies: historical purchase frequency, co-purchase pattern, and a backup of popular items. This algorithm achieved a MAP@12 score of 0.021 in the H&M personalized fashion recommendation competition.

Run online:https://go.hyper.ai/fiegd

6. Titanic Data Science Solutions: Learning Machine Learning from Scratch

Titanic – Machine Learning from Disaster is one of the most classic introductory competitions on the Kaggle platform. The task is to predict whether passengers would have survived the sinking of the Titanic based on their personal information (cabin class, gender, age, family size, etc.).

Run online:https://go.hyper.ai/gGUJO

Data analysis

7. Disaster Tweet Classification: Using BERT for NLP Text Classification

This tutorial, based on xhlulu's classic Notebook on Kaggle, demonstrates how to perform binary classification of disaster-related tweets using BERT. Through a complete NLP project, this tutorial guides you through building a BERT text classification pipeline from scratch: data loading → tokenization encoding → model building → training and fine-tuning → inference and prediction. You will learn how to apply pre-trained BERT to custom binary classification tasks and understand the principles behind key design choices such as the [CLS] token and the sigmoid output layer.

Run online:https://go.hyper.ai/ENFgs

8. Ternary-Bonsai-27B: 7.2GB Ternary Quantization Inference 27B

Ternary Bonsai 27B is an inference-based large language model developed by the Prism ML team in July 2026. It is based on Qwen3.6-27B with end-to-end ternary quantization. The model employs a hybrid attention mechanism (~75% linear attention / ~25% full attention), supports ultra-long contexts of up to 262K, and requires only ~14.7 GB of peak memory in a 100K token context. It provides a DSpark speculative decoding acceleration layer, achieving a 1.34x lossless decoding speedup.

Run online:https://go.hyper.ai/OfSLw

Demo Page

9. Lyft Motion Prediction: Autonomous Vehicle Trajectory Prediction Based on ResNeXt50 FPN

Lyft Motion Prediction is a Kaggle competition held by Lyft in 2020. The task is to predict the motion trajectory of an agent (vehicle, pedestrian, etc.) in the next 5 seconds (50 time steps, 10Hz) based on contextual information in an autonomous driving scenario.

Run online:https://go.hyper.ai/vJYkN

Demo Page

💡We have also established a Stable Diffusion tutorial exchange group. Welcome friends to scan the QR code and remark [SD tutorial] to join the group to discuss various technical issues and share application results~

Community article interpretation

1. Based on training with over 10,000 tumor samples, Harvard Medical School and others proposed COMPASS, a pan-cancer basic model, which outperforms 22 existing methods on average.

A research team from Harvard Medical School, Roche Pharmaceuticals, and Zhejiang University has proposed a pan-cancer baseline model, COMPASS. This model was trained on 10,184 tumor samples from 33 cancer types and evaluated in 16 clinical cohorts covering 7 cancer types and 6 immune checkpoint inhibitors. The results show that COMPASS outperforms 22 existing methods on average, improving prediction accuracy by an average of 8.51 TP3T across different cohorts.

View the full report:https://go.hyper.ai/FyRsC

Popular Encyclopedia Articles

1. Skills

2. World Action Model WAM

3. Scaling Law

4. Frames Per Second

5. Automatic Speech Recognition

Here are hundreds of AI-related terms compiled to help you understand "artificial intelligence" here:

https://go.hyper.ai/wiki