Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

Towards Predictive, Aligned, and Scalable Robot Learning

Flow Matching in Feature Space for Stochastic World Modeling

Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit

TRACE: TURN-LEVEL REWARD ASSIGNMENT VIA CREDIT ESTIMATION FOR LONG-HORIZON AGENTS

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

BadWAM: When World-Action Models Dream Right but Act Wrong

SearchOS-V1 : Towards Robust Open-Domain Information-Seeking Agent Collaboration

SEED: SELF-EVOLVING ON-POLICY DISTILLATION FOR AGENTIC REINFORCEMENT LEARNING

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

Deep Learning in Remote Sensing: A Review

A Regression Approach to Speech Enhancement Based on Deep Neural Networks

Deep Neural Networks for Acoustic Modeling in Speech Recognition

RoboTTT: Context Scaling for Robot Policies

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Efficient Estimation of Word Representations in Vector Space

Depth Map Prediction from a Single Image using a Multi-Scale Deep Network

TabNet: Attentive Interpretable Tabular Learning

AudioPaLM: A Large Language Model That Can Speak and Listen

SQuAD: 100,000+ Questions for Machine Comprehension of Text

DeepPose: Human Pose Estimation via Deep Neural Networks

Self-Improvements in Modern Agentic Systems: A Survey

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

OvisOCR2 Technical Report

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable

Qwen-Music Technical Report

Spectral Rewiring for Exploration, Purification, and Model Merging

Towards Predictive, Aligned, and Scalable Robot Learning

Flow Matching in Feature Space for Stochastic World Modeling

Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit

TRACE: TURN-LEVEL REWARD ASSIGNMENT VIA CREDIT ESTIMATION FOR LONG-HORIZON AGENTS

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

BadWAM: When World-Action Models Dream Right but Act Wrong

SearchOS-V1 : Towards Robust Open-Domain Information-Seeking Agent Collaboration

SEED: SELF-EVOLVING ON-POLICY DISTILLATION FOR AGENTIC REINFORCEMENT LEARNING

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

Deep Learning in Remote Sensing: A Review

A Regression Approach to Speech Enhancement Based on Deep Neural Networks

Deep Neural Networks for Acoustic Modeling in Speech Recognition

RoboTTT: Context Scaling for Robot Policies

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Efficient Estimation of Word Representations in Vector Space

Depth Map Prediction from a Single Image using a Multi-Scale Deep Network

TabNet: Attentive Interpretable Tabular Learning

AudioPaLM: A Large Language Model That Can Speak and Listen

SQuAD: 100,000+ Questions for Machine Comprehension of Text

DeepPose: Human Pose Estimation via Deep Neural Networks

Self-Improvements in Modern Agentic Systems: A Survey

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

OvisOCR2 Technical Report

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable

Qwen-Music Technical Report

Spectral Rewiring for Exploration, Purification, and Model Merging