Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
3,031 papers

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation.

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

Causal Foundation Models

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

WorldSculpt: Generating Compositional Worlds from Grounded Videos

Iris: Climbing to the Search Frontier

MOTION-OMNI: END-TO-END JOINT SPEECH AND FULL-BODY MOTION FOR SPOKEN DIALOGUE

The Atention Triangle in Audio-Video Models

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

Dr. Claw: An AI Scientist Workspace for Vibe Research

WorldReward: Reward Modeling for Camera-Conditioned World Models

VIBEVOICE-ASR-STREAMING Technical Report

Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency

ROBOTOK: AN INTERNET-SCALE DATA ENGINE FOR HUMAN DEMONSTRATION RETRIEVAL AND DEXTEROUS MANIPULATION LEARNING

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

Random Attention: Rethinking KV Cache Eviction for Eficient Reasoning

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

Language Models Can Control Their Own Attention

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

H3-World: Turning Language Understanding into World Control

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

UI-Venus-2: A Large Language Model for GUI Agents

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

StudentSim: Training LLM-based Student Simulators