Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences































Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences






























TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs
LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
Predicting LLM Safety Before Release by Simulating Deployment
FastContext: Training Efficient Repository Explorer for Coding Agents
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
DreamX-World 1.0: A General-Purpose Interactive World Model
Geometric Action Model for Robot Policy Learning
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
dots.tts Technical Report
Deterministic Video Depth Estimation with Generative Priors
Galaxy Image Deconvolution for Weak Gravitational Lensing with Unrolled Plug-and-Play ADMM
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
Agents of Chaos
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
Orchestra-o1: Omnimodal Agent Orchestration
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents
APPO: Agentic Procedural Policy Optimization
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
InterleaveThinker: Reinforcing Agentic Interleaved Generation
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
WEAVEBENCH: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
MiniMax Sparse Attention
TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs
LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
Predicting LLM Safety Before Release by Simulating Deployment
FastContext: Training Efficient Repository Explorer for Coding Agents
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
DreamX-World 1.0: A General-Purpose Interactive World Model
Geometric Action Model for Robot Policy Learning
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
dots.tts Technical Report
Deterministic Video Depth Estimation with Generative Priors
Galaxy Image Deconvolution for Weak Gravitational Lensing with Unrolled Plug-and-Play ADMM
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
Agents of Chaos
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
Orchestra-o1: Omnimodal Agent Orchestration
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents
APPO: Agentic Procedural Policy Optimization
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
InterleaveThinker: Reinforcing Agentic Interleaved Generation
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
WEAVEBENCH: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
MiniMax Sparse Attention