Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

Atomic Task Graph: A Unified Framework for Agentic Planning and Execution

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models































Atomic Task Graph: A Unified Framework for Agentic Planning and Execution

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models






























UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
Why Can’t I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Video-Oasis: Rethinking Evaluation of Video Understanding
Vidu S1: A Real-Time Interactive Video Generation Model
Measuring the Gap Between Human and LLM Research Ideas
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Infinite Worlds with Versatile Interactions
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
LAME M-VLA: DUAL LATENT MEMORY IN VISION-LANGUAGE-ACTION MODELS FOR ROBOTIC MANIPULATION
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
Vision as Unified Multimodal Generation
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
AlayaWorld: Long-Horizon and Playable Video World Generation
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs
Multi-Turn On-Policy Distillation with Prefix Replay
Gemma 4 Technical Report
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
Wan-Streamer v0.2: Higher Resolution, Same Latency
EVA-Client: A Unified Framework for Deployment, Evaluation, and Data Collection on Real Robots
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
FINAL Bench: Measuring Functional Metacognitive Reasoning in Large Language Models
SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes
TheoremGraph: Bridging Formal and Informal Mathematics
Always-On Agents: A Survey of Persistent Memory, State, and Governance in LLM Agents
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
Why Can’t I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Video-Oasis: Rethinking Evaluation of Video Understanding
Vidu S1: A Real-Time Interactive Video Generation Model
Measuring the Gap Between Human and LLM Research Ideas
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Infinite Worlds with Versatile Interactions
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
LAME M-VLA: DUAL LATENT MEMORY IN VISION-LANGUAGE-ACTION MODELS FOR ROBOTIC MANIPULATION
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
Vision as Unified Multimodal Generation
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
AlayaWorld: Long-Horizon and Playable Video World Generation
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs
Multi-Turn On-Policy Distillation with Prefix Replay
Gemma 4 Technical Report
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
Wan-Streamer v0.2: Higher Resolution, Same Latency
EVA-Client: A Unified Framework for Deployment, Evaluation, and Data Collection on Real Robots
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
FINAL Bench: Measuring Functional Metacognitive Reasoning in Large Language Models
SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes
TheoremGraph: Bridging Formal and Informal Mathematics
Always-On Agents: A Survey of Persistent Memory, State, and Governance in LLM Agents