Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection































WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection






























Autonomous Code Evolution Meets NP-Completeness
Reinforcement Learning Foundations for Deep Research Systems: A Survey
Reinforced Visual Perception with Tools
Does DINOv3 Set a New Medical Vision Standard?
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
Reverse-Engineered Reasoning for Open-Ended Generation
OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting
LuxDiT: Lighting Estimation with Video Diffusion Transformer
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
Set Block Decoding is a Language Model Inference Accelerator
Symbolic Graphics Programming with Large Language Models
Why Language Models Hallucinate
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
Recomposer: Event-roll-guided generative audio editing
Transition Models: Rethinking the Generative Learning Objective
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
Towards a Unified View of Large Language Model Post-Training
From Editor to Dense Geometry Estimator
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory
CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
Multi-View 3D Point Tracking
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
On the Theoretical Limitations of Embedding-Based Retrieval
Autonomous Code Evolution Meets NP-Completeness
Reinforcement Learning Foundations for Deep Research Systems: A Survey
Reinforced Visual Perception with Tools
Does DINOv3 Set a New Medical Vision Standard?
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
Reverse-Engineered Reasoning for Open-Ended Generation
OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting
LuxDiT: Lighting Estimation with Video Diffusion Transformer
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
Set Block Decoding is a Language Model Inference Accelerator
Symbolic Graphics Programming with Large Language Models
Why Language Models Hallucinate
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
Recomposer: Event-roll-guided generative audio editing
Transition Models: Rethinking the Generative Learning Objective
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
Towards a Unified View of Large Language Model Post-Training
From Editor to Dense Geometry Estimator
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory
CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
Multi-View 3D Point Tracking
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
On the Theoretical Limitations of Embedding-Based Retrieval