Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

Matrix-Game: Interactive World Foundation Model































Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

Matrix-Game: Interactive World Foundation Model






























AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models
Learning Approach to Efficient Vision-based Active Tracking of a Flying Target by an Unmanned Aerial Vehicle
TritonZ: A Remotely Operated Underwater Rover with Manipulator Arm for Exploration and Rescue Operations
ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
RLPR: Extrapolating RLVR to General Domains without Verifiers
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
Light of Normals: Unified Feature Representation for Universal Photometric Stereo
Predicting cellular responses to perturbation across diverse contexts with State
CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
Optimizing Multilingual Text-To-Speech with Accents & Emotions
VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
All is Not Lost: LLM Recovery without Checkpoints
Sundial: A Family of Highly Capable Time Series Foundation Models
ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
Show-o2: Improved Native Unified Multimodal Models
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
s1: Simple test-time scaling
Search-o1: Agentic Search-Enhanced Large Reasoning Models
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models
Learning Approach to Efficient Vision-based Active Tracking of a Flying Target by an Unmanned Aerial Vehicle
TritonZ: A Remotely Operated Underwater Rover with Manipulator Arm for Exploration and Rescue Operations
ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
RLPR: Extrapolating RLVR to General Domains without Verifiers
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
Light of Normals: Unified Feature Representation for Universal Photometric Stereo
Predicting cellular responses to perturbation across diverse contexts with State
CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
Optimizing Multilingual Text-To-Speech with Accents & Emotions
VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
All is Not Lost: LLM Recovery without Checkpoints
Sundial: A Family of Highly Capable Time Series Foundation Models
ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
Show-o2: Improved Native Unified Multimodal Models
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
s1: Simple test-time scaling
Search-o1: Agentic Search-Enhanced Large Reasoning Models
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale