Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

Step-Audio 2 Technical Report

Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning































Step-Audio 2 Technical Report

Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning






























Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
Uncertainty-Aware Knowledge Transformers for Peer-to-Peer Energy Trading with Multi-Agent Reinforcement Learning
NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling
WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
The Invisible Leash: Why RLVR May Not Escape Its Origin
GUI-G^2: Gaussian Reward Modeling for GUI Grounding
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Design of intrinsically disordered region binding proteins
An All-Atom Generative Model for Designing Protein Complexes
RedOne: Revealing Domain-specific LLM Post-Training in Social Networking Services
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models
The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
PrefPalette: Personalized Preference Modeling with Latent Attributes
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
π^3: Scalable Permutation-Equivariant Visual Geometry Learning
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
A Survey of Context Engineering for Large Language Models
Assessing adaptive world models in machines with novel games
Emotional Support with LLM-based Empathetic Dialogue Generation
DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
MOSPA: Human Motion Generation Driven by Spatial Audio
MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
Uncertainty-Aware Knowledge Transformers for Peer-to-Peer Energy Trading with Multi-Agent Reinforcement Learning
NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling
WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
The Invisible Leash: Why RLVR May Not Escape Its Origin
GUI-G^2: Gaussian Reward Modeling for GUI Grounding
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Design of intrinsically disordered region binding proteins
An All-Atom Generative Model for Designing Protein Complexes
RedOne: Revealing Domain-specific LLM Post-Training in Social Networking Services
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models
The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
PrefPalette: Personalized Preference Modeling with Latent Attributes
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
π^3: Scalable Permutation-Equivariant Visual Geometry Learning
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
A Survey of Context Engineering for Large Language Models
Assessing adaptive world models in machines with novel games
Emotional Support with LLM-based Empathetic Dialogue Generation
DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
MOSPA: Human Motion Generation Driven by Spatial Audio
MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding