Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

iLRM: An Iterative Large 3D Reconstruction Model

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models































iLRM: An Iterative Large 3D Reconstruction Model

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models






























C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
RecGPT Technical Report
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
The Outcome of the 2022 Landslide4Sense Competition: Advanced Landslide Detection from Multi-Source Satellite Imagery
Less is More for Synthetic Speech Detection in the Wild
Solution-aware vs global ReLU selection: partial MILP strikes back for DNN verification
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
BANG: Dividing 3D Assets via Generative Exploded Dynamics
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
MIRepNet: A Pipeline and Foundation Model for EEG-Based Motor Imagery Classification
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data
Toward long-range ENSO prediction with an explainable deep learning model
OmniArch: Building Foundation Model for Scientific Computing
VA-MoE: Channel-Adapted MoE for Incremental Weather Forecasting
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
DualSG: A Dual-Stream Explicit Semantic-Guided Multivariate Time Series Forecasting Framework
When Tokens Talk Too Much: A Survey of Multimodal Long-Context Token Compression across Images, Videos, and Audios
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
Reconstructing 4D Spatial Intelligence: A Survey
Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning
C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
RecGPT Technical Report
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
The Outcome of the 2022 Landslide4Sense Competition: Advanced Landslide Detection from Multi-Source Satellite Imagery
Less is More for Synthetic Speech Detection in the Wild
Solution-aware vs global ReLU selection: partial MILP strikes back for DNN verification
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
BANG: Dividing 3D Assets via Generative Exploded Dynamics
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
MIRepNet: A Pipeline and Foundation Model for EEG-Based Motor Imagery Classification
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data
Toward long-range ENSO prediction with an explainable deep learning model
OmniArch: Building Foundation Model for Scientific Computing
VA-MoE: Channel-Adapted MoE for Incremental Weather Forecasting
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
DualSG: A Dual-Stream Explicit Semantic-Guided Multivariate Time Series Forecasting Framework
When Tokens Talk Too Much: A Survey of Multimodal Long-Context Token Compression across Images, Videos, and Audios
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
Reconstructing 4D Spatial Intelligence: A Survey
Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning