Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

Visual Spatial Tuning

Too Good to be Bad: On the Failure of LLMs to Role-Play Villains































Visual Spatial Tuning

Too Good to be Bad: On the Failure of LLMs to Role-Play Villains






























DeepEyesV2: Toward Agentic Multimodal Model
Use of Continuous Glucose Monitoring with Machine Learning to Identify Metabolic Subphenotypes and Inform Precision Lifestyle Changes
Reusing Pre-Training Data at Test Time is a Compute Multiplier
NVIDIA Nemotron Nano V2 VL
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
Cambrian-S: Towards Spatial Supersensing in Video
Scaling Agent Learning via Experience Synthesis
V-Thinker: Interactive Thinking with Images
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
Recent Developments in Amber Biomolecular Simulations
UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
Step-Audio-EditX Technical Report
LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
Diffusion Language Models are Super Data Learners
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
Dynamic Population Distribution Aware Human Trajectory Generation with Diffusion Model
Text to Robotic Assembly of Multi Component Objects using 3D Generative AI and Vision Language Models
Kosmos: An AI Scientist for Autonomous Discovery
Shorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVR
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
DeepEyesV2: Toward Agentic Multimodal Model
Use of Continuous Glucose Monitoring with Machine Learning to Identify Metabolic Subphenotypes and Inform Precision Lifestyle Changes
Reusing Pre-Training Data at Test Time is a Compute Multiplier
NVIDIA Nemotron Nano V2 VL
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
Cambrian-S: Towards Spatial Supersensing in Video
Scaling Agent Learning via Experience Synthesis
V-Thinker: Interactive Thinking with Images
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
Recent Developments in Amber Biomolecular Simulations
UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
Step-Audio-EditX Technical Report
LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
Diffusion Language Models are Super Data Learners
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
Dynamic Population Distribution Aware Human Trajectory Generation with Diffusion Model
Text to Robotic Assembly of Multi Component Objects using 3D Generative AI and Vision Language Models
Kosmos: An AI Scientist for Autonomous Discovery
Shorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVR
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation