Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

MultiRef: Controllable Image Generation with Multiple Visual References

Prompt Orchestration Markup Language































MultiRef: Controllable Image Generation with Multiple Visual References

Prompt Orchestration Markup Language






























LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
HPSv3: Towards Wide-Spectrum Human Preference Score
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
Evaluating Identity Leakage in Speaker De-Identification Systems
Next Visual Granularity Generation
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning
An integrated microwave neural network for broadband computation and communication
GTool: Graph Enhanced Tool Planning with Large Language Model
Observation of dendrite formation at Li metal-electrolyte interface by a machine-learning enhanced constant potential framework
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing
DINOv3
SSRL: Self-Search Reinforcement Learning
Thyme: Think Beyond Images
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation
Puppeteer: Rig and Animate Your 3D Models
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
HPSv3: Towards Wide-Spectrum Human Preference Score
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
Evaluating Identity Leakage in Speaker De-Identification Systems
Next Visual Granularity Generation
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning
An integrated microwave neural network for broadband computation and communication
GTool: Graph Enhanced Tool Planning with Large Language Model
Observation of dendrite formation at Li metal-electrolyte interface by a machine-learning enhanced constant potential framework
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing
DINOv3
SSRL: Self-Search Reinforcement Learning
Thyme: Think Beyond Images
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation
Puppeteer: Rig and Animate Your 3D Models
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving