Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

An Image is Worth 32 Tokens for Reconstruction and Generation

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

FIFO-Diffusion: Generating Infinite Videos from Text without Training

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

OmniFusion Technical Report

Machine learning prediction errors better than DFT accuracy

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search

Representation Shift: Unifying Token Compression with FlashAttention

CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward

LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Automated Algorithmic Discovery for Gravitational-Wave Detection Guided by LLM-Informed Evolutionary Monte Carlo Tree Search

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report

CellForge: Agentic Design of Virtual Cell Models

SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution

Multimodal Referring Segmentation: A Survey

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding

SWE-Exp: Experience-Driven Software Issue Resolution

PixNerd: Pixel Neural Field Diffusion

Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models

Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training

Co-Producing AI: Toward an Augmented, Participatory Lifecycle

An Image is Worth 32 Tokens for Reconstruction and Generation

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

FIFO-Diffusion: Generating Infinite Videos from Text without Training

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

OmniFusion Technical Report

Machine learning prediction errors better than DFT accuracy

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search

Representation Shift: Unifying Token Compression with FlashAttention

CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward

LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Automated Algorithmic Discovery for Gravitational-Wave Detection Guided by LLM-Informed Evolutionary Monte Carlo Tree Search

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report

CellForge: Agentic Design of Virtual Cell Models

SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution

Multimodal Referring Segmentation: A Survey

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding

SWE-Exp: Experience-Driven Software Issue Resolution

PixNerd: Pixel Neural Field Diffusion

Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models

Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training

Co-Producing AI: Toward an Augmented, Participatory Lifecycle