Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Autonomous de novo protein binder design with Claude






























Agentic Transaction: Towards ACID-Compliant Agent Systems
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
ClawGym II: Exploring Black-Box RL on Agent Harness
MOSS-VL Technical Report
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
VIBEWORLDING: CAN MULTIMODAL AGENTS CONSTRUCT 3D OPEN WORLDS END-TO-END?
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation
Adversarial Learning of Classifier-Free Guidance Schedules
Agent-Orchestration in Autonomous Chip Design
HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation
Demonstration of Space Robot Teleoperation over a Lossy and Delayed Network using ATMOS
CAKE: Compiler–Agent Co-Design for Frontier Kernel Evolution
Training AI Scientists to Replicate Research
Small-Scale Experiments: Are We There Yet?
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Intern-S2-Preview: Scientific Agentic Foundation Model
DarwinX: Evolving Agent Harnesses Through Natural Selection
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
LLMROUTER: UNIFIED INFRASTRUCTURE FOR DEVELOPING, EVALUATING, AND DEPLOYING LLM ROUTERS
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
WizardLM: EMPOWERING LARGE PRE-TRAINED LANGUAGE MODELS TO FOLLOW COMPLEX INSTRUCTIONS
Indoor Segmentation and Support Inference from RGBD Images