Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends

Rethinking the Evaluation of Harness Evolution for Agents

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning































Rethinking the Evaluation of Harness Evolution for Agents

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning






























Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
Towards Autonomous and Auditable Medical Imaging Model Development
MUSCRIPTOR: AN OPEN MODEL FOR MULTI-INSTRUMENT MUSIC TRANSCRIPTION
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
The Role of Rigor in Artificial Intelligence
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional Pruning
LLM-Guided Program Evolution for Targeted Black-Box Attacks on Perceptual Hash Algorithms
Are LLMs ready for HARDCHOICES?
Prezta: Provable Remote Execution of Zero-Trust Authorization using SNARKs
Score-Only Distillation for Compact Dense Retrieval
FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
ManiScope: LLM-Assisted Visual Analytics of Cryptocurrency Manipulation Risk
Event-based Neural Decoding for Neuroprosthetic Motor Control
Unlocking Every Expert in Domain-Specific Training
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
PithTrain: A Compact and Agent-Native MoE Training System
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
KronQ: LLM Quantization via Kronecker-Factored Hessian
Trust Region Policy Distillation
Video Generation Models are General-Purpose Vision Learners
Scalable Visual Pretraining for Language Intelligence
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
Towards Autonomous and Auditable Medical Imaging Model Development
MUSCRIPTOR: AN OPEN MODEL FOR MULTI-INSTRUMENT MUSIC TRANSCRIPTION
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
The Role of Rigor in Artificial Intelligence
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional Pruning
LLM-Guided Program Evolution for Targeted Black-Box Attacks on Perceptual Hash Algorithms
Are LLMs ready for HARDCHOICES?
Prezta: Provable Remote Execution of Zero-Trust Authorization using SNARKs
Score-Only Distillation for Compact Dense Retrieval
FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
ManiScope: LLM-Assisted Visual Analytics of Cryptocurrency Manipulation Risk
Event-based Neural Decoding for Neuroprosthetic Motor Control
Unlocking Every Expert in Domain-Specific Training
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
PithTrain: A Compact and Agent-Native MoE Training System
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
KronQ: LLM Quantization via Kronecker-Factored Hessian
Trust Region Policy Distillation
Video Generation Models are General-Purpose Vision Learners
Scalable Visual Pretraining for Language Intelligence
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL