Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

Evaluating Language Models for Harmful Manipulation

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control






























Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
MENDEL GÖDEL MACHINE: RECURSIVE SELF-IMPROVING CODING AGENTS VIA COMPARATIVE EVOLUTION
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
DeepTutor: Towards Agentic Personalized Tutoring
Flow-by-Flow: Content-Judgment Bypass for Governing AI Output in High-Loss Domains
Towards an Argumentative Foundation for Evaluative AI
Data-Driven Fire-Zone Segmentation for Improved Short-Term Wildfire Prediction
Application of Artificial Intelligence for Fraudulent Banking Operations Recognition
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Evidence-RL: Towards Evidence-intensive Visual Reasoning
Skaling: Chinchilla’s Exponents Meet Kaplan’s Coupling
PROTECT-90: A Fault Dataset for Power System Protection
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning Paradigms for LLMs
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
ENVACE: INTERNALIZING ENVIRONMENT DYNAMICS VIA WORLD REHEARSAL FOR AGENTIC REINFORCEMENT LEARNING
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
WorldClaw: Agentic 3D Open-World Generation at Scale
INTERPRETABLE MEG DECODING OF PERCEIVED SPEECH: CORTICAL SOURCES AND THE STIMULUS FEATURES THAT DRIVE RETRIEVAL
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
AGENTOPSD: RECURSIVE SELF-DISTILLATION FOR AGENTIC REINFORCEMENT LEARNING