HyperAIHyperAI

Command Palette

Search for a command to run...

يوريكا: تقييم وفهم النماذج الأساسية الكبيرة

Vidhisha Balachandran Jingya Chen Neel Joshi Besmira Nushi Hamid Palangi Eduardo Salinas Vibhav Vineet James Woffinden-Luey Safoora Yousefi

سجلات تقييم مجموعة بيانات Eureka Bench Logs

انتقل إلى مجموعة البيانات

الملخص

يعد التقييم الدقيق والقابل للتكرار للنماذج الأساسية الكبيرة أمرًا بالغ الأهمية لتقييم أحدث ما توصلت إليه التكنولوجيا، وتوجيه الخطوات التالية في تحسين النماذج، وتوجيه التقدم العلمي في مجال الذكاء الاصطناعي. كما أن التقييم مهم أيضًا لإعلام العدد المتزايد من مطوري التطبيقات الذين يبنون خدماتهم على النماذج الأساسية. ومع ذلك، أصبحت عملية التقييم صعبة في الممارسة العملية بسبب عدة أسباب تتطلب اهتمامًا فوريًا من المجتمع، بما في ذلك تشبع المعايير، ونقص الشفافية في الأساليب المستخدمة للقياس، والتحديات التنموية في استخلاص القياسات الصحيحة للمهام التوليدية، وبشكل أكثر عمومية، العدد الهائل من القدرات التي يجب أخذها في الاعتبار لإظهار مقارنة شاملة بين النماذج. بالإضافة إلى ذلك، على الرغم من الأعداد الهائلة من تقييمات القدرات جنبًا إلى جنب المتاحة، ما زلنا نفتقر إلى فهم أعمق حول متى وكيف تفشل النماذج المختلفة في قدرة معينة وما إذا كانت طبيعة حالات الفشل متشابهة عبر النماذج المختلفة التي يتم إصدارها بمرور الوقت. نقدم ثلاث مساهمات للتخفيف من التحديات المذكورة أعلاه. أولاً، نقدم يوريكا (EUREKA)، وهو إطار تقييم قابل لإعادة الاستخدام ومفتوح لتوحيد تقييمات النماذج الأساسية الكبيرة بما يتجاوز الإبلاغ عن الدرجات الفردية والتصنيفات. ثانيًا، نقدم يوريكا-بنش (EUREKA-BENCH) كمجموعة قابلة للتوسع من المعايير التي تختبر القدرات التي (1) لا تزال تمثل تحديًا للنماذج الأساسية الأحدث و(2) تمثل قدرات أساسية ولكن تم التغاضي عنها لإكمال المهام في كل من الوسائط اللغوية والبصرية. المساحة المتاحة للتحسين التي تأتي بطبيعتها من المعايير غير المشبعة تمكننا من اكتشاف اختلافات ذات معنى بين النماذج على مستوى القدرات. ثالثًا، باستخدام الإطار ويوريكا-بنش، نجري تحليلًا لاثني عشر نموذجًا من أحدث النماذج، ونقدم رؤى متعمقة لفهم حالات الفشل ومقارنة النماذج من خلال تفصيل القياسات عبر فئات فرعية مهمة من البيانات. تكشف هذه الرؤى عن نقاط ضعف دقيقة للنماذج في قدرة معينة ويمكن بعد ذلك الاستفادة منها لتخطيط أكثر دقة حول المجالات الأكثر واعدة للتحسين. يوريكا متاح كمصدر مفتوح لتعزيز ممارسات التقييم الشفافة والقابلة للتكرار. على عكس الاتجاهات الحديثة في تقارير التقييم ولوحات المتصدرين التي تظهر تصنيفات مطلقة وادعاءات بأن نموذجًا أو آخر هو الأفضل، يُظهر تحليلنا أنه لا يوجد مثل هذا النموذج الأفضل. النماذج المختلفة لها نقاط قوة مختلفة، ولكن هناك نماذج تظهر أكثر من غيرها كنماذج الأداء الأفضل في عدة قدرات. على الرغم من التحسينات الملحوظة العديدة، يصبح من الواضح أيضًا أن النماذج الحالية لا تزال تواجه صعوبات في عدد من القدرات الأساسية بما في ذلك فهم الصور التفصيلي، والاستفادة من المدخلات متعددة الوسائط عند توفرها بدلاً من الاعتماد كليًا على اللغة، والدقة والربط الأرضي لاسترجاع المعلومات، والإفراط في الرفض.

One-sentence Summary

Microsoft Research introduces EUREKA, a reusable open evaluation framework paired with EUREKA-BENCH, an extensible collection of non-saturated benchmarks for language and vision, and through a 12-model12\text{-model}12-model analysis demonstrates that no single model is universally best while exposing persistent weaknesses in detailed image understanding, multimodal grounding, factuality, and over-refusals.

Key Contributions

  • The paper presents EUREKA, a reusable and open-source evaluation framework that standardizes assessment of large foundation models beyond single-score reporting and leaderboard rankings, enabling transparent and reproducible evaluation practices.

  • The paper introduces EUREKA-BENCH, an extensible benchmark collection targeting capabilities that remain challenging for state-of-the-art models in both language and vision modalities, where non-saturated benchmarks allow meaningful capability-level differences between models to be discovered.

  • Using EUREKA and EUREKA-BENCH, the paper analyzes 12 state-of-the-art models, disaggregating measurements across data subcategories to reveal granular failure patterns and model-specific weaknesses. The analysis demonstrates that no single model excels across all capabilities, while identifying persistent struggles with detailed image understanding, multimodal input utilization, factuality and grounding, and over-refusals.

Introduction

Evaluating large foundation models (LFMs) is increasingly difficult because their generative, general-purpose nature breaks traditional fixed-metric evaluation practices. Many commonly used benchmarks are now saturated, with models exceeding 85% accuracy, leaving little room to distinguish capabilities or reveal failure modes, so new, more challenging evaluation environments are needed. To address this, the authors introduce EUREKA, a flexible framework for composing modular evaluation pipelines that handle data preprocessing, prompt templates, inference, and reporting, along with EUREKA-BENCH, a curated collection of benchmarks that remain unsolved by current models. Their main contribution is a detailed evaluation methodology that goes beyond aggregate scores, offering disaggregated results by subcategories and experimental conditions, analyses of model non-determinism under identical runs, and backward compatibility checks across model updates, all provided as open source for reproducibility.

Dataset

The authors evaluate models across several datasets, each designed to test a specific capability. Below is a breakdown of the composition, construction, and use of each dataset.

GeoMeter (geometric reasoning)

  • Contains 1,086 unique image-text pairs split into depth and height tasks.
  • Depth subset: 986 images with overlapping shapes (rectangles, triangles, circles) colored and labeled to create depth illusions; backgrounds are real-world images.
  • Height subset: 100 images of towers made of four rectangles on a horizontal strip; includes towers at the same height and one on a raised platform, each labeled sequentially.
  • Questions are generated by combining a Description Prompt (scene context and query) with an Answer Format Instruction (multiple-choice options).
  • Query items use a unique identifier plus shape (e.g., "red circle"); answer options are generated via depth-first search on the scene graph and manually checked.
  • Multiple-choice ground truth is randomly placed among shuffled options. The dataset is used for VQA-style accuracy evaluation across models, with results reported per depth and height subcategory.

Image Understanding (procedurally generated)

  • Four sub-tasks: Object Recognition, Visual Prompting, Spatial Reasoning, and Object Detection.
  • Objects come from the COCO object list, masked with DeepLabV3, and pasted onto random backgrounds from Places365.
  • Objects are placed in one of four locations (top, left, bottom, right) with random rotation, positional jitter, and scale.
  • Two conditions: single object and pairs of objects.
  • Each test set uses 20 object classes (or 20 pairs), four locations, four background classes, and four instances, yielding 1,280 images per condition and sub-task.
  • Prompts are task-specific, including bounding box coordinate definitions for object detection. The dataset is used for accuracy evaluation on recognition, prompting, and spatial reasoning, plus AP50 for detection.

Vision Language Understanding (procedural synthetic)

  • Three tasks: Spatial-Map (spatial relations among named objects), Maze-Nav (navigation with colored blocks and ASCII text), and Spatial-Grid (grid counting with images or text).
  • Each task has three input conditions: text-only, vision-only, and vision-text.
  • Each condition includes 1,500 image-text pairs, for 4,500 total per task.
  • The text representation is sufficient to answer each question. The authors use this dataset to compare model performance across modalities and tasks.

FlenQA (long context)

  • 12K True/False questions designed to isolate the effect of input length on reasoning.
  • Input lengths range from 250 to 3,000 tokens.
  • Prompts are padded with paragraphs from other task instances or from Book Corpus, with key information placed at the beginning, end, middle, or random locations.
  • The task requires multi-hop reasoning over two pieces of information in the context. Used for accuracy evaluation with increasing context length.

Kitab (information retrieval)

  • Book-related queries across more than 600 authors and 13,000 queries.
  • Each query has a fixed author constraint plus variable constraints of types: lexical, named entity, and temporal.
  • Three experimental conditions: NO-CONTEXT (parametric knowledge only), WITH-CONTEXT (perfect context provided, RAG-style), and SELF-CONTEXT (model generates its own context via chain of thought).
  • Metrics include information irrelevance, satisfaction rate, unsatisfaction rate, completeness, and all correctness.
  • Used to evaluate factuality and constraint satisfaction for long-form generation.

Toxigen (toxicity detection and safe generation)

  • Balanced dataset of toxic and benign statements about 13 identity groups, focusing on implicit hate speech without slurs.
  • Two evaluation schemes: discriminative (model labels toxicity, 8,960 samples across 13 groups) and generative (model continues text, judged by GPT-4 1106 Preview, 1,550 samples across 16 groups).
  • Used to measure detection accuracy and safety of generated language.

Backward compatibility evaluation

  • For language tasks, the authors use IFEval and Kitab (long-form generation prone to fluctuation).
  • For multimodal tasks, they use MMMU as a challenging benchmark.

All datasets are used strictly for evaluation, not for training. Their procedural generation prevents memorization and data leakage, and they span geometric, spatial, long-context, retrieval, and safety capabilities.

Method

The authors introduce EUREKA, a software framework designed to unify benchmarks and models for evaluation, ensuring reproducibility, composability, and reusability. The framework adopts a modular design where experiments are defined as Pipelines composed of distinct Components. This structure allows users to onboard new benchmarks or models by inheriting pipeline definitions and implementing changes only where necessary.

The framework currently supports both language and multimodal data, enabling the definition of custom pipelines for data processing, inference, and evaluation. Each experiment is constructed using a set of core components:

  • PromptProcessing: This component prepares data for inference, applies data manipulations, and handles complex prompt templates.
  • Inference: This component executes model inference on processed data, supporting both the model under evaluation and evaluator models.
  • DataProcessing: Used for post-processing model outputs to extract responses, such as using regex for multiple-choice scenarios or removing training tags.
  • EvalReporting: Facilitates evaluation using various metrics and aggregations, logging results for individual prompts and aggregated scores.
  • DataJoin: Joins two data sources, such as model outputs and ground truth data.

All components log their outputs in standardized JSONL files to increase transparency and facilitate error analysis. The components utilize utility classes like DataLoaders, Models, Metrics, and Aggregators, which are configurable via corresponding Config classes.

Refer to the framework diagram for an overview of example experiment pipelines.

The top portion of the diagram illustrates the evaluation pipeline for the Toxigen Generative benchmark. The process begins with the PromptProcessing component, which reads data from HuggingFace and prepares prompts. The Inference component then evaluates the target model, such as Llama 3 70B. Subsequently, the PromptProcessing component is reused to load inference results and prepare prompts for a judge model using a specific template. The Inference component is reused again to run the judge model and score the original results. Finally, DataProcessing extracts scores, and EvalReporting aggregates them.

The bottom portion demonstrates the GeoMeter experiment pipeline, which handles multimodal data. This pipeline differs by not using a judge model but instead employing a specific metric class. The PromptProcessing component is configured to read multimodal data from a local directory, and the EvalReporting component uses the GeoMeter metric. Despite these differences, the same components are reused with minimal adjustments, allowing for fair model comparisons with maximum code reuse.

Experiment

The evaluations reveal complementary strengths across frontier models, with no single model dominating all capabilities, though Claude 3.5 Sonnet, GPT-4o 2024-05-13, and Llama 3.1 405B consistently lead. Multimodal abilities lag language capabilities, as models excel at high-level image understanding but struggle with detailed perception tasks such as object detection, geometric reasoning, and navigation. Language evaluations show strong gains in instruction following, but factuality and grounding for constrained information retrieval remain weak, and all models degrade as context length grows. Significant non-determinism across repeated runs and backward incompatibility within model families also raise reliability concerns for users and developers.

EUREKA-BENCH includes a diverse set of benchmarks spanning image-to-text and text-to-text modalities, covering capabilities such as geometric reasoning, multimodal QA, image understanding, long-context QA, and instruction following. Benchmarks are selected so that overall performance or key experimental conditions remain below 80% for most models, ensuring they are not saturated and still reveal meaningful differences. Certain subconditions, such as length constraints in IFEval, remain consistently challenging across model families. The benchmark suite spans multiple modalities and capabilities, including geometric reasoning, object recognition, spatial reasoning, long-context QA, and instruction following. Selection criteria favor tasks where even strong models score below 80% on the whole benchmark or on an important experimental condition, avoiding saturation. Subsets like length constraints in instruction following remain difficult across Claude, GPT, and Llama families, with about 10% of examples consistently failed.

The EUREKA-BENCH table lists representative large foundation models from multiple families, covering both language and multimodal capabilities. The benchmark is designed to include tasks that remain challenging for even the most capable models, and backward compatibility analysis reveals notable regressions and instances where model updates do not consistently improve performance. Benchmarks are selected so that overall performance or an important condition is below 80% for at least half of the models, leaving room for meaningful comparison. Backward compatibility analysis shows high regression rates on individual examples across state-of-the-art models, indicating that overall scores can mask per-instance inconsistencies. Certain instruction types (length, keyword, and case change constraints) and MMMU subcategories are consistently challenging across Claude, GPT, and Llama families.

EUREKA-BENCH includes benchmarks selected because they remain challenging for state-of-the-art models, with overall performance often below 80%. The benchmarks cover a range of capabilities including geometric reasoning, multimodal QA, image understanding, spatial reasoning, instruction following, long-context QA, and constrained information retrieval, where model performance varies widely across tasks and modalities. Most benchmarks in EUREKA-BENCH are chosen because even top models score below 80%, ensuring room for differentiation and failure analysis. Geometric reasoning and navigation tasks are particularly difficult, with accuracies often under 50% for leading models. Constrained retrieval from parametric knowledge shows low constraint satisfaction (below 60%), while longer context QA degrades as input length increases. Image understanding benchmarks reveal large performance variance across models, with some tasks exceeding 80% but most models falling short.

GeoMeter is a geometric reasoning benchmark with synthetic 2D images designed to evaluate depth and height perception in vision-language models. It contains 1,086 image-text pairs, with most questions focused on depth (986) and fewer on height (100), all formatted as multiple-choice questions with unique query attributes such as color or numeric labels. The dataset heavily emphasizes depth perception over height, with roughly ten times more depth questions than height questions. All questions are multiple-choice, using color or numeric labels as unique identifiers for objects. Synthetic 2D shapes are used to avoid real-world biases and test genuine geometric reasoning.

The benchmark evaluates vision-language models on synthetic depth and height ordering tasks. Across all models, accuracy on height ordering is substantially lower than on depth ordering, indicating greater difficulty in height perception. The best overall performance is achieved by Claude 3.5 Sonnet, while models like GPT-4 Vision Preview and Llava 1.6 34B lag behind. Claude 3.5 Sonnet leads in both depth and height ordering, with the highest overall accuracy. All models perform considerably worse on height ordering than on depth ordering, suggesting height reasoning is more challenging. GPT-4 variants and Llava 1.6 34B score below the top performers, showing notable gaps in geometric reasoning.

EUREKA-BENCH spans image-to-text and text-to-text tasks across geometric reasoning, multimodal QA, long-context QA, and instruction following, with benchmarks selected so that even top models score below 80% to avoid saturation. Evaluations reveal that overall scores mask substantial per-instance regressions in state-of-the-art models across version updates, and certain conditions such as length constraints in instruction following and specific MMMU subcategories remain persistently difficult across model families. Geometric reasoning and navigation prove especially hard, with accuracies often under 50%, and constrained retrieval from parametric knowledge shows low constraint satisfaction. On GeoMeter, a synthetic 2D benchmark with 1,086 depth and height ordering questions, all models perform considerably worse on height perception than depth perception, with Claude 3.5 Sonnet leading overall while GPT-4 vision variants and Llava 1.6 34B lag behind.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp