HyperAIHyperAI

Command Palette

Search for a command to run...

EvalVerse: 전문 시네마틱 영상 생성을 위한 파이프라인 인식 및 전문가 보정 기반 벤치마킹

초록

생성형 비디오 파운데이션 모델의 급속한 진보는 해당 분야를 전문급 시네마틱 합성으로 이끌었다. 이러한 높은 품질을 달성하기 위해 커뮤니티는 Reinforcement Learning (RL)과 agentic 워크플로우로 전환하고 있다. 그러나 신뢰할 수 있는 평가가 중요한 병목 현상으로 대두되었다. 기존 벤치마크는 주로 “whether it is right” (기본 프롬프트 준수)를 평가하는 데 치중하여, 근본적으로 “whether it is good” (시네마틱 품질, 연기, 미학)를 간과해 왔다. 또한, 현재 자동화 지표는 신뢰할 수 있는 신호를 제공하기 위해 필요한 도메인 특화 엄격성을 결여하고 있어, 인간의 미적 인식과 기계 점수 간에 심각한 신뢰성 격차를 초래한다. 이러한 격차를 해소하기 위해, 우리는 포괄적이고 파이프라인을 인지하며 전문가가 보정한 평가 프레임워크인 EvalVerse를 제안한다. 우리는 비디오 생성 평가를 단순한 공학적 과제로 취급하는 것을 넘어, 주관적인 시네마틱 전문 지식의 체계적인 디지털화라는 핵심 과학적 문제로 간주한다. 첫째, 우리는 도메인 지식을 전문 영화 제작 워크플로우(사전 제작, 제작, 사후 제작)와 일치하는 평가 분류 체계로 조직한다. 둘째, 우리는 인간 전문가의 판단을 대규모 인간 주석이 포함된 선별 데이터셋으로 정제한다. 셋째, 우리는 전문가 보정 파인튜닝 전략을 통해 이 지식을 Vision-Language Models (VLMs)에 주입하여, VLM이 명시적인 Chain-of-Thought 추론을 수행할 수 있도록 한다. 기존 연구와 비교하여, EvalVerse는 기본 “rightness” 지표와의 호환성을 유지할 뿐만 아니라 평가 기준을 “goodness”으로 크게 확장하고, 작업 범위를 복잡한 다중 샷 시퀀싱 및 오디오-비주얼 통합으로 확대한다. 결과적으로, EvalVerse는 세분화된 진단 신호를 제공함으로써 정적 리더보드를 넘어 보상 모델 및 evaluator agent와 같은 향후 연구를 위한 기반 인프라를 구축한다.

One-sentence Summary

EvalVerse is a pipeline-aware and expert-calibrated benchmarking framework that advances cinematic video assessment beyond basic prompt-following by systematically digitizing subjective expertise through a workflow-aligned taxonomy and a curated human-annotated dataset, thereby bridging the credibility gap between automated metrics and human aesthetic perception for professional cinematic video generation.

Key Contributions

  • EvalVerse introduces a pipeline-aware evaluation framework that structures professional filmmaking expertise into a comprehensive taxonomy spanning pre-production, production, and post-production phases. The framework utilizes a Real-to-Gen data engine to construct high-fidelity test pairs through proportional sampling from authentic professional video distributions.
  • A systematic human-machine calibration mechanism distills large-scale expert annotations into a curated dataset and injects this domain knowledge into vision-language models. This process translates subjective cinematic standards into scalable, expert-aligned chain-of-thought reasoning for automated assessment.
  • Comprehensive evaluations demonstrate strong human-machine alignment across complex multi-shot sequencing and audio-visual integration dimensions. The resulting metrics provide trustworthy diagnostic signals and dense reward vectors to support reinforcement learning optimization and autonomous agentic workflows in generative video synthesis.

Introduction

The rapid evolution of generative video models is driving the field toward professional cinematic synthesis, making fine-grained evaluation critical for training next-generation systems with reinforcement learning and autonomous agents. Existing benchmarks, however, primarily measure basic prompt-following and fail to assess nuanced cinematic quality, creating a significant credibility gap between human aesthetic judgment and automated scoring. To bridge this divide, the authors introduce EvalVerse, a pipeline-aware evaluation framework that maps professional filmmaking workflows into a structured diagnostic taxonomy. By distilling expert judgments into a large-scale annotated dataset and fine-tuning vision-language models with a structured reasoning process, they successfully translate subjective cinematic expertise into scalable, interpretable machine metrics. This methodology enables rigorous assessment of complex multi-shot and audio-visual generation while providing the reliable reward signals required to advance future generative pipelines.

Dataset

  • Composition and Sources: The authors curate a benchmark dataset drawn from a diverse collection of professional films and animations. The database is structured to evaluate video generation models across nine core cinematic dimensions, emphasizing technical fidelity, artistic rendering, and narrative continuity.

  • Subset Details and Distribution: The authors do not partition the data for training. Instead, they use the full collection as an evaluation set and apply a proportional sampling strategy across the nine dimensions to establish precise mixture ratios. Key subsets include the Aesthetics dimension, which covers visual quality, chromaticity, materiality, and lighting, and the Multi-Shot dimension, which assesses sequential logic and editing rhythm.

  • Data Usage and Workflow: The dataset serves exclusively as a testing benchmark rather than a training resource. The authors construct Real to Gen test pairs to drive downstream generation tasks and measure model capabilities. These pairs function as structured ground truth references and prompt targets for comparative evaluation across different video generation architectures.

  • Processing and Metadata Construction: The pipeline begins with a multi modal perception suite that extracts structured JSON metadata, capturing camera parameters, character attributes, and environmental details. After industrial grade processing and rigorous manual verification, the authors use Gemini 3.1 Pro to synthesize professional cinematic prompts from the metadata and raw captions. For reference based tasks, they extract keyframes and process them through Nano Banana Pro to create high fidelity reference images, while a ControlNet tuned model generates corresponding depth sequences.

Method

The framework of EvalVerse is structured around a comprehensive, pipeline-aware taxonomy that mirrors the traditional cinematic workflow, dividing video evaluation into three distinct stages: Pre-Production, Production, and Post-Production. This hierarchical structure organizes 18 main dimensions and 45 sub-dimensions, each designed to assess specific aspects of video quality from a professional filmmaking perspective. The taxonomy serves as the foundation for a systematic evaluation pipeline, guiding both human and machine assessment processes. Refer to the framework diagram to understand how the core dimensions are distributed across the three stages and how they relate to the overall evaluation process.

The evaluation pipeline is implemented in five steps. Step I involves the establishment of the taxonomy, defining the conceptual framework and its constituent dimensions. Step II focuses on dataset curation, where a large-scale, high-quality database of film and television content is assembled, augmented with industrial operators and human annotations. This step includes comprehensive sampling strategies and test pair construction, ensuring diverse and representative data for evaluation. Step III involves expert human evaluation, where a team of 14 video AIGC research scientists and engineers, along with 20 professional artists, perform both ranking-based and thoroughly considered scoring, with a workflow designed to ensure cross-checking and validation. Step IV constitutes the machine evaluation suite, which includes professional operator development and chain-of-thought evaluation. Step V enables application through benchmarking and the deployment of the trained models.

At the core of the machine evaluation is a two-stage VLM fine-tuning process. The first stage involves score calibration, where the model is trained on a pointwise dataset to generate both a detailed CoT rationale and the final absolute score. The model learns to autoregressively produce the rationale followed by the score, with optimal parameters obtained by minimizing a cross-entropy loss. The second stage involves a progressive, three-tiered calibration mechanism to align the model with human expert criteria. This includes prompt-level calibration, where abstract evaluation dimensions are replaced with more perceptually grounded ones; fusion-level calibration, which employs a lightweight MLP to optimize weights for different evidence sources and reasoning components; and parameter-level calibration, where fine-tuning injects cinematic domain knowledge directly into the model's parameters.

The machine evaluation pipeline operates in two steps. First, a suite of specialized operators extracts deterministic, objective evidence from the input video, audio, text prompt, and reference. These operators, including DINO for cross-frame identity tracking, YOLO for semantic anchoring, SyncNet for audio-visual synchronization, and Whisper for speech emotion recognition, provide a perception prior that mitigates hallucinations and ensures reliable contextual grounding. As shown in the figure below, this evidence is then fed into the fine-tuned VLM, which performs expert-guided chain-of-thought reasoning.

The VLM, denoted as Mθ\mathcal{M}_{\theta^*}Mθ, processes the multi-modal context, which includes the extracted evidence, the text prompt, the reference, and a set of expert-designed multi-questions for a specific cinematic dimension. Instead of producing a direct score, the model generates a detailed CoT, which includes a self-reflection mechanism to re-examine its own reasoning for potential hallucinations. A context-aware gating mechanism dynamically bypasses certain metrics if the narrative context does not warrant them. The final score for a dimension is computed by combining the VLM's output with this gating indicator. This approach ensures that the evaluation is not only accurate but also transparent and interpretable, providing a clear rationale for the final judgment.

Experiment

The evaluation framework assesses video generation models across pre-production asset logic, post-production multi-shot sequencing, and affective storytelling through a rigorous three-stage human expert pipeline alongside automated benchmarking. This comprehensive testing validates how effectively models preserve visual concepts, maintain cinematic aesthetics, and synchronize audio-visual elements across different generation settings. Qualitative analysis reveals a clear performance hierarchy where leading models consistently excel in cinematography, emotional progression, and multimodal coherence, while others demonstrate specialized strengths or notable gaps in complex narrative and sound design. Alignment experiments further confirm that the automated evaluation system closely mirrors professional human judgment, demonstrating that parameter-level calibration is essential for accurately assessing abstract and temporally entangled cinematic criteria.

The authors evaluate the alignment between automated metrics and human expert judgments across various video generation models and evaluation dimensions. Results show strong correlation between automated predictions and human preferences, particularly for dimensions grounded in visual and cinematic criteria, with higher agreement observed for abstract and temporally complex dimensions after task-specific calibration. Automated metrics show strong alignment with human expert evaluations across multiple video generation dimensions. Higher correlation is observed for abstract and temporally entangled dimensions after task-specific calibration. Pixel-grounded dimensions achieve strong alignment with human judgments, indicating reliable performance of prompt-level reasoning.

The authors evaluate various video generation models using a comprehensive benchmarking framework that assesses multiple dimensions of video quality, including visual design, emotional resonance, and post-production coherence. The framework, EvalVerse, is designed to align with human expert evaluations through a calibrated pipeline, achieving high consistency with professional judgments across different modalities and evaluation criteria. EvalVerse demonstrates high alignment with human expert evaluations across various video generation models and evaluation dimensions. The benchmarking framework covers diverse task modalities, including text-to-video, reference-to-video, video with sound, and multi-shot sequences. EvalVerse achieves high interpretability and expert-guided evaluation, with strong performance in both pixel-grounded and abstract, temporally-entangled dimensions.

The authors evaluate the alignment between automated model predictions and human expert judgments across various video generation dimensions. The results show strong correlation between automated and human evaluations, particularly for dimensions that are grounded in visual content and those further refined with task-specific calibration. This indicates that the evaluation framework effectively captures human preferences, with higher agreement observed for both pixel-based and abstract, temporally complex attributes after calibration. Automated predictions closely align with human expert judgments across all evaluated dimensions. Dimensions grounded in visual content show strong alignment, indicating reliable performance of prompt-level reasoning. Abstract and temporally complex attributes achieve the highest agreement after task-specific calibration.

The authors evaluate a range of video generation models across multiple dimensions spanning pre-production, production, and post-production stages, using a multi-stage human evaluation protocol and an automated evaluation framework. Results show that the leading models exhibit strong performance in visual and cinematic aspects, while differences emerge in affectivity and sound-related dimensions, with the evaluation framework demonstrating high alignment with human expert judgments. Seedance 2.0 achieves the highest overall performance across most evaluation dimensions, particularly in aesthetics, cinematography, and sound-related criteria. Models exhibit varying strengths across different aspects, with some excelling in visual and camera control while showing weaknesses in affectivity and audio-visual synchronization. The automated evaluation framework shows strong alignment with human expert preferences, especially for dimensions grounded in visual and cinematic attributes, and further improves for abstract and temporally complex aspects through task-specific calibration.

The authors evaluate multiple video generation models across various cinematic dimensions, including visual design, affectivity, and sound design, using a multi-stage human evaluation protocol. Results show a clear hierarchy among models, with some demonstrating strong overall performance and others showing specialized strengths or weaknesses in specific areas. Seedance 2.0 achieves the highest overall performance across most evaluation dimensions. Kling-v3-Omni and Happy Horse 1.0 show strong and consistent results in visual and cinematic aspects, with varying strengths in sound-related and affective dimensions. Models exhibit uneven performance profiles, with some excelling in specific areas like visual consistency or cinematography while lagging in others such as affectivity or audio-visual synchronization.

The evaluation employs a comprehensive benchmarking framework paired with a multi-stage human assessment protocol to compare state-of-the-art video generation models across diverse modalities and cinematic dimensions. This experimental setup validates the alignment between automated scoring systems and professional human judgments while establishing a performance hierarchy among competing models. Qualitative results indicate that automated metrics reliably capture human preferences, particularly for visual and cinematic attributes, with task-specific calibration further strengthening agreement on abstract and temporally complex criteria. Overall, model capabilities reveal distinct specializations, with leading systems excelling in aesthetics and cinematography while demonstrating uneven strengths in emotional resonance and audio-visual synchronization.


AI로 AI 구축

아이디어에서 출시까지 — 무료 AI 코코딩, 즉시 사용 가능한 환경, 최적의 GPU 가격으로 AI 개발을 가속화하세요.

AI 협업 코딩
바로 사용 가능한 GPU
최적의 가격

HyperAI Newsletters

최신 정보 구독하기
한국 시간 매주 월요일 오전 9시 에 이번 주의 최신 업데이트를 메일로 발송합니다
이메일 서비스 제공: MailChimp
EvalVerse: 전문 시네마틱 영상 생성을 위한 파이프라인 인식 및 전문가 보정 기반 벤치마킹 | 문서 | HyperAI초신경