Command Palette
Search for a command to run...
Vera-Layered-Video Video Layered Diffusion Dataset
Vera-Layered-Video is a hierarchical diffusion dataset for video editing, released in 2026 by Caltech in collaboration with Netflix. It aims to support content-preserving video editing based on a hierarchical diffusion model. Related research papers include... Vera: A Layered Diffusion Model for Content-Preserving Video Editing . This dataset consists of a training set and a test set. The training set contains 11,996 49-frame clips and 6,005 81-frame clips, while the test set contains 141 clips covering two types of editing tasks: background replacement and object addition, involving both compositing and real-world scenes. The training set data comes from Pexels, Mixkit, and VideoMatte240K, while the test set also includes footage from DAVIS and VACEBench.
Dataset composition:
- 49 training set frames: each frame is 3 seconds long and contains 914 realistic-set1-bg-change, 470 realistic-set1-obj-add, 770 realistic-set2-obj-add, 4,994 synthetic-bg-change, and 4,848 synthetic-obj-add frames.
- 81 training set frames: each frame is 5 seconds long and contains 457 realistic-set1-bg-change, 235 realistic-set1-obj-add, 385 realistic-set2-obj-add, 2,497 synthetic-bg-change, and 2,431 synthetic-obj-add frames.
- Test set snippet: Contains 69 bg-change and 72 obj-add statements.
Citation
@article{zheng2026vera,
title = {Vera: A Layered Diffusion Model for Content-Preserving Video Editing},
author = {Zheng, Hongkai and Cheng, Ta-Ying and Klein, Benjamin and Yue, Yisong and Yuan, Zhuoning},
journal = {arXiv preprint arXiv:2606.23610},
year = {2026}
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.