Command Palette
Search for a command to run...
How2QA Video + Language Dataset
Date
Publish URL
Paper URL

How2QA is a video + language learning framework dataset. The dataset presents the same set of selected video clips to another set of AMT workers for multiple-choice question-answer annotation. Each worker is assigned a video clip and asked to write a question based on four prepared responses (one correct answer and three distracting answers). The video narration is hidden from the workers to ensure that the collected question-answer pairs are not affected by subtitles. The dataset contains 22,000 60-second clips and 44,007 question-answer pairs selected from 9,035 videos.
Citation
@inproceedings{li2020hero,
title={HERO: Hierarchical Encoder for Video+ Language Omni-representation Pre-training},
author={Li, Linjie and Chen, Yen-Chun and Cheng, Yu and Gan, Zhe and Yu, Licheng and Liu, Jingjing},
booktitle={EMNLP},
year={2020}
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.