Command Palette
Search for a command to run...
MS MARCO: A Human Generated MAchine Reading COmprehension
Date
Paper URL
MS MARCO is a dataset focused on deep learning for search, released by Microsoft in 2016, with the related paper "[MS MARCO: A Human Generated MAchine Reading COmprehension Dataset", designed to advance research in question answering, natural language generation, and information retrieval tasks using real user queries and human-generated answers.
The dataset includes two main versions, v1.1 and v2.1, where v1.1 contains approximately 100,000 examples, while v2.1 expands to over 1 million queries, with a massive data scale covering training, validation, and test sets, widely applied in scenarios such as voice answer generation for smart speakers and web search ranking.
Dataset Composition
- answers: List of answers
- passages: Dictionary structure containing passage_text, is_selected, and url
- query: User query text
- query_id: Query ID
- query_type: Query type
- wellFormedAnswers: List of well-formed answers
Citation
@article{DBLP:journals/corr/NguyenRSGTMD16,
author = {Tri Nguyen and
Mir Rosenberg and
Xia Song and
Jianfeng Gao and
Saurabh Tiwary and
Rangan Majumder and
Li Deng},
title = {{MS} {MARCO:} {A} Human Generated MAchine Reading COmprehension Dataset},
journal = {CoRR},
volume = {abs/1611.09268},
year = {2016},
url = {http://arxiv.org/abs/1611.09268},
archivePrefix = {arXiv},
eprint = {1611.09268},
timestamp = {Mon, 13 Aug 2018 16:49:03 +0200},
biburl = {https://dblp.org/rec/journals/corr/NguyenRSGTMD16.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.