HyperAIHyperAI

Command Palette

Search for a command to run...

Microsoft Research Sequential Question Answering 微软研究顺序问答数据集

日期

数据集组织

微软研究院

许可证

Other

Microsoft Research Sequential Question Answering 是由 Microsoft Research 于 2017 年发布的一个面向顺序问答任务的数据集,旨在探索更自然的对话式问答场景,解决传统语义解析中过于复杂且不自然的长问题。

该数据集包含 17,553 个问题,分为 6,066 个序列,每个问题都关联了表格中的单元格位置作为答案。 数据基于 WikiTableQuestions 数据集,由众包工人对原始问题进行分解生成。

数据集组成

数据集划分为以下三个部分:

  • 训练集(train):12,276 个样本。
  • 验证集(validation):2,265 个样本。
  • 测试集(test):3,012 个样本。

每条样本主要包含以下字段:

  • id:问题序列的唯一标识。
  • annotator:标注者标识。
  • position:问题在序列中的位置。
  • question:问题文本。
  • table_file:关联的表格文件。
  • table_header:表格列名。
  • table_data:表格数据,以二维数组形式存储。
  • answer_coordinates:答案所在单元格的行列坐标。
  • answer_text:答案文本内容。

Citation

@inproceedings{iyyer-etal-2017-search,
    title = "Search-based Neural Structured Learning for Sequential Question Answering",
    author = "Iyyer, Mohit  and
      Yih, Wen-tau  and
      Chang, Ming-Wei",
    booktitle = "Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2017",
    address = "Vancouver, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/P17-1167",
    doi = "10.18653/v1/P17-1167",
    pages = "1821--1831",
}

用 AI 构建 AI

从创意到上线——通过免费 AI 协同编码、开箱即用的环境和最优惠的 GPU 价格,加速您的 AI 开发。

AI 协同编码
开箱即用的 GPU
最优定价

HyperAI Newsletters

订阅我们的最新资讯
我们会在北京时间 每周一的上午九点 向您的邮箱投递本周内的最新更新
邮件发送服务由 MailChimp 提供