HyperAIHyperAI

Command Palette

Search for a command to run...

MsrTextCompression: Abstractive Sentence Compression

Date

Organization

Microsoft Research

License

Other

MsrTextCompression is a dataset released by Microsoft Research in 2016 for evaluating sentence and short paragraph summarization and compression, designed to provide a standard corpus containing compressed texts along with their quality judgments to advance the development of automatic summarization technology.

The dataset contains approximately 6,000 source texts from the Open American National Corpus (OANC1), totaling about 26,000 source text and compressed text pairs, covering domains such as newswire, business letters, journals, and technical documents. Each source text is accompanied by up to five compressed versions generated by crowdworkers, along with quality scores regarding meaning preservation and grammatical correctness.

Dataset Composition

The dataset is divided into three parts:

  • Training set (train): Contains 4,936 samples, with a data size of approximately 5 MB.
  • Validation set (validation): Contains 447 samples, with a data size of approximately 449 KB.
  • Test set (test): Contains 785 samples, with a data size of approximately 804 KB.

Each data record contains the following fields:

  • source_id: The index of the original text in the source dataset.
  • domain: The domain classification of the source text (e.g., Newswire).
  • source_text: The uncompressed original text.
  • targets: A list containing compression results and related evaluation information, with each element corresponding to one compressed version:
  • compressed_text: The compressed text.
  • judge_id: The anonymous ID of the crowdworker who provided this compressed version.
  • num_ratings: The number of ratings received by this compressed version.
  • ratings: The specific list of ratings, based on meaning preservation and grammatical correctness (e.g., 6 indicates the most important meaning with perfect grammar, while 24 indicates minimal meaning with ungrammatical output).

Citation```bibtex

@inproceedings{Toutanova2016ADA, title={A Dataset and Evaluation Metrics for Abstractive Compression of Sentences and Short Paragraphs}, author={Kristina Toutanova and Chris Brockett and Ke M. Tran and Saleema Amershi}, booktitle={EMNLP}, year={2016} }

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp