HyperAIHyperAI

Command Palette

Search for a command to run...

MsrZhenTranslationParity: Human-Level Chinese-English Translation Dataset

Date

Organization

Microsoft Corporation

License

Other

MsrZhenTranslationParity is a Chinese-English translation dataset released by Microsoft in 2018, containing human evaluation results of machine translation and translation outputs, designed to allow external verification of its claim of achieving human-level translation and to facilitate future research.

The dataset contains 6 additional English translations of the WMT17 Chinese-English language pair, including Reference-HT based on human translation from scratch, Reference-PE based on human post-editing, and translation outputs from research systems Combo-4, Combo-5, Combo-6, and the online machine translation service Online-A-1710. The dataset training set contains 2,001 samples with a total size of approximately 1.79 MB, and all fields are different translation versions of the same Chinese source sentence.

Dataset Composition

The dataset mainly contains the following fields and structure:

  • Reference-HT: Human translation from scratch.
  • Reference-PE: Human post-edited translation.
  • Combo-4, Combo-5, Combo-6: Three translation versions from research systems.
  • Online-A-1710: Translation version from an anonymous online machine translation service.
  • train: Training set, containing 2,001 samples, with a size of 1,797,033 bytes.

Citation

Citation information is available at this link Achieving Human Parity on Automatic Chinese to English News Translation

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp