Command Palette
Search for a command to run...
MsrZhenTranslationParity: Human-Level Chinese-English Translation Dataset
MsrZhenTranslationParity is a Chinese-English translation dataset released by Microsoft in 2018, containing human evaluation results of machine translation and translation outputs, designed to allow external verification of its claim of achieving human-level translation and to facilitate future research.
The dataset contains 6 additional English translations of the WMT17 Chinese-English language pair, including Reference-HT based on human translation from scratch, Reference-PE based on human post-editing, and translation outputs from research systems Combo-4, Combo-5, Combo-6, and the online machine translation service Online-A-1710. The dataset training set contains 2,001 samples with a total size of approximately 1.79 MB, and all fields are different translation versions of the same Chinese source sentence.
Dataset Composition
The dataset mainly contains the following fields and structure:
- Reference-HT: Human translation from scratch.
- Reference-PE: Human post-edited translation.
- Combo-4, Combo-5, Combo-6: Three translation versions from research systems.
- Online-A-1710: Translation version from an anonymous online machine translation service.
- train: Training set, containing 2,001 samples, with a size of 1,797,033 bytes.
Citation
Citation information is available at this link Achieving Human Parity on Automatic Chinese to English News Translation
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.