Command Palette
Search for a command to run...
MeetingBank-LLMCompressed Meeting Minutes Compression Training Dataset
Date
Paper URL
License
Other
MeetingBank-LLMCompressed is a dataset released by Microsoft in 2024 for building training data for the LLMLingua-2 compressor, with the associated paper titled "LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression", aimed at providing efficient and faithful task-agnostic prompt compression training samples.
The dataset contains 5,169 samples from the MeetingBank training set, each including the original meeting transcript text and its version compressed independently by GPT-4 in chunks, with a data size of approximately 246 MB.
Dataset Composition
The dataset includes the following 6 fields:
- idx: The index number of the instance
- prompt: The original text of the meeting transcript
- prompt_list: The list of original text chunks corresponding to the prompt
- compressed_prompt_list: The list of texts after each chunk is independently compressed by GPT-4
- compressed_prompt: The complete compressed text formed by concatenating all compressed chunks
- summary: The meeting transcript summary from MeetingBank
Citation
@inproceedings{pan2024llmlingua2,
title={LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression},
author={Zhuoshi Pan and Qianhui Wu and Huiqiang Jiang and Menglin Xia and Xufang Luo and Jue Zhang and Qingwei Lin and Victor Rühle and Yuqing Yang and Chin-Yew Lin and H. Vicky Zhao and Lili Qiu and Dongmei Zhang},
year={2024},
booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics},
publisher = {Association for Computational Linguistics}
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.