Command Palette
Search for a command to run...
CantTalkAboutThis Topic Control
Dataset Overview
The CantTalkAboutThis Topic Control Dataset, released by NVIDIA in 2024, is designed for training language models to maintain topic focus during task-oriented dialogues. The related research paper can be found at [insert link or title here].
(Note: If there's additional information about the dataset such as its size, usage examples, etc., it would follow this introductory paragraph.)CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues”, aimed at enhancing the model’s robustness against distractors and improving alignment in instruction-following and safety tasks.
This dataset comprises 1,080 synthetic conversations spanning nine domains: health, banking, travel, education, finance, insurance, law, real estate, and computer troubleshooting.
Each conversation includes distractors designed to test and improve the model’s performance when handling irrelevant topics.
The dataset contains no personally sensitive information and is generated entirely from synthetic data.
Dataset Composition
The dataset consists of the following fields:
domain: The category or field to which the conversation belongs.scenario: The specific context or task being discussed.system_instruction: Dialogue strategy provided to the model, typically comprising complex instructions regarding permitted and prohibited topics.conversation: The complete dialogue, including both main-topic exchanges and distractor interactions.distractors: A list of distractor items, encompassing bot responses within the conversation as well as user inputs intended as replies to those bot responses.conversation_with_distractors: The full dialogue incorporating all distractor elements.
The dataset is divided into two subsets: training (Mixtral) and testing (Human Test Set). The training subset was generated using the Mixtral-8x7B-Instruct model, while the test set features human-curated, more sophisticated, and realistic distractors for evaluating model performance.
Citation```bibtex
@inproceedings{sreedhar2024canttalkaboutthis, title={CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues}, author={Sreedhar, Makesh and Rebedea, Traian and Ghosh, Shaona and Zeng, Jiaqi and Parisien, Christopher}, booktitle={Findings of the Association for Computational Linguistics: EMNLP 2024}, pages={12232--12252}, year={2024}, organization={Association for Computational Linguistics} }
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.