Command Palette
Search for a command to run...
CantTalkAboutThis Topic Control - Non Commercial
Dataset Overview
The CantTalkAboutThis Topic Control Dataset - Non-Commercial is a dataset released by NVIDIA in 2024, focusing on dialogue topic control and alignment of large language models for safety. The related paper results can be found in [insert citation or link here].CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues”, which aims to enhance large language models’ ability to maintain topic focus in task-oriented conversations and improve their robustness against distractions.
The dataset comprises 1,080 synthetic dialogue samples spanning nine domains: health, banking, travel, education, finance, insurance, law, real estate, and computer troubleshooting. Each dialogue includes distracting turns designed to test the model’s resistance to topic drift. Fine-tuning on this dataset significantly boosts performance in instruction-following and safety-related tasks, enabling more effective identification of sensitive topics and handling of restricted content.
Dataset Composition
The dataset primarily contains the following fields:
- domain: The category or field to which the conversation belongs.
- scenario: The specific context or task being discussed.
- system_instruction: Dialogue strategies assigned to the model, typically including complex sets of instructions specifying allowed and prohibited discussion topics.
- conversation: Full transcript of the conversation, encompassing both main-topic exchanges and distracting turns.
- distractors: A list of distracting turns, comprising bot-generated responses as well as user-side replies intended as counter-responses to those bot turns.
- conversation_with_distractors: Complete dialogues that incorporate all distraction elements.
The dataset is split into training and testing subsets. The training set consists of synthetically generated examples produced by the GPT-4 Turbo model, while the evaluation (testing) subset features human-labeled data containing more sophisticated and realistic distractors for assessing model performance.
Citation```bibtex
@inproceedings{sreedhar2024canttalkaboutthis, title={CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues}, author={Sreedhar, Makesh and Rebedea, Traian and Ghosh, Shaona and Zeng, Jiaqi and Parisien, Christopher}, booktitle={Findings of the Association for Computational Linguistics: EMNLP 2024}, pages={12232--12252}, year={2024}, organization={Association for Computational Linguistics} }
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.