HyperAIHyperAI

Command Palette

Search for a command to run...

a month ago

SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models

Cheng-Han Chiang Xiaofei Wang Linjie Li Chung-Ching Lin Kevin Lin Shujie Liu Zhendong Wang Zhengyuan Yang Hung-yi Lee Lijuan Wang

SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models

Abstract

Current large language models (LLMs) and spoken language models (SLMs) beginthinking and taking actions only after the user has finished their turn. Thisprevents the model from interacting during the user's turn and can lead to highresponse latency while it waits to think. Consequently, thinking afterreceiving the full input is not suitable for speech-to-speech interaction,where real-time, low-latency exchange is important. We address this by notingthat humans naturally "think while listening." In this paper, we proposeSHANKS, a general inference framework that enables SLMs to generate unspokenchain-of-thought reasoning while listening to the user input. SHANKS streamsthe input speech in fixed-duration chunks and, as soon as a chunk is received,generates unspoken reasoning based on all previous speech and reasoning, whilethe user continues speaking. SHANKS uses this unspoken reasoning to decidewhether to interrupt the user and to make tool calls to complete the task. Wedemonstrate that SHANKS enhances real-time user-SLM interaction in twoscenarios: (1) when the user is presenting a step-by-step solution to a mathproblem, SHANKS can listen, reason, and interrupt when the user makes amistake, achieving 37.1% higher interruption accuracy than a baseline thatinterrupts without thinking; and (2) in a tool-augmented dialogue, SHANKS cancomplete 56.9% of the tool calls before the user finishes their turn. Overall,SHANKS moves toward models that keep thinking throughout the conversation, notonly after a turn ends. Animated illustrations of Shanks can be found athttps://d223302.github.io/SHANKS/

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models | Papers | HyperAI