HyperAIHyperAI

Command Palette

Search for a command to run...

VoxMem: 大規模オーディオ言語モデルにおけるマルチモーダル記憶のベンチマーク評価

Yang Xiao Vidhyasaharan Sethu Eun-Jung Holden Ting Dang

概要

音声対話システムは過去のやり取りから情報を復元する必要がある(すなわち記憶)。しかし、音声における関連情報は、何が話されたかだけでなく、誰が話したか、どのように話されたか、何が聞こえたかにも及び、これらの情報は音声信号にのみ存在し、書き起こしからは復元できない。何を記憶するかに加えて、記憶は多様な操作も要求する。単一の事実の検索、ターンをまたぐ証拠の統合、変化する状態の追跡などである。実世界のやり取りはさらにセッションをまたいで展開するため、情報は単一の連続した録音ではなく、別個のエピソードにわたって蓄積される。既存のベンチマークはこれら3つの側面すべてで不十分である。主に語彙的内容に焦点を当て、限定的で場当たり的な記憶操作を採用し、記憶を単一セッションの問題として扱っている。我々は、原理的な記憶評価には、保持すべき音響的証拠とそれに適用される操作を統合的に特徴づける必要があると主張し、これら2つの軸に沿った分類体系を導入する。この分類体系に基づき、VoxMemを提示する。VoxMemは、4種類の音響的証拠タイプ(発話意味、話者識別、パラ言語的手がかり、環境音)と4種類の記憶操作(情報抽出、複数セッション推論、時間的追跡、回答拒否)を交差させ、34,743件の音声セッション(177時間)にわたる3,196件の評価インスタンスから成る。マルチセッション履歴に基づき、8Kから64Kトークンまでの文脈予算にわたって層別されている。15の大規模オーディオ言語モデル(LALM)を評価したところ、32Kで40%を超えるモデルはなかった。モデルは、何が話されたかを、誰が話したか、どのように話されたか、何が聞こえたかよりもはるかによく保持する。この差は複雑な操作で拡大し、履歴が長くなるにつれて大きくなり、証拠タイプごとに質的に異なる失敗モードとして現れる。VoxMemは、音声対話記憶の全範囲にわたる進歩を測定し推進するための基盤を提供することを目指す。

One-sentence Summary

Researchers from the University of Melbourne and the University of New South Wales introduce VoxMem, a benchmark of 3,196 evaluation instances over 34,743 spoken sessions that applies a taxonomy jointly characterizing acoustic evidence and memory operations to evaluate 15 large audio language models across 8K to 64K token contexts, finding that none exceeds 40% at 32K and that models retain lexical content far better than speaker identity, paralinguistic cues, and environmental sound, exposing gaps beyond lexical memory.

Key Contributions

  • Introduces a two-dimensional taxonomy for spoken conversational memory that jointly characterizes acoustic evidence to be retained, including speech semantics, speaker identity, paralinguistic cues, and environmental sound, alongside memory operations such as information extraction, multi-session reasoning, temporal tracking, and answer refusal.
  • Presents VoxMem, a multi-session benchmark with 3,196 quality-controlled instances over 34,743 spoken sessions totaling 177 hours, stratified across context budgets from 8K to 64K tokens and built from evidence, competing, and unrelated sessions across 20 topic families.
  • Reports a systematic evaluation of 15 LALMs showing that no model exceeds 40% at 32K context, that lexical content is retained much better than speaker identity, paralinguistic cues, or environmental sounds, and that failure modes differ qualitatively by evidence type, such as misattribution for speaker identity and lost acoustic cues for paralinguistic information.

Introduction

Large audio language models (LALMs) are enabling longer spoken interactions across separate sessions, making long-term memory a critical requirement. However, existing spoken-memory benchmarks lack a principled taxonomy of what acoustic evidence must be retained and which memory operations are needed, and they mostly evaluate isolated recordings or continuous dialogues rather than multi-session histories. They also tend to confound history length with question difficulty and evidence type. The authors introduce VoxMem, a multi-session benchmark organized along two axes: acoustic evidence, encompassing speech semantics, speaker identity, paralinguistic cues, and environmental sound, and memory operation, comprising information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal. VoxMem contains 3,196 quality-controlled instances over 34,743 spoken sessions, spanning four context budgets from 8K to 64K tokens, and supports controlled evaluation of how spoken memory degrades as context length scales.

Dataset

VoxMem Benchmark Dataset

  • Scope and taxonomy. VoxMem is a spoken conversational memory benchmark organized around two dimensions: acoustic evidence and memory operations. Acoustic evidence covers speech semantics, speaker identity, paralinguistic cues, and environmental sound. Memory operations cover information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal. The benchmark includes 15 evaluation scenarios, leaving out information extraction over speech semantics.

  • Data composition. Each multi-session history contains three session types. Evidence sessions hold the information required to answer a question. Haystack sessions add plausible but misleading content on the same topic, preventing the model from answering by topic matching alone. Filler sessions contain unrelated conversation and extend the history to the target length. Every item begins with a structured plan specifying the question, gold answer, and required evidence before dialogue or audio is written.

  • Text and audio sources. Evidence and haystack dialogues are written as text first. Filler text is taken from InstructS2S and trimmed to match VoxMem session lengths. User turns are synthesized with Higgs-TTS-3, using a fixed VCTK voice per user. Paralinguistic cues are introduced through style controls, and environmental sounds from ESC-50 are mixed in at 10dB SNR. Questions are rendered into natural language using Gemini-3.7-Flash and GPT-5.6-Luna. In the final input, user turns are audio, assistant turns are text, and session boundaries and timestamps are included.

  • Scale. The construction pipeline generates more than 30,000 plans. The final answerable set contains 669 questions, with answer-refusal variants derived from existing answerable questions. Each answerable question is embedded in histories at four lengths: 8K, 16K, 32K, and 64K tokens measured with a Whisper encoder, approximately 2.5 to 20 minutes of audio. Longer histories strictly extend shorter ones, and evidence and haystack sessions are distributed uniformly except where temporal order matters for temporal evolution tracking.

  • Filtering and quality control. Session-level checks verify transcript fidelity, speaker consistency, target-cue perceptibility, and deterministic consistency of paired paralinguistic and environmental renditions. Question-family checks verify that evidence uniquely determines the answer, the question requires the intended operation, audio-native questions cannot be answered from the transcript, and full histories contain no leaked or alternative answers. Items that fail are revised and rechecked or discarded. A text-only classifier also cannot reliably distinguish evidence from haystack sessions, confirming that surface patterns do not reveal the answer.

  • Usage. VoxMem is an evaluation-only benchmark, not a training set. It is designed to evaluate LALMs that accept audio input but do not generate speech. Every model receives the same history with audio user turns and text assistant turns, and performance is compared across the four context lengths so that changes can be attributed to growing context rather than changing evidence.

Method

The authors construct the VoxMem benchmark along a two-dimensional taxonomy, utilizing a multi-session conversation structure composed of three distinct session types to ensure both naturalism and controlled evaluation scenarios. Evidence sessions contain the specific information required to answer a query. Haystack sessions introduce plausible but misleading content on the same topic, forcing the model to reason over the full acoustic and semantic context rather than relying on topic matching. Filler sessions consist of ordinary conversations unrelated to the question, serving to extend the history to the target context length.

As shown in the figure below, this structure is applied across different memory types. For instance, in an Information Extraction scenario involving speaker identity, the evidence session contains the answer spoken by the querying user, while a haystack session features a different user discussing the same topic with different content.

The construction of every item follows a rigorous pipeline, as illustrated in the framework diagram.

The process begins with Task and Evidence Design, where a structured plan specifies the question, gold answer, and required evidence. Evidence is distributed across sessions based on the operation type: a single session for Information Extraction, multiple for Multi-Session Reasoning, and an ordered sequence for Temporal Evolution Tracking. During the Dialogue and Audio Realization phase, evidence sessions are written as user-assistant dialogues. To prevent answer leakage, the two sides are authored independently. User turns are synthesized using a TTS system with fixed voices, and paralinguistic cues or environmental sounds are introduced.

In the Variants and Context Assembly stage, the authors create Answerable and Answer Refusal variants. Haystack and filler sessions are added as distractors. Histories are assembled at four lengths (8K, 16K, 32K, and 64K tokens) by strictly extending shorter histories with more distractors, ensuring the question and evidence remain identical. Finally, the Validation and Final Benchmark stage involves multi-stage quality control.

The authors implement a two-level quality control process. Session-level validation targets failures within a single session, verifying transcript fidelity, speaker consistency, and the perceptibility of acoustic cues. Question-family validation targets failures emerging in the full history. This includes checking answer and operation validity, ensuring acoustic necessity by rejecting questions answerable from text alone, and verifying full-history validity to ensure no leaked or alternative answers exist. Items failing these checks are revised and rechecked.

Experiment

The VoxMem benchmark evaluates spoken conversational memory across a two-dimensional taxonomy of acoustic evidence and memory operations, testing 15 large audio-language models on histories from 8K to 64K tokens with LLM-judged short answers. The experiments show that current models remain unreliable, with no model exceeding 40% overall accuracy at 32K, and that non-lexical acoustic information such as speaker identity, paralinguistic cues, and environmental sounds is substantially harder to retain than speech semantics. Memory operation difficulty depends on the evidence type, while refusal patterns indicate that models often abstain because of weak acoustic processing rather than genuine awareness of missing evidence. Longer histories reduce access to the same evidence, with speaker and environmental memory degrading fastest and error profiles differing across evidence types and operations.

The comparison shows that existing spoken-memory benchmarks cover acoustic evidence and memory operations only partially. Most emphasize linguistic content or selected acoustic cues and evaluate only retrieval or integration, while temporal evolution tracking and answer refusal are absent. The benchmarks also rely on single-session monologues or dialogues, with controlled scaling and evidence provenance rarely reported. Prior benchmarks rarely test speaker identity, paralinguistic cues, or environmental sound; several focus only on lexical content. Memory operation coverage is narrow: all listed benchmarks include information extraction, a few add multi-session reasoning, and none include temporal evolution tracking or answer refusal. All listed benchmarks use single-session histories, with no multi-session evaluation. Evidence provenance is missing for most listed benchmarks, with only one dialogue benchmark providing it.

The evaluation taxonomy is uneven, with paralinguistic and semantic questions dominating and answer refusal less common. Experimental results show that longer histories reduce access to previously available evidence, especially for speaker and environmental sound memory, while paralinguistic understanding remains low at all context lengths. Error attribution also reveals distinct failure modes across operations and evidence types, such as binding errors for speaker information and localization errors for temporal evolution. The taxonomy is unevenly distributed: paralinguistic and semantic questions are the most common evidence categories, with answer refusal less frequent across all operations. As history grows, models retain less access to evidence, and speaker identity and environmental memory degrade fastest; paralinguistic cues remain difficult at every context length.

The benchmark includes several hundred questions and several thousand instances drawn from tens of thousands of sessions, with a smaller subset of evidence sessions. Sessions are relatively compact, averaging about nine turns and a few clips per session. Controlled context expansion from 8K to 64K shows accuracy declines as history grows, with the steepest relative drops for speaker identity and environmental sound. The dataset contains 799 questions and 3,196 instances, but only a small fraction of sessions serve as evidence sessions. Longer context windows reduce accuracy across evidence types, with speaker and environmental memory degrading faster than speech semantics and paralinguistic cues.

For Gemini-3.7-Flash on answerable acoustic-memory questions, replacing full audio with transcript-only turns causes a large overall accuracy drop, confirming reliance on non-text acoustic cues. The impact is highly uneven: speech semantics remains relatively stable, while speaker identity, paralinguistic cues, and environmental sounds lose most of their accuracy. Overall full-audio accuracy falls from about 55 percent to below a quarter with transcript-only input. Speech semantics is least affected, while speaker identity drops roughly 60 points to just over 10 percent, and paralinguistic and environmental scores fall near zero.

Spoken-memory accuracy declines as reference history grows, with both proprietary and open-weight models losing accuracy between 8K and 32K. The drop is steeper for speaker identity and environmental sound than for speech semantics and paralinguistic cues. Paralinguistic cues show the lowest absolute accuracy but decline at a rate similar to speech semantics, indicating that baseline difficulty and history sensitivity are distinct failure modes. At 64K, speech semantics and paralinguistic cues retain about 70% of their 8K accuracy, while speaker identity and environmental sound retain about 63% to 67%. Proprietary models score above open-weight models at 8K and 32K, but both groups lose accuracy as history grows. Speaker errors are most often binding failures, whereas paralinguistic errors are dominated by localization failures and environmental errors split between localization and unsupported answers.

Existing spoken-memory benchmarks cover acoustic evidence and memory operations only partially, often missing speaker identity, paralinguistic and environmental cues, temporal evolution tracking, answer refusal, multi-session histories, and evidence provenance. The benchmark's controlled context expansion shows that both proprietary and open-weight models lose access to earlier evidence as history grows, with speaker identity and environmental sound degrading more steeply than speech semantics, while paralinguistic cues remain difficult across all context lengths. Replacing audio with transcripts causes a large overall drop and leaves speaker, paralinguistic, and environmental performance near zero, confirming that models rely on acoustic cues beyond text. Error attribution reveals distinct failure modes, including binding errors for speaker information and localization errors for paralinguistic and temporal evolution questions.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています