Command Palette
Search for a command to run...
OneStreamer: Vereinheitlichung von Wahrnehmung, Gedächtnis und proaktiver Reaktion in der Streaming-Video-Interaktion
OneStreamer: Vereinheitlichung von Wahrnehmung, Gedächtnis und proaktiver Reaktion in der Streaming-Video-Interaktion
Zusammenfassung
Streaming-Video-LLMs müssen Belege festhalten, bevor deren Relevanz für künftige Aufgaben bekannt ist, und antworten, sobald ausreichende Belege verfügbar sind. Die Herausforderung besteht darin, ein wiederverwendbares faktisches Gedächtnis aufzubauen, ohne die Echtzeitwahrnehmung zu beeinträchtigen. Wir stellen OneStreamer vor, das abfrageunabhängige Evidenzaufzeichnung und Aufgabenbeantwortung gemeinsam über einen geteilten proaktiven Generierungsprozess lernt. Sein Proactive Hierarchical Caption Memory (PHCM) erzeugt zeitlich verankerte, lokaldetaillierte Captions und Zusammenfassungen abgeschlossener Ereignisse. Streaming-Caption-Ziele überwachen während des Trainings die Interpretation beobachteter Videopräfixe. Bei der Inferenz ergänzen modellgenerierte Aufzeichnungen ein aktuelles visuelles Fenster und liefern wiederverwendbaren faktischen Kontext, ohne auf frühere visuelle Merkmale zurückzugreifen. Proactive State Transition Learning (PSTL) reduziert die Dominanz wiederholter Wartezustände, indem es die Supervision an allen Ausgabeankern erhält und repräsentative Zustandswechselund Zustandspersistenz-Token auswählt. Darüber hinaus entwickeln wir eine Pipeline zur Synthese von Streaming-Daten, die Ausgabeinhalt und -zeitpunkt mit der verfügbaren Evidenz in Einklang bringt. Die Kombination der resultierenden Streaming-Captions und QA mit bereinigten Open-Source-Daten ergibt OneStreamer-1M, einen breit abdeckenden Datensatz für die Streaming-Video-Interaktion mit über einer Million Einträgen zu unterschiedlichen Aufgaben. Unser 4B-Modell erzielt unter den verglichenen Methoden die besten Ergebnisse über alle acht evaluierten Benchmarks zum Streaming-Videoverständnis hinweg. Ablationsstudien zeigen, dass das Beibehalten generierter Captions die historische QA verbessert, ohne die Echtzeitwahrnehmung zu verschlechtern. PSTL übertrifft zudem dichte Zustands-Supervision, während es nur 27,5 % der annotierten Zustands-Token supervidiert. Insgesamt stützen diese Ergebnisse die proaktive Generierung als gemeinsame Lernschnittstelle, die Wahrnehmung, Gedächtnisbildung und zeitgerechte Reaktion in der Streaming-Video-Interaktion verbindet.
One-sentence Summary
Researchers from NJU, PJLAB, JD, and other affiliated institutions propose OneStreamer, a streaming video interaction model that jointly learns query-independent evidence recording and task response through shared proactive generation using Proactive Hierarchical Caption Memory and Proactive State Transition Learning, and whose OneStreamer-1M dataset and 4B model achieve state-of-the-art results across all eight evaluated streaming video understanding benchmarks.
Key Contributions
- OneStreamer jointly learns query-independent evidence recording and task response through proactive generation, with Proactive Hierarchical Caption Memory producing time-grounded local-detail captions and summaries of completed events that complement a recent visual window without revisiting historical visual features.
- Proactive State Transition Learning preserves supervision at all output anchors and selects representative state-change and state-persistence tokens to reduce the dominance of repeated waiting states; it outperforms dense state supervision while supervising only 27.5% of annotated state tokens.
- A streaming data synthesis pipeline aligns caption and QA targets with available evidence in content and timing, yielding OneStreamer-1M with over one million records. The resulting 4B model achieves the best results among compared methods across all eight evaluated streaming video understanding benchmarks, and ablations show that retained generated captions improve historical QA without degrading real-time perception.
Introduction
Streaming video LLMs are needed in live assistance, wearable agents, and security monitoring, where frames arrive continuously and only later queries reveal which earlier observations mattered. Prior approaches compress or retrieve historical visual tokens, but those tokens compete with current frames for limited context and can weaken real-time perception. Streaming reasoning traces are not necessarily reusable time-grounded records, and dense state-token supervision can overemphasize silence labels relative to sparse response decisions. The authors introduce OneStreamer, a model that unifies evidence recording and task response through proactive generation. It uses Proactive Hierarchical Caption Memory to produce time-aligned local-detail captions and semantic event summaries as query-independent factual records, and Proactive State Transition Learning to supervise response timing without dense silence labels. A streaming data synthesis pipeline builds OneStreamer-1M, and the 4B model reports the best results across eight online video understanding and proactive-response benchmarks.
Dataset
The authors construct OneStreamer-1M by combining synthesized streaming examples with cleaned open-source data. The synthesis pipeline has two branches: streaming caption synthesis and streaming QA synthesis. Figure 5 summarizes the pipeline and task composition.
Dataset sources and composition
- Video sources: a curated candidate pool built through semantic retrieval, scene detection, and visual-richness assessment.
- Annotation sources: Gemini and Seed generate timestamp-grounded captions and QA pairs; selected annotations from existing datasets are also used.
- Final dataset: synthesized streaming caption examples, synthesized streaming QA examples, and cleaned open-source data.
Streaming caption branch
- Filters out videos with low visual quality, prolonged static periods, black frames, excessive shot fragmentation, or decoding failures.
- Generates multi-granularity captions at frame/clip, segment, and video levels.
- Verifies all captions against the source video for factual and temporal consistency.
- Discards unsupported or temporally misaligned captions.
- Converts verified captions into causal streaming sequences:
- Frame/clip captions become
</Observe>targets at the end of their supporting intervals. - Segment captions become
</Summary>targets at segment boundaries. - After the stream ends, a full-video summary instruction uses the video-level caption as the answer target after
</Response>.
- Frame/clip captions become
Streaming QA branch
- Uses verified streaming caption annotations and selected annotations from existing datasets as metadata sources.
- Defines target capabilities for streaming interaction, then prompts Gemini and Seed with task-specific templates to generate grounded question-answer pairs.
- Localizes a coarse evidence interval for each QA pair.
- Verifies the interval by asking a VLM to answer from only that clip.
- Discards examples where the model cannot produce the correct answer.
- Uses fine-grained response-time calibration within each coarse interval:
- Reasoning-oriented QA is scored by conditional likelihood of the reference answer given the video prefix.
- Perception-oriented tasks are scored by visual-text similarity between each window and the target event.
- Selects the earliest timestamp meeting the reliability criterion, producing a localized evidence interval and a calibrated response timestamp.
How the data is used
- The synthesized examples are combined with cleaned open-source data to form OneStreamer-1M.
- The dataset provides training supervision for streaming video interaction, with captions, summaries, and QA responses aligned to observed video prefixes.
- Exact subset sizes, training split ratios, and mixture ratios are not specified in the provided excerpt.
Method
The authors propose OneStreamer, which unifies query-independent evidence recording and task response through a shared causal generation process. As shown in the figure below, the model processes a live video stream incrementally as temporally ordered clips. A vision encoder extracts visual tokens from each incoming clip, and a projector maps them into the LLM's embedding space. A Recent-N FIFO sliding window retains only the visual tokens corresponding to the latest N observed frames. These tokens are interleaved with a text history that includes accumulated caption records and dialogue. The resulting causal visual-language sequence contains only information available at the current time. The LLM predicts task-specific control tokens from this sequence and generates the associated text when the predicted state initiates an output.
The model uses task-specific control tokens to indicate whether to record evidence, continue observing, or produce a user-visible response. For memory formation, </Observe> introduces a local-detail caption of observed objects, actions, scenes, and state changes. </Summary> introduces a semantic summary of a completed event or segment. Both caption types are aligned with their source intervals and appended to the text history as factual context for subsequent predictions. Proactive QA uses </Standby> to indicate that relevant evidence is emerging but the model is not yet ready to answer. The model generates user-visible answers or other task outputs only after </Response>. The shared </Silence> token indicates continued observation without generating caption or response.
Under limited context capacity, retaining historical visual tokens can compete with the fine-grained visual evidence needed for current perception, whereas a Recent-N window alone discards distant visual history. Therefore, the authors introduce Proactive Hierarchical Caption Memory (PHCM), which complements recent visual tokens with time-aligned textual captions. As the stream unfolds, OneStreamer proactively generates these records from already observed content and retains them after their source frames leave the visual window. This heterogeneous representation extends access to past events through compact text while keeping the Recent-N visual window unchanged.
PHCM organizes observed evidence into two types of time-aligned records. </Observe> introduces dense local-detail captions describing directly observed objects, actions, scenes, and state changes over short intervals. </Summary> introduces sparser semantic summaries of completed events or segments at a coarser temporal granularity. Each record is associated with its source interval, while the two granularities form a hierarchy of local details and event-level summaries. At inference time, OneStreamer incrementally builds PHCM using only the currently available visual-language context. Outside user-visible answer generation, a predicted </Observe> or </Summary> initiates the corresponding record, whereas </Silence> continues observation without writing one. Each generated record is appended to the text history in generation order and becomes available to subsequent predictions.
Proactive streaming interaction requires the model to decide when the available evidence warrants a memory record or a task response. Dense state-token supervision can overemphasize waiting when repeated silence tokens outnumber output-initiating tokens, yet supervising only state changes omits direct supervision of when the current state should persist. To address this, the authors introduce Proactive State Transition Learning (PSTL) to preserve supervision at all output anchors and select representative state tokens associated with state changes or persistence.
PSTL is implemented through selective state-token supervision guided by task-specific output anchors. These anchors are control states that initiate textual outputs. Within each training sequence, state tokens are grouped by the transition from the preceding control state to the current one. The size of the largest group targeting an output anchor defines a supervision quota shared across all groups in that sequence. Full supervision is retained for groups within the quota, and larger groups are uniformly subsampled without replacement to match the quota. This preserves supervision for all output-initiating tokens while selecting examples of both state changes and state persistence. Unselected state tokens remain in the causal sequence but are excluded from the state loss.
Formally, let Aτ denote the output-anchor set for task τ. Let nXY be the number of occurrences of transition XY between consecutive annotated control states in a training sequence. Here X and Y range over the task's control states. The supervision quota shared across all transition groups in this sequence is defined as:
q=Xmaxa∈AτmaxnX→a.From each transition group XY, state-token indices are sampled uniformly without replacement to form TXY of size:
∣IX→Y∣=min(nX→Y,q).Only state tokens indexed by T=⋃X,YTX→Y contribute to the state loss. By preserving all output anchors and capping frequent transition groups, PSTL limits the influence of repeated silence tokens on the state objective.
To support this training process, the authors develop a reusable streaming data synthesis pipeline with two complementary branches: streaming caption synthesis and streaming QA synthesis. As shown in the figure below, this pipeline produces verified streaming captions with evidence-aligned release times and evidence-grounded streaming QA sequences with calibrated response times.
The caption branch aligns both caption content and release time with the available evidence. Local descriptions are released only after their supporting visual intervals have been observed, whereas segment-level summaries are released only after the corresponding events or segments are complete. The QA branch follows a timing principle where each response is assigned to an earlier point provided that the available evidence is sufficient to support the answer. Within each coarse evidence interval, candidate timestamps are evaluated using sliding windows. Reasoning-oriented QA is scored by the conditional likelihood of the reference answer given each video prefix, and perception-oriented tasks are scored by the visual-text similarity between each window and the target event. The earliest timestamp satisfying the reliability criterion is selected as the calibrated response time.
Experiment
OneStreamer is initialized from Qwen3-VL-4B-Instruct and fine-tuned on the OneStreamer-1M streaming video interaction dataset with additional offline data, then evaluated across streaming perception, memory, proactive-response, efficiency, and qualitative benchmarks. The main results show consistent gains over size-matched and larger baselines in both real-time perception and long-range online understanding. Ablations indicate that proactively generated caption memory provides reusable temporal evidence beyond the recent visual window, jointly improving perception and memory rather than trading one for the other, while selective supervision of state transitions and persistence strengthens proactive responses. Efficiency analysis and qualitative examples further confirm that the approach preserves distant evidence with substantially lower context and latency than retaining full visual history, and enables timely responses as relevant evidence appears.
OneStreamer, with only 4B parameters, achieves the highest scores among the compared methods on all eight online video benchmarks spanning perception, memory, and proactive response. It consistently outperforms a size-matched Qwen3-VL base model and a larger 11B competitor, with especially large gains on several perception and proactive-response metrics. The results indicate strong real-time perception and long-range online understanding without requiring a larger model. OneStreamer leads all compared methods on every benchmark in the suite despite having only 4B parameters. Compared with the size-matched Qwen3-VL base model, OneStreamer shows double-digit improvements on several perception and proactive-response benchmarks. OneStreamer also surpasses the larger 11B MOSS-VL-Realtime across both perception and memory benchmarks and proactive-response benchmarks. On ViSpeak, OneStreamer records a higher overall score than both Qwen3-VL and MOSS-VL-Realtime.
PHCM augments a recent visual window with hierarchical caption memory and improves backward memory over using the same visual window alone. It also yields small real-time perception gains over FIFO and surpasses full visual history on backward ASI, though its EPM is slightly lower than full visual history. The results indicate that caption memory provides reusable long-term context without sacrificing current perception. Retaining generated caption records improves backward ASI from 63.5 to 71.6 and backward EPM from 62.0 to 62.6 compared with FIFO. With the same recent visual window as FIFO, PHCM improves real-time perception and outperforms Full on backward ASI by 4.0 points despite using caption records instead of distant visual history.
This ablation compares state-token supervision strategies for proactive response learning. Dense focal supervision improves over dense cross-entropy, but PSTL's selective supervision of state transitions and persistence achieves the best results on all three benchmarks while using only about a quarter of annotated state tokens. Random sparse selection with the same supervision ratio performs substantially worse, suggesting that the specific token selection strategy is responsible for the gains. Dense focal loss substantially outperforms dense cross-entropy, reflecting the benefit of addressing state-label imbalance. Transition-only supervision stays competitive on ProactiveVQA but trails PSTL on OmniMMI and OVO-Timing, indicating the value of also supervising state persistence. PSTL supervises roughly a quarter of state tokens yet achieves the highest scores on ProactiveVQA, OmniMMI, and OVO-Timing, while random sparse selection at the same ratio performs much worse.
Retaining the full visual history incurs the highest answer-stage cost, with the largest context, GPU memory, and time-to-first-token. PHCM uses a recent visual window plus compact caption memory, which reduces context and memory use substantially relative to Full while adding only a small overhead over FIFO. It therefore remains much more efficient than Full and only modestly more costly than FIFO at answer time. PHCM cuts context length and GPU memory by large margins relative to retaining the full visual history, and sharply lowers time-to-first-token. Compared with FIFO, PHCM adds only a small amount of context, GPU memory, and latency while retaining access to older visual evidence through caption memory.
The evaluation covers an online video benchmark suite spanning perception, memory, and proactive response, where OneStreamer with only 4B parameters outperforms both a size-matched Qwen3-VL base model and a larger 11B competitor, with particularly strong gains in perception and proactive response. A memory ablation shows that augmenting a recent visual window with hierarchical caption memory improves backward memory over FIFO and remains competitive with or better than full visual history on backward understanding while greatly reducing context, GPU memory, and latency. A proactive-response ablation further indicates that selective supervision of state transitions and persistence achieves the best results across all proactive benchmarks, outperforming dense and random sparse strategies while using only a fraction of annotated state tokens.