OneStreamer Fixes AI Video Memory and Response
Based on research by Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li
Imagine watching a live sports broadcast where the AI narrator doesn’t just describe what is happening right now, but remembers a crucial foul from ten minutes ago to explain a current penalty. Current streaming video models struggle with this balance. They often forget the past when trying to react to the present, or they get stuck waiting for more data instead of speaking up when they have enough. This disconnect makes real-time interaction feel disjointed and unintelligent.
Researchers have introduced OneStreamer, a new approach that unifies perception, memory, and proactive response in a single system. Instead of treating video analysis as a series of isolated snapshots, OneStreamer uses a shared proactive generation process. It creates a Proactive Hierarchical Caption Memory that records time-grounded details and summaries of completed events as they happen. This allows the model to build a reusable factual history without needing to re-process old visual data. When a question arises, the model can draw on these generated captions to provide accurate answers, effectively bridging the gap between immediate observation and long-term context.
The real breakthrough lies in how the system handles timing. Traditional models often waste time in repeated waiting states, unsure if they have enough information to respond. OneStreamer uses Proactive State Transition Learning to break this cycle. It preserves supervision at every output anchor, encouraging the model to speak up when it has sufficient evidence rather than waiting indefinitely. This method not only speeds up responses but also improves accuracy. The researchers also created OneStreamer-1M, a massive dataset with over one million records, to train the model on diverse streaming tasks.
The results are striking. The four-billion-parameter model achieves the best performance across all eight evaluated streaming video understanding benchmarks, outperforming existing methods. Crucially, retaining these generated captions improves historical question-answering without degrading the model’s ability to perceive real-time events. Proactive State Transition Learning also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. This proves that proactive generation can serve as a shared interface for learning, allowing AI to remember the past while staying sharp in the present. As streaming video becomes more common in real-time applications, this ability to seamlessly blend memory and reaction will be essential for creating truly intelligent interactive systems.