Publications

Memory in Vision-Language Models: Taxonomy, Mechanisms, and Applications

Abstract

Vision-language models (VLMs) have achieved remarkable progress in multimodal understanding and generation by integrating visual and linguistic representations. However, most current VLMs lack explicit mechanisms for persistent memory, limiting their ability to maintain contextual coherence, accumulate knowledge over time, and support long-term reasoning across extended interactions. To address these limitations, a diverse set of memory mechanisms has emerged, including latent memory, key-value caches, external memory stores, retrieval-augmented memory, and hybrid memory systems. Despite rapid advances, the design space of memory in VLMs remains fragmented, and a unified understanding of its design and application remains lacking. In this survey, we provide a comprehensive review of memory mechanisms in VLMs from a system-oriented perspective. We introduce a novel four-dimensional (4D) taxonomy that organizes existing approaches along four orthogonal aspects: when memory is maintained (temporal scope), where memory is stored (storage location), what memory encodes (information stored), and how memory is accessed, updated, and utilized (memory operations). Using this taxonomy as a unifying framework, we systematically analyze representative memory-enhanced VLM architectures, review evaluation protocols for memory capabilities, and summarize key application areas, including long-video understanding, multimodal dialogue, embodied agents, and robotic reasoning. We further discuss key challenges such as scalability, memory efficiency, continual updating and forgetting, and multimodal …

Date
2026
Authors
Shao-Jun Xia, Yizhuo He, Jiashen Liu, Yuner Zhang, Yifan Jiang, Xiaoyang Chen, Liangxi Liu, Chenwen Luo, Jinbao Wang
Publisher
Preprints