Randomized trial demonstrates accurate user recall using persistent KV cache in large language models, suggesting advancements in AI memory design.
We demonstrate persistent KV cache across sessions in large language models: the internal attention state of a transformer is serialized to disk, reloaded in a new process, and used directly for inference without reprocessing the original context. Using Mistral 7B (bfloat16) on a Tesla T4 GPU, an 84-token conversational context encoded into a 10.52 MB file enables accurate recall of user identity and research topics across session boundaries — with no text history passed to the model. The model correctly identifies the user by name using only the deserialized KV cache as input. This establishes a proof-of-concept for stateful AI memory: persistent, model-native, and free of retrieval overhead. Seventh paper in a series on KV cache optimization for consumer GPU inference.
No takes yet. Share an insight, caveat, or question.
Andrew Gaveta (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: