SciGroveBeta
Machine Learning

Addressable Memory for Video World Models

Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taixé, Despoina Paschalidou, Jonathan Lorraine, Aljosa Osep

Featured August 16, 2026

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Video models usually forget old scenes because their memory system gets confused by new positions or mixes up old memories; WorldTrace fixes this by giving old memories special, understandable 'virtual' spots and storing them clearly, so the model can always find and use them correctly.

In depth
The paper introduces WorldTrace, a training-free memory framework that enables video world models to maintain visual persistence over long generation horizons. It addresses two critical issues: out-of-distribution temporal Rotary Positional Embeddings (RoPE) that make distant memories unaddressable, and naive compression methods that corrupt memory by averaging incompatible positional phases. WorldTrace assigns each summary slot a distinct, in-distribution virtual position and compresses memory in a canonical (unrotated) key domain, ensuring both addressability and informativeness.

Key Takeaways

  • 1
    Identifies that long-horizon failure in video world models stems from both memory addressability (due to out-of-distribution RoPE offsets) and memory informativeness (due to RoPE phase cancellation during naive compression).
  • 2
    Proposes WorldTrace, a training-free framework that assigns in-distribution virtual positions to memory slots and uses canonical key compression, realized by WorldTrace-Field for temporal coherence and WorldTrace-Landmark for episodic recall.
  • 3
    Introduces LoopBench, a novel benchmark for evaluating episodic recall in compressed caches, demonstrating that WorldTrace significantly extends visually persistent generation without requiring model retraining.

Conceptual Flow

HIGH LEVEL
1
Methodology: How WorldTrace Works

The new method keeps a small, recent memory and a special summary of older memories, making sure all memories are easy to find and understand, even very old ones.

New Video Frame
Add to Recent Memory
Recent Memory
Old Memories
2
Results: Better Memory for Video Models

The new method helps video models remember past scenes much better, making generated videos look more consistent and realistic over long periods.

Old Method: Forgets Scene
New Method: Remembers Scene
Compare Scene Recall
Inconsistent Video
Consistent Video

This breakdown was generated by SciGrove. Get the same analysis — intuition, storyboard, peer review, a runnable prototype and a glossary — on any paper you upload or paste a DOI for.