Ir al contenido principalSaltar al contenido

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

Abstract

arXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to

Más info: Política de Cookies