🚀NEW LABGetting Started with Claude AgentsStart lab
Memory · Agents · Multimodal

OCR-Memory

First page
OCR-Memory
Paper summary

Most agent memory systems compress trajectories into text summaries and hope the model remembers what matters, which is exactly where the information loss hides. OCR-Memory renders the agent's interaction history as images with indexed visual anchors, then retrieves via a locate-and-transcribe pipeline: the model scans visual memory, predicts the index of the relevant region, and the original text is fetched verbatim from a database. Older trajectories are stored as low-resolution thumbnails with active-recall up-sampling, and the method reaches SOTA on Mind2Web and AppWorld under strict context limits.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack