EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory

AuthorsXuanyu Meng, Xing Fan, Xinyi Fan et al.

arXiv 20262026

TL;DR

EnSIMem uses entity structured indexing with requirement aware retrieval to reach 92.8% accuracy on LongMemEval and 90.6% on LoCoMo.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Long term agents lose entity specific evidence

EnSIMem notes that existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult to identify the correct entity, property, and evidence.

This harms long term conversational agents, where missing temporal qualifiers or provenance causes wrong answers on benchmarks like LoCoMo and LongMemEval despite having the information in history.

HOW IT WORKS

EnSIMem — Entity structured indexing for episodic memory

EnSIMem combines Theme-coherent episodic memory construction, Dialogue-grounded entity-property indexing, Granularity-controlled property alignment, and Requirement-aware evidence localization into a reusable long term memory substrate.

Think of EnSIMem like a card catalog for an agent’s life: episodes are books, entity property records are catalog cards, and queries pull only the needed pages instead of the whole library.

This design lets EnSIMem localize evidence via structure while still reasoning over original dialogue episodes, something a plain context window or lossy summaries cannot provide.

DIAGRAM

Online requirement aware retrieval flow

This diagram shows how EnSIMem processes a user request into evidence requirements, retrieves episodes, and generates an answer during online interaction.

DIAGRAM

Evaluation and ablation pipeline for EnSIMem

This diagram shows how EnSIMem is evaluated on LoCoMo and LongMemEval and how ablations on episodes, properties, and retrieval are run.

PROCESS

How EnSIMem Handles a Long term conversational request

  1. 01

    Theme-coherent episodic memory construction

    EnSIMem segments long dialogue into theme coherent episodes, preserving contiguous turns, temporal order, and provenance to support later reasoning.

  2. 02

    Dialogue-grounded entity-property indexing

    EnSIMem extracts entities and salient events from each episode, building records of the form [entity][entity type][property : value] linked to source turns.

  3. 03

    Granularity-controlled property alignment

    EnSIMem normalizes properties to an intermediate granularity so that corpus and query properties align across paraphrases without collapsing distinct events.

  4. 04

    Requirement-aware evidence localization

    EnSIMem decomposes the new request into evidence requirements, classifies query type, retrieves episodes via structured matching and dense fallback, and builds an evidence rich short context buffer.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Theme-coherent episodic memory construction

    EnSIMem introduces theme coherent, contiguous episodes that preserve local context and provenance, improving accuracy by 2.27 points over per session and 5.68 over per turn units on LoCoMo.

  • 02

    Dialogue-grounded entity-property indexing

    EnSIMem builds dialogue grounded entity property records linked to original turns and multimodal fields, providing precise addresses into episodic memory rather than anonymous chunks.

  • 03

    Requirement-aware evidence localization

    EnSIMem decomposes requests into explicit evidence requirements and uses query type aware adaptive retrieval, reaching 96.59% accuracy versus 86.36% for dense retrieval only in ablations.

RESULTS

By the Numbers

LoCoMo Overall

90.6%

+4.3 over MEMORA (P)

LoCoMo Single-hop

94.2%

+2.4 over MEMORA (P)

LongMemEval Average

92.8%

+5.4 over MEMORA (P)

LoCoMo Context tokens

7470

1029 fewer than MEMORA (P)

LoCoMo and LongMemEval are long term conversational memory benchmarks testing multi hop, temporal, and multi session recall. EnSIMem’s 92.8% and 90.6% accuracies show that entity structured indexing plus requirement aware retrieval yields higher accuracy with compact, evidence focused contexts than MEMORA (P).

BENCHMARK

By the Numbers

LoCoMo and LongMemEval are long term conversational memory benchmarks testing multi hop, temporal, and multi session recall. EnSIMem’s 92.8% and 90.6% accuracies show that entity structured indexing plus requirement aware retrieval yields higher accuracy with compact, evidence focused contexts than MEMORA (P).

BENCHMARK

LoCoMo overall accuracy comparison

Overall accuracy on LoCoMo under the Memora evaluation protocol.

BENCHMARK

LongMemEval average accuracy comparison

Average accuracy on LongMemEval across categories.

KEY INSIGHT

The Counterintuitive Finding

EnSIMem’s theme coherent episodes reach 96.59% accuracy, beating both per session (94.32%) and per turn (90.91%) units despite being a middle granularity.

This is surprising because many builders assume finer per turn segmentation or full session context is always better, but EnSIMem shows a carefully chosen intermediate unit yields higher accuracy.

WHY IT MATTERS

What this unlocks for the field

EnSIMem unlocks long term agents that can reliably answer entity specific, temporal, and aggregative questions from compact, evidence rich buffers instead of full transcripts.

Builders can now design agents that reuse a persistent, query independent memory index while serving many future queries with grounded answers and inspectable evidence paths.

~12 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory

Answers use this explainer on Memory Papers.

Checking…