SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

AuthorsFengrong Wan, Chengcan Wu, Ningtao Lyu

arXiv 20262026

TL;DR

SodaMem uses an evidence-grounded temporal graph memory with FactEvents and connection-density fusion to reach 92.8% accuracy on LongMemEval-S at $0.00161 per question.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Flat RAG diaries fail on currency and temporal reasoning

LLM agents assisting users over weeks must remember what is currently true, not merely what was once said in append-only logs.

Flat RAG diaries and Markdown logs leave currency, provenance, and ordered temporal reasoning unresolved, causing incoherent personal-memory QA over multi-session histories.

HOW IT WORKS

SodaMem: Evidence-Grounded Temporal Graph Memory

SodaMem centers on FactEvents, a temporal graph store, hybrid BM25–dense indexes, and a planner–reader loop to keep user memory coherent and citable.

You can think of SodaMem like a card catalog plus timeline: every fact is a card with timestamps and links, and retrieval walks the catalog instead of a single long notebook.

This evidence-grounded temporal graph lets SodaMem enforce supersession, soft time windows, and citation-backed answers that a plain context window or flat RAG store cannot provide.

DIAGRAM

SodaMem Query-Time Memory Retrieval Pipeline

This diagram shows how SodaMem performs multi-tunnel retrieval and connection-density fusion when answering a user question.

DIAGRAM

LongMemEval-S Evaluation Setup for SodaMem

This diagram shows how SodaMem is evaluated on LongMemEval-S, including ingest, answering, and cost measurement.

PROCESS

How SodaMem Handles a LongMemEval-S Question

  1. 01

    IngestSession

    SodaMem runs IngestSession over each dialogue session, extracting FactEvents with provenance spans and normalized modality and dates before writing facts and edges.

  2. 02

    Store and maintain

    SodaMem persists FactEvents in SQLite with hybrid BM25–dense indexes, maintaining SUPERSEDES, CONTRADICTS, and UPDATES edges plus optional dreaming and timeline resolution.

  3. 03

    MultiTunnelRetrieve

    SodaMem executes MultiTunnelRetrieve, using graph, BM25, and embedding tunnels with validity gates and soft time bonuses to build a fused evidence pool ranked by connection density.

  4. 04

    Planner–Reader Loop

    SodaMem runs a planner–reader loop where the planner calls memory tools to expand evidence and the reader composes an evidence-grounded answer with mandatory citations.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Problem framing

    SodaMem isolates currency, temporal structure, provenance, and association as failure modes of Markdown and flat RAG for long-horizon personal assistants, guiding the FactEvent schema and graph design.

  • 02

    System

    SodaMem introduces an ingest–store–planner–reader pipeline with typed FactEvents, hybrid multi-signal retrieval, supersession semantics, and a timeline-resolution layer targeting temporal-reasoning errors.

  • 03

    Cost–accuracy evaluation

    SodaMem reports 92.8% accuracy on LongMemEval-S at mean $0.00161 per question and compiles a cost–accuracy map showing SodaMem strictly dominates several higher-cost, lower-accuracy systems.

RESULTS

By the Numbers

Accuracy

92.8%

+15.0 percentage points over MemOS (77.8%)

Questions

464/500 correct

best of N=3 runs on LongMemEval-S

Mean cost per 103Q

$1.61

Flash planner+reader spend excluding ingest and judge

Median cost per 103Q

$1.11

≈14.6k tokens per question, 25% cheaper than mean

On LongMemEval-S, which probes multi-session recall, updates, temporal reasoning, preference, and abstention, SodaMem reaches 92.8% accuracy at Flash-tier cost. This MAIN_RESULT shows SodaMem can maintain a coherent, updatable, and citable personal memory while staying near the accuracy frontier without Opus or GPT-4o-level spend.

BENCHMARK

By the Numbers

On LongMemEval-S, which probes multi-session recall, updates, temporal reasoning, preference, and abstention, SodaMem reaches 92.8% accuracy at Flash-tier cost. This MAIN_RESULT shows SodaMem can maintain a coherent, updatable, and citable personal memory while staying near the accuracy frontier without Opus or GPT-4o-level spend.

BENCHMARK

LongMemEval-S Accuracy vs Cost for Memory Systems

Accuracy on LongMemEval-S and relative position of SodaMem among systems with estimable API cost.

KEY INSIGHT

The Counterintuitive Finding

SodaMem achieves 92.8% accuracy on LongMemEval-S at a mean cost of only $0.00161 per question with deepseek-v4-flash.

This is surprising because systems like agentmemory V4 and Mem0 2026 use Claude Opus and GPT-4o at roughly 10–40× higher estimated cost to gain only 1.6–3.4 percentage points of accuracy.

WHY IT MATTERS

What this unlocks for the field

SodaMem unlocks a practical way to keep long-horizon personal agents evidence-grounded, temporally coherent, and cost-efficient using FactEvents and connection-density fusion.

Builders can now deploy multi-week memory with explicit supersession, soft temporal reasoning, and auditable citations at Flash-tier spend, instead of relying on opaque long-context prompts or expensive high-end models.

~12 min read← Back to papers

Related papers

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

Answers use this explainer on Memory Papers.

Checking…