Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents

AuthorsChidera Biringa, Lucas Yannul, Xiaowen Wang et al.

arXiv 20262026

TL;DR

Stashbird links episodic provenance to structured graph memory, cutting LoCoMo ingestion prompt tokens by 76.4× vs Graphiti while keeping accuracy within 5.5 points of Hindsight.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Long histories break context-only agents with billion-token ingestion costs

Long-term benchmarks like LongMemEval-S require roughly 115K tokens per question, making full-context prompting unreliable and extremely expensive.

Systems such as Hindsight and Graphiti need up to 700M and 3.57B ingestion prompt tokens respectively, making persistent memory, updates, and deletion impractical at scale.

HOW IT WORKS

Stashbird — episode-grounded graph memory with provenance

Stashbird’s core mechanism organizes memory into episodic records, semantic relations, preference traces, community summaries, and persisted graph state with explicit episode-level provenance.

You can think of Stashbird like a database-backed hippocampus: episodes are raw experiences, while entities, relations, and communities are indexed summaries and schemas.

By tying every semantic fact back to episodes, Stashbird enables precise incremental updates and episode-level deletion that a plain context window or flat vector store cannot support.

DIAGRAM

Retrieval and query decomposition pipeline in Stashbird

This diagram shows how Stashbird decomposes a user query, runs hybrid search over episodes, entities, relations, and communities, then fuses and reranks evidence for the answer backbone.

DIAGRAM

Benchmark evaluation and efficiency measurement for Stashbird

This diagram shows how Stashbird is evaluated across four benchmarks while tracking ingestion and retrieval workload against baselines like Hindsight and Graphiti.

PROCESS

How Stashbird Handles a Conversation Lifecycle

  1. 01

    Ingestion and Update

    Stashbird ingests timestamped episodes, builds a global summary, and uses a rolling entity context to populate episodic records, entities, and relations.

  2. 02

    Semantic Consolidation

    Stashbird reconciles entities using cosine and Levenshtein thresholds, updates semantic relations, and optionally forms community summaries over the persisted graph state.

  3. 03

    Deletion

    Stashbird applies episode anchored memory deletion, updating provenance sets P(e) and P(r), performing meta forgetting, and reconnecting episode neighbors.

  4. 04

    Retrieval

    Stashbird decomposes queries, runs hybrid search over episodes, entities, relations, and communities, fuses with reciprocal rank fusion, reranks, and sends top evidence to the backbone.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Episode grounded memory system

    Stashbird integrates episodic records, entities, relations, preference traces, community summaries, and persisted graph state with provenance, enabling structured views over the same conversation history.

  • 02

    Lifecycle operations for updates and deletion

    Stashbird introduces episode anchored memory deletion with entity meta forgetting and neighbor adoption, plus incremental updates via persisted checkpoints and rolling context restoration.

  • 03

    Efficient long term memory benchmarks

    Stashbird achieves 81.0% on LongMemEval S with GPT 4.1 mini while using 164M ingestion prompt tokens versus 700M for Hindsight and 3.57B for Graphiti.

RESULTS

By the Numbers

LongMemEval-S Overall

81.0%

+2.7 points over Hindsight with GPT-4.1-mini

LoCoMo Overall

87.5%

-1.6 points vs Hindsight with GPT-4.1-mini

EverMemBench Avg

55.4%

+0.6 points over Hindsight with GPT-4.1-mini

GroupMemBench Avg

51.1%

+26.2 points over Hindsight with GPT-5.2-chat

Across LongMemEval-S, LoCoMo, EverMemBench, and GroupMemBench, Stashbird maintains competitive or better accuracy while drastically reducing ingestion and retrieval prompt tokens compared to Hindsight and Graphiti.

BENCHMARK

By the Numbers

Across LongMemEval-S, LoCoMo, EverMemBench, and GroupMemBench, Stashbird maintains competitive or better accuracy while drastically reducing ingestion and retrieval prompt tokens compared to Hindsight and Graphiti.

BENCHMARK

LLM-as-a-judge accuracy on LongMemEval-S with GPT-4.1-mini

Overall accuracy (%) on LongMemEval-S using GPT-4.1-mini as the backbone.

BENCHMARK

LLM-as-a-judge accuracy on LoCoMo with GPT-4o-mini

Overall accuracy (%) on LoCoMo using GPT-4o-mini as the backbone.

KEY INSIGHT

The Counterintuitive Finding

Stashbird uses 164M ingestion prompt tokens on LongMemEval-S, versus 700M for Hindsight and 3.57B for Graphiti, yet reaches higher overall accuracy than Hindsight.

This is surprising because many builders assume more LLM processing and larger ingestion budgets always buy better memory quality, but Stashbird shows structured provenance can beat brute force context.

WHY IT MATTERS

What this unlocks for the field

Stashbird unlocks scalable, speaker aware, provenance tracked memory that survives updates and deletions across user, dyadic, and group conversations.

Builders can now design agents that maintain long lived, editable memories with explicit episode links, instead of fragile, ever growing prompts or opaque vector stores.

~12 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents

Answers use this explainer on Memory Papers.

Checking…