Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents

AuthorsDonghua Cai, Yongheng Deng, Yifei Wang et al.

arXiv 20262026

TL;DR

Threader uses incremental topic-coherent segmentation plus multi-view, multi-signal retrieval to reach 90.60% overall accuracy on LongMemEval, +5.80 points over EverMemOS on Knowledge-Update.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Memory construction collapses in long-horizon dialogue: Mem0 drops to 5.00% accuracy at 8 concatenated sessions

As conversations grow to 8 concatenated sessions, Mem0’s QA accuracy falls from 55.00% to just 5.00%, while constructed memory units shrink sharply.

In these high-entropy settings, LLM-based memory construction like Mem0 loses query-relevant evidence, causing downstream agents to miss facts and answer incorrectly.

HOW IT WORKS

Threader: Incremental segmentation plus multi-view, multi-signal access over raw interactions

Threader’s core mechanism combines Incremental Topic-Coherent Segmentation, Multi-View Segment Representation, and Multi-Signal Evidence Retrieval to turn raw interaction logs into usable long-term memory.

You can think of Threader like a card catalog for dialogue: it files each conversation turn into topic threads, then looks up cards using multiple semantic and lexical views.

This design lets Threader retrieve complete, coherent evidence from 26K-token histories in PersonaMem, something a plain context window or fixed-size chunking cannot reliably achieve.

DIAGRAM

Threader Query-Time Evidence Retrieval Flow

This diagram shows how Threader performs two-stage multi-signal retrieval over segmented conversations when answering a new query.

DIAGRAM

Evaluation Pipeline and Cost Comparison for Threader

This diagram shows how Threader is evaluated on LongMemEval and PersonaMem, including construction and query-time cost measurement against baselines.

PROCESS

How Threader Handles a Long-Horizon Conversational Query

  1. 01

    Incremental Topic-Coherent Segmentation

    Threader streams turns through Incremental Topic-Coherent Segmentation, using a lightweight BERT segmenter to decide CONTINUE or SEGMENT for each new turn and build topic-aligned segments.

  2. 02

    Multi-View Segment Representation

    For every segment, Threader constructs Multi-View Segment Representation with user-only, assistant-only, and full views, encoding each via a shared encoder to avoid collapsing heterogeneous evidence.

  3. 03

    Multi-Signal Evidence Retrieval

    Given a query, Threader runs Multi-Signal Evidence Retrieval, combining dense multi-view similarity, BM25 lexical scores, and cross-encoder reranking in a two-stage coarse-to-fine pipeline.

  4. 04

    Response Generation with Retrieved Segments

    Threader feeds the top-K temporally consistent segments to GPT-5-mini or GPT-5-nano, letting the response LLM synthesize answers while relying on Threader’s evidence-complete context.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Threader memory framework for long-horizon, high-entropy dialogue

    Threader reframes agent memory from LLM rewriting to structure-aware access, using Incremental Topic-Coherent Segmentation to avoid construction-time information loss and handle 26K-token PersonaMem sessions.

  • 02

    Unified system with segmentation, multi-view representation, and multi-signal retrieval

    Threader integrates Multi-View Segment Representation and Multi-Signal Evidence Retrieval, achieving 90.60% overall accuracy on LongMemEval and 69.95% on PersonaMem with GPT-5-mini.

  • 03

    Cost-efficient memory access without LLM construction

    Threader eliminates construction-time LLM calls, building indices for 100K tokens in 18.33s and serving queries in 8.66s, compared to 2032.45s construction for A-Mem and 2013.97s for EverMemOS.

RESULTS

By the Numbers

Overall

90.60%

+6.84 over EverMemOS on LongMemEval with GPT-5-mini

Knowledge-Update

93.59%

+5.80 over EverMemOS on LongMemEval with GPT-5-mini

Temporal-Reasoning

90.98%

+9.03 over HippoRAG2 and +11.02 over EverMemOS with GPT-5-mini

PersonaMem Overall

69.95%

+4.75 over EverMemOS and +5.26 over MemOS with GPT-5-mini

On LongMemEval, which tests information extraction, multi-session reasoning, temporal reasoning, and knowledge updates, Threader’s 90.60% overall accuracy shows that structure-aware access over raw interactions can beat EverMemOS and agentic RAG baselines. On PersonaMem’s 26K-token personalization tasks, Threader’s 69.95% exact match demonstrates robust recall of user-specific facts across long, entangled histories.

BENCHMARK

By the Numbers

On LongMemEval, which tests information extraction, multi-session reasoning, temporal reasoning, and knowledge updates, Threader’s 90.60% overall accuracy shows that structure-aware access over raw interactions can beat EverMemOS and agentic RAG baselines. On PersonaMem’s 26K-token personalization tasks, Threader’s 69.95% exact match demonstrates robust recall of user-specific facts across long, entangled histories.

BENCHMARK

Overall Accuracy on LongMemEval with GPT-5-mini

Overall accuracy (%) on LongMemEval across memory and agentic RAG systems.

BENCHMARK

Overall Accuracy on PersonaMem with GPT-5-mini

Overall exact match (%) on PersonaMem personalization benchmark.

KEY INSIGHT

The Counterintuitive Finding

Threader achieves 82.60% accuracy and 82.55% full evidence recall with Top-5 segments, while fixed 2K-token chunking drops to 63.60% accuracy.

This is surprising because larger chunks should provide more context, yet Threader’s topic-coherent segmentation shows that semantically aligned units beat brute-force context expansion.

WHY IT MATTERS

What this unlocks for the field

Threader unlocks long-horizon conversational agents that can reliably recall sparse, early evidence without paying 21× token amplification for LLM-based memory construction.

Builders can now deploy cost-efficient, evidence-complete memory for 100K-token histories, using Threader’s segmentation and multi-signal retrieval instead of brittle summaries and expensive agentic RAG loops.

~12 min read← Back to papers

Related papers

RAG

A Dynamic Retrieval-Augmented Generation System with Selective Memory and Remembrance

Okan Bursa

· 2026

Adaptive RAG Memory (ARM) augments a standard retriever–generator stack with a Dynamic Embedding Layer and Remembrance Engine that track usage statistics and apply selective remembrance and decay to embeddings. On a lightweight retrieval benchmark, ARM achieves NDCG@5 ≈ 0.9401 and Recall@5 = 1.000 with 22M parameters, matching larger baselines like gte-small while providing the best efficiency among ultra-efficient models.

RAG

Are We Ready For An Agent-Native Memory System?

Wei Zhou, Xuanhe Zhou et al.

arXiv 2026 · 2026

Are We Ready For An Agent-Native Memory System? analyzes Memory Representation and Storage, Memory Extraction, Memory Retrieval and Routing, and Memory Maintenance across 12 real systems like Mem0, Zep, MemTree, LightMem, MemOS, MemoryOS, and A-MEM. The study’s main result is that structured systems such as Zep reach 48.0 LLM Judge Accuracy on LongMemEval while Long Context reaches 19.0, and that localized maintenance strategies like LightMem achieve 48.3% normalized utility at only 3.67 s per query.

Questions about this paper?

Paper: Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents

Answers use this explainer on Memory Papers.

Checking…