Grounding Memory Summarization in Utility Intent

AuthorsZhenyu Lei, Mingjia Shi, Xingbo Fu et al.

arXiv 20262026

TL;DR

MemSuit uses utility-aware self-distillation plus entry-aware retrieval calibration to raise average F1 on LoCoMo to 35.56 vs 30.09 for SimpleMem (+5.47).

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Memory summaries drop key evidence for temporal and open-domain queries

Existing memory summarizers optimize for human-facing faithfulness, not downstream utility, so they often discard the fine-grained facts future queries need.

On LoCoMo with Qwen2.5-3B-Instruct, vanilla summarization yields only 19.1 F1 overall, and loses substantial evidence for temporal and open-domain questions.

HOW IT WORKS

MemSuit: Adaptive Entry Decomposition + Utility-Aware Self-Distillation + Retrieval Calibration

MemSuit’s core mechanism combines Adaptive Entry Decomposition, Utility-Aware Self-Distillation, and Entry-Aware Retrieval Calibration to align memory with query-answer utility.

Think of MemSuit as turning conversations into a card catalog of fact-dense entries, then training the retriever to find the right cards for each future question.

This KEY_MECHANISM lets MemSuit preserve multi-faceted, query-relevant evidence in compact entries that a plain context window or vanilla summarizer would erase.

DIAGRAM

MemSuit Query-Time Memory Retrieval Pipeline

This diagram shows how MemSuit retrieves and uses student-generated memory entries to answer a query at inference time.

DIAGRAM

MemSuit Training Loop with Distillation and Calibration

This diagram shows how MemSuit trains the student summarizer and retriever using teacher entries and query-answer supervision.

PROCESS

How MemSuit Handles a Conversational Query Session

  1. 01

    Adaptive Entry Decomposition

    MemSuit uses Adaptive Entry Decomposition to let the teacher split each conversation block into multiple self-contained entries, avoiding collateral erasure across topics.

  2. 02

    Utility-Aware Self-Distillation

    MemSuit applies Utility-Aware Self-Distillation so the student summarizer learns to reproduce teacher entries from raw blocks without seeing query-answer pairs.

  3. 03

    Entry-Aware Retrieval Calibration

    MemSuit runs Entry-Aware Retrieval Calibration, using teacher entries and LLM-judged positives to train the embedding model with a contrastive objective.

  4. 04

    Inference with Student and Calibrated Retriever

    MemSuit stores student entries, retrieves top K entries per query using the calibrated retriever, and passes them to the fixed reader to generate answers.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Downstream-Utility-Oriented Memory Summarization

    MemSuit reframes memory summarization around downstream utility, showing utility-aware summaries nearly double overall F1 from 19.1 to 37.4 in oracle analysis on LoCoMo.

  • 02

    Adaptive Entry Decomposition and Self-Distillation

    MemSuit introduces Adaptive Entry Decomposition plus Utility-Aware Self-Distillation so a query-agnostic student learns to emit multiple fact-dense entries per block.

  • 03

    Entry-Aware Retrieval Calibration

    MemSuit adds Entry-Aware Retrieval Calibration, improving Recall@1 of grounding entries from 25.77 to 35.05 for Llama-3.1-8B-Instruct and from 21.78 to 26.73 for Qwen2.5-3B-Instruct.

RESULTS

By the Numbers

Avg F1

35.56

+5.47 over SimpleMem

Avg BLEU

27.73

+5.49 over SimpleMem

Temporal F1

40.77

+5.66 over SimpleMem

Open F1

25.42

+3.01 over SimpleMem

These numbers come from LoCoMo, a long-horizon conversational benchmark with multi-hop, temporal, open-domain, and single-hop queries. MemSuit’s MAIN_RESULT shows that grounding summarization in utility intent yields higher answer quality than SimpleMem and other memory systems.

BENCHMARK

By the Numbers

These numbers come from LoCoMo, a long-horizon conversational benchmark with multi-hop, temporal, open-domain, and single-hop queries. MemSuit’s MAIN_RESULT shows that grounding summarization in utility intent yields higher answer quality than SimpleMem and other memory systems.

BENCHMARK

Main Results on LoCoMo with Llama-3.1-8B-Instruct

Macro-average F1 on LoCoMo across multi-hop, temporal, open-domain, and single-hop queries.

BENCHMARK

Ablation Study of MemSuit Components (Llama-3.1-8B-Instruct)

Macro-average F1 on LoCoMo for MemSuit and ablated variants.

KEY INSIGHT

The Counterintuitive Finding

MemSuit shows that utility-aware summarization nearly doubles overall F1 on LoCoMo, jumping from 19.1 with vanilla summaries to 37.4 when conditioned on query-answer pairs.

This is surprising because memory summarization is usually optimized for human-style coverage and coherence, yet those objectives leave large amounts of query-supporting evidence unused.

WHY IT MATTERS

What this unlocks for the field

MemSuit unlocks memory systems that store compact, fact-dense entries explicitly tuned to support future questions, rather than generic narrative summaries.

Builders can now design LLM agents whose memories are trained against downstream utility, enabling more reliable long-horizon reasoning without exploding token budgets.

~12 min read← Back to papers

Related papers

Memory Architecture

A Control Architecture for Training-Free Memory Use

Yanzhen Lu, Muchen Jiang et al.

· 2026

TAG routes low-confidence steps to uncertainty-based routing, filters them with guarded acceptance with rollback, chooses between bank selection across rule and exemplar memory, and prunes via evidence-based retirement inside a unified control loop. On SVAMP and ASDiv, TAG reaches 81.0% and 85.2% accuracy, improving over the 74.0% and 77.5% no-memory baselines while a compute-matched Retry baseline stays flat.

Memory Architecture

ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

Song-Li Wu, Jingyi Wang et al.

arXiv 2026 · 2026

ActiveMem organizes experiences with a Hierarchical Latent Memory Tree, Dual-Head Memory Controller, Latent Injection Head, and Tree Action Head to build dependency-aware memory paths. On ALFWorld with Qwen3-8B, ActiveMemGRPO reaches 95.57% vs MemGenGRPO’s 90.60%, while also boosting TriviaQA from 80.65% to 87.46%.

Questions about this paper?

Paper: Grounding Memory Summarization in Utility Intent

Answers use this explainer on Memory Papers.

Checking…