Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

AuthorsJiahe Geng, Jinpeng Wang, Kun Yuan

arXiv 20262026

TL;DR

RSM-full uses a cosine-gated max-member merge rule and atom-aware grouped packing to reach 0.311 (83% of Full-Context) at 32% token cost on AMA-Bench.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Compact-memory agents under 5k tokens lose quality fast

Full-context prompting preserves quality but its token cost grows linearly with interaction length, making it impractical beyond compact budgets ≤5k tokens.

Under tight prompt budgets, long-horizon LLM agents must drop past context, which hurts answer quality, causal reasoning, and state updates on benchmarks like AMA-Bench.

HOW IT WORKS

Cosine-gated clustering plus atom-aware packing

RSM-full centers on Cosine-gated max-member merge for streaming writes and Atom-aware grouped context assembly for compact prompt construction.

You can think of RSM-full like a smart card catalog: it clusters related experiences into atoms, then pulls grouped folders instead of scattered pages when answering.

This max-member gate and grouped packer let RSM-full keep high-quality, structured memories within 2k–5k tokens, something a plain context window or flat retrieval cannot do.

DIAGRAM

Streaming write and read flow in RSM-full

This diagram shows how RSM-full writes incoming chunks into atoms with cosine-gated max-member merge and later reads them with atom-aware grouped assembly.

DIAGRAM

AMA-Bench ablation design for clustering and packing

This diagram shows how AMA-Bench evaluates RSM-full against Online K-Means, DP-means, and flat concatenation in a 2×3 factorial ablation.

PROCESS

How RSM-full Handles an AMA-Bench Episode

  1. 01

    Cosine-gated max-member merge

    RSM-full encodes each trajectory chunk h, finds m* = arg max cos(h, μm), and uses the cosine-gated max-member merge rule with threshold τ≈Q0.70.

  2. 02

    Atom-aware grouped context assembly

    At query time, RSM-full retrieves top K atoms, ranks members within each atom, and packs chunks grouped by atom with headers and temporal ordering.

  3. 03

    |v⊤₁,m q| retrieval rule

    RSM-full scores each atom m using |v⊤₁,m q|, where v₁,m is the top singular vector of the atom’s member matrix, equivalent to normalized-centroid cosine on BGE.

  4. 04

    AMA-Bench compact-memory operating point

    RSM-full runs GPT-4o-mini agents under a 4k-token budget, then llmmelon GPT-5.4 judges all 2,496 QAs to measure average correctness and quality–token trade-offs.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Strong compact-memory operating point

    RSM-full reaches 0.311 average correctness on AMA-Bench, delivering 83% of Full-Context quality at 32% of the 12,519-token cost at a 4k budget.

  • 02

    Mechanism account via merging and packing

    RSM-full’s cosine-gated max-member merge rule adds +5.74 pp over Online K-Means and +5.45 pp over DP-means, while atom-aware packing adds +5.02 pp over flat concatenation.

  • 03

    Replication on RealMem persona memory

    RSM-full achieves 0.4684 accuracy on RealMem, improving Budget-RAG by +0.69 pp and Streaming-Proto by +2.97 pp under a 4k-token budget.

RESULTS

By the Numbers

Avg (per ep) Recall

0.311

+1.9 pp over Budget-RAG at ~4k tokens

TotalTok

4,001 tokens

32% of Full-Context’s 12,519 tokens on AMA-Bench

Avg correctness RealMem

0.4684

+0.69 pp over Budget-RAG and +2.97 pp over Streaming-Proto

Clustering ablation gain

5.74 pp

RSM-full vs Online K-Means with |v1 q| retrieval on AMA-Bench

On AMA-Bench, which tests long-horizon tool-use and reasoning, RSM-full achieves 0.311 average correctness at 4,001 tokens versus Full-Context’s 0.373 at 12,519 tokens. On RealMem, which probes persona memory over 135–276 sessions, RSM-full’s 0.4684 accuracy confirms that compact clustered memory can preserve quality under 4k-token budgets.

BENCHMARK

By the Numbers

On AMA-Bench, which tests long-horizon tool-use and reasoning, RSM-full achieves 0.311 average correctness at 4,001 tokens versus Full-Context’s 0.373 at 12,519 tokens. On RealMem, which probes persona memory over 135–276 sessions, RSM-full’s 0.4684 accuracy confirms that compact clustered memory can preserve quality under 4k-token budgets.

BENCHMARK

AMA-Bench end-to-end comparison at ~4k tokens

Average correctness on AMA-Bench at the compact-memory 4k-token operating point.

BENCHMARK

RealMem compact-memory comparison across methods

Three-seed mean accuracy on RealMem under a 4k-token budget.

KEY INSIGHT

The Counterintuitive Finding

On BGE-normalized AMA-Bench, swapping |v⊤₁,m q| retrieval for centroid cosine on RSM-full clusters changes average correctness by only +0.33 pp (p=0.40).

This is surprising because a more sophisticated singular-vector retrieval rule seems strictly better, yet RSM-full’s gains come almost entirely from merging and packing rather than retrieval scoring.

WHY IT MATTERS

What this unlocks for the field

RSM-full shows that carefully designed streaming clustering and atom-aware packing can define strong quality–token Pareto points for long-horizon agents.

Builders can now deploy LLM agents that retain rich, structured memories over many sessions while staying within 2k–5k prompt tokens, avoiding the costs of full-context prompting.

~14 min read← Back to papers

Related papers

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Answers use this explainer on Memory Papers.

Checking…