MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

AuthorsRuike Cao, Fanyu Zhao, Fugen Yao et al.

arXiv 20262026

TL;DR

MemCalib-RL uses ordered bidirectional counterfactual credit assignment to balance memory over-use and under-use, reaching 81.12 SCS on MemCalib with Qwen3.5-35B-A3B.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

LLM agents miscalibrate memory influence with directional skew (e.g. GPT-5.6-SOL SCS 46.25 and Exact 28.40)

MemCalib shows frontier models like GPT-5.6-SOL reach only 46.25 SCS and 28.40 Exact, with widespread mismatches between ideal and actual memory use levels.

Across models, MemCalib reveals strong directional skew: most over-use memory more severely than they under-use it, causing biased, low-quality responses in LLM-based agents.

HOW IT WORKS

MemCalib-RL — ordered bidirectional counterfactual credit assignment

MemCalib uses MemCalib benchmark, LLM-as-a-Judge, ordered reward channels, bidirectional counterfactual localization, and channel-wise advantage redistribution to calibrate each memory atom’s influence.

You can think of MemCalib-RL like a CPU profiler for memory: it ablates specific memory atoms, measures how they change token likelihoods, and routes reward to the exact response tokens they support or suppress.

This design lets MemCalib-RL adjust over-use and under-use at the token level, something a plain context window or uniform GRPO advantage cannot achieve.

DIAGRAM

MemCalib evaluation pipeline with atom-level use judgments

This diagram shows how MemCalib evaluates memory use by mapping each atom to Ignore, Bound, or Control and computing over-use and under-use metrics.

DIAGRAM

MemCalib-RL training loop with counterfactual localization

This diagram shows how MemCalib-RL samples rollouts, builds reward channels, runs atom ablations, and redistributes advantages during training.

PROCESS

How MemCalib Handles a MemCalib training example

  1. 01

    Benchmark Construction

    MemCalib extracts atomic propositions from eight source datasets, assembles composite memory blocks, and assigns ideal Ignore, Bound, or Control labels.

  2. 02

    Rubric Guided Judge

    MemCalib uses rubric guided LLM as a Judge to classify each atom’s actual use level in the generated response for that training example.

  3. 03

    Ordered Reward Channels

    MemCalib-RL groups atoms into nine ordered reward channels based on ideal and actual levels, assigning channel rewards for each rollout response.

  4. 04

    Bidirectional Counterfactual Localization

    MemCalib-RL performs exact atom ablation, computes token level log likelihood differences, and redistributes channel advantages with channel wise mean preserving multipliers.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    MemCalib benchmark

    MemCalib introduces a 15,000 example benchmark with 13,500 training and 1,500 test examples, evaluating atom level Ignore, Bound, and Control memory use across health, general assistance, and coding.

  • 02

    MemCalib-RL algorithm

    MemCalib-RL defines nine ordered reward channels, uses bidirectional counterfactual localization via exact atom ablation, and applies channel wise advantage redistribution for GRPO style training.

  • 03

    Directional skew analysis

    MemCalib reveals calibration seesaw behavior in GRPO and OPSD, and shows MemCalib-RL is the only method that reduces both over use and under use across Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B.

RESULTS

By the Numbers

Sample Calibration Score

81.12

+1.73 over GDPO on Qwen3.5-35B-A3B

Exact Calibration

70.16

+2.65 over GDPO on Qwen3.5-35B-A3B

Sample Calibration Score

79.54

+7.29 over GDPO on Qwen3-8B

Exact Calibration

67.89

+9.98 over GDPO on Qwen3-8B

On the 1,500 example MemCalib test set, which measures over use and under use via SCS and Exact, MemCalib-RL consistently improves both metrics over GRPO and GDPO. These results show MemCalib-RL can balance memory influence rather than shifting errors from one direction to the other.

BENCHMARK

By the Numbers

On the 1,500 example MemCalib test set, which measures over use and under use via SCS and Exact, MemCalib-RL consistently improves both metrics over GRPO and GDPO. These results show MemCalib-RL can balance memory influence rather than shifting errors from one direction to the other.

BENCHMARK

Main results across model families and scales (Sample Calibration Score, SCS)

Sample Calibration Score on the MemCalib test set for Qwen3.5-35B-A3B.

KEY INSIGHT

The Counterintuitive Finding

MemCalib shows that Qwen3-8B achieves higher SCS and Exact than Qwen3.5-35B-A3B in the base setting, despite having far fewer parameters.

This breaks the common assumption that larger LLMs automatically use memory more appropriately, highlighting that memory calibration is not simply a function of model scale.

WHY IT MATTERS

What this unlocks for the field

MemCalib-RL gives practitioners a way to explicitly tune how much each memory atom influences responses, balancing Ignore, Bound, and Control behavior.

With MemCalib and MemCalib-RL, builders can train LLM agents that respect constraints, avoid over personalization, and maintain consistent long horizon behavior in realistic memory heavy environments.

~12 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

Answers use this explainer on Memory Papers.

Checking…