COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents

AuthorsHongji Pu, Ruixiang Tang, Yongfeng Zhang

arXiv 20262026

TL;DR

COUNTERMEM uses world-model-verified counterfactual corrections plus a learned selector to boost ReAct from 39.0 to 57.5 average score (+18.5 points) while cutting tokens.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Agents ignore counterfactual experience from failed decisions

Existing agent memory mainly stores factual trajectories, never asking what would have happened under alternative actions at the same state.

Without counterfactual checks, a coding or SQL agent may repeat the same failed decision, missing corrections that could generalize across tasks and domains.

HOW IT WORKS

COUNTERMEM: World-model-verified counterfactual memory

Module I: Propose Alternatives uses a generator to produce up to K = 4 local edits after a failed action in a given state.

Module II: Check Alternatives and Store Corrections evaluates edits with executable world models and saves records in a bounded memory store with applicability conditions.

Module III: Retrieve and Select Records uses a DQN-based selector policy to choose one record or skip memory, enabling verified counterfactual guidance beyond a plain context window.

DIAGRAM

COUNTERMEM decision-time flow for a single task

This diagram shows how COUNTERMEM handles a failed decision by generating counterfactual edits, verifying them, and optionally reusing stored corrections on the same task.

DIAGRAM

COUNTERMEM evaluation and ablation pipeline

This diagram shows how COUNTERMEM builds memory, trains the selector, and runs held-out evaluation plus ablations like removing world-model verification or storage.

PROCESS

How COUNTERMEM Handles a Task Decision Unit

  1. 01

    Module I: Propose Alternatives to a Failed Action

    COUNTERMEM uses Module I to take the current state and failed action, then the generator produces up to K = 4 local alternatives tailored to the observed failure.

  2. 02

    Module II: Check Alternatives and Store Corrections

    COUNTERMEM runs each alternative through the world model, computes the improvement ∆, and admits records into the bounded memory store when ∆ exceeds ϵadm.

  3. 03

    Module III: Retrieve a Record and Decide Whether to Use It

    COUNTERMEM filters records by applicability conditions, ranks candidates using situation and reuse statistic u, and lets the selector choose one record or skip memory.

  4. 04

    Evaluation: Freeze Memory and Selector

    COUNTERMEM freezes the memory store and selector policy, then solves held-out tasks so that success and token cost reflect reuse of previously verified counterfactual corrections.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Constructing memory beyond factual trajectories

    COUNTERMEM turns checked local alternatives into records m = (q, a−, a+, e, c, u), storing action contrasts, evidence, and applicability conditions for reuse across up to B = 200 records.

  • 02

    Learning when counterfactual memory is useful

    COUNTERMEM trains a DQN selector over retrieved records and skipping, using rewards that balance task success, evaluator calls, failed attempts, and token use across 2,000 decisions.

  • 03

    Testing memory quality as well as task performance

    COUNTERMEM evaluates gains, accumulation, transfer, and ablations, showing average +12.6 points over base agents with gpt-oss-120b and token reductions of 7.7–42.0% across backbones.

RESULTS

By the Numbers

Math

57.5 score

+27.7 over Generative Agents (22.6) for ReAct gpt-oss-120b

Coding

66.6 score

+17.9 over ReAct None (48.7) with gpt-oss-120b

Text-to-SQL

47.0 score

-1.9 vs MemGPT (48.9) but +19.8 over ExpeL (27.2) for ReAct gpt-oss-120b

Tokens(M)

4.4M tokens

-1.6M vs ReAct None (6.0M) on gpt-oss-120b

On the four-domain suite (Math, Coding, Text-to-SQL, SAT) with gpt-oss-120b and ReAct, COUNTERMEM achieves 57.5, 66.6, 47.0, and 58.9 scores respectively. This MAIN_RESULT shows that COUNTERMEM converts verified counterfactual corrections into substantial accuracy gains while reducing task-run token consumption compared to baselines like ReAct, ExpeL, Generative Agents, Voyager, and MemGPT.

BENCHMARK

By the Numbers

On the four-domain suite (Math, Coding, Text-to-SQL, SAT) with gpt-oss-120b and ReAct, COUNTERMEM achieves 57.5, 66.6, 47.0, and 58.9 scores respectively. This MAIN_RESULT shows that COUNTERMEM converts verified counterfactual corrections into substantial accuracy gains while reducing task-run token consumption compared to baselines like ReAct, ExpeL, Generative Agents, Voyager, and MemGPT.

BENCHMARK

Four-domain comparison of memory mechanisms across agents and backbones (ReAct, gpt-oss-120b, Coding)

Average Coding score on the four-domain suite for ReAct with gpt-oss-120b under different memory mechanisms.

BENCHMARK

Component necessity ablation: gains over base agent

Gain over base agent on the four-domain suite when removing world-model verification or persistent counterfactual storage.

KEY INSIGHT

The Counterintuitive Finding

Best-of-N uses 8.2–10.0M tokens yet still scores lower than COUNTERMEM, which uses only 2.9–4.8M task-run tokens.

This is counterintuitive because more candidate generations and compute are often assumed to match or beat memory reuse, but COUNTERMEM shows verified counterfactual memory can be both more accurate and cheaper.

WHY IT MATTERS

What this unlocks for the field

COUNTERMEM unlocks agents that can systematically learn from failed actions by storing world-model-verified counterfactual corrections with explicit applicability conditions.

Builders can now deploy language agents that reuse precise, tested fixes across tasks and domains, reducing interaction cost while avoiding harmful misapplication of past corrections.

~12 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents

Answers use this explainer on Memory Papers.

Checking…