Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

AuthorsYefan Zhou, Yang Li, Zeyu Leo Liu et al.

arXiv 20262026

TL;DR

Just-in-Time Memory (JITMEM) defers memory curation to read time, letting a GRPO-trained curator synthesize task-adaptive payloads that beat SkillOS by up to 16.3 SR points on WebShop.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Write-time memory discards useful experience before future tasks are known

Existing agentic memory systems curate at write time, producing query-independent summaries that must serve many future tasks. This prematurely and irreversibly discards trajectory details needed later.

This breaks long-horizon LLM agents, forcing them to guess what matters before seeing downstream tasks and creating a difficult credit-assignment problem over many future interactions.

HOW IT WORKS

JITMEM: Read-time curation over a streaming memory bank

JITMEM’s core mechanism combines a Memory Bank, Retriever, Memory Curator, and frozen Agent Executor so only raw trajectories are stored and curated at read time. The UPDATE gate uses the executor-as-judge to keep only successful trajectories.

You can think of JITMEM like a card catalog plus librarian: the Memory Bank keeps full books, the Retriever finds shelves, and the Memory Curator summarizes exactly the chapters relevant to the current reader’s question.

This read-time design lets JITMEM tailor payloads to each task, avoid premature information loss, and learn curation directly from immediate task rewards in a way a plain context window cannot.

DIAGRAM

Inference Flow: Just-in-Time Memory Read Pipeline

This diagram shows how JITMEM retrieves raw trajectories and synthesizes a task-adaptive payload at inference time before executing the current task.

DIAGRAM

Training Loop with GRPO over a Fixed Training Bank

This diagram shows how JITMEM trains the Memory Curator with GRPO using a fixed training bank and a frozen Agent Executor.

PROCESS

How JITMEM Handles a Streaming Task Sequence

  1. 01

    Retrieve

    At each task xt, JITMEM uses the Retriever R over the Memory Bank Mt to fetch the top k raw trajectories most relevant to the task description.

  2. 02

    Curate

    The Memory Curator πϕ reads xt and the retrieved trajectories, then synthesizes a compact task adaptive payload pt that extracts strategies and guidance tailored to xt.

  3. 03

    Execute

    The frozen Agent Executor πL receives xt and pt in its prompt, uses the curated payload as guidance, interacts with the environment, and produces a trajectory ξt and task reward rt.

  4. 04

    Update

    JITMEM runs an executor-as-judge to assess correctness and uses UPDATE to append ξt to the Memory Bank only if the trajectory passes the quality gate.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Read-time curation enables task-adaptive memory

    JITMEM defers curation to read time so the Memory Curator sees the current task xt and can distill different payloads from the same raw trajectory for different tasks.

  • 02

    Immediate credit assignment via GRPO

    JITMEM trains the Memory Curator with GRPO, using the same task’s reward rt to update curation without task grouping, collapsing credit assignment to a single interaction.

  • 03

    Empirical gains over write-time memory

    On ALFWorld, WebShop, and τ2-bench, JITMEM improves success rate by up to +16.2 points over SkillOS and reduces executor steps by 28.4%–31.4% compared to write-time methods.

RESULTS

By the Numbers

ALFWorld SR

77.4%

+16.2 over SkillOS with Qwen3-8B executor (61.2%)

WebShop SR

32.8%

+16.3 over SkillOS with Qwen3-8B executor (16.5%)

WebShop SR Gemini

50.5%

+9.2 over SkillOS with Gemini-2.5-Pro executor (41.3%)

τ2-bench Macro SR

73.4%

+3.0 over ReasoningBank with GPT-5.4 curator (70.4%)

These results come from ALFWorld, WebShop, and τ2-bench, which test embodied control, web shopping, and multi-domain tool-use. The gains show JITMEM’s read-time curation improves task success across diverse agent benchmarks compared to ReasoningBank, MemP, and SkillOS.

BENCHMARK

By the Numbers

These results come from ALFWorld, WebShop, and τ2-bench, which test embodied control, web shopping, and multi-domain tool-use. The gains show JITMEM’s read-time curation improves task success across diverse agent benchmarks compared to ReasoningBank, MemP, and SkillOS.

BENCHMARK

ALFWorld Success Rate with Qwen3-8B Executor

Success rate on ALFWorld comparing JITMEM to write-time memory baselines using Qwen3-8B as executor.

BENCHMARK

WebShop Success Rate with Qwen3-8B Executor

Success rate on WebShop comparing JITMEM to write-time memory baselines using Qwen3-8B as executor.

KEY INSIGHT

The Counterintuitive Finding

Even an untrained JITMEM curator is competitive with or surpasses RL-trained write-time methods, reaching 61.0% SR on WebShop with Gemini-2.5-Pro versus 41.0% for SkillOS.

This is surprising because SkillOS uses a stronger curator model and RL training, yet JITMEM’s simple task-adaptive read-time curation alone yields a +20.0 point success-rate advantage.

WHY IT MATTERS

What this unlocks for the field

JITMEM unlocks reconstructive, task-conditioned episodic memory for LLM agents, letting them reuse full past trajectories without committing to fixed write-time summaries.

Builders can now plug a trained Memory Curator into different executors and get immediate, task-specific guidance from streaming experience, without designing task groups or complex write-time policies.

~14 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Answers use this explainer on Memory Papers.

Checking…