ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

AuthorsSong-Li Wu, Jingyi Wang, Zhaocheng Du, Weinan Gan

arXiv 20262026

TL;DR

ActiveMem uses a Hierarchical Latent Memory Tree with latent prefix injection and RL-trained tree actions to reach 95.57% on ALFWorld vs 90.60% for MemGenGRPO.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Flat memory causes context fragmentation and cross task interference

Existing memory systems retrieve historical trajectories or summaries as independent context fragments, overlooking procedural dependencies in multi step execution.

Flat memory retrieval introduces context fragmentation and cross task interference, causing structurally inconsistent reasoning trajectories and workflow mismatch in long horizon LLM agents.

HOW IT WORKS

ActiveMem — Hierarchical Latent Memory Trees with RL

ActiveMem builds a Hierarchical Latent Memory Tree, controlled by a Dual-Head Memory Controller with a Latent Injection Head and Tree Action Head, plus TreeMaintenance for pruning and merging.

You can think of ActiveMem like a file system with folders and subfolders, where frequently useful procedures stay pinned while redundant branches are merged or deleted.

This design lets ActiveMem retrieve dependency consistent execution paths and inject them as latent prefixes, enabling long horizon reasoning that a plain context window cannot sustain.

DIAGRAM

ActiveMem Inference and Memory Update Flow

This diagram shows how ActiveMem processes each step with fenc, fretrieve, the Dual-Head Memory Controller, TreeUpdateImmediate, and TreeMaintenance.

DIAGRAM

Training and Evaluation Pipeline for ActiveMem

This diagram shows how ActiveMem collects trajectories, updates the Hierarchical Latent Memory Tree, and optimizes ϕ with GRPO across benchmarks.

PROCESS

How ActiveMem Handles a Long Horizon Agent Trajectory

  1. 01

    Hierarchical Latent Memory Tree

    ActiveMem initializes the Hierarchical Latent Memory Tree with a virtual root node nr and stores node embeddings ei for compressed historical states.

  2. 02

    fretrieve and Greedy Traversal

    Using fenc and fretrieve, ActiveMem performs recursive greedy traversal over HLMT, selecting children by cosine similarity until the τstop threshold.

  3. 03

    Dual Head Memory Controller

    The Dual Head Memory Controller mean pools the retrieved path Pt, builds ut via MLPshared, and feeds it into the Latent Injection Head and Tree Action Head.

  4. 04

    TreeUpdateImmediate and TreeMaintenance

    TreeUpdateImmediate applies insert or update actions online, while TreeMaintenance prunes low SEMA nodes and merges similar siblings using τdel and τmerge.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Hierarchical Latent Memory Tree

    ActiveMem introduces the Hierarchical Latent Memory Tree that organizes experiences as dependency aware latent nodes, improving ALFWorld from 90.60% with MemGenGRPO to 95.57%.

  • 02

    Dual Head Memory Controller

    ActiveMem designs a Dual Head Memory Controller with a Latent Injection Head and Tree Action Head to jointly modulate reasoning and memory topology.

  • 03

    GRPO based Memory Optimization

    ActiveMem applies Group Relative Policy Optimization with hierarchical advantage to learn dynamic expansion, retrieval, and pruning policies directly from trajectory rewards.

RESULTS

By the Numbers

ALFWorld

95.57%

+4.97 over MemGenGRPO

TriviaQA

87.46%

+6.81 over MemGenGRPO

PopQA

73.20%

+10.90 over MemGenGRPO

BigCodeBench

84.39%

+8.83 over MemGenGRPO

On Qwen3-8B across ALFWorld, TriviaQA, PopQA, and BigCodeBench, ActiveMemGRPO consistently exceeds MemGenGRPO, showing that hierarchical latent memory and RL trained topology control materially improve long horizon reasoning.

BENCHMARK

By the Numbers

On Qwen3-8B across ALFWorld, TriviaQA, PopQA, and BigCodeBench, ActiveMemGRPO consistently exceeds MemGenGRPO, showing that hierarchical latent memory and RL trained topology control materially improve long horizon reasoning.

BENCHMARK

Table 1: Results on Qwen3-8B

Accuracy on ALFWorld for Qwen3-8B backbone.

BENCHMARK

Table 2: Ablation Study on Qwen3-8B

ALFWorld accuracy for ActiveMemGRPO ablations.

KEY INSIGHT

The Counterintuitive Finding

ActiveMem with a structured HLMT and frozen backbone reaches 95.57% on ALFWorld, surpassing MemGenGRPO’s 90.60% and even large parametric baselines.

This is surprising because ActiveMem keeps LLM weights fixed, yet hierarchical memory and RL trained control compensate enough to beat heavier finetuned systems.

WHY IT MATTERS

What this unlocks for the field

ActiveMem enables compact open weight models to match or exceed much larger agents on long horizon tasks by externalizing structure into a latent memory tree.

Builders can now bolt ActiveMem onto frozen LLMs to gain reusable procedural memories, efficient pruning, and dependency aware retrieval without retraining the backbone.

~14 min read← Back to papers

Related papers

Memory Architecture

A Control Architecture for Training-Free Memory Use

Yanzhen Lu, Muchen Jiang et al.

· 2026

TAG routes low-confidence steps to uncertainty-based routing, filters them with guarded acceptance with rollback, chooses between bank selection across rule and exemplar memory, and prunes via evidence-based retirement inside a unified control loop. On SVAMP and ASDiv, TAG reaches 81.0% and 85.2% accuracy, improving over the 74.0% and 77.5% no-memory baselines while a compute-matched Retry baseline stays flat.

Questions about this paper?

Paper: ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

Answers use this explainer on Memory Papers.

Checking…