LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

AuthorsXianglong Shi, Ruijie Yang, Sirui Zhao et al.

arXiv 20262026

TL;DR

LSTMem uses hierarchical cell and hidden matrix states with cross-layer feedback to boost Qwen3-4B HotpotQA F1 from 60.47 to 68.12 (+7.65 over δ-Mem MSW).

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Long-horizon agents couple retention and exposure, hurting memory tasks (full history in context degrades performance)

Compact online memories typically use a single persistent state both to accumulate history and to serve readout, so retention and exposure are coupled.

Multi-session dialogue and multi-step agent tasks on benchmarks like MemoryAgentBench and LoCoMo then suffer: useful information stays exposed when irrelevant, and lower-layer memory must be redundantly stored in limited states.

HOW IT WORKS

LSTMem: Hierarchical Long Short-Term Online Memory

LSTMem introduces cell memory C, hidden memory D, memory readout and attention correction, gated accumulation and expression, and cross-layer feedback at each memory layer of a frozen LLM.

Like RAM plus cache, cell memory C persistently accumulates history while hidden memory D acts as a fast, controllable view that corrects attention and propagates across depth.

This design lets LSTMem selectively express stored content, build deeper memories on shallower ones, and refine states via feedback, enabling long-horizon reasoning beyond what a fixed context window can support.

DIAGRAM

Token and Depth Memory Flow in LSTMem

This diagram shows how LSTMem reads temporal and depth memories and updates cell and hidden states across layers for each token.

DIAGRAM

Training and Evaluation Pipeline for LSTMem

This diagram shows how LSTMem is trained on QASPER and the Long dataset, then evaluated on memory and general benchmarks.

PROCESS

How LSTMem Handles a History Block and Answer

  1. 01

    Memory readout and attention correction

    LSTMem projects x t l to q mem, k mem, v mem and uses memory readout and attention correction from hidden memory D to compute Δq t and Δo t that guide attention.

  2. 02

    Gated accumulation and expression

    LSTMem computes input, forget, output, and depth gates and updates cell memory C and hidden memory D via gated accumulation and expression using residuals and lower-layer states.

  3. 03

    Cross-layer feedback

    At history block boundaries, cross-layer feedback uses reconstruction gradients from higher layers to correct lower-layer cell memory C and rebuild hidden memory D from shallow to deep.

  4. 04

    Training and inference procedure

    During training and inference, LSTMem processes histories in blocks, applies feedback, then uses fixed cell memory C and hidden memory D to influence answer generation without changing backbone weights.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    LSTMem online memory with gated C and D

    LSTMem proposes an online memory for frozen LLMs that uses separately gated cell memory C and hidden memory D to supply low-rank attention corrections, improving Qwen3-4B HotpotQA F1 from 60.47 to 68.12.

  • 02

    Layer-hierarchical memory structure

    LSTMem introduces a layer-hierarchical memory where hidden memory D connects successive layers and cross-layer feedback uses higher-layer reconstruction gradients to revise lower-layer cell memory C.

  • 03

    Extensive experiments on memory benchmarks

    LSTMem demonstrates effectiveness on MemoryAgentBench, LoCoMo, and HotpotQA, achieving a 45.08 MAB average on Qwen3-4B-Instruct versus 38.85 for δ-Mem (MSW) and retaining 92.10% of δ-Mem decoding speed.

RESULTS

By the Numbers

HotpotQA F1

68.12 F1

+7.65 over δ-Mem (MSW) on Qwen3-4B-Instruct

MemoryAgentBench Avg

45.08 score

+6.23 over δ-Mem (MSW) on Qwen3-4B-Instruct

LoCoMo Avg F1

53.60 F1

+4.48 over δ-Mem (MSW) on Qwen3-4B-Instruct

Decoding speed retention

92.10 %

LSTMem retains 92.10% of δ-Mem equivalent single-step decoding speed

On Qwen3-4B-Instruct, LSTMem is evaluated on HotpotQA, MemoryAgentBench, and LoCoMo, which test multi-hop QA, long-horizon agents, and conversational memory. The 68.12 HotpotQA F1 and 45.08 MAB average show that LSTMem’s hierarchical gated memory substantially strengthens long-range retrieval and test-time learning compared to δ-Mem (MSW).

BENCHMARK

By the Numbers

On Qwen3-4B-Instruct, LSTMem is evaluated on HotpotQA, MemoryAgentBench, and LoCoMo, which test multi-hop QA, long-horizon agents, and conversational memory. The 68.12 HotpotQA F1 and 45.08 MAB average show that LSTMem’s hierarchical gated memory substantially strengthens long-range retrieval and test-time learning compared to δ-Mem (MSW).

BENCHMARK

Main results across three base models on HotpotQA F1

F1 on HotpotQA for LSTMem versus δ-Mem (MSW) and Context2LoRA on each backbone.

BENCHMARK

LoCoMo F1 on Qwen3-4B-Instruct

Average F1 on LoCoMo for LSTMem compared to δ-Mem (MSW), Context2LoRA, and the unadapted backbone.

KEY INSIGHT

The Counterintuitive Finding

Removing hidden-memory propagation across layers lowers MemoryAgentBench by only 1.92 and 1.77 points, but test-time learning drops by 6.21 and 8.00 points.

This is counterintuitive because depth coupling might seem like a minor architectural detail, yet LSTMem shows that hierarchical memory pathways are crucial specifically for learning new tasks from history.

WHY IT MATTERS

What this unlocks for the field

LSTMem unlocks long-horizon assistants that can retain specific facts and task instructions across very long histories without bloating the context window or retraining the backbone.

Builders can now attach LSTMem to frozen LLMs to get hierarchical, controllable online memory that improves benchmarks like HotpotQA, MemoryAgentBench, and LoCoMo while preserving general capabilities such as IFEval accuracy.

~12 min read← Back to papers

Related papers

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

Answers use this explainer on Memory Papers.

Checking…