FlowState: Execution State as Memory for Long-Horizon LLM Agents

AuthorsMinghao Li, Bangyan Li, Zifan Wang et al.

arXiv 20262026

TL;DR

FlowState uses persistent execution state with Incremental State Update and Progressive State Access to boost MemoryArena success by 4.55pp while cutting tokens by 43.2%.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Long-horizon agents lose crucial earlier decisions and evidence

Long-horizon tasks need earlier information, but full histories raise context costs and compression risks losing details needed later, especially when relevance emerges gradually.

When travel planning or policy-constrained service interactions forget earlier constraints, decisions, or tool observations, agents mis-handle later subtasks and fail benchmarks like MemoryArena and τ3-Bench.

HOW IT WORKS

FlowState — execution state as reusable memory

FlowState’s core mechanism links a Persistent State Repository, Historical State Index, Active State, Incremental State Update, and Progressive State Access into a single execution loop.

Think of FlowState like RAM plus disk: the Active State is fast working memory, while the Persistent State Repository and Historical State Index act as structured long-term storage and an index.

This design lets FlowState revisit semantically typed execution states and their evidence on demand, something a plain context window or naive compression cannot achieve.

DIAGRAM

State-driven execution loop with ISU and PSA

This diagram shows how FlowState runs Algorithm 1, combining Incremental State Update and Progressive State Access within each ReAct step.

DIAGRAM

Evaluation pipeline and ablation design for FlowState

This diagram shows how FlowState is evaluated on MemoryArena and τ3-Bench, including baselines and ablation variants.

PROCESS

How FlowState Handles a Sequential Interaction Task

  1. 01

    Persistent Execution State

    FlowState initializes the Persistent State Repository Mk, builds the Historical State Index Gk, and sets an empty Active State Sk,1 before processing request qk.

  2. 02

    State Driven Execution Loop

    FlowState repeatedly serializes qk, the latest observations, Gk, and Sk,t, then uses Incremental State Update and Progressive State Access to update Sk,t and choose zk,t.

  3. 03

    Incremental State Update

    FlowState applies Add, Update, and Remove operations to current request nodes in the Active State, committing a validated state delta ΔSk,t under execution rules P.

  4. 04

    Progressive State Access

    FlowState validates access requests Rk,t against Gk and Sk,t, discloses selected historical states from Mk, and merges them into Sk,t+1 to inform subsequent decisions.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Unified memory representation centered on execution state

    FlowState extends execution states into reusable memory units in the Persistent State Repository, linked via the Historical State Index and Active State to decisions and tool evidence.

  • 02

    Incremental maintenance and progressive access in one loop

    FlowState combines Incremental State Update and Progressive State Access so memory accumulates during execution and informs decisions without reconstructing full state each step.

  • 03

    Improved success rates and token efficiency on long horizon tasks

    FlowState raises MemoryArena success rate by 4.55 percentage points and τ3-Bench pass rate by 13.95 percentage points while cutting tokens by over 40% versus Full Context.

RESULTS

By the Numbers

Average Success Rate

FlowState +4.55 pp

+4.55 over Full Context DeepSeek V4 Flash

Average Pass Rate

FlowState +13.95 pp

+13.95 over Full Context DeepSeek V4 Flash

Token Consumption MemoryArena

−43.2%

relative to Full Context DeepSeek V4 Flash

Token Consumption τ3 Bench

−40.6%

relative to Full Context DeepSeek V4 Flash

On MemoryArena and τ3-Bench, which test interdependent shopping, travel, formal reasoning, and policy constrained service interactions, FlowState shows higher task success with substantially fewer tokens. These MAIN_RESULT numbers demonstrate that execution state as memory can beat full context baselines in both performance and efficiency for long horizon agents.

BENCHMARK

By the Numbers

On MemoryArena and τ3-Bench, which test interdependent shopping, travel, formal reasoning, and policy constrained service interactions, FlowState shows higher task success with substantially fewer tokens. These MAIN_RESULT numbers demonstrate that execution state as memory can beat full context baselines in both performance and efficiency for long horizon agents.

BENCHMARK

MemoryArena Bundled Web Shopping Progress Score

Progress Score (%) on Bundled Web Shopping comparing FlowState to key baselines.

BENCHMARK

τ3 Bench Airline Pass Rate with DeepSeek V4 Flash

Pass Rate (%) on τ3-Bench Airline comparing FlowState to Full Context.

KEY INSIGHT

The Counterintuitive Finding

FlowState reduces total token consumption by 43.2% on MemoryArena and 40.6% on τ3-Bench while still increasing success and pass rates.

This is surprising because many assume trimming context or externalizing memory always harms accuracy, yet FlowState’s structured execution state and PSA avoid that trade off.

WHY IT MATTERS

What this unlocks for the field

FlowState unlocks long horizon agents that can revisit specific past decisions and evidence without replaying or compressing entire interaction histories.

Builders can now design multi session tools, planners, and policy agents whose memory grows across requests while keeping context windows compact and decision making grounded in linked execution state.

~12 min read← Back to papers

Related papers

Memory Architecture

A Control Architecture for Training-Free Memory Use

Yanzhen Lu, Muchen Jiang et al.

· 2026

TAG routes low-confidence steps to uncertainty-based routing, filters them with guarded acceptance with rollback, chooses between bank selection across rule and exemplar memory, and prunes via evidence-based retirement inside a unified control loop. On SVAMP and ASDiv, TAG reaches 81.0% and 85.2% accuracy, improving over the 74.0% and 77.5% no-memory baselines while a compute-matched Retry baseline stays flat.

Memory Architecture

ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

Song-Li Wu, Jingyi Wang et al.

arXiv 2026 · 2026

ActiveMem organizes experiences with a Hierarchical Latent Memory Tree, Dual-Head Memory Controller, Latent Injection Head, and Tree Action Head to build dependency-aware memory paths. On ALFWorld with Qwen3-8B, ActiveMemGRPO reaches 95.57% vs MemGenGRPO’s 90.60%, while also boosting TriviaQA from 80.65% to 87.46%.

Questions about this paper?

Paper: FlowState: Execution State as Memory for Long-Horizon LLM Agents

Answers use this explainer on Memory Papers.

Checking…