ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

AuthorsCai Ke, Xin Liu, Han Zhang et al.

arXiv 20262026

TL;DR

ThinkFlow uses probabilistic latent memory skills with gated consolidation and hyper-alignment to reach 2.18 BLEU-4 and 77.20 Mauve on CC with Qwen3-8B, beating LD-Agent by +0.91 BLEU-4 and +19.01 Mauve.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Explicit Text Memory Causes Information Bottleneck and Semantic Dilution

Explicit textual memory pipelines in lifelong agents suffer from a severe information bottleneck, losing subtle behavioral patterns and emotional shifts.

In long-term multi-session chat, this bottleneck and semantic dilution make agents misread evolving preferences and emotions, breaking deep personalization and empathetic companionship.

HOW IT WORKS

ThinkFlow Latent Memory Skills and Test-Time Evolution

ThinkFlow introduces Probabilistic Latent Memory Skills, a Gated Latent Consolidator, and a Context-Aware Hyper-Aligner trained with Self-Supervised Test-Time Evolution.

You can think of ThinkFlow as a brain-like memory system: skills are specialized latent “circuits,” the gate is attention, and the hyper-aligner is a dynamic adapter between memory and current thoughts.

This KEY_MECHANISM of probabilistic, disentangled latent skills plus predictive evolution lets ThinkFlow track user states far beyond what a static context window or explicit text summaries can maintain.

DIAGRAM

Turn-by-Turn Self-Supervised Test-Time Evolution Flow

This diagram shows how ThinkFlow uses teacher-guided latent alignment and next-user-utterance prediction to evolve its probabilistic memory skills during conversations.

DIAGRAM

Evaluation Pipeline Across Long-Term Conversation Benchmarks

This diagram shows how ThinkFlow is evaluated on CC, MSC, GC, and PersonaMem with baselines and metrics.

PROCESS

How ThinkFlow Handles a Lifelong Multi-Session Conversation

  1. 01

    Probabilistic Latent Memory Skills

    ThinkFlow uses Probabilistic Latent Memory Skills to compress hidden states Et into K disentangled skill vectors via cross attention, preserving uncertainty with µ_t and σ_t.

  2. 02

    Gated Latent Consolidator

    ThinkFlow applies the Gated Latent Consolidator, computing an information gain gate g_t with an MLP and updating skills via a GRU while filtering conversational noise.

  3. 03

    Context-Aware Hyper-Aligner

    ThinkFlow runs the Context-Aware Hyper-Aligner, generating W_adapt from the current query and projecting St to Saligned, then prepending aligned skills as soft prompts.

  4. 04

    Self-Supervised Test-Time Evolution

    ThinkFlow performs Self-Supervised Test-Time Evolution, first distilling from a teacher distribution and then predicting the next user utterance to refine skills with L_CE and L_Conf.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    End-to-End Latent Memory Paradigm

    ThinkFlow introduces an end-to-end latent memory paradigm using Probabilistic Latent Memory Skills and the Gated Latent Consolidator, bypassing explicit text and improving Mauve to 77.20 on CC.

  • 02

    Disentangled Skill Vectors

    ThinkFlow dynamically compresses conversational flow into disentangled skill vectors via Probabilistic Latent Memory Skills, enabling separation of facts, emotions, and preferences without semantic interference.

  • 03

    Self-Supervised Test-Time Evolution

    ThinkFlow proposes Self-Supervised Test-Time Evolution combining Teacher-Guided Latent Alignment and Next-User-Utterance Prediction, achieving label-free lifelong personalization across CC, MSC, GC, and PersonaMem.

RESULTS

By the Numbers

BLEU-4

2.18

+0.91 over LD-Agent on CC with Qwen3-8B

ROUGE-L

17.83

vs 15.13 for LD-Agent on CC with Qwen3-8B

BertScore

47.77

context: Qwen3-8B backbone on CC long-term generation

Mauve

77.20

+19.01 over LD-Agent on CC with Qwen3-8B

On the CC benchmark for long-term open-domain conversation with Qwen3-8B, ThinkFlow’s 2.18 BLEU-4 and 77.20 Mauve show more coherent, human-like responses than LD-Agent and other baselines. These numbers demonstrate that ThinkFlow’s latent memory and evolution mechanisms materially improve multi-session personalization quality.

BENCHMARK

By the Numbers

On the CC benchmark for long-term open-domain conversation with Qwen3-8B, ThinkFlow’s 2.18 BLEU-4 and 77.20 Mauve show more coherent, human-like responses than LD-Agent and other baselines. These numbers demonstrate that ThinkFlow’s latent memory and evolution mechanisms materially improve multi-session personalization quality.

BENCHMARK

Generation Performance on CC with Qwen3-8B

BLEU-4 scores on CC for Qwen3-8B-based memory systems.

BENCHMARK

PersonaMem Accuracy at 1M Context (Qwen3-8B Comparable Methods)

Average accuracy on PersonaMem with 1M-token context.

KEY INSIGHT

The Counterintuitive Finding

Scaling the number of latent memory skills K from 1 to 10 steadily increases PersonaMem accuracy, with ThinkFlow-8B reaching 41.94% at 1M context.

This is surprising because many assume more memory slots would cause interference, yet ThinkFlow’s disentangled probabilistic skills instead improve fine-grained tracking of evolving user states.

WHY IT MATTERS

What this unlocks for the field

ThinkFlow unlocks lifelong conversational agents that maintain rich, evolving user models purely in latent space without explicit text memories or manual labels.

Builders can now deploy agents that continuously personalize via next-utterance prediction, achieving high-quality multi-session companionship with minimal token cost and real-time scalability.

~14 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

Answers use this explainer on Memory Papers.

Checking…