Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation

AuthorsDac Duy Anh Nguyen, Zhangchi Qiu, Shigeng Chen, Alan Wee-Chung Liew

arXiv 20262026

TL;DR

Graph-Based Personalized Memory for LLM Agents uses a lifecycle taxonomy over representation, evolution, retrieval, and evaluation to unify fragmented graph-memory designs.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Personalized Agents Need More Than Raw Logs

LLM agents are evolving into long-term assistants, but dialogue logs and flat summaries cannot capture how user facts interrelate or change over time.

This breaks sustained personalization for tasks like dynamic preference tracking, where agents misinterpret evolving goals and constraints, degrading user modeling and decision quality.

HOW IT WORKS

Lifecycle-Oriented Graph-Based Personalized Memory

Graph-Based Personalized Memory for LLM Agents structures personalized memory into memory representation, memory evolution, memory retrieval, and memory evaluation as explicit lifecycle stages.

Think of it like a card catalog plus history log for an LLM assistant, where each user fact is linked, versioned, and reachable instead of buried in a giant text archive.

This lifecycle lets Graph-Based Personalized Memory for LLM Agents support multi-hop reasoning, preference drift, and provenance-aware updates that a plain context window or vector store cannot manage.

DIAGRAM

Personal Memory Dimensions and Graph Structures

This diagram shows how Graph-Based Personalized Memory for LLM Agents organizes personal memory dimensions and graph representation patterns.

DIAGRAM

Benchmark Landscape for Personalized Graph Memory

This diagram shows how Graph-Based Personalized Memory for LLM Agents groups benchmarks by what they test in personalized memory.

PROCESS

How Graph-Based Personalized Memory for LLM Agents Handles the Memory Lifecycle

  1. 01

    Memory Representation

    Graph-Based Personalized Memory for LLM Agents defines node roles, relation types, and lifecycle metadata to encode semantic, temporal, relational, experiential, and affective user information.

  2. 02

    Memory Evolution

    Graph-Based Personalized Memory for LLM Agents analyzes admission, integration, conflict resolution, consolidation, and removal to maintain an accurate active user state over time.

  3. 03

    Memory Retrieval

    Graph-Based Personalized Memory for LLM Agents categorizes similarity-based, structure-based, and adaptive retrieval to map queries to the right personalized graph context.

  4. 04

    Memory Evaluation

    Graph-Based Personalized Memory for LLM Agents surveys benchmarks and metrics to assess answer quality, evidence retrieval, user-state modeling, evolution, structure, and efficiency.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Lifecycle Taxonomy

    Graph-Based Personalized Memory for LLM Agents introduces a lifecycle taxonomy over memory representation, memory evolution, memory retrieval, and memory evaluation for personalized graph memory.

  • 02

    Evaluation Survey

    Graph-Based Personalized Memory for LLM Agents summarizes benchmarks like LoCoMo, PersonaMem-v2, EngramaBench, and EvoMemBench, highlighting gaps in graph quality and user control.

  • 03

    Future Directions

    Graph-Based Personalized Memory for LLM Agents outlines challenges in scalable lifelong memory, multimodal personal graphs, and causal and counterfactual user modeling.

RESULTS

By the Numbers

Benchmark Count

12 benchmarks

covers more settings than single-benchmark surveys

Lifecycle Stages

4 stages

representation, evolution, retrieval, evaluation

Memory Dimensions

5 dimensions

semantic, temporal, relational, experiential, affective

Representation Patterns

5 patterns

flat, hierarchical, hypergraph, hybrid, multiple graphs

Graph-Based Personalized Memory for LLM Agents aggregates at least twelve named benchmarks across long-horizon recall and personalization. This breadth shows how the taxonomy spans answer quality, retrieval, user modeling, evolution, and structure for LLM agents.

BENCHMARK

By the Numbers

Graph-Based Personalized Memory for LLM Agents aggregates at least twelve named benchmarks across long-horizon recall and personalization. This breadth shows how the taxonomy spans answer quality, retrieval, user modeling, evolution, and structure for LLM agents.

BENCHMARK

Coverage of Personalized Memory Lifecycle Components

Count of core lifecycle components discussed by Graph-Based Personalized Memory for LLM Agents.

KEY INSIGHT

The Counterintuitive Finding

Graph-Based Personalized Memory for LLM Agents notes that adopting a graph alone does not guarantee better retrieval or personalization quality.

This is counterintuitive because many assume structured graphs automatically improve memory, but the survey shows lifecycle design and evaluation practices are equally critical.

WHY IT MATTERS

What this unlocks for the field

Graph-Based Personalized Memory for LLM Agents gives builders a clear blueprint for designing personalized memory across representation, evolution, retrieval, and evaluation.

With this structure, developers can systematically compare graph-based memories, choose appropriate benchmarks, and design agents that maintain trustworthy long-term personalization rather than ad hoc context hacks.

~12 min read← Back to papers

Related papers

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation

Answers use this explainer on Memory Papers.

Checking…