LifeSide: Benchmarking Agents as Lifelong Digital Companions

AuthorsYuqian Wu, Zhijie Deng, Wei Chen et al.

arXiv 20262026

TL;DR

LifeSide + multi-session Memory Emotion Environment loops + reveals that even saturated memory benchmarks still fail at long horizon companionship across 2,000 personas and 111K tasks.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Lifelong companions fail across 2,000 personas and 111K tasks

LifeSide shows that existing evaluations miss lifelong behavior, even though LifeSide tests 2,000 personas and 111K tasks across key capabilities.

When systems that saturate current memory benchmarks face LifeSide, they still lose accurate user understanding and true companionship over long horizons.

HOW IT WORKS

LifeSide — multi session Memory Emotion Environment loops

LifeSide builds multi session Memory Emotion Environment loops by modeling users as persistent worlds with layered profiles and event trajectories across dialogues.

You can think of LifeSide like a simulated life log, where each persona is a long running operating system with changing processes and hidden states.

This lets LifeSide probe what a plain context window cannot capture, namely the gap between latent thoughts and observable expressions over many evolving sessions.

DIAGRAM

Multi session interaction in LifeSide

This diagram shows how LifeSide simulates repeated user agent sessions using Memory Emotion Environment loops while preserving latent user states.

DIAGRAM

LifeSide evaluation pipeline over 2,000 personas

This diagram shows how LifeSide constructs personas, simulates trajectories, and generates 111K tasks for evaluating lifelong companionship.

PROCESS

How LifeSide Handles a Memory Emotion Environment loop

  1. 01

    Layered profiles construction

    LifeSide first builds layered profiles for each persona, defining stable traits and hidden preferences that drive Memory Emotion Environment dynamics.

  2. 02

    Event trajectories generation

    LifeSide then simulates event trajectories, creating long term life events that interact with the layered profiles across many sessions.

  3. 03

    Multi agent simulation

    LifeSide runs multi agent simulation, projecting environmental dynamics into dialogue while keeping a gap between latent thoughts and observable expressions.

  4. 04

    Memory Emotion Environment evaluation

    LifeSide evaluates memory tracking, user understanding, privacy control, and emotional companionship over the evolving Memory Emotion Environment loops.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    LifeSide benchmark design

    LifeSide introduces a benchmark centered on multi session Memory Emotion Environment loops over 2,000 personas and 111K tasks for lifelong companionship.

  • 02

    Persistent world user modeling

    LifeSide models users as persistent worlds with layered profiles and event trajectories, preserving the gap between latent thoughts and observable expressions.

  • 03

    Four capability evaluations

    LifeSide jointly evaluates memory tracking, user understanding, privacy control, and emotional companionship under realistic long horizon multi session conditions.

RESULTS

By the Numbers

Personas count

2000 personas

+2000 over prior single user tests

Tasks total

111K tasks

vs prior small scale evaluations

Evaluation areas

4 areas

memory tracking, user understanding, privacy, companionship

Legacy benchmarks

saturated

yet still fail at long horizons

LifeSide evaluates 2,000 personas and 111K tasks across four areas, testing long horizon behavior. LifeSide proves that saturating existing memory benchmarks does not guarantee sustained user understanding or companionship.

BENCHMARK

By the Numbers

LifeSide evaluates 2,000 personas and 111K tasks across four areas, testing long horizon behavior. LifeSide proves that saturating existing memory benchmarks does not guarantee sustained user understanding or companionship.

BENCHMARK

Scale of LifeSide benchmark components

Relative scale of personas, tasks, and evaluation areas used in LifeSide.

KEY INSIGHT

The Counterintuitive Finding

LifeSide reveals that even systems that saturate current memory benchmarks still fail to sustain accurate user understanding and true companionship over long horizons.

This is surprising because many assumed strong short term memory and empathy scores would transfer directly to lifelong digital companionship, but LifeSide shows that assumption breaks under multi session dynamics.

WHY IT MATTERS

What this unlocks for the field

LifeSide unlocks a way to stress test lifelong digital companions under realistic Memory Emotion Environment loops with persistent user worlds and hidden states.

Builders can now design and debug agents that respect evolving privacy boundaries and emotional needs across many sessions, instead of optimizing only for single turn memory recall.

Related papers

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Agent Memory

ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents

Xiaohui Zhang, Zequn Sun et al.

· 2026

ActMem transforms dialogue history into atomic facts via Memory Fact Extraction, groups them with Fact Clustering, links them through a Memory KG Construction module, and uses Counterfactual-based Retrieval and Reasoning for action-aware answers. On ActMemEval, ActMem reaches 76.52% QA accuracy with DeepSeek-V3, beating LightMem’s 63.97% by 12.55 points and NaiveRAG’s 61.54%.

Questions about this paper?

Paper: LifeSide: Benchmarking Agents as Lifelong Digital Companions

Answers use this explainer on Memory Papers.

Checking…