EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents

AuthorsFengzhou Sun, Yuan Zhang, Xintong Yu, Jinyao Yan

arXiv 20262026

TL;DR

EP-Mem uses an evolvable privacy engine aligned with token-level memory to cut audience-differentiated privacy leakage PB from 0.738 to 0.136 on EP-Bench.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Long-term agents leak private memories in human–agent–human chats (PB=0.738 without EP-Mem)

LLM agents acting as delegates in human–agent–human communication can leak sensitive long-term memories, with NoEngine_evid showing privacy permeability breadth PB=0.738.

This breaks safe social disclosure in multi-party dialogues, causing unauthorized sharing of personal events, work details, and relational secrets across evolving social relationships.

HOW IT WORKS

EP-Mem: Elastic Privacy Memory with a Pluggable Privacy Engine

EP-Mem centers on a User Pre-configured Privacy Policy File, Elastic Privacy Memory Unit, EP-Mem Sidecar Overlay, and Evolvable EP-Mem Privacy Engine attached to explicit memory systems.

Think of EP-Mem like a card catalog plus access-control gates in front of a library: memories stay on the shelves, but each card encodes who may read which story.

This alignment of policy-aware memory and a privacy engine lets EP-Mem enforce elastic, role-conditioned disclosure decisions that a plain context window or naive retrieval pipeline cannot express or audit.

DIAGRAM

Elastic Disclosure Flow in EP-Mem for a Single Query

This diagram shows how EP-Mem processes a receiver query through privacy judgment and context desensitization before answering.

DIAGRAM

EP-Mem Evaluation Pipeline on EP-Bench

This diagram shows how EP-Mem is evaluated across EP-Bench tasks and retrieval benchmarks.

PROCESS

How EP-Mem Handles a Human–Agent–Human Interaction Session

  1. 01

    Memory Material Preprocessing

    EP-Mem converts heterogeneous sources into standardized sessions and writes interpersonal labels into metadata so the Elastic Privacy Memory Unit and EP-Mem Sidecar Overlay can reason over roles.

  2. 02

    Memory Unit Construction and Reprocessing

    EP-Mem builds multi-granular units F0–F5, then the Evolvable EP-Mem Privacy Engine annotates F2 event-level facts with categories, levels, and whitelist or blacklist rules.

  3. 03

    Memory Retrieval

    EP-Mem combines native retrieval scores with event-level proposition scores inside the Elastic Privacy Memory Unit, keeping SimpleMem or MemGAS intact while adding privacy-aware ranking.

  4. 04

    Context Management via EP-Mem Privacy Engine

    The Evolvable EP-Mem Privacy Engine performs disclosure-permission judgment, context desensitization, and audience-differentiated response generation before the agent replies.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Elastic Privacy Memory Unit

    EP-Mem introduces a token-level Elastic Privacy Memory Unit that uses Disc(M | Q,S,R; P) to select Mauth, achieving 94.0% event-level privacy classification accuracy on EP-Bench Task 1.

  • 02

    Evolvable EP-Mem Privacy Engine

    EP-Mem designs an Evolvable EP-Mem Privacy Engine with categorization, whitelist and blacklist filling, disclosure judgment, and context desensitization aligned to F2–F5 memory layers.

  • 03

    EP-Bench Long-term Relational Disclosure Benchmark

    EP-Mem is evaluated on EP-Bench, a 10-script long-term multi-party dataset with 752 event-level facts and 427 sessions, plus noisy retrieval releases and cross-benchmark tests on LoCoMo and CIMemories.

RESULTS

By the Numbers

PB (Permeability Breadth)

0.136

-0.602 vs NoEngine_evid (0.738) on Task 3 audience-differentiated answering

MIQ

4.32

+1.02 vs NoEngine_evid (3.30) showing stronger memory integration with privacy control

Event-level Accuracy

94.0%

EP-Mem achieves 94/100 correct privacy levels on EP-Bench Task 1

Exact Judgment Accuracy

82%

+60 percentage points over NoEngine_full (22%) on Task 2 person-level disclosability

On EP-Bench, which tests policy-conditioned relational disclosure in long-term multi-party dialogues, EP-Mem and EP-Mem+Evo drastically reduce privacy leakage while improving memory integration and maintaining retrieval performance. These numbers show EP-Mem can enforce fine-grained social disclosure boundaries without sacrificing answer quality.

BENCHMARK

By the Numbers

On EP-Bench, which tests policy-conditioned relational disclosure in long-term multi-party dialogues, EP-Mem and EP-Mem+Evo drastically reduce privacy leakage while improving memory integration and maintaining retrieval performance. These numbers show EP-Mem can enforce fine-grained social disclosure boundaries without sacrificing answer quality.

BENCHMARK

EP-Bench Task 3 Audience-Differentiated Answering Privacy PB

Permeability breadth PB on EP-Bench Task 3 (lower is better).

BENCHMARK

EP-Bench Task 2 Person-level Disclosability Exact Accuracy

Exact disclosure-permission accuracy on EP-Bench Task 2 (higher is better).

KEY INSIGHT

The Counterintuitive Finding

EP-Mem+Evo reduces PB from 0.738 to 0.136 while MIQ rises from 3.30 to 4.32, and Style stays at 3.63.

This is surprising because stricter privacy filters usually hurt response richness, yet EP-Mem’s privacy engine simultaneously tightens disclosure and strengthens memory integration and audience differentiation.

WHY IT MATTERS

What this unlocks for the field

EP-Mem unlocks long-term agents that can remember rich personal histories yet elastically respect user-defined social boundaries across roles and domains.

Builders can now deploy human–agent–human delegates that share the right facts with the right people over months of interaction, without manual vetting of every memory retrieval.

~12 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents

Answers use this explainer on Memory Papers.

Checking…