CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations

AuthorsYulin Hu, Yanyan Zhao, Zimo Long et al.

arXiv 20262026

TL;DR

CUE-Mem benchmarks long-term user memory from recurring implicit multimodal cues, revealing explicit–implicit accuracy gaps up to 42.5 points in textualized systems.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Long-term agents fail on implicit multimodal cues with a 42.5 point accuracy drop

Across textualized memory systems on CUE-Mem, Qwen3.6-35B-A3B shows EI-Gaps of 42.5 points on Entity Recall and 36.5 on Long Pattern.

These gaps mean multimodal assistants misplace recurring background cues, breaking user memory for preferences and habits and undermining personalized assistance over long histories.

HOW IT WORKS

CUE-Mem: Explicit vs implicit multimodal memory across four tasks

CUE-Mem builds synthetic histories via Script Synthesis, Multimodal Dialog Synthesis, and Data Review, then probes memory with Entity Recall and Long Pattern tasks.

Think of CUE-Mem as a stress-test where explicit cues are like labeled files, while implicit cues are faint traces scattered across a giant card catalog of images and audio.

By separating explicit and implicit evidence, CUE-Mem exposes how long-term memory systems fail to preserve and retrieve subtle multimodal signals that a plain context window cannot reliably capture.

DIAGRAM

CUE-Mem dataset construction pipeline from profiles to multimodal conversations

This diagram shows how CUE-Mem constructs synthetic text-image-audio histories through three stages before task generation.

DIAGRAM

CUE-Mem evaluation setup for textualized and native multimodal memory systems

This diagram shows how CUE-Mem feeds histories into textualized and native multimodal memory paradigms and measures explicit–implicit gaps.

PROCESS

How CUE-Mem Handles a Multimodal Memory Evaluation Session

  1. 01

    Script Synthesis

    CUE-Mem expands persona seeds into profiles and event timelines, assigning modalities and evidence types that later drive Entity Recall and Long Pattern tasks.

  2. 02

    Multimodal Dialog Synthesis

    CUE-Mem generates text-image-audio turns, placing explicit cues in foreground content and implicit cues as recurring background signals across sessions.

  3. 03

    Data Review

    CUE-Mem applies automatic checks and human review to ensure multimodal coherence, long-term consistency, and valid implicit cues before evaluation.

  4. 04

    Task Design

    CUE-Mem builds Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal questions with annotated clues to probe long-term user memory.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Text-image-audio benchmark for implicit cues

    CUE-Mem introduces 2,674 QA pairs over 20 users, 648 sessions, 496 images, and 2,228 audio clips to test long-term memory from subtle multimodal evidence.

  • 02

    Four-task memory decomposition

    CUE-Mem defines Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal, each with explicit and implicit variants to isolate different memory capabilities.

  • 03

    Analysis of textualized and native multimodal memory

    CUE-Mem benchmarks Full Memory, NaiveRAG, Generative Agents, Reflexion, MemGPT, A-Mem, MemoryOS, and multimodal RAG, revealing persistent explicit–implicit gaps and retrieval noise.

RESULTS

By the Numbers

Memory Avg. Qwen3.6-35B-A3B Oracle Evidence

80.8%

+42.5 points over Full Memory implicit (38.3%)

Memory Avg. GPT-5.4-mini Oracle Evidence

78.4%

+40.9 points over Full Memory implicit (37.5%)

Human Full implicit Memory Avg.

82.8%

+44.5 points over Qwen3.6-35B-A3B Full Memory implicit (38.3%)

EI-Gap Qwen3.6-35B-A3B Full Memory

38.6 points

difference between explicit 76.9% and implicit 38.3% Memory Avg.

These metrics come from CUE-Mem’s main textualized memory results table, which tests long-term user memory across explicit and implicit evidence. The numbers show CUE-Mem exposes a large gap between oracle evidence and what current memory systems can recover from real multimodal histories.

BENCHMARK

By the Numbers

These metrics come from CUE-Mem’s main textualized memory results table, which tests long-term user memory across explicit and implicit evidence. The numbers show CUE-Mem exposes a large gap between oracle evidence and what current memory systems can recover from real multimodal histories.

BENCHMARK

Textualized Memory Avg. on CUE-Mem (Qwen3.6-35B-A3B)

Memory Avg. accuracy (%) across Entity Recall, Long Pattern, and Personalized Recommendation for different memory systems under implicit evidence.

KEY INSIGHT

The Counterintuitive Finding

On CUE-Mem, Answer Refusal accuracy jumps from 63.2 to 97.2 for Qwen3.6-35B-A3B when moving from explicit to implicit evidence.

This is surprising because better refusal scores usually signal improved calibration, but in CUE-Mem they instead reveal that systems simply fail to access implicit memories and default to abstaining.

WHY IT MATTERS

What this unlocks for the field

CUE-Mem gives researchers a controlled way to stress-test whether long-term memory systems can form user memories from recurring background visual and acoustic cues.

With CUE-Mem, builders can now quantify preservation, retrieval, and use of implicit multimodal evidence, guiding designs for selective memory retention beyond naive context-window scaling.

~12 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations

Answers use this explainer on Memory Papers.

Checking…