PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use

AuthorsMingfei Lu, Mengjia Wu, Yi Zhang

arXiv 20262026

TL;DR

PairPref pairs fixed preferences and requests with contrasting situations to test contextual applicability, revealing selectivity up to 64.9 points but free-generation success as low as 3.6–18.3%.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Models keep applying preferences even when context says they should not (Use both up to 92.0%)

PairPref shows that models apply user preferences in both situations for 71.0%–92.0% of generated pairs, even when one situation requires withholding.

This means assistants using PairPref’s setup often ignore contextual constraints, leading to over-personalization, misaligned recommendations, and failures in situational judgment.

HOW IT WORKS

PairPref — contextual preference pairs with dual tracks

PairPref’s core mechanism combines Stored preferences, a Companion memory pool, Paired instance construction, and Multi-stage validation to build controlled situation pairs for both Selection and Free generation.

You can think of PairPref like a lab experiment where the preference and request are held constant, while only the “environment” changes, isolating when memory should or should not influence behavior.

This design lets PairPref test contextual applicability decisions that a plain context window cannot, because it jointly scores both sides of a pair and separates preference availability from appropriate use.

DIAGRAM

Paired situation flow for selection and free generation

This diagram shows how PairPref routes a fixed preference and request through two contrasting situations into selection and free-generation evaluations.

DIAGRAM

PairPref construction and validation pipeline

This diagram shows how PairPref builds and filters 1,227 validated pairs through rule checks, semantic verification, and context-free screening before evaluation.

PROCESS

How PairPref Handles a Paired Contextual Preference Task

  1. 01

    Task formulation

    PairPref defines Mi, qi, C+ i, C− i, and Ri so that a single preference and request are shared while applicability flips across situations.

  2. 02

    Paired instance construction

    PairPref uses GPT-5.5 to generate the shared request, two situations, and four replies rM, rsup, robl, and rirr while building the Companion memory pool.

  3. 03

    Multi-stage validation

    PairPref applies rule-based checks, DeepSeek V4 Flash semantic verification, and Qwen context-free choice screening to filter candidates.

  4. 04

    Selection and free generation

    PairPref runs Selection with fixed replies and Free generation without replies on the same 1,227 pairs, then scores Use+, Use−, selectivity, Use both, and Pair success.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Contextual preference benchmark

    PairPref introduces 1,227 pairs across 45 preferences and eight situation categories, using a Companion memory pool and Paired instance construction to isolate applicability changes.

  • 02

    Paired evaluation protocol

    PairPref combines Selection and Free generation on the same pairs, jointly scoring both sides to measure Use+, Use−, selectivity ∆, Use both, and Gen-OK.

  • 03

    Empirical findings on persistent use

    PairPref shows that eight models achieve selection selectivity between 51.3 and 64.9 points but free-generation Pair success of only 3.6%–18.3%, with Use both as high as 92.0%.

RESULTS

By the Numbers

Selectivity ∆

64.9 points

+13.6 over Qwen 3.8 Max

Selection Pair success

66.4%

Claude Opus 5 vs Grok 4.6 at 23.6%

Use both

92.0%

Highest persistent use for Grok 4.6

Gen-OK Pair success

3.6%–18.3%

Range across eight models on PairPref

PairPref evaluates eight models on 1,227 pairs, measuring contextual preference use via Use+, Use−, selectivity ∆, Use both, and Gen-OK. These results show that PairPref exposes a gap where models like Claude Opus 5 reach 64.9-point selectivity in selection but only 9.3% Pair success in free generation, while Grok 4.6 reaches 92.0% Use both and just 3.6% Gen-OK.

BENCHMARK

By the Numbers

PairPref evaluates eight models on 1,227 pairs, measuring contextual preference use via Use+, Use−, selectivity ∆, Use both, and Gen-OK. These results show that PairPref exposes a gap where models like Claude Opus 5 reach 64.9-point selectivity in selection but only 9.3% Pair success in free generation, while Grok 4.6 reaches 92.0% Use both and just 3.6% Gen-OK.

BENCHMARK

Main results on PairPref selection and generation

Selectivity ∆ on PairPref selection track for eight models.

KEY INSIGHT

The Counterintuitive Finding

PairPref shows that models apply the preference in both generated answers for 71.0%–92.0% of retained pairs, even when one side requires setting it aside.

This is surprising because the same models achieve selection selectivity between 51.3 and 64.9 points on PairPref, suggesting they can recognize contextual differences but fail to carry that judgment into free-form answers.

WHY IT MATTERS

What this unlocks for the field

PairPref gives researchers a controlled way to test when memory should guide an answer, not just whether preferences are stored and retrieved.

With PairPref, builders can design and debug systems that separate topical relevance from situational applicability, making personalized assistants that know when to apply or withhold user preferences.

~12 min read← Back to papers

Related papers

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use

Answers use this explainer on Memory Papers.

Checking…