VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

AuthorsLiyang Fan, Yingcheng Shi, Yongbin Li et al.

arXiv 20262026

TL;DR

VibeMemBench uses the SIEVE construction pipeline and a matched memory-on vs memory-off protocol to show frozen verified experience yields up to +4.5pp Resolved while four existing memory systems fail in 11 of 12 pairings to beat the no-memory baseline.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Memory systems rarely improve executable coding work

VibeMemBench shows that in eleven of twelve solver and system pairings, existing memory systems fail to exceed the matched memory-off baseline.

This means coding agents with Mem0, SimpleMem, MemoryOS, or A-MEM often carry persistent memory that does not translate into better executable repository repairs, leaving useful experience latent.

HOW IT WORKS

VibeMemBench and the SIEVE construction pipeline

VibeMemBench centers on the SIEVE construction pipeline, Frozen Verified Experience, a Matched Intervention Protocol, and two evaluation layers over 3,634 history trajectories.

You can think of SIEVE like a careful lab experiment: it first proves a history patch helps in one controlled setting, then reuses that “labeled” experience across different solvers and memory systems.

This design lets VibeMemBench test whether memory actually changes executable outcomes, something a plain context window or recall-only benchmark cannot reveal.

DIAGRAM

Matched memory-on vs memory-off intervention for coding agents

This diagram shows how VibeMemBench runs paired agent trajectories under fixed configuration A, changing only the declared memory condition.

DIAGRAM

VibeMemBench evaluation layers and memory systems

This diagram shows how VibeMemBench first tests frozen verified experience transfer, then evaluates Mem0, SimpleMem, MemoryOS, and A-MEM retrieval on the same targets.

PROCESS

How VibeMemBench Handles a Repository Coding Task

  1. 01

    SIEVE Construction

    VibeMemBench uses SIEVE to validate SWE-rebench V2 tasks, screen same-repository histories, execute them, and distill one experience per history-target pair.

  2. 02

    Verify Uplift and Freeze

    VibeMemBench runs MiniSWEAgent with deepseek-v4-flash across candidate memories, computes uplift gt,k, and freezes the best k⋆t as frozen verified experience.

  3. 03

    Matched Intervention Protocol

    VibeMemBench defines a fixed agent configuration A and runs each target-seed pair twice, once with memory off and once with the declared memory condition on.

  4. 04

    Frozen and System Evaluation

    VibeMemBench first tests frozen verified experience transfer on five solvers, then evaluates Mem0, SimpleMem, MemoryOS, and A-MEM retrieval under the same protocol.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    VibeMemBench construction and SIEVE

    VibeMemBench builds 111 coding targets from 90 repositories and 3,634 trajectories using SIEVE, which controls leakage and verifies that each target has a useful experience by execution uplift.

  • 02

    Matched intervention protocol

    VibeMemBench formulates a matched intervention protocol that changes only the declared memory condition, separating frozen verified experience transfer from memory-system retrieval under an auditable manifest.

  • 03

    Gap between availability and use

    VibeMemBench shows frozen verified experience yields up to +4.5pp Resolved and fewer steps, while existing memory systems fail to exceed the memory-off baseline in eleven of twelve solver–system pairings.

RESULTS

By the Numbers

Resolved deepseek-v4-pro

67.1%

+0.0 over Memory off

Resolved glm-5

57.4%

+2.4 over Memory off

Resolved kimi-k2.7-code

74.3%

+4.5 over Memory off

Resolved qwen3.8-max

80.2%

+1.1 over Memory off

These Resolved scores come from the frozen verified experience transfer layer on the 111 outcome-selected VibeMemBench targets, each averaged over 444 target-seed runs. They show that VibeMemBench’s frozen verified experience can improve executable repository repair success by up to +4.5 percentage points for kimi-k2.7-code compared to its memory-off baseline.

BENCHMARK

By the Numbers

These Resolved scores come from the frozen verified experience transfer layer on the 111 outcome-selected VibeMemBench targets, each averaged over 444 target-seed runs. They show that VibeMemBench’s frozen verified experience can improve executable repository repair success by up to +4.5 percentage points for kimi-k2.7-code compared to its memory-off baseline.

BENCHMARK

Frozen verified experience transfer on VibeMemBench

Resolved on 111 targets with frozen verified experience vs memory-off baselines.

KEY INSIGHT

The Counterintuitive Finding

VibeMemBench finds that in 42.0 percent of retrieval pairings, a memory-off ceiling of four resolved seeds leaves injected records able only to lose seeds.

Even more surprisingly, the same record often helps a weaker solver and harms a stronger one, so the value of experience depends on solver headroom rather than any intrinsic record quality.

WHY IT MATTERS

What this unlocks for the field

VibeMemBench gives researchers a way to measure whether persistent memory actually improves executable repository work, not just recall or token savings.

Builders can now design memory systems that are evaluated on paired Resolved, tokens, and steps, tuning record form, retrieval, and gating policies for real coding agents instead of proxy QA benchmarks.

~12 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

Answers use this explainer on Memory Papers.

Checking…