CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems

AuthorsChengxin Yu, Zhaoxin Fan, Faguo Wu et al.

arXiv 20262026

TL;DR

CoMem uses Private Experience Sedimentation, Collective Wisdom Curation, and Parallel Dual-Stream Retrieval to boost MacNet ALFWorld success from 79.85% to 89.55% (+9.70pp).

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Memory Pollution Drives MAS Toward Mediocrity Over Time

CoMem targets memory pollution, where flat shared repositories suffer noise accumulation and behavioral homogenization, eventually driving the collective toward mediocrity.

In conventional MAS, indiscriminate cross-agent writing into a single shared pool causes execution noise dilution and loss of specialization, degrading success rates below even No-memory baselines.

HOW IT WORKS

Collective-Individual Memory Synergy in CoMem

CoMem centers on Private Experience Sedimentation, Collective Wisdom Curation, and Parallel Dual-Stream Retrieval, all tied together by a unified metadata schema and an empirical promotion gateway.

You can think of CoMem like a team with personal notebooks and a curated handbook: agents first refine notes privately, then only proven strategies are promoted into the shared handbook.

This design lets CoMem maintain diverse, role-specific expertise while still building a high-quality group memory, something a plain context window or flat shared log cannot provide.

DIAGRAM

CoMem Task-Time Memory Retrieval Flow

This diagram shows how CoMem retrieves and combines private and collective memories in parallel when an agent faces a new task.

DIAGRAM

CoMem Evaluation and Ablation Pipeline

This diagram shows how CoMem is plugged into different MAS frameworks, evaluated on ALFWorld and PDDL, and ablated over key hyper-parameters.

PROCESS

How CoMem Handles a Collaborative Task Episode

  1. 01

    Private Experience Sedimentation

    CoMem lets each agent extract its execution trace and use the Reflection Distiller to convert raw trajectories into concise insights stored in Private Experience Sedimentation slots.

  2. 02

    Collective Wisdom Curation

    CoMem applies Collective Wisdom Curation with the empirical promotion gateway, checking usage count and above-baseline rewards before moving insights into the Collective Memory Pool.

  3. 03

    Parallel Dual-Stream Retrieval

    During new tasks, CoMem runs Parallel Dual-Stream Retrieval, pulling top k items from both private memory and the collective pool, then applying clustering-based diversity filtering.

  4. 04

    Hybrid Utility Tracking and Pruning

    CoMem updates hybrid utility scores via EMA, increments miss counters, and performs rolling pruning in private memory while purging low-score entries from the collective pool.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    CoMem Architecture for Memory Pollution

    CoMem introduces a two-tier group memory with Private Experience Sedimentation and Collective Wisdom Curation, lifting MacNet ALFWorld success from 79.85% to 89.55% (+9.70pp).

  • 02

    Collective-Individual Memory Synergy

    CoMem formalizes collective-individual memory synergy by combining isolated private stores, a curated Collective Memory Pool, and Parallel Dual-Stream Retrieval for balanced individual and group learning.

  • 03

    Systematic Evaluation Across MAS Frameworks

    CoMem is evaluated on ALFWorld and PDDL with AutoGen, DyLAN, MacNet, and CARD, showing average absolute gains of 8.02 percentage points on ALFWorld and 4.70 on PDDL over No-memory baselines.

RESULTS

By the Numbers

ALFWorld MacNet Success Rate

89.55%

+9.70pp over No-memory MacNet (79.85%)

PDDL MacNet Success Rate

70.19%

+9.41pp over No-memory MacNet (60.78%)

ALFWorld AutoGen Success Rate

88.81%

+9.71pp over No-memory AutoGen (79.10%)

PDDL AutoGen Success Rate

74.44%

+4.92pp over No-memory AutoGen (69.52%)

These metrics come from ALFWorld and PDDL, which test embodied control and symbolic planning respectively. The gains show that CoMem’s collective-individual memory synergy reliably improves multi-agent success over strong No-memory baselines.

BENCHMARK

By the Numbers

These metrics come from ALFWorld and PDDL, which test embodied control and symbolic planning respectively. The gains show that CoMem’s collective-individual memory synergy reliably improves multi-agent success over strong No-memory baselines.

BENCHMARK

Success Rates on ALFWorld with MacNet

Success rate (%) on ALFWorld for MacNet with different memory mechanisms.

BENCHMARK

Success Rates on PDDL with AutoGen

Success rate (%) on PDDL for AutoGen with different memory mechanisms.

KEY INSIGHT

The Counterintuitive Finding

On DyLAN with PDDL, G-Memory drops success from 67.78% (No-memory) to 59.53%, while CoMem raises it to 69.50% (+9.97pp over G-Memory).

This is surprising because adding shared memory is usually expected to help, but flat cooperative memory actually harms performance until CoMem’s curated collective-individual design is used.

WHY IT MATTERS

What this unlocks for the field

CoMem unlocks stable, long-term group memory where agents keep specialized private knowledge while still benefiting from rigorously curated collective wisdom.

Builders can now deploy LLM-based multi-agent systems that learn across tasks without collapsing into noisy, homogenized behavior, enabling sustained specialization and reliable shared strategies.

~13 min read← Back to papers

Related papers

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Agent Memory

ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents

Xiaohui Zhang, Zequn Sun et al.

· 2026

ActMem transforms dialogue history into atomic facts via Memory Fact Extraction, groups them with Fact Clustering, links them through a Memory KG Construction module, and uses Counterfactual-based Retrieval and Reasoning for action-aware answers. On ActMemEval, ActMem reaches 76.52% QA accuracy with DeepSeek-V3, beating LightMem’s 63.97% by 12.55 points and NaiveRAG’s 61.54%.

Questions about this paper?

Paper: CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems

Answers use this explainer on Memory Papers.

Checking…