RuleMem: Active Rule Memory for Long-Term Conversational Agents

AuthorsXingyuan Zeng, Zuohan Wu, Quanming Yao et al.

arXiv 20262026

TL;DR

RuleMem uses Rule Perplexity Consistency filtered natural language Horn clause rules to guide memory, reaching 78.05% accuracy on LoCoMo (+5.39 points over GraphRAG).

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Long-term QA suffers recall and reasoning failures across semantic gaps

RuleMem targets long-term QA where existing memory designs cause recall failures and reasoning failures, leaving semantic gaps between queries and evidence.

These failures break conversational agents on datasets like LoCoMo, causing missing evidence and broken logic chains that lead directly to incorrect answers.

HOW IT WORKS

RuleMem: Bottom-Up Induction and Top-Down Rule-Guided QA

RuleMem organizes memory via a Fact Memory Base, Rule Memory Base, Reasoning Path Mining, and RPC Validation to construct and curate natural language rules.

Think of RuleMem like a card catalog plus a rulebook: facts are stored as cards, while induced rules are reusable instructions for how to combine cards.

This rule-centric design lets RuleMem bridge semantic gaps and guide deduction in ways that a plain context window or similarity-only retrieval cannot.

DIAGRAM

Top-Down Rule-Guided Question Answering Flow

This diagram shows how RuleMem uses rules to activate retrieval and perform explicit reasoning at query time.

DIAGRAM

RuleMem Evaluation and Ablation Pipeline

This diagram shows how RuleMem is evaluated on LoCoMo with ablations and baseline comparisons.

PROCESS

How RuleMem Handles a Long-Term Conversation QA Session

  1. 01

    Fact Memorization

    RuleMem parses dialogues into temporal quadruples and stores them in the Fact Memory Base as the foundation for later reasoning and rule induction.

  2. 02

    Reasoning Path Mining

    RuleMem organizes facts into a temporal knowledge graph and uses Reasoning Path Mining to sample and reconstruct coherent multi-hop reasoning paths.

  3. 03

    Rule Induction and RPC Validation

    RuleMem abstracts mined paths into Horn clause rules in the Rule Memory Base, then filters them with RPC Validation using perplexity reductions.

  4. 04

    Top-Down Rule-Guided QA

    At query time, RuleMem activates rules by matching heads, performs Guided Recall from the Fact Memory Base, and uses Explicit Reasoning to generate answers.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    RuleMem: Active Rule Memory Framework

    RuleMem introduces a Rule Memory Base alongside a Fact Memory Base, turning memory into an active guide for retrieval and deduction, reaching 78.05% accuracy on LoCoMo.

  • 02

    Rule Perplexity Consistency

    RuleMem proposes RPC Validation, combining ∆self and ∆fact perplexity reductions with α and τ to filter unreliable induced rules before committing them to memory.

  • 03

    Guided Recall and Explicit Reasoning

    RuleMem adds Guided Recall using rule bodies and Explicit Reasoning with rule scaffolds, raising average recall from 0.56 to 0.79 and reducing reasoning failures by 12.0%.

RESULTS

By the Numbers

Accuracy

78.05%

+9.04 points over GraphRAG

BLEU

36.90

+23.36 over Mem0 average BLEU 13.54

F1

32.96

-0.92 vs MemInsight average F1 33.88

Recall Rate

0.79

+0.23 over default recall 0.56 with Guided Recall

On the LoCoMo benchmark, which tests single-hop, multi-hop, open-domain, and temporal questions, RuleMem achieves 78.05% average accuracy. This proves that RuleMem's rule-guided memory substantially improves long-term conversational QA compared to baselines like GraphRAG and Mem0.

BENCHMARK

By the Numbers

On the LoCoMo benchmark, which tests single-hop, multi-hop, open-domain, and temporal questions, RuleMem achieves 78.05% average accuracy. This proves that RuleMem's rule-guided memory substantially improves long-term conversational QA compared to baselines like GraphRAG and Mem0.

BENCHMARK

Performance comparison on the LoCoMo benchmark

Accuracy on LoCoMo averaged across Single-hop, Multi-hop, Open-domain, and Temporal questions.

KEY INSIGHT

The Counterintuitive Finding

RuleMem's Guided Recall boosts average recall from 0.56 to 0.79, a 41.1% improvement, even though it starts from the same fact memory as baselines.

This is surprising because many assume better retrieval requires new storage structures, yet RuleMem shows that rule-level guidance alone can dramatically improve recall.

WHY IT MATTERS

What this unlocks for the field

RuleMem unlocks rule-based memory that actively guides retrieval and reasoning, using Horn clauses and RPC to maintain reliable, reusable abstractions.

Builders can now create conversational agents that learn generalizable rules from interactions, enabling robust multi-hop reasoning over massive dialogue histories that were previously too noisy and dispersed.

~12 min read← Back to papers

Related papers

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: RuleMem: Active Rule Memory for Long-Term Conversational Agents

Answers use this explainer on Memory Papers.

Checking…