Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 18 of 19

Memory Architecture

Hierarchical Neural Memory Network for Low Latency Event Processing

Ryuhei Hamaguchi, Yasutaka Furukawa et al.

arXiv 2023 · 2023

Hierarchical Neural Memory Network (HMNet) stacks multi-level latent memories z1–z3 with Event-write, Up-write, Down-write, Update, and Readout operations driven by Event Sparse Cross Attention. On DSEC-Semantic, HMNet-L3 reaches 57.4 mIoU with event–RGB fusion, improving over a ResNet-50 baseline at 54.1 mIoU while also reducing latency, and on GEN1 HMNet-B1 matches AED at similar mAP with 57% lower latency.

Benchmark

In-context Autoencoder for Context Compression in a Large Language Model

Tao Ge, Jing Hu et al.

arXiv 2023 · 2023

In-context Autoencoder (ICAE) adds a LoRA-adapted encoder, memory tokens and slots, and combined autoencoding plus language modeling pretraining on top of an unchanged Llama decoder to compress long contexts into short, reusable memory spans. On Pile autoencoding with Llama-7B, ICAE compresses 512-token contexts into 128 memory slots with 99.1% BLEU and 0.017 loss, and on the PWC dataset ICAE with 128 slots reaches a 74.2% win+tie rate against GPT-4 when built on Llama-2-7B-chat.

Long-Term Memory

In search of dispersed memories: Generative diffusion models are associative memory networks

Luca Ambrogioni

· 2023

In search of dispersed memories connects associative memory networks, modern Hopfield networks, generative diffusion models, and the denoising loss to show that diffusion score networks encode Hopfield-like energy landscapes in their weights. In search of dispersed memories demonstrates that exact diffusion dynamics reach Pearson correlations of 0.995–0.996 with modern Hopfield iterations on denoising and completion tasks, while classical Hopfield networks reach only 0.700–0.741.

Long-Term Memory

LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination

Kai Zhang, Yangyang Kang et al.

· 2023

MaLP combines a Dual-Process enhanced Memory (DPeM), Working Memory, Short-Term Memory (STM), Long-Term Memory (LTM), a Coordinator C, and Retriever R to coordinate short- and long-term personalization around a PEFT-tuned LLM. On MaLP’s medical dialogue benchmark with LLaMA-7B, MaLP achieves 69.95% preference classification accuracy and a 91.53% win rate in response generation, improving ROUGE-L profile QA from 29.66 to 33.91 over the LoRA baseline.

Memory Architecture

MBPTrack: Improving 3D Point Cloud Tracking with Memory Networks and Box Priors

Tian-Xing Xu, Yuan-Chen Guo et al.

arXiv 2023 · 2023

MBPTrack combines a Decoupling Feature Propagation Module, BPLocNet, box-prior sampling, and point-to-reference aggregation to track 3D objects from point clouds using temporal memory and size-aware localization. On KITTI, MBPTrack achieves 70.3% Success and 87.9% Precision, improving over CXTrack’s 67.5%/85.3% by +2.8/+2.6.

Benchmark

MemoryBank: Enhancing Large Language Models with Long-Term Memory

Wanjun Zhong, Lianghong Guo et al.

· 2023

MemoryBank combines Memory Storage, Memory Retrieval, and a Memory Updating Mechanism to maintain daily conversations, event summaries, and user portraits for long-term personalization. On a 10-day, 15-user simulated dialog benchmark with 194 probing questions, MemoryBank-powered SiliconFriend ChatGPT achieves 0.716 correctness and 0.912 contextual coherence, surpassing SiliconFriend ChatGLM and SiliconFriend BELLE.

Benchmark

Recurrent Neural Networks and Long Short-Term Memory Networks: Tutorial and Survey

Benyamin Ghojogh, Ali Ghodsi

arXiv 2023 · 2023

Recurrent Neural Networks and Long Short-Term Memory Networks explains how Backpropagation Through Time, LSTM gates and cells, Gated Recurrent Units, bidirectional RNN, and ELMo fit into one dynamical-systems view. Recurrent Neural Networks and Long Short-Term Memory Networks mainly contributes a structured tutorial and survey rather than new benchmark numbers.

Benchmark

Working Memory Capacity of ChatGPT: An Empirical Study

Dongyu Gong, Xingchen Wan, Dingmin Wang

· 2023

Working Memory Capacity of ChatGPT probes ChatGPT with verbal and spatial n-back blocks, adding noise, feedback, and chain-of-thought variants to stress working memory. Working Memory Capacity of ChatGPT finds d′ falls to around 1 at n=3 in most conditions and that GPT-4’s verbal n-back capacity far exceeds Bloomz, ChatGLM, and Vicuna baselines.

Memory Architecture

BayesPCN: A Continually Learnable Predictive Coding Associative Memory

Jason Yoo, Frank Wood

arXiv 2022 · 2022

BayesPCN combines predictive coding, conjugate Bayesian updates over W 0:L, sequential importance sampling, and a diffusion-based forget mechanism to build a hierarchical associative memory that supports continual one-shot writes. On CIFAR10 and Tiny ImageNet hetero-associative tasks, BayesPCN matches offline GPCN with MSE as low as 0.0000 while online GPCN rises to 0.0791 MSE on CIFAR10 mask at sequence length 1024.

Benchmark

Episodic Memory Question Answering

Samyak Datta, Sameer Dharur et al.

· 2022

Episodic Memory Question Answering combines allocentric top-down semantic features, spatiotemporal memory, and a LingUNet-based question-answering model to localize answers on scene floorplans from egocentric RGB-D tours. On the EMQA benchmark built from Matterport3D tours, Episodic Memory Question Answering with temporal features achieves 29.11 IoU and 62.27 recall on top-down maps, improving over SMNetDecoder by 2.19 IoU and 18.41 recall.

Long-Term Memory

Evaluating Long-Term Memory in 3D Mazes

Jurgis Pasukonis, Timothy Lillicrap, Danijar Hafner

· 2022

Memory Maze combines an online reinforcement learning environment, an offline dataset, and an offline probing protocol to stress-test long-term memory using Dreamer, Dreamer (TBTT), IMPALA, and supervised probe networks. On Memory 9x9, Dreamer (TBTT) achieves a return of 33.2 compared to 23.4 for IMPALA, while humans reach 26.4 and the oracle reaches 34.8.

Benchmark

Large Language Models with Controllable Working Memory

Daliang Li, Ankit Singh Rawat et al.

· 2022

Knowledge Aware Finetuning (KAFT) augments QA data with relevant context, irrelevant context, counterfactual context, and empty context built from SQuAD 2.0, T-REx, QASC, and TriviaQA, and uses pretrained model’s answer to label irrelevant slices. KAFT achieves up to 24× higher controllability than Noisy Finetuning and up to 6× higher robustness than Noisy Finetuning on PaLM 540B while matching TriviaQA validation accuracy.

Memory Architecture

Training Language Models with Memory Augmentation

Zexuan Zhong, Tao Lei, Danqi Chen

arXiv 2022 · 2022

TRIME aligns token embeddings with in-batch contextualized representations using a contrastive training objective, combined with local memory, long-term memory, and external memory constructed via specialized batching strategies. On WIKITEXT-103, TRIME with external memory (TRIMELMext) reaches 15.37 test perplexity using a 247M Transformer, a 0.75 improvement over kNN-LM with continuous cache.

Memory Architecture

Universal Hopfield Networks: A General Framework for Single-Shot Associative Memory Models

Beren Millidge, Tommaso Salvatori et al.

· 2022

Universal Hopfield Networks decompose single-shot associative memory into similarity, separation, and projection, instantiating Hopfield networks, sparse distributed memories, dense associative memories, and modern continuous Hopfield networks within one energy-based framework. Universal Hopfield Networks then compare dot product, Euclidean, Manhattan, and other similarity functions, finding that Manhattan and Euclidean similarity often yield higher retrieval capacity and robustness than dot-product-based modern continuous Hopfield networks on MNIST, CIFAR10, and Tiny ImageNet.

RAG

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Alex Mallen, Akari Asai et al.

arXiv 2022 · 2022

Adaptive Retrieval combines parametric memory, non-parametric memories, and popularity-based thresholds to decide when to augment questions with BM25 or Contriever passages on POPQA and EntityQuestions. On POPQA, Adaptive Retrieval with GPT‑3 davinci‑003 plus GenRead and Contriever reaches 46.5% accuracy, 5.3 points above always-retrieving systems while cutting GPT‑3 API costs by about 50%.

Benchmark

Generalizable Episodic Memory for Deep Reinforcement Learning

Hao Hu, Jianing Ye et al.

arXiv 2021 · 2021

Generalizable Episodic Memory (GEM) combines a parametric network Mθ, implicit memory-based planning, twin back-propagation process, and conservative estimation on single step to generalize episodic returns across trajectories. GEM achieves higher average returns than TD3, SAC, DDPG, and TD3+SIL on MuJoCo tasks like Ant-v2, HalfCheetah-v2, and Humanoid-v2 within 1M environment steps.

Memory Architecture

Hierarchical Associative Memory

Dmitry Krotov

arXiv 2021 · 2021

Hierarchical Associative Memory organizes fully recurrent Modern Hopfield Networks into layered architectures using Lagrangian functions, hierarchical time scales, and symmetric feedforward–feedback weights as core components. Hierarchical Associative Memory theoretically extends Dense Associative Memories to arbitrary depth and local connectivity, deriving explicit dynamical and energy functions rather than reporting benchmark numbers.

Benchmark

Offline Reinforcement Learning with Value-based Episodic Memory

Xiaoteng Ma, Yiqin Yang et al.

arXiv 2021 · 2021

Value-based Episodic Memory (VEM) uses Expectile V-Learning (EVL), Implicit Memory-based Planning, and Generalized Advantage-weighted Learning to learn conservative V-values and plan along offline trajectories. On the D4RL benchmark, VEM attains 87.5 on antmaze-umaze and 128.3 on adroit-hammer-expert, improving over BAIL and CQL on most AntMaze and Adroit tasks.

Benchmark

Solving Continuous Control with Episodic Memory

Igor Kuznetsov, Andrey Filchenkov

arXiv 2021 · 2021

Episodic Memory Actor Critic (EMAC) augments DDPG with a Memory Module, Episodic-based Experience Replay Prioritization, and a modified critic objective that blends Bellman targets with retrieved Monte Carlo returns. On OpenAI Gym continuous control, EMAC reaches 2236.88 ± 808 average return on Walker2d-v3, beating TD3 (1008.32) by +1228.56 and SAC (1787.28) by +449.6 in the 200000-step small-data regime.

Benchmark

Working Memory Connections for LSTM

Federico Landi, Lorenzo Baraldi et al.

arXiv 2021 · 2021

Working Memory Connections for LSTM (LSTM-WM) augments LSTM gates, Working Memory Connections, peephole connections, and the memory cell so gates see a protected projection of the cell state. On PTB character-level language modeling, LSTM-WM achieves 1.299 BPC with TPTB=150 compared to 1.334 BPC for vanilla LSTM.

Memory Architecture

Emergent Symbols through Binding in External Memory

Taylor W. Webb, Ishan Sinha, Jonathan D. Cohen

arXiv 2020 · 2020

Emergent Symbol Binding Network (ESBN) combines an LSTM controller, a shared image encoder fe, temporal context normalization, and a two column key value memory to bind abstract variables to concrete image embeddings. ESBN achieves ≥95% test accuracy on same different, RMTS, distribution of three, and identity rules tasks while generalizing to withheld Unicode characters, unlike LSTM, NTM, MNM, Relation Net, Transformer, and PrediNet baselines.