Bounded recurrent memory · routed banks · measured results

Mixture of Memories versus Gated DeltaNet

We did not find a reliable Mixture of Memories advantage over Gated DeltaNet in our tested recipes. One original-MoM run reached 84.29% on the primary recall test, near its paired GDN result of 85.85%. The other two MoM runs reached 8.06% and 0.74%. A separate language-model profile measured MoM training updates about 38% longer.

The published MoM paper reports higher recall scores than GDN after much larger language-pretraining runs. Our work uses small models and a synthetic overwrite-recall task. It is not a reproduction of that campaign and does not disprove its result.

Measured conclusionGated DeltaNet is the working baseline for these local recipes. MoM can learn the recall task well, but it did not do so consistently across the three paired runs.
THREE RECALL SEEDSFIXED 8,000 UPDATESALL ENDPOINTS RETAINEDONE FILTER CONTROL FAILED

What changes when memory is split into banks

Gated DeltaNet (GDN) stores associations in recurrent matrix state and uses a learned gate to control retention. Mixture of Memories (MoM) routes each token to a subset of several memory banks and also updates an always-active shared bank. A router chooses where each token writes and reads.

Gated DeltaNet

Every token updates the same multi-head recurrent memory. The model does not need to learn a bank assignment before it can use the state.

shared recurrent state

Mixture of Memories

The tested MoM uses four routed banks, activates the top two for each token, and keeps one shared bank active for every token.

bank 1
bank 2
bank 3
bank 4
shared
Stored state is not active state. MoM stores several banks but updates only selected routed banks plus the shared bank for a token. More stored state does not by itself prove more useful recall.

The paper reports a positive result at larger scale

The MoM paper trains 24-layer models on SlimPajama and averages six recall tasks with inputs up to 2,048 tokens. Its models use four routed memories, top-2 activation, a shared memory, and GDN updates. These scores are not percentages from our task.

380M label · 15B tokens

Paper Table 2. Nominal scale labels are not exact equal-parameter matches.

GDN
24.78
MoM
28.16

1.3B label · 100B tokens

Paper Table 2. MoM leads GDN by 3.74 score points.

GDN
32.30
MoM
36.04

Expanded GDN control

Paper Table 4 changes value width and reports 550M versus 444M parameters.

GDN
26.32
MoM
28.16
Paper-reported recall averages. Higher is better.
paper settingGDNMoMMoM minus GDN
nominal 380M, 15B training tokens24.7828.16+3.38
nominal 1.3B, 100B training tokens32.3036.04+3.74
400M activated parameters24.7826.51+1.73
expanded GDN versus MoM26.3228.16+1.84

Source: paper Tables 2, 4, and 9. The expanded control matches activated memory capacity, not our total stored-state comparison. The paper's main efficiency comparison is against Transformer inference, not GDN training-update time.

Local latest-value recall varies sharply by seed

The synthetic task writes key/value pairs, can overwrite a value, inserts unrelated tokens, and asks for the latest value of a queried key. Recall accuracy is the fraction of exact answer tokens. Each model has two layers and width 256. Each job runs 8,000 optimizer updates and scores 8,192 answers in the primary 64-key cell.

Seed 912011

Original MoM does not acquire the primary task. Original GDN reaches 77.14%.

MoM
8.06%
GDN
77.14%

Seed 912013

Original MoM reaches 84.29%, 1.56 percentage points below paired GDN.

MoM
84.29%
GDN
85.85%

Seed 913019

Original MoM remains near chance while paired GDN reaches 93.44%.

MoM
0.74%
GDN
93.44%
original MoMoriginal GDN
Primary 64-key recall. Exact counts out of 8,192 answers.
seedoriginal MoMMoM + causal filteroriginal GDNGDN + causal filter
912011660 / 8192 · 8.06%43 / 8192 · 0.52%6319 / 8192 · 77.14%1182 / 8192 · 14.43%
9120136905 / 8192 · 84.29%59 / 8192 · 0.72%7033 / 8192 · 85.85%7602 / 8192 · 92.80%
91301961 / 8192 · 0.74%1380 / 8192 · 16.85%7655 / 8192 · 93.44%7036 / 8192 · 85.89%
Easy 8-key acquisition check. Exact counts out of 1,024 answers.
seedoriginal MoMMoM + causal filteroriginal GDNGDN + causal filter
912011649 / 102414 / 10241024 / 1024882 / 1024
9120131024 / 102427 / 10241024 / 10241024 / 1024
91301927 / 1024908 / 10241024 / 10241024 / 1024
The causal-filter comparison failed its control requirement. GDN with the filter scored 882/1024 on the easy test for seed 912011. The required floor was 973/1024. The formal comparison is therefore unassessed. The endpoints still describe the completed runs, but they are not a qualified negative mechanism result.

The filter is a learned residual causal convolution. It lets each token use its current input and two previous inputs before routing. It starts at zero. It adds 1,536 parameters and 4,096 logical FP32 history bytes per sequence. The original models use 716,800 logical state bytes; the filter variants use 720,896. These are state-size calculations, not measured memory traffic.

A separate profile measures slower MoM training updates

This profile uses a small language-model workload on one AMD Radeon Pro W7900 GPU. It is not timing from the H100 recall cohort. The metric is profiler-off median time for a complete training update. Lower is better.

Seed 41011
1.3778×

GDN: 295.66 ms. Normalized top-2 MoM: 407.35 ms.

Seed 41013
1.3745×

GDN: 297.93 ms. Normalized top-2 MoM: 409.50 ms.

Profiler-off complete-update timing.
seedGDN median updatetop-2 MoM median updateMoM / GDN
41011295.66 ms407.35 ms1.3778×
41013297.93 ms409.50 ms1.3745×
What the profile does not show. The padded projection layout held about 1.90 times as many rows as the valid routed set, but recurrence runs after compaction. The profile did not isolate backward traffic or prove an HBM bottleneck. It does not support an inference-speed claim.

Small repairs did not produce a reliable challenger

These development studies test different interventions and often reuse development seeds. They are not one pooled statistical experiment.

interventionsupported observation
Normalize selected top-2 weightsRouting engages, but the development screen does not pass.
Change routed convolution spansNo nominated variant; one favorable five-seed mean remains exploratory.
Put queries on the original timelineMoM loss remains above GDN by +0.0145, +0.0023, and +0.0328.
Change the output gateNo qualifying gain; MoM remains above its GDN control.
Use two active banks plus a shared bankBeats a tied sparse reference, but does not reliably beat GDN.
Bank-specific output mapsNo consistent mechanism or challenger signal across seeds.
Causal context filterPoor MoM endpoints; one required GDN filter control fails.

Methods and evidence scope

GenerateSample feasible latest-value recall worlds with overwrites, interference, and controlled query distance.
TrainRun 8,000 fixed updates, 32 prefix pairs and 64 query rows per update, with paired data inside each seed.
ScoreUse exact answer-token accuracy on 128 prefixes for each terminal in-distribution cell.
AuditRetain every endpoint and independently reconstruct 675,840 prediction rows and 228 cell summaries.

Recall runtime

One NVIDIA H100 80 GB GPU; Python 3.11.12; PyTorch 2.9.1 with CUDA 12.8; Triton 3.8.0; Flash Linear Attention 0.5.2; FP32 with TF32 disabled. AdamW uses peak LR 0.001, 200 warmup updates, betas 0.9/0.95, weight decay 0.1, and gradient clipping at 1.0. MoM adds a 0.01 balance objective.

Public evidence

The machine-readable aggregate bundle contains exact counts, configs, paper anchors, and timing rows. No W&B tracker was used for the recall cohort. The exact runtime source snapshot, checkpoints, and row-level predictions remain outside this public repository, so the original runs cannot yet be reproduced from the public tree alone.

What these experiments establish—and what they do not

Established here

Learning is inconsistent

Original MoM can acquire the 64-key task, but only one of three paired runs reaches GDN-like recall at the fixed endpoint.

Not established

The cause of weak runs

Initialization, sampled training examples, evaluation examples, routing, normalization, and optimization all vary or remain plausible. Endpoint scores do not isolate a cause.

Different regime

The paper's scale result

The published models, data, objectives, and training budgets differ substantially. Our result neither reproduces nor disproves the paper's positive finding.

Supported conclusion: the tested MoM recipes have not produced a reliable advantage over GDN. Recall varies sharply across seeds, and the measured language-model implementation has slower updates. This result is limited to the tested configurations.