Mixture of Memories versus Gated DeltaNet
We did not find a reliable Mixture of Memories advantage over Gated DeltaNet in our tested recipes. One original-MoM run reached 84.29% on the primary recall test, near its paired GDN result of 85.85%. The other two MoM runs reached 8.06% and 0.74%. A separate language-model profile measured MoM training updates about 38% longer.
The published MoM paper reports higher recall scores than GDN after much larger language-pretraining runs. Our work uses small models and a synthetic overwrite-recall task. It is not a reproduction of that campaign and does not disprove its result.
What changes when memory is split into banks
Gated DeltaNet (GDN) stores associations in recurrent matrix state and uses a learned gate to control retention. Mixture of Memories (MoM) routes each token to a subset of several memory banks and also updates an always-active shared bank. A router chooses where each token writes and reads.
Gated DeltaNet
Every token updates the same multi-head recurrent memory. The model does not need to learn a bank assignment before it can use the state.
Mixture of Memories
The tested MoM uses four routed banks, activates the top two for each token, and keeps one shared bank active for every token.
The paper reports a positive result at larger scale
The MoM paper trains 24-layer models on SlimPajama and averages six recall tasks with inputs up to 2,048 tokens. Its models use four routed memories, top-2 activation, a shared memory, and GDN updates. These scores are not percentages from our task.
380M label · 15B tokens
Paper Table 2. Nominal scale labels are not exact equal-parameter matches.
1.3B label · 100B tokens
Paper Table 2. MoM leads GDN by 3.74 score points.
Expanded GDN control
Paper Table 4 changes value width and reports 550M versus 444M parameters.
| paper setting | GDN | MoM | MoM minus GDN |
|---|---|---|---|
| nominal 380M, 15B training tokens | 24.78 | 28.16 | +3.38 |
| nominal 1.3B, 100B training tokens | 32.30 | 36.04 | +3.74 |
| 400M activated parameters | 24.78 | 26.51 | +1.73 |
| expanded GDN versus MoM | 26.32 | 28.16 | +1.84 |
Source: paper Tables 2, 4, and 9. The expanded control matches activated memory capacity, not our total stored-state comparison. The paper's main efficiency comparison is against Transformer inference, not GDN training-update time.
Local latest-value recall varies sharply by seed
The synthetic task writes key/value pairs, can overwrite a value, inserts unrelated tokens, and asks for the latest value of a queried key. Recall accuracy is the fraction of exact answer tokens. Each model has two layers and width 256. Each job runs 8,000 optimizer updates and scores 8,192 answers in the primary 64-key cell.
Seed 912011
Original MoM does not acquire the primary task. Original GDN reaches 77.14%.
Seed 912013
Original MoM reaches 84.29%, 1.56 percentage points below paired GDN.
Seed 913019
Original MoM remains near chance while paired GDN reaches 93.44%.
| seed | original MoM | MoM + causal filter | original GDN | GDN + causal filter |
|---|---|---|---|---|
| 912011 | 660 / 8192 · 8.06% | 43 / 8192 · 0.52% | 6319 / 8192 · 77.14% | 1182 / 8192 · 14.43% |
| 912013 | 6905 / 8192 · 84.29% | 59 / 8192 · 0.72% | 7033 / 8192 · 85.85% | 7602 / 8192 · 92.80% |
| 913019 | 61 / 8192 · 0.74% | 1380 / 8192 · 16.85% | 7655 / 8192 · 93.44% | 7036 / 8192 · 85.89% |
| seed | original MoM | MoM + causal filter | original GDN | GDN + causal filter |
|---|---|---|---|---|
| 912011 | 649 / 1024 | 14 / 1024 | 1024 / 1024 | 882 / 1024 |
| 912013 | 1024 / 1024 | 27 / 1024 | 1024 / 1024 | 1024 / 1024 |
| 913019 | 27 / 1024 | 908 / 1024 | 1024 / 1024 | 1024 / 1024 |
The filter is a learned residual causal convolution. It lets each token use its current input and two previous inputs before routing. It starts at zero. It adds 1,536 parameters and 4,096 logical FP32 history bytes per sequence. The original models use 716,800 logical state bytes; the filter variants use 720,896. These are state-size calculations, not measured memory traffic.
A separate profile measures slower MoM training updates
This profile uses a small language-model workload on one AMD Radeon Pro W7900 GPU. It is not timing from the H100 recall cohort. The metric is profiler-off median time for a complete training update. Lower is better.
GDN: 295.66 ms. Normalized top-2 MoM: 407.35 ms.
GDN: 297.93 ms. Normalized top-2 MoM: 409.50 ms.
| seed | GDN median update | top-2 MoM median update | MoM / GDN |
|---|---|---|---|
| 41011 | 295.66 ms | 407.35 ms | 1.3778× |
| 41013 | 297.93 ms | 409.50 ms | 1.3745× |
Small repairs did not produce a reliable challenger
These development studies test different interventions and often reuse development seeds. They are not one pooled statistical experiment.
| intervention | supported observation |
|---|---|
| Normalize selected top-2 weights | Routing engages, but the development screen does not pass. |
| Change routed convolution spans | No nominated variant; one favorable five-seed mean remains exploratory. |
| Put queries on the original timeline | MoM loss remains above GDN by +0.0145, +0.0023, and +0.0328. |
| Change the output gate | No qualifying gain; MoM remains above its GDN control. |
| Use two active banks plus a shared bank | Beats a tied sparse reference, but does not reliably beat GDN. |
| Bank-specific output maps | No consistent mechanism or challenger signal across seeds. |
| Causal context filter | Poor MoM endpoints; one required GDN filter control fails. |
Methods and evidence scope
Recall runtime
One NVIDIA H100 80 GB GPU; Python 3.11.12; PyTorch 2.9.1 with CUDA 12.8; Triton 3.8.0; Flash Linear Attention 0.5.2; FP32 with TF32 disabled. AdamW uses peak LR 0.001, 200 warmup updates, betas 0.9/0.95, weight decay 0.1, and gradient clipping at 1.0. MoM adds a 0.01 balance objective.
Public evidence
The machine-readable aggregate bundle contains exact counts, configs, paper anchors, and timing rows. No W&B tracker was used for the recall cohort. The exact runtime source snapshot, checkpoints, and row-level predictions remain outside this public repository, so the original runs cannot yet be reproduced from the public tree alone.
What these experiments establish—and what they do not
Learning is inconsistent
Original MoM can acquire the 64-key task, but only one of three paired runs reaches GDN-like recall at the fixed endpoint.
The cause of weak runs
Initialization, sampled training examples, evaluation examples, routing, normalization, and optimization all vary or remain plausible. Endpoint scores do not isolate a cause.
The paper's scale result
The published models, data, objectives, and training budgets differ substantially. Our result neither reproduces nor disproves the paper's positive finding.
Supported conclusion: the tested MoM recipes have not produced a reliable advantage over GDN. Recall varies sharply across seeds, and the measured language-model implementation has slower updates. This result is limited to the tested configurations.