Measurement records

Soup's repository publishes its measurement records in a benchmarks/ directory. They are the working records kept while each result was built and checked, not a report written afterwards, so they keep the failures, the readings that turned out wrong and the numbers that were measured and then discarded, in the order those things happened. This page indexes every record in that directory at the v0.75.2 tag: the question each one asked, the hardware and software it states, its headline result with the caveat it carries, and whether it passed, failed or was never a gate.

The site says its performance numbers are measured rather than claimed, and that is only worth something if a reader can find the measurement. This page is the map. It is not a second copy of the results: the record is the authority, and where this page and a record differ, the record wins. Upstream keeps its own index in benchmarks/README.md. The layer streaming records are the evidence behind the layer streaming preprint, whose concept DOI is 10.5281/zenodo.21771064.

The directory holds fourteen records, that README, and seven scripts. Every layer streaming figure published up to v0.72.4 came from one laptop. The later records add hardware the maintainer does not own: a borrowed 8x H100 box, a free Colab T4, and machines belonging to contributors.

How to read this page

  • Each record is linked at the tag, at the top of the file. Read it to the end before quoting a number from the middle. Several leave a wrong passage standing with its correction beside it or further down, on purpose, so a figure lifted from the middle may be one the same file later withdraws. Where a record does this, its entry below says so, and a figure a record later withdraws is not quoted here.
  • The status words are the record's own. "Not a gate", "partial pass" and "gate failed" are quoted from the files, not softened or sharpened.
  • Units are as the record prints them. Most gate records use decimal GB, the H100 record prints MiB, the pinned-memory record prints KiB, the MLX record uses GB as MLX reports them, and the QuEST record reports loss in nats. Nothing here is converted.
  • "Bit-exact" is always two claims. The forward pass (logits, compared with torch.equal) and the backward pass (every LoRA gradient tensor) are measured independently, and above roughly 165 MiB per NF4 layer they disagreed before the repair. Wherever this page says exact, it names the half, and the quantisation where the record states one. The per-model ledger has all four cells for every row.
  • A correctness reference always matches the numerics under test. A streamed NF4 run is compared with a resident NF4 run, never with resident bf16, whose quantisation error would hide a real defect.
  • A throughput figure is only comparable at the clock it was taken at. The laptop card's boost clock varies about 13% between sessions, so the records quote the SM clock beside throughput and compare against a GEMM ceiling measured in the same session.
  • Derived figures are labelled as arithmetic. "One million tokens takes about 2.3 hours" is division, not a wall-clock run.

At a glance

RecordBelongs toMachineStatus as the record states it
Streaming pathv0.72.0RTX 3050 Laptop 4 GBBoth gates pass
NF4 streamingv0.72.2RTX 3050 Laptop 4 GBGate 1 passes (6 of 6), gate 2 complete
Breadthv0.72.3RTX 3050 Laptop 4 GBSix gates, all pass
Preference lossesv0.72.4RTX 3050 Laptop 4 GBGate 1 passes 14 of 14, gates 2 and 3 pass 10 of 11
What bounds a streamed stepv0.73.0RTX 3050 Laptop 4 GBTask 1 answered, task 2 answered for the kernel and blocked for the shipped flag
Measuring the VRAM fitv0.73.1RTX 3050 Laptop 4 GBFound the pre-flight under-predicting at long sequence
Leg 2 of soup shipv0.73.2RTX 3050 Laptop 4 GB, CPU for model runsWorking record, no verdict line
8x H100The v0.72.4 line, repair in v0.73.08x H100 80 GBExternal-hardware gate, defect repaired and re-gated, dated corrections
Free Colab T4v0.73.1Colab Tesla T4Run completed, correctness not gated, not a gate
A10G#395, Soup 0.73.3AWS A10G 23 GBUnder-prediction does not reproduce, hypothesis not supported
L4#655NVIDIA L4 containerMeasurement supporting a design decision
M4 Maxv0.74.0MacBook Pro M4 Max 128 GiBPartial pass, optimizer step not validated
8 GB M1#23Apple M1 8 GBFive runs completed, one fixture failed, not a gate
QuEST W4A4#674RTX 5090 Laptop 24 GB classGate failed

Read the last column first. One record is a failed gate, one is a partial pass, and two say outright that they are not gates. One is a negative result for a hypothesis, and one is the gate that found a shipped formula failing its own contract. They are in this index because an index of the flattering ones would be worth nothing.

Layer streaming, as it was built

The first four records are the gates that built the feature, all measured on one machine: an RTX 3050 Laptop with 4 GB of VRAM (4.29 GB usable), 16.9 GB of host RAM, NVMe and Windows 11. Windows matters when reading them. Its driver spills into shared host memory instead of raising an out-of-memory error, so a run completing is not evidence that its configuration fits, which is why peak VRAM is reported beside every throughput figure.

v0.72.0: the streaming path

  • File: gate-v0.72.0-layer-streaming.md. Belongs to v0.72.0, and has been in a tagged release since v0.72.4.
  • Question: Does streaming the frozen base one decoder layer at a time give the same numbers as holding it resident, and what throughput does a 4 GB card sustain doing it?
  • Hardware and stack: RTX 3050 Laptop 4 GB (compute capability 8.6, driver 591.44), 16.9 GB RAM, NVMe, Windows 11. torch 2.5.1+cu121, peft 0.18.1, accelerate 1.12.0, Python 3.10.
  • Result: On SmolLM2-135M in bf16 the streamed forward is bit-exact against the resident model (logits, maximum absolute difference 0.0) and a 100-step loss curve is identical. Qwen2.5-3B in bf16 trains in 2.15 GB where the same model trained resident runs out of memory on that card.
  • Status: "Both gates PASS", in the record's words.
  • Caveats: The 3B row used a pageable store, so it is a lower bound. The numbers are Windows and pessimistic against Linux, and nothing was measured above 3B. The only valid streaming-versus-resident overhead is 1.43x, at 0.5B, because the resident 1.5B row spilled into host memory and its number is discarded.
  • Read to the end: Near the end, "After the gates" records what a passing gate did not catch: no test had built a real trainer from the streamed model, so every real run would have died at trainer construction, and a shard cache could silently stream stale weights. A paragraph that first said the transfer was fully hidden was corrected in place on 2026-07-27 and then explained row by row. Treat those row-level explanations as the author's readings at the time, and read the probe record below, where a mechanism was actually measured, before reusing one. The file also records that three third-party figures drafted for the write-up were wrong when checked against source, which is why the defensible claim is 1.3B to 3B on a 4 GB card and not "nobody has done this".
  • On this site: Layer streaming, measured numbers and the paper.

v0.72.2: NF4 streaming

  • File: gate-v0.72.2-nf4.md. Belongs to v0.72.2 (its own heading says "now the v0.72.2 slot", because the gates ran while NF4 was planned as v0.72.1), and has been in a tagged release since v0.72.4.
  • Question: Can the base be quantised to NF4 offline and streamed, still matching a resident NF4 model, and how large a model does that bring onto 4 GB?
  • Hardware and stack: RTX 3050 Laptop 4 GB, 16.9 GB RAM, NVMe, Windows 11. torch 2.5.1+cu121, bitsandbytes 0.49.2, transformers 4.57.6, peft 0.18.1, trl 0.19.1, accelerate 1.12.0, Python 3.10.8.
  • Result: On SmolLM2-135M the streamed NF4 forward is bit-exact against a resident NF4 model (logits, maximum difference 0.0), gradients are non-zero on all 30 layers and a 25-step loss curve matches. Llama-3.1-8B in NF4 then runs on the 4 GB card from a 3.60 GB page-locked store.
  • Status: "GATE 1 PASS (6/6, bit-exact vs resident NF4). GATE 2 complete", in the record's words.
  • Caveats: The file carries two 8B rows, a first measurement from a throwaway spike and a later reproduction through the shipped code, and the reproduction is the figure the release notes, the README and this site quote: 119.6 tok/s and 3.32 GB, on a 3.60 GB page-locked store. Two caveats travel with the published figure. It was measured before the NF4 gradient repair of v0.73.0 and has not been re-run on a 4 GB card since, and the 3.32 GB peak predates the large-layer streaming change of v0.74.0, which shards an untied embedding and output head separately and so moves the resident share. The H100 cross-check, a median 113.00 tok/s at a 3,397 MiB peak (the H100 record prints MiB, where the laptop records print decimal GB), is a reproduction on other hardware and not a re-measurement of the laptop figure.
  • Read to the end: The gate found the v0.72.1 defect, where streamed adapters were saved with an extra key segment and reloaded as the untuned base. Two of its own checks were nearly vacuous until fixed, because an adapter initialised to zero makes the adapter path do nothing. A first TFLOPS reading was held back and corrected: the ceiling tracks the GPU clock, so a fraction of a ceiling means nothing without the clock it was taken at.
  • On this site: Layer streaming, measured numbers and why NF4 also made the 3B row faster.

v0.72.3: breadth

  • File: gate-v0.72.3-breadth.md. Belongs to v0.72.3, and has been in a tagged release since v0.72.4.
  • Question: Can streaming widen to more architectures, batches above 1, gradient accumulation, resume and a disk tier without silently training the wrong thing, and can the VRAM pre-flight predict a peak well enough to refuse a run that will not fit?
  • Hardware and stack: RTX 3050 Laptop 4.29 GB, 16.9 GB RAM, NVMe, Windows 11, Python 3.10. torch 2.5.1+cu121, transformers 4.57.6, peft 0.18.1, trl 0.19.1, bitsandbytes 0.49.2.
  • Result: Each of the six added architectures (mistral, gemma, gemma2, gemma3_text, phi, phi3) streams with forward logits equal to the same checkpoint loaded resident, in an unquantised and an NF4 arm, but on tiny from-config CPU float32 fixtures, not real checkpoints and not the RTX 3050. The peak-VRAM predictor came within 0.85% of ten real runs on two models and never under-predicted on that grid. Raising the batch beat accumulating at the same effective batch by a measured 2.52x, and accumulation is neutral per token.
  • Status: All six gates are recorded as PASS, gate 1 as 14 of 14.
  • Caveats: The predictor was fitted at sequences of 256 and 512 only, and the v0.73.1 record later found it under-predicting at long sequence. Refusing a run that is too big is policy applied to a measured demand: Windows spills instead of raising, so the record does not claim it was verified by observing an out-of-memory error. End-to-end soup train --resume could not be demonstrated on that box, because torch 2.5.1 makes transformers refuse every resume and a stream_layers: false control fails identically, so only the streaming-specific half is verified. The speed difference between the RAM tier and the disk tier is deliberately unmeasured and no figure is claimed. Only the disk tier's correctness is shown, with logits equal to both the RAM tier and a resident run. No tok/s exists for the six added architectures.
  • Read to the end: The record keeps its own over-claims. A first reading of the soup doctor --disk cost was wrong and is corrected, and a GEMM probe that swung by a factor of two at the same reported clock was diagnosed rather than accepted.
  • On this site: Scaling a streaming run.

v0.72.4: preference losses

  • File: gate-v0.72.4-preference-losses.md. Belongs to v0.72.4, and has been in a tagged release since v0.72.4.
  • Question: Can DPO, ORPO, SimPO and KTO run over a streamed base without a second copy of the weights? DPO's reference model has to be the same streamed base with adapters disabled, because a second instance would defeat the feature, and a passing loss curve cannot detect that.
  • Hardware and stack: RTX 3050 Laptop 4.29 GB, 16.9 GB RAM, NVMe, Windows 11, Python 3.10.8. torch 2.5.1+cu121, transformers 4.57.6, peft 0.18.1, trl 0.19.1. The record states that CI resolves a much newer stack, so it asserts the properties the shipped tests depend on rather than TRL's internals.
  • Result: On a synthetic 365.2M-parameter fixture (24 layers, vocabulary 260, sequence 64, batch 1), streamed DPO peaked at 0.914x the streamed SFT peak, while a control that forced a real second model instance added +730.44 MB against 730.44 MB of weights, which is exactly one copy.
  • Status: Gate 1 (DPO) is recorded as PASS 14 of 14, and gates 2 and 3 (ORPO, SimPO, KTO) as PASS 10 of 11. The one miss is a pre-flight check on the SFT path at -1.7%, which the record attributes, with evidence, to a fixture far outside the range the formula was fitted on.
  • Caveats: Read 0.914x as "no second copy", not as DPO being cheaper than SFT: the SFT arm computes its loss over all 64 positions while the DPO arm splits them into prompt and completion, so its logits tensor is smaller. The control is the half that carries the result. The cost is 1.52x the layer reads per step, free in memory and not in time. The fixture is not a real checkpoint and not Llama-3.1-8B, DPO's loss matches a resident DPO run exactly (difference 0.0, on a 4-layer CPU float32 model), and no tok/s was measured for any preference loss. KTO is not reference-free and needs a batch size of 2 or more.
  • Read to the end: The file opens with a table of TRL versions that is wrong, and a correction directly beneath it says so and was left in place on purpose. The lesson the record draws is that a version bound derived by reading source is a hypothesis, and the experiment that tests it is constructing the object. It also lists three measurement attempts it marks INVALID so they are not repeated.
  • On this site: Preference losses over layer streaming, and why 0.914x means no second copy.

The probe that retracted an explanation

One record measured what had only been inferred, and its result changed what the preprint and this site may say about why a streamed step takes as long as it does.

v0.73.0: what bounds a streamed step

  • File: probe-v0.73.0-what-bounds-streaming.md. Belongs to v0.73.0 (its header names the v0.73.0 working tree at commit 57c18e5), and has been in a tagged release since v0.73.1.
  • Question: What actually bounds a streamed step, and what happens when Cut Cross-Entropy is switched on beside streaming?
  • Hardware and stack: RTX 3050 Laptop 4 GB, 16.9 GB RAM, NVMe, Windows 11. torch 2.5.1+cu121, bitsandbytes 0.49.2, transformers 4.57.6, peft 0.18.1, Python 3.10.8.
  • Result: On Llama-3.1-8B in NF4 at batch 1 and sequence 512, deleting every host-to-device byte (6.864 GB per step) makes the step 1.44% faster. The compute stream is blocked on a copy for 8.4 ms of a 4190 ms step, and the step runs at 71.3% of the same-session, shape-matched GEMM ceiling at 952 MHz. Removing the NF4 dequantisation instead buys 9.80%. These are the measurements that retracted the explanation version 2 of the preprint gave for the H100 replication, and the record limits that explanation to very short steps, where the fixed transfer volume dominates. Its own figures put that boundary near 128 tokens per step, and the published configuration is not there.
  • Status: "TASK 1 ANSWERED", and task 2 "ANSWERED for the kernel, BLOCKED for the shipped flag".
  • Task 2, Cut Cross-Entropy: A hand-wired Cut Cross-Entropy kernel raised the usable microbatch from 1 to 3 for +9.6% throughput on that card, but Soup's own training.use_cut_ce could not engage on the stack measured, for three separate import and packaging reasons, and the kernel's backward is not reproducible against itself, so the bit-exactness standard cannot apply to it.
  • Caveats: Those blockers were measured on the v0.73.0 stack (transformers 4.57.6). Soup has required transformers 5.x since v0.74.0, and the record does not speak to that stack. It is one card on Windows. The step decomposition is Llama-3.1-8B in NF4 only, and the Cut Cross-Entropy correctness check ran on SmolLM2-135M in NF4. bf16 streaming, sequences beyond 512 and the disk tier were not measured, and the H100's own bottleneck was never instrumented, so the record makes no claim about it.
  • Read to the end: The premise was corrected before the first measurement, because the brief had assumed a 128-token step. Section 10 lists what was not measured. The record gives its own one-sentence reading of the laptop result in section 6; this page quotes the measurements, not that reading, and offers no mechanism of its own.
  • On this site: What version 3 of the paper adds and the laptop number reproducing on other hardware.

The pre-flight and the release gate

Two records test the instruments rather than the feature: the VRAM pre-flight that decides whether a streaming run may start, and the scorers behind the soup ship verdict.

v0.73.1: measuring the VRAM fit

  • File: gate-v0.73.1-measured-vram-fit.md. Belongs to v0.73.1 (issue #349), and has been in a tagged release since v0.73.1.
  • Question: The streaming pre-flight predicts peak VRAM from a formula and promises it never under-predicts. Can that prediction be replaced by a measurement, and does the promise hold?
  • Hardware and stack: RTX 3050 Laptop 4 GB (4.294 GB total, 3.45 GB free at rest), Windows 11, NVMe. torch 2.5.1+cu121, transformers 4.57.6, trl 0.19.1, peft 0.18.1, Python 3.10.
  • Result: Measured through the real soup train on SmolLM2-135M (streamed, bf16, batch 1), the formula over-predicts at sequence 4352 (1.081x) and under-predicts at 5120 (0.934x) and at 6144 (0.787x, under by 21%). The ten-run grid behind the "never under-predicts" contract varied batch and never sequence length, so it could not have seen this.
  • Status: This is the gate that found a contract failing, and its headline is "the formula under-predicts at long sequence".
  • Caveats: The mechanism is "NOT established" in the record's words. The release added an opt-in measured probe, training.stream_vram_probe, which applies to task: sft only. The probe is not validated for preference losses, where one shape is not a validation.
  • Read to the end: Three readings were withdrawn during the work, and two of them looked like the headline result. They are left in with their refutations, and none of them is a result. The record also corrects a draft claim about the probe on a preference loss, which had compared two different shapes.
  • On this site: The v0.73.1 release page and the long-sequence under-prediction in the pre-flight.

v0.73.2: leg 2 of soup ship

  • File: gate-v0.73.2-leg2-scoring.md. Belongs to v0.73.2, and has been in a tagged release since v0.73.2.
  • Question: Do four reported defects in the scorers behind soup ship's regression leg reproduce (#357, #346, #355 and #317), and does the repair hold?
  • Hardware and stack: RTX 3050 Laptop 4 GB on Windows 11 with Python 3.10, but every model run was on the CPU (a 135M pair at 256 new tokens), and the record says no GPU was needed and none is claimed. The baseline is shipped v0.73.1.
  • Result: A stub answering every item correctly in a boxed-letter style scored 0.000 on two multiple-choice suites. A stub naming the right tool 40 of 40 times, one closing brace short, scored 0.000 on the tool-call suite. Two models with byte-identical leg-2 scores, one of them refusing every benign request, got the same SHIP verdict.
  • Status: A working record with no gate verdict line.
  • Caveats: The 0.423 to 0.731 and 0.225 figures in the issues are the H100 record's, on an 8B, and are not re-measured here. The noise floor measured on this box is 0.0000 because CPU greedy decode is deterministic, and the H100's 0.015 and 0.020 spreads are explicitly not re-claimed. The source issue's floor is n=3 on one model and one dataset, so it sizes an effect and does not calibrate a threshold. The record also explains why the preprint is scoped and not amended: its leg-2 numbers were measured with the earlier scorer, are correct as measured, and should not be compared across the boundary.
  • Read to the end: A "66 failures" scare turned out to be a read of a half-written file and was withdrawn. A control that varied nothing is recorded as such, and a review finding was checked against three implementations and partly rejected.
  • On this site: The v0.73.2 release page and soup ship.

Hardware the maintainer does not own

Everything above was measured on the one laptop. These records are the exceptions, and each says which of its claims depend on that.

8x H100: external validation

  • File: gate-h100-validation.md. Measured on a checkout made just after the v0.72.4 tag, which still reported itself as 0.72.4 (the defect it found is present across v0.72.0 to v0.72.4), with the repair shipped in v0.73.0, and has been in a tagged release since v0.73.0. It is a long file, about five thousand lines.
  • Question: Does the method hold on someone else's hardware, at model sizes where a resident reference exists to compare against, which a 4 GB card cannot hold?
  • Hardware and stack: 8x NVIDIA H100 80 GB HBM3 over PCIe with no NVLink, 503 GB RAM, Ubuntu 24.04.3, driver 590.48.01. torch 2.13.0+cu130, bitsandbytes 0.50.0, transformers 4.57.6, trl 0.26.2, peft 0.20.0.
  • Result: The forward (logits, torch.equal) matched a resident reference of the same numerics at 0.5B, 8B, 14B, 32B and 72B. The backward (every LoRA gradient tensor) matched at 0.5B, 8B and 14B in NF4, and above about 165 MiB per NF4 layer it was wrong until the repair: 8 of 256 tensors on 32B and 8 of 320 on 72B. After the repair it re-gated at 256 of 256 on 32B and 320 of 320 on 72B, each against a control arm that reproduced the defect in the same process.
  • Status: An external-hardware gate, with its status block saying the defect is repaired and re-gated.
  • Caveats: Nowhere does the record call 72B exact in the backward before the repair, and "bit-exact at 72B" on its own is a forward statement.
  • Also in the record: The laptop headline reproduced at a median 113.00 tok/s with a 3,397 MiB peak (n=5), with no explanation attached on this site; a DeepSpeed ZeRO-3 comparison, which the record frames as not faster than DeepSpeed but built for a different situation, one card that is too small; and a convergence comparison against resident runs. The external validation page carries those figures with their caveats.
  • Read to the end: The opening summary explains the H100 replication by host-to-device transfer. That was measured and refuted afterwards, by the probe record above, and the file carries dated corrections (2026-08-13) beside each place it appeared, with the original text left standing. It opens with a per-model exactness ledger giving the forward, the backward, the quantisation, MiB per layer and the reference for every row, with unmeasured cells marked "not tested" instead of left blank, because two readers took an unqualified "bit-exact" to cover both halves. The defect is also filed upstream, as bitsandbytes issue 2034.
  • On this site: Validation on hardware we do not own, the paper and what to re-run after v0.73.0.

Free Colab T4: an 8B run

  • File: run-t4-colab-free-tier.md. Belongs to v0.73.1 (the pre-Ampere fp16 repair, #385 and #387), and has been in a tagged release since v0.73.1.
  • Question: Does the streaming path run to completion on a pre-Ampere (Turing) card? The repair had only been verified using fp16 on an Ampere card, which establishes the plumbing and says nothing about Turing kernels.
  • Hardware and stack: A free-tier Google Colab Tesla T4 (sm_75, 15.6 GB), with the process capped at 4.00 GB through torch.cuda.set_per_process_memory_fraction. Driver and library versions were not recorded. Llama-3.1-8B-Instruct in NF4, fp16, batch 1, max_length: 256, LoRA r=8.
  • Result: Peak VRAM was 2.91 GB against a predicted 3.02 GB, so the pre-flight over-predicted by 3.8%, the safe direction. Seven steps completed and the adapter was written with 128 of 128 tensors non-zero. The cap was shown to bite: a deliberate 4.29 GiB allocation was refused.
  • Status: "RUN COMPLETED. Correctness NOT gated", and the record's own index calls it not a gate.
  • Caveats: It makes no throughput claim, and the pre-flight panel's forecast is a compute bound derived from a GEMM probe, not a measurement. Neither forward nor backward exactness on Turing is shown, because non-zero gradients prove gradients flowed and not that they were correct. The notebook's streamed-versus-resident section produced no output and is recorded as unrun. It is one run on one seed with no repeats, and the library versions are missing, so it is not exactly reproducible. The seven-step loss series is not evidence of learning.
  • Read to the end: The run noticed by accident that the pre-flight read the device's free VRAM (15.10 GB), not the 4.00 GB cap the run was held to. On capped hardware the pre-flight is therefore not what enforces the real budget, which is the case training.stream_vram_override exists for. The record says its list of what it does not establish is longer than its results on purpose.
  • On this site: An 8B on a free Colab T4 and running it yourself on a free Colab T4.

A10G: the VRAM fit on a second stack

  • File: gate-395-second-stack-vram.md. Filed against issue #395 and measured against Soup 0.73.3, and has been in a tagged release since v0.74.0.
  • Question: Does the long-sequence under-prediction from the v0.73.1 record reproduce on a second GPU and software stack, and is the leading hypothesis, a switch to the attention math path that materialises the full attention matrix, the cause?
  • Hardware and stack: NVIDIA A10G 23.0 GB (AWS g5.xlarge), Ubuntu 22.04, driver 580.126.09, torch 2.13.0+cu130, transformers 5.16.1, trl 0.29.1, peft 0.20.0, Python 3.10. The record says every number was measured in that configuration and nothing is extrapolated. Its transformers and peft versions are exactly the floor Soup now requires and its trl is one patch above that floor, whereas the laptop's are below it.
  • Result: It does not reproduce. The prediction over-estimates by a flat 1.162x to 1.169x from sequence 2048 to 6144 on SmolLM2-135M and Qwen2.5-0.5B (bf16, batch 1, LoRA r=8), against 0.934x at 5120 and 0.787x at 6144 on the RTX 3050.
  • Status: A negative result for a hypothesis, published with a positive control.
  • Caveats: The math path is exactly as expensive as the hypothesis needs, but SDPA does not select it on this stack. The hypothesis is therefore not supported here, and it is not refuted for the RTX 3050, whose Windows and torch 2.5.1 combination is where a math fallback is plausible and untested. The record says nothing about the RTX 3050 itself. Its accounting finds the retained-copy charge in the formula absent on this stack, which narrows issue #327 by one stack without explaining it, and every row is batch 1.
  • Read to the end: A back-solved coefficient from an earlier revision was withdrawn once the instrument to measure it directly was run, and a method note records the harness trap behind the first attempt.
  • On this site: The long-sequence under-prediction in the pre-flight and what was measured in v0.74.0.

L4: pinned host memory accounting

  • File: gate-655-pinned-store-accounting.md. Filed against issue #655, measured by Tristan Grech on 2026-09-10 UTC, and has been in a tagged release since v0.75.0.
  • Question: Does a pinned CUDA host allocation consume /dev/shm, that is, does the streaming pinned-store path need a /dev/shm free-space pre-flight?
  • Hardware and stack: NVIDIA L4, driver 570.195.03, PyTorch 2.9.1+cu128, CUDA runtime 12.8, Python 3.12.3, an Ubuntu 24.04.3 container on a RunPod host with kernel 6.14.0-33-generic.
  • Result: Two fresh-process pinned 4 GiB allocations each added exactly 4 GiB to RssShmem with no change to RssAnon or /dev/shm usage, and the pageable control added exactly 4 GiB to RssAnon with no change to RssShmem.
  • Status: A measurement record whose stated conclusion is a design decision: keep the /dev/shm free-space check out of the pinned-store path.
  • Caveats: It is one stack. It does not establish the driver's internal implementation, the behaviour on other drivers or kernels, or reclaimability, and it does not guarantee that a full streaming run fits. Host-wide counters can include other tenants, so the small background changes in those columns are not attributed to the allocation. The optional streaming test near 55% of host RAM was not run and no out-of-memory condition was induced.
  • Read to the end: The original issue comment is reproduced without edits, including the host-counter drift, at the maintainer's request. The tables are in KiB, and the exact probe script is included in the file.
  • On this site: What was measured in v0.75.0.

M4 Max: Qwen4-Exp, a partial pass

  • File: gate-qwen4-ple-m4-max.md. Dated 2026-08-31, belongs to v0.74.0 (the qwen4_exp streaming architecture), and has been in a tagged release since v0.74.0.
  • Question: Do the Qwen4-Exp and oQ decoder mapping and the read-only external N-gram (PLE) table work on Apple Silicon, and can the full production checkpoint be trained there?
  • Hardware and stack: MacBook Pro M4 Max with 128 GiB unified memory and an external SSD. macOS, Python 3.12, PyTorch 2.13.0. The checkpoint is an oMLX/oQ affine bundle (model_type=qwen4_exp, 48 decoder layers, 106.29 GB of source safetensors).
  • Result: All 1,167 expected tensors mapped, with none missing, unexpected or mis-shaped. The oQ affine decoder vectors matched MLX at 4, 5, 6 and 8 bits. The tiny native Qwen4 resident-versus-streamed gate was bit-exact on CPU for rows, logits, loss and LoRA gradients, and both MPS variants passed. The production checkpoint then reached "Training started!", but the one-step run stopped without completing an optimizer step.
  • Status: A partial pass. The record marks checkpoint discovery, oQ decoding, expert mapping, cache construction, tiny CPU and MPS parity and production setup as PASS, a complete optimizer step on the 176.9B-parameter checkpoint as NOT VALIDATED, and Qwen4 resident-versus-streamed bf16 parity on CUDA as PENDING. Its own line: "Do not market this record as proof that the full model is trainable on a 128 GiB Mac."
  • Caveats: It also says it must not be read as a throughput, peak-memory or production-trainability claim. It does not attribute the stop to host-RAM exhaustion, and the low-level cause is unresolved.
  • On this site: What was measured in v0.74.0 and which architectures stream.

8 GB M1: MLX SFT

  • File: run-m1-8gb-mlx-sft.md. Filed against issue #23, which asked for exactly this and had been open since v0.25.0 because CI has no Apple Silicon, and has been in a tagged release since v0.75.0.
  • Question: Does MLXSFTTrainerWrapper run end to end against a real MLX runtime, and where is the memory ceiling on an 8 GB machine?
  • Hardware and stack: Apple M1, 8 GB unified memory, 8 cores, macOS 26.6.2 (arm64). Python 3.12.14, mlx 0.32.2, mlx-lm 0.31.3. The config is backend: mlx, task: sft, LoRA r=8, batch 1, max_length: 512, on 48 synthetic chat rows for one epoch.
  • Result: The model the shipped llama3.1-8b-sft-mlx recipe names (mlx-community/Llama-3.1-8B-Instruct-4bit) trained on an 8 GB M1 under the smaller config above: 48 iterations in 71 s at a 5.154 GB peak, which is mlx-lm's own figure and the one the record says to quote, with the adapter written and reloadable. The recipe's own settings (batch size 2, max_length: 2048, LoRA r=16, 3 epochs) were not run, and its catalog description says M2+ 16GB, so the record's wording that the recipe fits an 8 GB M1 is stronger than what was measured. A model too large for memory does not fail there, it pages, so an allocation-failure pre-flight would not fire.
  • Status: "FIVE RUNS COMPLETED, ONE FIXTURE FAILED. This is not a gate."
  • Caveats: There is no correctness claim: the loss falls and the adapter reloads, and on 48 repeated rows a falling loss is memorisation. The LoRA is query and value only under target_modules: auto. Throughput is an order of magnitude and not a benchmark, since re-running one model with the same config gave a rate 3.5 times different, so this page quotes none. It is one box with one run per rung, and DPO and GRPO are untouched.
  • Read to the end: The failed fixture and the corrections are left in. The small model the issue recommends ships legacy weight file names and cannot be loaded. A tok/s column that mixed two metrics was caught in review, and the author's own prediction that the 8B would not fit was wrong, which is why the row exists.
  • On this site: The MLX backend and what was measured in v0.75.0.

A failed experiment

One record is a contributor's research experiment, not a Soup feature. It failed its own preregistered criterion and is published as written.

QuEST W4A4: the gate failed

  • File: gate-674-quest-w4a4-sft.md. A research record for issue #674, not a release and not shipped code, measured by @Shutaru (the main experiment on 2026-09-09, the last bounded diagnostic on 2026-09-11), and in a tagged release since v0.75.0. It leaves #674 open.
  • Question: Can an experimental QuEST-style W4A4 supervised fine-tune, with 4-bit weights and 4-bit activations simulated by fake quantization, come within a preregistered 0.1 nat of the strongest full-precision run on the same held-out panel?
  • Hardware and stack: NVIDIA GeForce RTX 5090 Laptop GPU (24,463 MiB reported), Windows, Python 3.12.10, PyTorch 2.11.0+cu128, Transformers 5.16.1. The model is ahxt/LiteLlama-460M-1T and the data is databricks/databricks-dolly-15k, with revisions pinned in the record. The NVIDIA driver was not captured at the time, and the record refuses to backfill it from a later query of the installed driver.
  • Result: The gate failed. The group128 candidate scored 2.349768 against 2.235367 for the strongest retained full-precision endpoint, a gap of 0.114 nat with a paired 95% interval of [0.099, 0.130], where the preregistered criterion required both the gap and its upper bound to be at most 0.1.
  • Status: The record's opening sentence is "Gate failed: 0.114 nat [0.099, 0.130], one training seed, one model, fake quantization. Not parity, not an integration, not an efficiency claim." The separate requirement of a demonstrated improvement over the non-group128 control also fails, since that interval crosses zero. No alternate arm was promoted and the reserved confirmation panel was never evaluated.
  • Caveats: The scripts and full records remain local, so the record publishes descriptions and SHA-256 bindings rather than a runnable integration, and the result cannot be re-run from the repository. In its descriptive desktop timings the reference implementation was slower and used more memory than the full-precision run; those are not benchmarks. There is no QuEST option in Soup and none is implied.
  • Read to the end: The record retains its failed directions, its cost table and a history of every evaluation panel it spent, and it says it can stand as a negative result without another training run.
  • On this site: What was measured in v0.75.0.

The scripts

The harness/ directory holds measurement scripts that can be run against a released Soup, so a claim in a record can be re-measured instead of taken on trust. It starts small on purpose: most of the original session's harnesses lived only in a scratchpad on a machine that no longer exists, which upstream tracks as issue #379.

ScriptRecordWhat it asksWhat it needs
issue331_qlora_scope.py8x H100Does the wrong-gradient defect reach ordinary QLoRA? Three arms in one process with a positive control. The answer on one Linear4bit is no: 0.0 against a control that diverges by 3.77e-01About 15 s on a 4 GB card, no downloads
bnb_repro.py8x H100, bitsandbytes issue 2034A minimal standalone reproducer of the NF4 stale-gradient mechanism. Private buffers are the reference and bf16 is the controlA CUDA GPU, no downloads, about a minute. Historical environment: torch 2.13.0+cu130, bitsandbytes 0.50.0
bitexact.py8x H100, its Reproducing sectionShard, stream, then compare logits, LoRA gradients and the loss curve against a resident reference of matching numerics. NF4 uses a resident NF4 referenceA CUDA GPU, model weights, and room for the resident reference as well
fsdp_sharding_probe.py8x H100, step 20 and issue #373Is the base actually sharded, or was the flag merely accepted? Each rank's local count must equal the total divided by the world size, with a single-GPU control expecting the full countThe accounting is pure Python and CPU-tested, but a real reading needs the multi-GPU box
issue395_second_stack_vram.pyA10GDoes the long-sequence under-prediction reproduce on a second stack? Three modes: the sweep, --measure and --sdpaA 24 GB card, downloads of two small models, and jinja2 3.1.0 or newer
mlx_sft_smoke.py8 GB M1, issue #23Does the MLX wrapper's training loop run on a real MLX runtime, and what does it peak at? It asserts MLX dispatch before measuring, so a transformers-path number cannot be published as an MLX oneApple Silicon and the [mlx] extra, and it downloads the chosen model

A seventh script sits beside the records rather than in harness/: bench_validator.py. It is a timing script, not a record. It compares a preserved copy of the previous validate_and_stats against the current one over 20,000 replicated fixture rows, takes the median of 15 interleaved runs, and asserts the two return identical output. The upstream README does not list it and no result is recorded in the directory, so this site quotes no figure from it.

See also

Soup is free and Apache-2.0. If it saved you a training run, starring the repo costs nothing and helps most. You can also fund the GPU time behind the work a 4 GB laptop cannot reach.