Tiered local-AI NAS · Minisforum N5 MAX · final campaign 2026-10

An AI box first.
A fast NAS second. RAIDZ1 by default.

Five NM790 NVMe drives (four on Gen4 ×1, one on ×4), 128 GB unified memory and five unbenchmarked 3.5″ SATA bays. btrfs raid5 and mdadm raid5 read fastest when healthy. RAIDZ1 is the default: it keeps 93.51 % of its reads with a drive out, against 5.87 % for btrfs raid5, and it rebuilds in 29 minutes. Once a model is resident, inference speed is the same on every pool.

Stages 1–7 sealed Protocol B · steady rails 5× NM790 4TB · SN1–SN5 21 AI cells 8 rebuild arms + 2 controls mdadm 8× · resolved (link fault)

FASTEST MEASURED (TIE)

btrfs raid5 ≈ mdadm raid5

9.034 / 9.035 GB/s seqRD · 14.43 / 14.41 TiB

Cold FS controls · stage5 · STO-01_seq_read_1M_j4 · Pareto frontier on capacity × cold read

RECOMMENDED DEFAULT

ZFS RAIDZ1 · 5-wide

14.11 TiB · 7.83 GB/s seqRD · 93.51 % degraded

Cold ZFS stage1 + stage7 STO-09 · rebuild 1,740 s @ 1.1656 GB/s

Every layout, one decision.

Cold regime (primarycache=none or dropped caches) is the layout story. Non-steady cases are hatched; PROVISIONAL cases carry ⚠. ZFS random/sync IO ran on zvols and btrfs/mdadm on files, so compare random IO within a stack only.

Interactive · dual pick + weighted score

Driven offline from data/sealed-metrics.json. Scores are normalized A↔B under your weights — not a sealed verdict. Default filter = cold only.

capacity30
seqRD35
seqWR15
rnd4kW15
sync5

—

0.0

—

0.0

Recommended pick

—

Set weights and pick two layouts.

Paired bars · A (top) / B (bottom) · normalized

mdadm 8× · resolved

The 1.14 GB/s pin was SN1 link-trained at Gen1 ×1. Gen4 reruns read 9.035 GB/s, tying btrfs raid5.

stage5 mdadm-ext4-raid5 cold · STO-01_seq_read_1M_j4

Write gap = drive state

The fs controls put btrfs raid5 at 4.066 vs RAIDZ1 0.718 GB/s, but only ZFS was preconditioned. Matched STO-09 S1: RAIDZ1 4.499, btrfs 2.78.

stage7 S1 · STO-01_seq_write_1M_j4

Lane tax is asymmetric

×4 vs ×1: seq read 3.27×, 4K random read 1.18×, sync 1.09×. Flush-bound, not link-bound.

stage1-single-x1 vs stage2-single-x4

Warm DRAM ceiling

Warm FS reads land at 35.685–36.74 GB/s on every layout. Never sell warm numbers as the array.

stage5 warm · STO-01_seq_read_1M_j4

mdadm parity RMW

mdadm raid5 has the fastest seq write (5.04 GB/s) and only 2,489 IOPS on 16K random write (btrfs raid5: 22,981).

stage5 mdadm-ext4-raid5 · STO-02_rand_write_16k_j4
Tradeoff compass · capacity × cold seq read · bubble ∝ seq write
Usable TiB vs cold sequential read for every layout
Small multiples · every cold layout, six workloads
Bars per layout for seq read, seq write, random read, random write, mixed and sync
Cold vs metadata ARC · ZFS
Slopegraph cold to metadata
Lane tax · Gen4 ×1 vs ×4
Single-drive x4 over x1 ratios
Write amplification · physical rail vs geometry
Measured write amplification per layout
Steadiness heat strip · every published case
Steady and non-steady cases

Baseline charts · from sealed-metrics.json

RAIDZ1 5-wide
14.11 TiB
seqRD 7.83 · sync 415 IOPS · degraded 93.51 %
ZFS RAID10 4×x1
6.71 TiB
seqRD 7.024 · sync 1,675 IOPS · mixed 104,485
btrfs raid5 5-wide
14.43 TiB
seqRD 9.034 · degraded 5.87 % · write hole
×4 / ×1 seqRD
3.27×
1.8 → 5.888 GB/s · sync 1.09×

Degraded reads decide the default.

Each arm was filled to 50 % with incompressible data, then one member was failed and replaced. ZFS reconstructs from parity at near-full speed and resilvers only allocated blocks. btrfs raid5/6 reads collapse to about 0.53 GB/s, whichever member fails. The S4 write shortfall is partly drive write state: a no-failure control returned 66.58 % and 67.71 % of S1 (n=2).

Degraded (S2), rebuilding (S3) and healthy-again (S4) per arm, with rebuild time
Degraded and rebuild chart per arm
armrebuildGB/sS2 seqRD %S4 seqWR %verdict
ZFS RAIDZ11,740 s1.165693.5139.32FAIL (write return)
ZFS RAIDZ1 · loaded1,550 s1.382897.1147.04FAIL (write return)
ZFS RAIDZ21,690 s1.18659.5546.95FAIL (write return)
btrfs raid102,790 s0.731671.1992.75PASS
btrfs raid53,720 s0.54985.87109.28PASS
btrfs raid5 · loaded3,480 s0.61635.8786.04FAIL (write return)
btrfs raid5 · x4 victim2,590 s0.78915.8675.45FAIL (write return)
btrfs raid66,380 s0.32075.8632.67FAIL (write return)

FAIL = S4 sequential write returned more than 10 % below S1 (the harness rule). Every arm’s reads came back (93.17–112.7 %).

Pick by workload, then by OS.

Decision graphic: AI-first RAG box, AI + VM server, media archive

The ecosystem decides the details.

Proxmox VE

RAIDZ1 (or RAID10 + RAIDZ1 split)

ZFS-native. VM disks as zvols. Run llama.cpp on the host or in an LXC with the iGPU. Check negotiated PCIe link speeds after every boot: SN1 downtrained to Gen1 twice during the campaign.

stage1-raidz1-full · 14.11 TiB · 7.83stage7 · rebuild 29 min

TrueNAS SCALE

RAIDZ1 or RAIDZ2

A ZFS appliance. RAIDZ2 for archives (two failures, 59.55 % degraded reads), RAIDZ1 for capacity. ROCm on TrueNAS was not tested.

stage1-raidz2-full · 10.21 TiB · 8.911

Fedora / Ubuntu

OpenZFS RAIDZ1 · btrfs raid1/10 only

btrfs raid5 is fast when healthy (9.034 GB/s), but degraded reads drop to 5.87 % and the data write hole is unmitigated upstream (-m raid1c3 protects metadata only). If you want btrfs, use raid1/raid10 profiles. mdadm raid5 suits bulk sequential work only.

stage7-btrfs-raid5 · degraded 5.87 %write-hole trip test · NOT YET MEASURED

21 sealed cells. Software moves tok/s more than storage.

llama.cpp llama-bench (HIP and Vulkan), models loaded from a cold btrfs raid5 pool with paired ZFS RAIDZ1 cells. Across all 16 comparisons we land at 0.7128–0.9715× of the Minisforum-published floor; the vendor states no conditions beyond ngl 99 and FA on.

Ours vs the Minisforum vendor floor
Dumbbell of ours vs vendor floor
Cold load: backend and placement beat layout
Load time pairs
Prefill to 32k · Qwen3.8-27B depth to 131k vs Gufo
Prefill curves and depth sweep
Decode retention to 55k tokens
Decode retention lines
model (HIP · btrfs cold)pp512tg128vs floor pp / tgcold load
gpt-oss-20b MXFP41,488.0275.490.9608 / 0.90645.04 s
Gemma 4 26B-A4B Q4_K_M1,130.8750.150.957 / 0.71285.035 s
gpt-oss-120b MXFP4564.0153.120.8562 / 0.897611.256 s
Qwen3.5-122B-A10B Q4_K_M308.2122.970.9646 / 0.725212.111 s
Qwen3.8-27B Q4_K_XL (dense)332.9412.59—4.635 s
Llama-2-7B Q4_01,366.2754.16—2.025 s
DeepSeek V4 Flash IQ2_XXS27.3515.26—13.65 s
Kimi K2 IQ1 · streamed (btrfs / ZFS)3.57 / 4.691.09 / 1.55—38.334 / 181.147 s

External references are quoted as “X reports Y” in the final report’s competitive section (ITPro, ServeTheHome, Level1Techs, Gufo, kyuz0, Lin/llm-tracker, MLPerf v6.1). KV-cache quantisation was not run.

How a number earns the page.

Matrix live status → matrix/ · Issues live on the private lab tracker.

What the full campaign changed.

The default is now earned, not asserted

Before stage 7, “RAIDZ1 by default” rested on capacity, ecosystem and caution. Now it rests on measurement. With a drive out, RAIDZ1 reads 93.51 % of healthy while btrfs raid5 reads 5.87 %, and RAIDZ1 resilvers at 1.1656 GB/s against 0.5498. stage7 STO-09 · zfs-raidz1 vs btrfs-raid5

The capacity-matched write race was a drive-state race

The fs controls showed btrfs raid5 writing 5.66× faster than RAIDZ1, but the ZFS cells had fully preconditioned their NM790s and the btrfs cells had not. Under one matched STO-09 protocol, RAIDZ1 wrote 4.499 GB/s and btrfs raid5 2.78. On DRAM-less, no-PLP drives, write numbers mean little without the drive state. stage5 vs stage7 S1 · STO-01_seq_write_1M_j4

The fastest pool is a tie

With the SN1 link fault removed, mdadm raid5 reads 9.035 GB/s, level with btrfs raid5 at 9.034. It also writes sequentially fastest (5.04 GB/s), but its 16K random writes collapse to 2,489 IOPS. stage5 mdadm-ext4-raid5 reruns

AI and storage meet only at load time

Resident decode is identical across pools. Cold load differs modestly for resident models (5.035 vs 6.758 s for Gemma) and sharply for streamed ones (Kimi K2 38.334 vs 181.147 s). The backend matters more than the layout: Vulkan loads gpt-oss-120b in 49.375 s against 11.256 s for HIP. stage3 AI-01 load rail

What the matrix whispers.

  1. Every 5-wide parity layout reads ~9.0 GB/s because SN5 sits on Gen4 ×4; the four ×1 members set the ceiling
    stage1/5 STO-01_seq_read_1M_j4
  2. One ×1 drive reads 1.8 GB/s (91 % of the lane); ZFS RAID10 over four reads 3.90× that
    stage1-single-x1 vs stage1-raid10-full
  3. Write amplification tracks geometry: mirrors 2.0×, 5-wide single parity 1.25×; double parity sits slightly above 5/3 on btrfs (1.764×) and mdadm (1.716×)
    STO-01_seq_write_1M_j4 physical rail
  4. A no-failure control loses a third of sequential write after the S1–S2 write volume alone (66.58 % / 67.71 %)
    stage7 control replicates
  5. The fastest decoders lose the largest share with context: gpt-oss-120b keeps 59.0 % of tg at 55k, Qwen3.5-122B 86.8 %
    stage3 depth.tg128
  6. Speculative decoding (DFlash2) lifts dense Qwen3.8-27B tg from 12.34 to 22.68 tok/s on prose
    gufo-model-bench · USB live boot · post-campaign
  7. A one-hour mixed soak stays steady at 0.886 GB/s; drives peak at 59.85 °C, the SoC at 52.0 °C
    stage4 SOAK-02 raidz1
  8. btrfs raid10 reads only 3.616 GB/s cold against mdadm raid10’s 7.223 on the same drives (unexplained)
    stage5 cold STO-01_seq_read_1M_j4

What this campaign did not measure.

Each item is tracked as a follow-up on the private lab tracker. None of the figures above stands in for these.

itemstatus
btrfs raid5/6 write-hole trip test (unclean shutdown)NOT YET MEASURED harness built, stage 8 deferred
ZFS RAIDZ1 + SLOG on the Gen4 ×4 slot (sync A/B)NOT YET MEASURED
5×3.5″ SATA bay tier (JMB58x, shared Gen3 ×2)NOT YET MEASURED
2×10GbE end-to-end client throughput (SMB / NFS)NOT YET MEASURED
KV-cache quantisation (q8_0) A/BNOT YET MEASURED every cell f16 KV
Prefill beyond 32k in the stage-3 benchPARTIAL gufo-model-bench rows reach 131k for one model
Prefill-under-decode scheduling probeNOT YET MEASURED
In-campaign USB ingest (STO-13)LOST post-campaign STO-13b: 905.6 / 376.2 MB/s r/w
File-based ZFS random IO (removes the zvol-vs-file confound)NOT YET MEASURED
Drive-state-matched seq write across ZFS / btrfs / mdadmNOT YET MEASURED
Attended rebuild arm with per-member io_ticksNOT YET MEASURED
RAIDZ1 cold mixed 70/30 (PROVISIONAL, 19× spread)RE-MEASURE
Streamed-model load on ZFS with primarycache=metadataNOT YET MEASURED
Why btrfs raid10 cold reads are lowOPEN
Per-member NM790 endurance (TBW) over the campaignNOT YET SEALED

Campaign complete. Follow-ups queued.

The full write-up, with evidence tables for every figure, is the final report (also as a PDF). The matrix keeps the cell history.