Skip to content

PTQTP — Post-Training Quantization to Trit-Planes

Design record. Data-free post-training quantization of weight matrices into K ∈ {1,2,3} ternary planes over the TQ2_0 machinery (docs/TERNARY.md). Method: arXiv:2509.16989 (Xiao et al.), implemented from the paper's formulas; K = 3 is a Fucina capacity extension beyond the paper's dual decomposition. No calibration data, no retraining, no gradients: weights in, packed trit-planes out.

The method

Each weight matrix decomposes as W ≈ Σₖ diag(αₖ)Tₖ with trit planes Tₖ ∈ {-1,0,+1} (topology) and one scale per plane per length-256 column group (magnitude). Each group is an independent problem, solved by alternating:

  1. scales — closed-form K×K ridge regression α = (SᵀS + λI)⁻¹Sᵀw (Gram entries are integer trit counts). λ escalates ×10 from lambda0 (1e-8) while the Frobenius condition estimate exceeds kappa_max (1e6), clamped at lambda_max (1.0) — required at init, where all planes start at sign(w) and the unregularized system is singular.
  2. topology — per-element exhaustive 3ᴷ-way argmin of (w − Σₖ αₖcₖ)² over trit tuples.

A group converges when its scale vector moves less than epsilon (1e-4), capped at max_iterations (50). The candidate order is pinned — zero tuple first, sparser before denser, finer planes first, strictly-less keeps-first — so exact ties prefer sparser trits, the symmetric init breaks deterministically, and the whole solve is bitwise reproducible for any thread count (rows fan out with disjoint outputs; per-row stats reduce in row order).

Deliberate deltas from the paper:

  • G = 256, packed. The paper uses G = 128 unpacked. Fucina's group size is the TQ2_0 block width, so each plane is a byte-valid standalone TQ2_0 tensor whose per-block fp16 d IS the group scale: inference is K stock ternary matmuls plus adds, no new kernels, and every plane is individually llama.cpp-dequantizable. Measured cost of the coarser groups: ~1.3% relative error (reconstructReference runs any G for such studies).
  • fp16 scale rounding at pack time. |α| is rounded to fp16 first, then one final topology pass runs against the rounded scales — stored trits are elementwise-optimal for exactly the scales inference multiplies by (measured: packed error equals the f32-scale reference to 4 decimals). Packing |α| loses nothing: the candidate set is sign-symmetric.
  • Non-finite weights are excluded from the scale regression and forced to trit 0 in every plane — one NaN degrades only itself.
  • K = 1 is a least-squares upgrade over the blind absmean b1.58 encoder; K = 3 adds a residual plane (27 levels per group, error bound ~1/3 of dual) at +2.06 bpw where applied. Planes are separable by construction: serving fewer planes than were solved is valid.

Surfaces

  • src/ptqtp.zig (fucina.ptqtp): solveGroup (pure, allocation-free), quantizeMatrix → PlanePair (owns up to three []BlockTQ2_0 planes, borrowed rhs(plane) matmul views, reconstructInto, MatrixStats — rel Frobenius error of the served reconstruction, per-plane zero fractions, iteration/convergence counts), reconstructReference (arbitrary G, f32 scales, unpacked; the fidelity-study path). Sibling tests pin: exact ternary recovery, byte parity with quantizeRowTQ2_0ScaledInto (the layout contract), error ordering K3 < K2 < K1 < absmean, RHS-view matmul equivalence, NaN benignity, determinism, all-zero packing, option/shape validation.
  • src/weights.zig: the LinearWeight union has a ptqtp: WeightPtqtp arm (up to three plane tensors). LinearWeight.toPtqtp dequantizes rows in chunks through getRowsAs — any loadable source dtype quantizes through one code path (f32, f16, bf16, K-quants, legacy, cold formats) — packs the planes, and drops the original storage. ptqtpEligible gates on the 256-block contract (in-dim % 256 == 0). getRowsAs on the arm returns the dequantized plane sum, so toResidentF16 doubles as un-decorate. Both LLM trainers' frozen-dot dispatch handles the arm (per-plane frozen dots + adds).
  • Fused inference entry (linearSeqPtqtpFused, the linearSeq fast path for the arm): one Q8_K activation quantization and ONE worker-team dispatch per decorated linear; column-partitioned tasks compute every plane and sum in the fixed plane order. Bitwise identical to a per-plane facade dot chain (pinned by test). Falls back to facade dots for gradient-tracking or non-contiguous inputs. Dispatch granularity is the decode bottleneck this design addresses: per-plane fork-joins cost more than the ternary kernel itself at 1.7B decode shapes.
  • src/models/qwen3/ptqtp.zig: qwen3.ptqtp.decorate(model, ctx, options) walks attention q/k/v (split or fused), o_proj, and dense FFN projections in place. DecorateOptions: solver (plane count etc.), skip_first_layers/skip_last_layers (edge layers stay in source precision — pure configuration, still data-free), down_planes/ o_planes (per-projection plane-count overrides for the sensitive residual-writing projections). Embeddings, lm_head, and norms are not walked; MoE FFNs are counted skipped.
  • GGUF persistence (src/ptqtp_gguf.zig): a decorated model saves as one byte-valid standalone TQ2_0 tensor per plane — <name>.ptqtp0/1/2 replaces <name>, each plane individually llama.cpp-dequantizable — plus a fucina.ptqtp.version metadata key; every other tensor and metadata entry passes through byte-verbatim. Loading pair-detects per tensor (metadata gate first, so undecorated files pay one map lookup; skip-layer tensors inside a decorated file fall through to their base) and rebuilds the .ptqtp arm bitwise. Fused in-memory weights persist under their SOURCE tensor names — the solver's per-group independence makes the fused rows' planes byte-identical to solo decoration, so they row-slice out losslessly — and re-fuse at load through fuseLinear's ptqtp arm. qwen3.ptqtp.save walks the projections decorate covers plus the output head; save→load→save is byte-stable. The format invariants — plane replacement, fused row-slicing vs solo decoration, 3-part re-fusion, resave stability, save/load validation errors — are pinned by sibling tests (ptqtp_gguf_tests.zig). Tie-fitted saves additionally stamp fucina.ptqtp.tie_scales = 1 (all-or-nothing over the file's decorated linears; absent on free-fit and legacy files, preserving their byte stability) and the loader rebuilds folded serving from it — a round-trip pin asserts the loaded weight carries the folded pack. Decoration thus runs once (--save): the saved file serves through the ordinary qwen3 runners — chat CLI, speculation, batch — with no re-decoration. Pair-detection is wired in the qwen3 loaders (dense + MoE) and the deepseek4 MoE expert trio (resident and streamed); other families do not read decorated files yet.
  • Scale-tied fit (Options.tie_scales, --tie-scales): locks the plane scales to the exact ratio 3 ([3s, s] at K=2, [9s, 3s, s] at K=3), making the K planes one uniform symmetric 3^K-level quantizer — the precondition for folding all K planes into a single dot pass (c = 3t₁+t₂, exact algebra). Measured on Qwen3-0.6B (512-token teacher-forced NLL vs the f16 baseline ppl 71.98): K=2 free ppl 210.5 vs tied 203.9; K=3 free ppl 80.7 (35.7 s fit, 10,497 unconverged groups) vs tied 75.9 (7.0 s, 0 unconverged) — the tied fit reconstructs slightly worse (uniform levels, fewer degrees of freedom: rel err .1887 vs .1784 at K=2) yet measures no downstream quality loss, fits 1.5-5x faster, and always converges. Single-eval-text caveat; a second model/eval confirmation is the gate for making it the default. At K=2 the tie is served FOLDED: WeightPtqtp.init packs both planes into one 4-bit code plane (BlockTQ2_0Foldedx4, cu = 3t₁+t₂+4 in {0..8} — same total bits as the two 2-bit packs, half the pack memory) and the fused linear runs ONE dot pass (matmulTQ2_0FoldedX4RhsTile, pinned bitwise vs its scalar reference in ternary_tests.zig and x86dot-check). Measured (M1, bench-ternary interleaved pair): 2.09-2.37x over the two-pass x4 path at every m — a K=2 tied linear costs what K=1 costs — and folded serving measures the best K=2 quality yet (ppl 195.97 on the gate eval). An in-register fold (combining crumbs per granule) was built first and measured 1.01-1.07x — the combine ops cost what the saved dots cost; pack-time folding is the real form. K=3's 27 levels exceed a nibble, so tied K=3 serves through the two-pass path. The per-plane f16 scales round independently, so folded execution derives the coarse scale from the finest in f32 (exact) — the folded pack stores only the fine scale.
  • Metal prefill offload (-Dgpu=metal): WeightPtqtp.init also copies each plane into GPU-resident bytes, and prefill-sized fused linears (m ≥ 32, work ≥ FUCINA_GPU_MIN_WORK_DENSE_TQ2, default 2^25) dispatch each plane through the ternary dequant-in-kernel mul_mm (fucina_mul_mm_tq2_0_f32), summing the K plane outputs on the CPU. Not bitwise vs the CPU chain — the same accepted numerics stance as the Q4_K/Q6_K/Q8_0 dense offload; provider parity tests pin the kernel. Measured (M1 Max, Qwen3-0.6B, pp1001, same binary, FUCINA_GPU=0 as the CPU arm): prefill 2830 → 1291 ms at K=2 with an f16 head (2.2×), and 2841 → 956 ms fully ternary with --head-planes 2 (2.97×, ~1047 prefill tok/s). Decode never dispatches (m ≥ 32 gate). Tie-fitted K=2 serves the GPU folded too: residency holds ONE buffer of row-major folded blocks (BlockTQ2_0Folded, half the bytes of two plane copies) and the fused linear issues a single tq2_0_folded dispatch (fucina_mul_mm_tq2_0_folded_f32, dedicated parity test) whose output returns async with no CPU plane-sum sync — measured pp1001: per-plane 1317 ms → folded 873 ms (1.51×), 2.29× over the same binary's CPU folded path.
  • Runtime speed path: WeightPtqtp.init packs every plane into the x4 column-interleaved form at construction (BlockTQ2_0x4, docs/TERNARY.md — same bytes rearranged, zero per-block reduces), and the fused linear runs all K planes in ONE worker-team dispatch on the x4 kernels, the accumulating twin folding each extra plane straight into the output with no scratch pass. Bitwise equal to the per-plane facade dot chain (pinned in weights_tests.zig). Measured on M1 Max (Qwen3-0.6B, paired same-window decode runs): +14% at K=2 and +30% at K=3 over the pre-pack fused path. Odd n or unreadable plane storage falls back to the row kernels — identical bits, just slower.
  • MoE expert stacks (MoeRhs.ptqtp, src/moe/expert_ffn.zig): expert stacks quantized at K=2/3 run through the fused MoE ops — the tile dot runs the ternary kernel once per plane and SUMS per element in fixed plane order before the gated nonlinearity, so a PTQTP expert equals the dense fused linear bitwise on the same weights (decode GEMV and batched prefill both; pinned against host-side per-plane K=1 sums in expert_store_tests.zig). Persistence reuses the dense convention — <name>.ptqtpK siblings with the base stack's 3D shape, plane-major on disk — and the streamed tier gathers one expert's K plane row-blocks into a contiguous cache-slab section (ProjSpec.plane_count/ plane_offsets), so K-plane expert models stream out-of-core with no new kernels. Loaders: ptqtp_gguf.maybeLoadMoeRhs / maybeStreamedMoeProjSpec, wired into the qwen3 and deepseek4 MoE load paths. Tie-fitted K=2 stacks serve folded here too: the resident loader builds an expert-major BlockTQ2_0Foldedx4 pack next to the planes and the expert dot runs the one-pass folded kernel — the same matmulTQ2_0FoldedX4RhsTile the dense fused linear uses, so the dense 2.09-2.37x kernel measurement carries over per expert dot. The streamed tier folds at fill instead (ExpertStore.readExpert): both plane reads bounce through a scratch and the 4-bit pack lands in the slab section (same section budget — 520 vs 528 bytes per 4-column group; disk bytes and the two preads per expert are unchanged, the win is the halved dot on every cache hit). Requires out_dim % 4 == 0; anything else serves the per-plane path. Both arms pinned bitwise against the direct folded kernel in expert_store_tests.zig ("tie-folded ptqtp experts"). FUCINA_PTQTP_FOLD=0 serves tie-fitted MoE plane sets through the per-plane path instead (resident AND streamed, so the tiers agree): the expert-store L2 tier stripes primary-file plane bytes and never covers fold-served layers (a fold-time slab is not primary bytes), so on a two-drive box the knob trades the halved cache-hit dot for L2 striping of every miss. Guards: l2Open drops tier coverage recorded for a layer now served folded, and an --moe-l2-build-gb build over an all-folded store refuses (L2NoStripeableLayers) instead of truncating the tier to zero coverage — both pinned in expert_store_tests.zig ("fold-mode flip").
  • Examples: zig build ptqtp-spirals (self-verifying acceptance demo: float-trains an MLP, decorates post-training, PASSes only if dual planes hold accuracy on the deployed int8 path) and zig build ptqtp-qwen3 (decorate a GGUF in place: --planes 1|2|3, --down-planes/--o-planes N, --skip-first/--skip-last N, --head-planes N to decorate the lm_head, --nll FILE teacher-forced perplexity before/after, --save FILE to persist the decorated model, greedy completion + decode timing).

Shard-streaming quantizer (zig build export-gguf -- --ptqtp)

The copy-paste walkthrough with measured end-to-end numbers (quantize → run resident → run streamed) is the Recipe section below; this section is the tool reference.

ptqtp-qwen3 --save decorates a loaded model — it needs the model in RAM and knows only the qwen3 family walk. The export tool's --ptqtp mode is the scale path: it quantizes GGUF→GGUF one source tensor at a time, so a hundreds-of-GB model (DeepSeek/GLM class) quantizes on a 64 GB machine.

zig build export-gguf -Doptimize=ReleaseFast -- \
  --from-gguf big-model-BF16.gguf --out big-model-ptqtp3.gguf --ptqtp=3
Flag Meaning
--ptqtp[=K] enable the mode; K = plane count 1–3 (default 2)
--ptqtp-planes K plane count as a separate knob (implies --ptqtp)
--ptqtp-tie scale-tied fit (Options.tie_scales; needs K ≥ 2): stamps fucina.ptqtp.tie_scales = 1 so the loaders serve tied K=2 tensors — dense linears AND MoE expert stacks — through the folded one-pass kernel
--ptqtp-include SUB[,SUB] quantize only tensors whose name contains a substring (replaces the default embeddings/head-stay name policy); repeatable
--ptqtp-exclude SUB[,SUB] never quantize matching tensors; always subtracts; repeatable
--dry-run print the per-tensor plan (name, shape, source dtype → target, bytes before/after) and exit without writing

Mechanics and policy:

  • Streaming both ways. The source is mmap-loaded (split GGUFs load via loadMmapAuto; the single-file output drops the split.* markers), and the output uses the writer's streaming path (gguf.Writer.declareTensor
  • beginStream): header first, then per tensor decode → solve → write → release. The tool never holds more than one tensor's f32 buffer plus its packed planes (reported as peak tensor working set; a 0.6B run peaks at ~14 MiB heap). Source pages get MADV.DONTNEED after each tensor — released immediately on Linux; Darwin ignores the hint for file-backed maps and evicts only under pressure, so macOS peak RSS includes clean evictable mmap pages (annotated in the summary). MoE expert stacks are the one deliberate exception to one-tensor residency: their K plane stacks accumulate in RAM for the stack's duration (see below).
  • Same on-disk format as ptqtp_gguf.zig: eligible matrices are replaced by <name>.ptqtp0..K-1 byte-valid TQ2_0 plane tensors plus the fucina.ptqtp.version stamp, so outputs load through the existing pair-detection with zero loader changes (wired in the qwen3 loaders and the deepseek4 expert trio today — see the persistence bullet above; other families read the format once their loaders adopt the same seam).
  • Eligibility: 2D matrix or 3D *_exps expert stack, name ends .weight, no norm, contract dim divisible by 256, source dtype decodable (f32/f16/bf16/legacy/K-quants; quantized sources dequantize first — the from-quantized path degrades gracefully, see above). Default name policy keeps embeddings (token_embd) and output.weight in source precision; --ptqtp-include replaces it (e.g. to decorate the head).
  • MoE expert stacks (3D) quantize per expert slice (ptqtp_gguf.quantizeMoeStack): each expert's [out x in] matrix of the expert-major [in, out, n_expert] stack runs through ptqtp.quantizeMatrix independently — group independence makes every expert row-block byte-identical to decorating that expert alone — and the K plane tensors keep the base 3D shape, plane-major (the exact MoE convention the qwen3 loaders pair-detect, resident and streamed). Memory: the K accumulating plane stacks stay resident for the whole stack plus one expert's f32 slice (~550 MiB per plane for a 4096 x 2048 x 256-expert stack, ~1.7 GiB at K=3); both the dry-run plan and the run summary report the figure, and source pages still release expert-by-expert.
  • Per-tensor solver diagnostics print as it runs (rel_err, mean iterations, unconverged groups) — the same fp16-rounded-scale reconstruction error MatrixStats measures, so a bad tensor is visible immediately, not after a full pass. Expert stacks fold their per-expert stats into one line (mean + max rel_err) instead of printing one row per expert.

Validated on Qwen3-0.6B f16: K=2 full decoration (196 matrices → 392 planes, 840→217 MiB linears) loads through the qwen3 runners and generates fluent text; a K=3 partial run (--ptqtp-include blk.0.attn) reproduces the expected error ladder (rel_err ~0.068 vs dual's ~0.18) and the mixed decorated/undecorated file serves correctly. MoE validated on Qwen3-30B-A3B Q5_K_M: quantizing one layer's three 128-expert stacks at K=2 (--ptqtp-include blk.0.ffn_gate_exps,...) peaks at a 105.8 MiB working set, and the mixed file loads through the qwen3 MoE pair-detection and generates correct text; per-expert plane bitwise identity against standalone quantizeMatrix is pinned by ptqtp_gguf_tests.

Native folded expert format (tq2_0_fx4)

Tie-fitted K=2 expert stacks can skip the sibling-plane container entirely: GGUF type tq2_0_fx4 stores the one-pass 4-bit pack (BlockTQ2_0Foldedx4, 520 bytes per 4 columns × 256 elements, 4.0625 bpw) under the base tensor name, expert-major for stacks. Compared with .ptqtp0/1 siblings this is the same served numbers (bit-identical — the pack algebra is exactly what the fold produces) in a strictly better container. No metadata stamps, no pair-detection: the tensor loads through the ordinary quant switches. Constraints: contract dim % 256, out dim % 4, K=2 tied only (27-level tied K=3 exceeds a nibble). Fucina-custom type (enum 43) — not parseable by stock GGUF tooling; the sibling-plane format remains the interchange form.

  • Streamed MoE experts (StreamedQuant.tq2_0_fx4): one pread per projection per miss instead of two, no fill-time fold, and slab bytes == file bytes, so the L2 tier stripes these layers like any plain quant — the fold-vs-striping exclusion disappears by construction. Measured on DeepSeek-V4-Flash-0731 streamed from disk (M1 Max, striped NVMe L2 tier, matched fresh tiers both arms): decode 3.11–3.12 tok/s vs 2.91–2.95 for the q4_k experts file at bit-identical output per format; 30.31 vs 33.89 GB read per 32-token round (−10.6%); 139.2 vs 153.3 GiB file bytes (−9.2%). Absolute rates vary with tier build state and disk headroom; the byte and pin advantages are invariant. The format serves identically on x86-64 (AVX2/AVX-VNNI kernel arms, bitwise-pinned against the aarch64 path; i9-13950HX/NVMe: 2.6 tok/s evicted-cache, 3.3 page-cache-warm).
  • Resident MoE experts (loadMoeRhs): the folded-only ptqtp arm — the pack borrows straight from the mapping (zero heap for experts, no load-time fold; the sibling path holds borrowed planes AND a heap fold).
  • Dense linears (WeightPtqtpFx4): pack-only weight serving bitwise the tied .ptqtp pfold path at every m; qkv/gate-up fusion is a plain pack concat (byte-identical to folding the fused matrix); Metal prefill residency comes from a pure byte relayout (packMatmulRhsTQ2_0FoldedRowsFromX4). Holds ~4.06 bpw vs the sibling arm's planes + x4 repacks + fold (~12.3 bpw) — ~3× less RAM and no pack-building at load. Not a training arm: autograd/frozen-training walks reject it (sibling planes remain the trainable form).

Producers: export-gguf --ptqtp-native quantizes any model straight to fx4 (tied K=2 implied; tensors with out % 4 != 0 — none in the qwen3 family — fall back to sibling planes, and only then do fucina.ptqtp.* stamps appear); convert-ds4-fp4 --repack-native transforms an existing tied sibling-plane MoE GGUF without re-solving. Oracle: sibling-tied and native exports of the same source produce bit-equal NLL and byte-identical greedy completions (the solver is bitwise-deterministic and both serve the same pack bytes).

DeepSeek-V4 fp4-expert converter (zig build convert-ds4-fp4)

The DeepSeek-V4-Flash checkpoint releases its routed experts in MXFP4 (fp4 e2m1 codes packed two per byte, one e8m0 power-of-two scale per 32 elements along K) — export-gguf --ptqtp cannot see them (safetensors, no GGUF dtype). tools/convert_ds4_fp4.zig closes the gap for the deepseek4 family: it dequantizes each expert exactly (both factors are dyadic, so fp4 × 2^k is lossless in f32 — a single-hop requant of the released weights, unlike requantizing an intermediate GGUF), solves tied-K2 planes per expert via ptqtp.quantizeMatrix, and streams a single-file GGUF whose metadata and every non-expert tensor pass through byte-verbatim from a trunk GGUF (identical attention/router/shared-expert bytes — a clean A/B against that file).

zig build convert-ds4-fp4 -Doptimize=ReleaseFast -- \
  --src-dir deepseek-v4-fp8/ --trunk-gguf TRUNK.gguf --out OUT.gguf [--planes K] [--no-tie]

--smoke N cross-checks layer N's fp4 dequant against the trunk's Q4_K experts (both derive from the same release, so cos ≈ 0.9985 validates the name map, nibble order, scale layout, and expert-major slicing in one shot); --verify N re-solves sample experts post-conversion and byte-compares against the stored planes (the solver is bitwise- deterministic, so equality is exact). The tool refuses already-decorated trunks and hard-errors on 0xFF (NaN) e8m0 scale bytes. Measured on the 284B-A13B 0731 checkpoint: 129 stacks × 256 experts in ~64 min on an M1 Max, 157.0 → 141.2 GiB, per-expert rel err ≈ 0.18, 0 unconverged groups.

Configuration guidance

  • Source precision: quantize from bf16/f16 originals when available. The from-quantized path (e.g. a Q4_K_M file) works through the same code and degrades gracefully (~15–18% worse ppl), never collapses. f16 vs f32 is immaterial (source rounding sits orders below the ternary error floor).
  • Plane count: K=3 for accuracy (parity-class, see below); K=2 only with selective K=3 overrides on down_proj/o_proj — model sensitivity to the K=2 error grows sharply with model size (0.6B tolerates it at ×3.2 ppl; 1.7B collapses ×17) and concentrates in the residual-writing projections, with a tail spread over q/k/v/gate/up.
  • lm_head (--head-planes): quality-free at K=3 (measured Δppl ≈ 0), and with the x4 column-interleaved fused path (docs/TERNARY.md) the ternary head now also WINS on ARM: paired same-thermal-window decode runs on M1 Max (Qwen3-0.6B) measure the K=2 head ~7-8% faster end-to-end than the f16 head at both an f16 body and a K=2 body — the pre-x4 verdict ("keep the bf16 head on ARM; the ternary GEMV is ALU-bound") is overturned. Decorate the head for speed and memory alike (fully-ternary 1.7B = 1.20 GiB of weights at baseline-parity ppl); x86-VNNI already favored it via the ~2× per-plane kernel margin.
  • Edge-layer skip: subsumed by K=3 (buys ~nothing on top); useful as a cheap quality lever for K=2-budget deployments (first layers matter more than last).

Recipe: quantize a model and run it (resident or streamed)

The end-to-end walkthrough: take a GGUF, quantize its weight matrices — including every routed MoE expert — into PTQTP trit planes, and run the result with the ordinary runners, resident or streamed from disk. Every number below was measured on this exact sequence (M1 Max 64 GB, source model on an external USB-3 SSD, 2026-07-12). The method itself, the on-disk format, and the full quality ladder are the rest of this document.

Pick the source

Any GGUF works as input. For best quality use the highest-precision release available (f16 / bf16 / q8_0): quantizing from an already-K-quantized file (as in the example below) stacks its error on top of PTQTP's. Quantized sources are dequantized tensor-by-tensor first — deliberate, and documented as graceful degradation under Configuration guidance above.

Preview the plan (optional, fast)

zig build export-gguf -Doptimize=ReleaseFast -- \
  --from-gguf models/Qwen3-30B-A3B-Instruct-2507-Q5_K_M.gguf \
  --ptqtp=2 --ptqtp-include ffn_gate_exps,ffn_up_exps,ffn_down_exps \
  --dry-run

Prints one line per tensor (shape, source dtype → target, bytes before/after), the totals, and the peak buffering estimate. No output file is written.

Quantize

zig build export-gguf -Doptimize=ReleaseFast -- \
  --from-gguf models/Qwen3-30B-A3B-Instruct-2507-Q5_K_M.gguf \
  --out       models/Qwen3-30B-A3B-ptqtp2-experts.gguf \
  --ptqtp=2 --ptqtp-include ffn_gate_exps,ffn_up_exps,ffn_down_exps

Knobs:

  • --ptqtp=K — trit planes per matrix. K=1 ≈ 2.06 bpw (maximum compression), K=2 ≈ 4.1 bpw (the sweet spot; reconstruction rel_err ≈ 0.17), K=3 ≈ 6.2 bpw (near-parity; rel_err ≈ 0.067). §Measured below has the NLL ladder.
  • --ptqtp-tie — scale-tied fit (K ≥ 2): plane scales locked to the exact ratio 3, making the K planes one uniform 3^K-level quantizer. Fits 1.5-5× faster, always converges, and measures better downstream quality than the free fit (Qwen3-1.7B, 512-token NLL vs BF16 ppl 56.15: K=2 ppl 132.8 tied vs 261.7 free; K=3 58.1 vs 61.4 — §Scale-tied below has the ladder). At K=2 the stamped file additionally serves FOLDED everywhere — dense linears, resident expert stacks, and streamed expert fills alike run the one-pass 4-bit folded kernel (2.09-2.37× the two-pass dot), so tied K=2 is the recommended shape for streamed MoE experts.
  • --ptqtp-include SUB[,SUB] — name-substring filter. The value shown quantizes only the routed expert stacks; each expert's matrix is quantized independently and persists as plane-major <name>.ptqtp0/1/… siblings with the base 3D shape. Omit the flag to also decorate attention/dense/shared-expert matrices under the default policy (embeddings, norms, router, and the output head always stay in source precision). --ptqtp-exclude subtracts from either.
  • Re-running on an already-decorated file quantizes nothing (idempotent).

Cost, measured on the 30B (48 layers × 3 stacks × 128 experts = the full ~28 B-parameter expert mass): 8.5 minutes wall, 105.8 MiB peak tensor working set — the tool streams one tensor at a time, so source size does not matter: a 300 GB model quantizes on a 64 GB machine. Every expert converged (0 unconverged trit groups out of 786 432 per stack; per-stack rel_err mean/max is printed as it goes).

Result for this example: file 20.7 → 15.3 GiB (quantized linears 19.6 → 14.3 GiB). From an f16 source the expert mass shrinks ~4× at K=2.

Run it

The output is a normal GGUF; the family runner picks up the plane tensors automatically (pair-detection at load — nothing to configure).

# resident
zig build qwen3 -Doptimize=ReleaseFast -- models/Qwen3-30B-A3B-ptqtp2-experts.gguf \
  --prompt "The capital of France is" --gen 32

# streamed: only dense weights stay resident; routed experts page from
# disk through the pinned-set + LRU tier under the given budget
zig build qwen3 -Doptimize=ReleaseFast -- models/Qwen3-30B-A3B-ptqtp2-experts.gguf \
  --prompt "The capital of France is" --gen 32 --moe-stream --moe-cache-mb=6144

Measured on the file produced above, both modes: byte-identical output, coherent and factual ("… Paris. What is the capital of Spain? The capital of Spain is Madrid. Madrid is the largest city in Spain and serves as the country's political, economic …"). The streamed run hit 80.1 % in the expert cache at a 6 GiB budget (6.94 GB read for prefill + 32 tokens); all --moe-stream companions (--moe-pin-mb, --moe-pilot, the .experts learning-cache sidecar, --kv-save) apply unchanged — see RUNNING-MODELS.md §Streaming.

Recipe notes

  • Bit-exactness contract: a PTQTP expert computes the same sum-of-plane-dots the dense PTQTP linear does (tie-fitted K=2 stacks: the same one-pass folded dot), and the streamed path is bit-identical to resident (both pinned by tests in src/store/expert_store_tests.zig).
  • Mixed files are fine: quantize one layer, a subset of projections, or everything; decorated and undecorated tensors serve side by side.
  • K per use case: K=1 for maximum-compression experiments, K=2 as the daily-driver size/quality point, K=3 when the goal is parity with the source at ~1.9× decode speed (see TERNARY.md for the kernel story). With --ptqtp-tie: K=2 tied is the speed shape (folded one-pass serving, streamed or resident), K=3 tied the quality shape (near-parity, two-pass serving — 27 levels exceed the nibble).

Measured (M1 Max, ReleaseFast; NLL = teacher-forced over 512 held-out tokens)

Two-spirals MLP (2→256→256→2 tanh, float-trained to 1.000, w2/w3 decorated post-training):

variant acc (exact-f32 path) acc (deployed int8 path) rel err w2
float 1.000 (CE 0.0042) — —
absmean b1.58, 1 plane 0.670 0.670 —
PTQTP K=1 0.644 0.644 0.474
PTQTP K=2 1.000 (CE 0.0127) 1.000 (CE 0.0126) 0.190

Single ternary planes collapse post-hoc; the dual decomposition holds full accuracy on the deployed int8 path.

Perplexity (all data-free; decoration takes seconds at 0.6B, ~90 s at 1.7B K=3, multi-threaded):

model / source baseline K=2 K=2 +down3+o3 K=3
0.6B f16 24.96 80.47 51.69 27.36
1.7B Q4_K_M 19.41 330.96 — 21.71
1.7B BF16 18.57 184.45 45.70 18.43

K=3 from the bf16 original matches the full-precision baseline at 1.7B (statistical parity; the greedy completion is flawless) and lands within 10% at 0.6B, with weight-space rel err 0.067 on the 27-level bound (~1/3 of dual's 0.179). Solver diagnostics at these scales: mean ~17 iterations, unconverged groups ≤ 0.2%. 0.6B data-free edge-skip curve (K=2 base): 1/1 → 69.0, 2/2 → 65.0, 3/3 → 51.6, 4/4 → 46.6.

Decode speed (1.7B vs its BF16 original, interleaved runs; fused entry):

config t/s vs baseline ppl
bf16 baseline 20.2–21.0 1× 18.57
K=2 47.1–48.3 2.3× 184
K=3 39.5–39.8 1.9× 18.43
K=3 + K=3 head (fully ternary) 33.9 1.65× 18.40
K=2 +down3+o3 49.3 2.4× 45.7

Against a Q4_K_M baseline (~44–48 t/s at 1.7B) K=2 is speed-parity and K=3 ~0.75×; at 0.6B K=2 is ~1.9× the f16 source and parity with Q4_K_M. Weights: 1.7B linears 3.2 GiB bf16 → 693 MiB (K=2) / 1040 MiB (K=3). Kernel truth is zig build bench-ternary: per plane the TQ2_0 kernel is ~2.1× Q4_K on ARM and ~4.8× on x86-VNNI, so multi-plane configs win outright on x86-class VNNI hardware. The ARM kernel is verified at ~86% of the NEON ALU roofline (LLVM emits the fully-folded 10-ops-per-64-weights sequence; hand asm has ≤14% headroom), and the ARM↔x86 per-instruction gap is instruction-set density (sdot 16 weights/instr vs vpdpbusd 32; i8mm smmla would close it on ARMv8.6+ targets).

Limits and future work

  • Measured dead ends, recorded so they are not re-chased: sparse exact outlier carry cannot rescue K=2 (carrying the top 1.56% |w| per group exactly improves rel err only 0.184→0.146 vs K=3's 0.069 — the K=2 gap is bulk 9-level resolution, not tails); per-128 packed scales (~1.3% fidelity, needs a block-layout change); E-core threads (even column splits go straggler-bound on heterogeneous cores); custom NEON asm (≤14% under the verified roofline).
  • Decode dispatch: ~140 fork-joins/token remain at 1.7B (linears + attention
  • norms), worth ~3–5 ms against a ~64 t/s ARM ceiling at K=3. The dependency-respecting fix is a per-layer phase chain (thread.parallelChained, moe/chain.zig precedent, family-local per the placement policy) — a second decode path with a parity burden; cheaper spin/dispatch-cost tuning should be measured first.
  • Per-row selective third plane (plane3 over a chosen row subset + row index, selection data-free) would shape the K=2↔K=3 frontier finer than the per-projection overrides; unbuilt.
  • The generic facade tq2_0 dot (non-PTQTP consumers) still pays one fork-join per call; the PTQTP path bypasses it via the fused entry.
  • MoE expert stacks SERVE from persisted planes (resident and streamed, see Surfaces) but no in-tree walker decorates them yet — producing a K-plane expert GGUF takes an external converter running ptqtp.quantizeMatrix per expert and writing the <name>.ptqtpK siblings.
  • Calibration-based extensions (activation-weighted solve, mean-correction bias, sample repair, self-calibration) live on the feat/ptqtp-repair branch, deliberately out of the data-free mainline.
  • PTQTP-as-init for STE fine-tuning (dotTernarySte) is unexplored.

Provenance

Method: arXiv:2509.16989v3 (PTQTP), reimplemented from Algorithm 1 / Eq. 3–10; no reference code existed (the paper's repository link was empty at porting time). Substrate: the TQ2_0 kernels, encoders, and block layout of docs/TERNARY.md (ggml/llama.cpp lineage — MIT, see docs/THIRD-PARTY-NOTICES.md).