Fucina Zig Architecture¶
This document describes the current Zig implementation in this tree (src/,
examples/, apps/, bench/, tools/). It is derived from the actual
source layout and behavior, not from historical design notes. Structure lives
here; the command cheat sheet lives in AGENTS.md, and per-model recipes live
in the per-target examples/<name>/README.md and apps/<name>/README.md
guides indexed by RUNNING-MODELS.md. Last reconciled against the tree:
2026-08-24.
Status¶
Fucina is an eager, close-to-metal CPU tensor/autograd runtime plus an LLM/ASR
inference stack, written in Zig 0.16. It is multi-dtype: bool/integer scalar
dtypes, f16/bf16/f32/f64, and the GGML block-quantized formats (see
src/dtype.zig). The public API is the tagged autograd Tensor facade
exposed from src/fucina.zig. Model families (Qwen3 dense + MoE, Qwen3.5,
Gemma 4, DiffusionGemma, DeepSeek V2 + V4 Flash, GLM-4.5, Inkling, Parakeet
ASR, plus the OmniVoice TTS, LocateAnything VLM, and NAM ports in
apps/) run from GGUF weights through the sibling fucina_models module
(src/models.zig) and the runners in examples/ and apps/. Execution is CPU-first with optional
Metal/CUDA callable-accelerator offload (-Dgpu=metal|cuda,
src/backend/{metal,cuda}.zig).
The core architecture is internally coherent: production dependencies are
acyclic and direction-banded (machine-enforced — see Layering And
Enforcement), execution is eager and explicit, buffers are owned
deterministically, and the autograd surface is unified around public Tensor.
What it is not: a general ML framework product with a stable, versioned
external API (see Current Production Gaps).
Layer Stack¶
Top-down; a band may depend only on bands at or below it:
| Band | Contents |
|---|---|
| apps | examples/**, apps/**, tools/**, bench/**, src/bench_raw.zig, src/x86dot_check.zig |
| serving | src/serving.zig, src/serving/** (the fucina_serving module) |
| models | src/models.zig, src/models/** (the fucina_models module) |
| facade | src/fucina.zig (the fucina module root) |
| ag + training/serialization | src/ag.zig, src/ag/**, src/optim.zig, src/optim/**, src/es.zig, src/es/**, src/ptqtp.zig, src/gguf.zig, src/gguf/**, src/lora.zig, src/rnn.zig, src/safetensors.zig, src/state_dict.zig, src/training_checkpoint.zig, src/param_registry.zig, src/weights.zig, src/weights/**, src/gguf_meta.zig, src/ptqtp_gguf.zig (model I/O) |
| tagged | src/tag_ops.zig (tag-ops library) |
| moe | src/moe.zig, src/moe/** (the MoE band: the MoeRhs expert-stack container, the decode and batched-prefill expert FFN engines, the phase-chain scheduling every family shares, the decode-scratch views family engines carve) |
| exec | src/exec.zig, src/exec/** (eager runtime) |
| store | src/store/** (disk-streamed block stores: the out-of-core MoE expert tier — expert_store.zig the facade over the io/geometry/tiers/policy/persist concern files) |
| backend | src/backend.zig, src/backend/** (the one CPU kernel provider behind the conformance-checked backend.kernels set; the single-implementation fused kernels live beside their ops in exec/) |
| tags | src/tags.zig (comptime tag algebra) |
| tensor | src/tensor.zig (raw tensor) |
| primitives | src/thread.zig, src/parallel.zig, src/tuning.zig (the tuning table over env's readers) |
| core | src/dtype.zig, src/shape.zig, src/storage.zig, src/accelerator.zig, src/rng.zig, and the std-only leaves src/fpenv.zig, src/caching_allocator.zig, src/streamconv.zig, src/cpu_topology.zig, src/env.zig |
Public Surface¶
src/fucina.zig exports:
Tensor: the public tagged/autograd tensor constructor fromsrc/ag.zig.ExecContext(plusRhsLifetime, the MoE/RouterTopKOptions/Reduction/CrossEntropyOptions/StandardizeOptionsoption types,UnaryOp, andRopeMode/RopeTable) fromsrc/exec.zig, plus theexpert_storedisk-streamed MoE expert tier (src/store/expert_store.zig) and thefakequantlow-precision grid round-trips.DTypeand the GGML block types (BlockQ4_K,BlockIQ2_XS, ...) fromsrc/dtype.zig, plus the quantized RHS container types from the backend.BackendKindand the backend build/runtime constants (active_backend_kind,native_uses_blas, ...).- Autograd framework pillars:
checkpoint/checkpointWithContext,noGrad/isGradEnabled/NoGradScope,customVjp, andgradcheck/GradcheckOptions/GradcheckResult. einsumMany: N-ary multi-index contraction (comptime left-fold of the binaryTensor.einsum).- Training/persistence namespaces:
optim,es,lora,gguf,rng,parallel,ParamRegistry,state_dict,safetensors,training_checkpoint, plus theptqtptrit-plane PTQ namespace. - Layers over the facade:
rnn(the LSTM cell and stack on theTensor.lstmsequence op, for training and streaming).
The root intentionally does not export the raw tensor or raw autograd
internals. A comptime guard in src/fucina.zig makes re-exporting RawTensor
at the public root a compile error; in-tree code that genuinely needs the raw
type names it through the fucina.internal escape hatch
(internal.RawTensor, plus internal.backend_mod/tensor_mod/thread_mod
for exact type identity in the fucina_models module, and internal.gpu for the
Metal residency/tracing hooks). Microbenchmarks use the separate bench_raw
module (src/bench_raw.zig).
Tensor(tags_or_rank) is the user-facing tensor type. It supports named tags
(Tensor(.{ .batch, .hidden })), a numeric rank (Tensor(2), generating axis
tags ._0, ._1, ...), and dtype specs such as
Tensor(.{ .dtype = .u16, .tags = .{ .batch, .seq } }) or
Tensor(.{ .dtype = .i64, .rank = 2 }). Rank is part of the public tensor
type at comptime; dimension sizes remain runtime values. DType is the public
logical format tag, not a promise that every tensor has one scalar storage
element per logical element: scalar dtypes store []Scalar(dtype);
block-quantized dtypes store []Storage(dtype) blocks over the last logical
axis.
The .f32 public tensor branch is the differentiable autograd tensor.
Non-f32 scalar public tensors are typed tensors: storage, tags, views,
broadcasting, gather, narrow/slice, concat, and slice/row updates. Floating
non-f32 tensors also expose forward-only math (typed ops never record);
the 16-bit floats (f16, bf16) are additionally trainable leaves: a
variable of either dtype carries an f32 gradient slot, and gradients reach
it through the recorded to(.f32) cast (dtype.supportsGrad, the one
operation/dtype contract: docs/reference/03).
f64, integer and bool tensors are constants; integer and bool tensors do
not expose float math. Block-quantized public tensors are constant
inference tensors: loaded-block construction, to(.f32), embedding-style
getRows, and f32 x quantized-RHS tagged dot/matmul when the RHS is stored as
[free, contract] — no generic pointwise math, softmax, norms, or autograd.
These boundaries are enforced by the Zig type system through separate public
tensor branches, not by runtime dtype checks. A scalar-tag tensor,
Tensor(.{}), is represented internally as a rank-1 raw tensor with shape
{1} because the raw tensor layer has no rank-0 shape.
Forward float dtype policy is explicit per operation family
(computeDType/outputDType in src/dtype.zig):
- Pointwise ops preserve input/output dtype.
bf16computes throughf32because it is stored as bits;f16computes asf16;f64asf64. - Reductions on
f16/bf16compute inf32and returnf32;f64reductions compute and returnf64. - Dot/matmul on
f16/bf16/f32accumulates inf32and returns the input dtype;f64matmul computes and returnsf64. - Explicit casts are required when a caller wants a different output dtype.
Source Layout¶
Core value types and substrate:
src/dtype.zig: comptime dtype metadata, scalar storage mapping, block-quantized storage block definitions, float compute/output policy.src/storage.zig: refcounted typed buffer storage (BufferOf(dtype)), including borrowed-slice storage with an optional release hook (fromBorrowedSliceWithRelease) used for device-resident weight bytes; storage also owns optional submitted-writer/latest-reader accelerator fences and storage-lifetime mapping resources.src/accelerator.zig: backend-neutral lifetime tokens for already-submitted eager GPU work (Work) and per-storage mapping caches (Resource). They contain no operation description or compute graph.src/shape.zig: shape and stride arithmetic, the one home for it: theShapeinline-dims value and itsfromnormalization of every shape spelling, element counting (logical and block-storage), contiguous strides,dispatchRank, and the(outer, axis_dim, inner)axis geometry the axis-wise kernels share.src/tensor.zig: raw tensor value (TensorOf(dtype)) overshape.zig's metadata: views, broadcast, reshape, materialization, fixed-rank views.src/tags.zig: comptime tag/rank algebra (no runtime representation).src/rng.zig: repo-owned deterministic RNG; the (seed → values) mapping is a checkpoint contract (APOLLO projections, dropout masks).src/parallel.zig: the comptime dispatch thresholds and the worker-team size, over two std-only leaves it re-exports:src/cpu_topology.zig(parallel.topology: physical/performance cores, the cgroup budget, the schedulable count) andsrc/env.zig(parallel.env: the sanctionedFUCINA_*readers, libc and libc-free Linux arms).src/thread.zig: thread pool. The worker team stays hot between dispatches (spin-then-park); the dependency-chained fork-join mode carries a documented exactly-once/exact-count enqueue contract with safety-build-only accounting (duplicate detection, stall diagnostics — comptime-elided in ReleaseFast).
Execution runtime:
src/exec.zig:ExecContext, the public runtime boundary. One struct embeds the runtime substrate asrt: Runtime(thread-safe allocator, published worker team,BufferPool, exec-scope stack, tuning, fp env), declares the model/session execution state beside it (kernel pinning, MoE decode scratch), and carries every op as an alias line (pub const add = exec_elementwise.add;) grouped by domain; the struct body is the registry, the bodies live undersrc/exec/.src/exec/runtime.zig: theRuntimestruct and the substrate functions ofExecContext(lifecycle, exec scopes, worker team, tensor allocation primitives,replace), each taking*ExecContextfirst. Domain modules receive*ExecContextexplicitly (neverself: anytype), so their code is monomorphic;src/exec.zigandsrc/exec/*.zigform the one permitted root-anchored import cycle (see Layering And Enforcement).src/exec/buffer_pool.zig: the reusable transient-buffer pool leaf.src/exec/domain modules:attention.zig,matmul.zig,quant_matmul.zig,fakequant.zig(FP8/FP4/f16 grid round-trips),elementwise.zig,norm.zig,softmax.zig,loss.zig,reduce.zig,topk.zig,stats.zig,gather_scatter.zig,rope.zig,convert.zig,conv.zig,pool.zig. These are not public API;src/exec.zigremains the runtime boundary.src/moe/chain.zig(the moe band above exec): shared batched-MoE scheduling scaffolding (expert-grouped route plan, gather → gate/up → act → down phase-chain machinery, chunking helpers, profile timers). Consumed bymoe/expert_ffn.zigand, asfucina.moe.chain, by the gemma fused gate|up kernels (models/gemma/moe_gu.zig), so scheduler fixes land once for every family.
Backends:
src/backend.zig: backend facade;backend.kernelsis the one provider's kernel set and the shared kernel vocabulary (ops, block and RHS types) is re-exported from here.-Dbackend=scalaris not a provider swap: it setsbackend/isa.zig'sreferenceflag and the entries select their scalar arms internally.- The kernel interface is the declaration list of
native.zig'skernelsnamespace itself;backend.zig'sconformKernelsreads thepc-first/pool-free contract from the signatures alone. src/backend/vector/rows.zigandsrc/backend/vector/attention.zigare the two kernel seams: each is the root file of a directory holding the Task payloads, factories, adapters and helpers the exec domain modules take asbackend.rows/backend.attention, over a<seam>/kernels.zigchild that holds the bodies (imported privately by the root, reached by everyone else throughkernels).src/backend/offload.zig: the accelerator seam, the one module above the providers that namesgpu_impl. Capability queries, resident storage, tracing, and the offload entries with their decisions built in (quantized GEMM, attention, grouped MoE, the ES flat kernels); every band above the backend calls it and carries no GPU branch of its own, and on-Dgpu=noneeach entry folds to its refusal at comptime.src/backend/native.zig: the one CPU kernel provider (Zig@Vectorkernels plus optional CBLAS for GEMM), exportingpub const kernels; each entry carries its scalar reference arm (thescalarnamespaces inbackend/vector/), andsrc/backend/parity_test.zigkeeps entry and reference arm in numeric agreement.src/backend/blas.zig: the CBLAS provider seam (the onecblas_sgemmextern behindgemm/gemmStrided, the vendor thread setters, the MKL nested scope);backend.blasre-exports it on BLAS-backed builds.src/backend/vector.zig+src/backend/vector/: portable SIMD kernels, addressed by child module (vector.gemm.gemm):primitives.zig,gemm.zig,gemm_blocked.zig— the BLIS-style blocked packed f32 GEMM for the no-BLAS path,gemm_packed.zig,matmul_quant.zig,elementwise.zig,conv.zig,pool.zig— channel-last pool2d/upsample2x,winograd.zig— F(2×2,3×3) conv transforms for the no-BLAS conv route,rows.zig— the fused row kernels (softmax/logsumexp, layer/RMS norm, cross-entropy, dropout, scatter-add, gated activations, the inner-lane strided-axis family) with their Task payloads,attention.zig— the grouped-causal attention kernels (per-query units, query-tiled online-softmax prefill, tiled backward + BLAS strips, multi-stream decode) with their Task payloads and adapters,batched.zig;common.zigholdsParallelConfig, the vector-width aliases and the thread-count gates;tile.zigis the payload-generic range splitter (forRange/reduceRange) behind the kernels' parallel dispatch. Every pool-taking kernel takespcfirst.src/backend/quant.zig+src/backend/quant/: GGML-compatible block helpers, dequantization, quantized-RHS containers and dot kernels, addressed by child module (quant.q4_k.matmulQ4_Kx8RhsTile):q4_k.zig,q5_k.zig,q6_k.zig,q8_0.zig,q8k.zig,ternary.zig,mxfp4.zig,cold.zigfor the rare formats,types.zig(the block and RHS types;quant.zigforwards these outward),common.zig(the shared SIMD primitives). The f32 → quantized row encoders live in the format modules behindquantizeRowForDTypeinquant.zig(byte-exact ggml parity).src/backend/packed.zig(the dense f32 output-row panel; f16/bf16 weights widen once at pack time),src/backend/ops.zig(shared op enums),src/backend/quant_tables.zig(GGML lookup tables).src/backend/metal.zig+src/backend/metal/: the-Dgpu=metalGPU GEMM provider — Zig host (lazy init, persistent queue, eager-async f32/f16/dense-quant completion, work-threshold gates, device-owned weight storage, storage-lifetime page wrappers) plus the ObjC shim (shim.m) and vendored kernels (mlx_gemm.metalf32/f16,ggml_mul_mm.metaldequant-in-kernel).src/backend/cuda.zig+src/backend/cuda/: the Linux/NVIDIA provider — dlopen'd driver/cuBLAS, persistent upload/compute/download streams, a bounded reusable in-flight slot pool and storage-lifetime host registration for eager-async f32/f16/dense quant, managed weight residency, and vendored PTX quant/GEMV/attention kernels.src/x86dot_check.zig: standalone cross-ISA parity checker for the int8 dot primitives + Q4_K/Q8_0 dot kernels (per-arm coverage table in its header).
Autograd:
src/tag_ops.zig: tag-semantics op library over raw tensors (see Tagged Tensor Semantics).src/ag.zig: autograd module root, exporting the publicTensorand the framework pillars.src/tuning.zig: tuning policy — the typed table of every FUCINA_* route gate and numeric crossover (read-once env load, measured defaults, programmatic pins), and the per-contextOverridescarried byExecContext(setTuning).src/ag/tensor.zig: public tagged/autograd tensor facade — theTensordispatcher (normalizes the spec, then instantiates one of four branches: f32, typed float, typed scalar, block-quantized) and, per branch, one alias line per method onto the shared mixins insrc/ag/tensor/. Which branch aliases what is one positive capability table (Caps/caps): the alias lines are grouped by capability, each group is guarded against the table (requireCap), and a comptime audit (auditMixin) requires every mixin decl to be aliased or covered by a decl-group the dtype's row does not grant (the group'swhydocuments the absence); only thegrad_slotdtypes (f32/f16/bf16) carry the live?*GradStatefield — on f64 it isvoid. The mixins:common.zig(lifetime, raw access, tag/shape queries, every branch),views.zig(the dtype-generic views and data movement, every scalar dtype; differentiable on f32),elementwise.zig(the dtype-generic pointwise family: f32 differentiable, f16/bf16 constants through thewidenedpolicy, the integer arithmetic subset),autograd.zig(leaves, gradients, backward; f32 and the 16-bit leaves), the per-domain mixins insrc/ag/tensor/float/(matmul, reduce, softmax, stats, shape, norm, ...; f32 differentiable, and the ops whose exec entry takes a dtype are aliased on the 16-bit branch as constants), and the typed-branch mixins insrc/ag/tensor/typed/(creation.zig,math.zigfor the typed reductions, casts and the typeddot,int.zig);src/ag/tensor/plumbing.zig: shared result-finishing and dispatch helpers. Thesrc/ag/tensor/files never import the facade back — they receive it as a comptime parameter (Self.ag_root/Mod(ag_tensor)), keeping the import graph acyclic.src/ag/backward/: the per-domain VJP modules (mirroringsrc/exec/'s taxonomy, over a sharedcommon.zig); each facade mixin imports its own domain file directly, andsrc/ag/backward.zigis the test-forwarding root only;src/ag/core.zig: backward-only gradient state and scheduling engine;src/ag/checkpoint.zig: activation checkpointing (recompute-in-backward);src/ag/control.zig: no-grad scopes;src/ag/custom.zig: thecustomVjpadapter;src/ag/gradcheck.zig: the finite-difference gradient oracle.
Training and persistence (see Training And Persistence): src/optim.zig,
src/es.zig, src/ptqtp.zig, src/param_registry.zig, src/state_dict.zig,
src/safetensors.zig, src/training_checkpoint.zig, src/lora.zig,
src/gguf.zig. Layers over the facade: src/rnn.zig (the LSTM cell and
stack on the Tensor.lstm sequence op, one call per layer and block for
training and streaming).
LLM stack (see LLM Stack): src/models.zig + src/models/.
Apps band: examples/ (single-file teaching programs, one main.zig per
directory: smoke/, spirals/, es_spirals/, gemma4/, qwen35/,
engram/, and more), apps/ (product- and port-shaped programs with their
own tests, shims, and goldens: run/ (the registry runner), qwen3/,
deepseek4/, diffusion_gemma/, lmserve/, parakeet/, omnivoice/,
voiceagent/, facedetect/, locate_anything/, nanochat/,
finetune/, es_finetune/, cartridge/, cartridge_fleet/), tools/
(export_gguf.zig, check_import_graph.zig, check_doc_links.zig, plus the
benchmark/parity helper scripts), bench/ (microbenchmarks plus the shared
alloc.zig/timer.zig helpers).
Layering And Enforcement¶
The intended production dependency direction inside the fucina module:
fucina.zig
-> ag.zig, exec.zig, backend.zig, tag_ops.zig, tensor.zig, storage.zig,
dtype.zig, thread.zig, and the training/persistence + model-I/O
modules (gguf, weights, gguf_meta, ptqtp_gguf, optim, es, ptqtp,
lora, rng, parallel, tuning, param_registry, state_dict,
safetensors, training_checkpoint)
ag/tensor.zig
-> ag/tensor/ (common/views/autograd mixins, float/ and typed/ method
mixins, plumbing.zig),
ag/{core,backward,control}.zig, tags.zig, tag_ops.zig, exec.zig,
backend.zig, tensor.zig, dtype.zig
ag/tensor/*.zig, ag/tensor/float/*.zig
-> ag/{core,backward,control,elemental}.zig, tags.zig, tag_ops.zig,
exec.zig, backend.zig, tensor.zig, dtype.zig, rng.zig
(never ag/tensor.zig — the facade arrives as a comptime parameter)
ag/tensor/**, ag/elemental.zig
-> ag/backward/<domain>.zig (per-domain VJP modules + common.zig; the
ag/backward.zig root only forwards their tests)
ag/backward/*.zig
-> ag/core.zig, tags.zig, tag_ops.zig, exec.zig, backend.zig (ops),
tensor.zig, dtype.zig, parallel.zig
tag_ops.zig
-> exec.zig, tags.zig, tensor.zig, shape.zig
exec.zig
-> exec/*.zig, backend.zig, tensor.zig, thread.zig, tuning.zig, fpenv.zig
exec/runtime.zig, exec/<domain>.zig
-> exec.zig (the `ExecContext` type), exec/buffer_pool.zig, backend.zig,
dtype.zig, parallel.zig, shape.zig, storage.zig, tensor.zig, thread.zig
(the root-anchored cycle with exec.zig: struct body in the root,
function bodies in the children)
backend.zig
-> backend/{ops,packed,quant,cpu,native,gpu,vector}.zig, dtype.zig,
tensor.zig, thread.zig
backend/gpu.zig (comptime -Dgpu selector, a leaf so native.zig can reach it)
-> backend/{gpu_provider,gpu_none,metal,cuda}.zig
(the providers share backend/{gpu_policy,gpu_trace}.zig: gate
arithmetic over the tuning table, and the trace counter shell)
tags.zig -> tensor.zig
tensor.zig -> shape.zig, storage.zig, dtype.zig
shape.zig -> dtype.zig
storage.zig -> accelerator.zig, dtype.zig
The fucina_models module (src/models.zig + src/models/) sits above the facade:
its files import the fucina module (public surface plus fucina.internal),
never individual src/*.zig files. build.zig wires the models module against
the fucina module root only (the bench_raw/raw_backend microbench
modules are separate, apps-band roots).
Enforcement:
-
zig build arch-checkrunstools/check_import_graph.zigover the production (non-test) import graph ofsrc/,examples/,apps/,bench/, andtools/, and enforces four invariants: no forbidden strongly-connected components, zero band inversions, every sibling test file forwarded from a production file, and the backend door: a production file of the core bands (ag, tagged, moe, exec, store) imports nothing undersrc/backend/(onlysrc/backend.zig) and never names a per-format kernel child of the quant module (quant.q8k,quant.q4_k, ...; every kernel it needs is abackend.kernelsentry; the models band takes the children throughfucina.internaland the apps band throughraw_backend, the two documented escape hatches). An SCC is permitted only when every member is in the same band and one member is the directory root of another (P.zigwith a member underP/:src/exec.zigwithsrc/exec/*.zig, the struct-body-in-the-root shapestd.zigandstd/array_list.zigshare). Children cycling without their root, or any cycle that crosses a band, fail the build. The apps-band roots are scanned because theapps/entries are complete model ports rather than snippets, and an unforwarded test file there is just as silently dead as one insrc/. A file whose name ends in_tests.zigbut which declarespub fn mainis an executable root, not a suite (tools/gen_snippet_tests.ziggenerates tests), and is exempt from the forwarding rule. The checker is AST-based and test-aware:@imports insidetestdeclarations, and inside non-pub file-scope decls reachable only from tests, are excluded, so sibling-test forwarding stanzas and private test helpers do not count as production edges. The step prints the current file/edge count; the counts grow with the tree, so they are not pinned here; the root-anchored SCC count is printed too (root-with-children module shapes such asexec.zigwithsrc/exec/*.zig; like the file/edge counts it grows with the tree and is not pinned here). -
The direction bands are the Layer Stack table, encoded as
band_tablein that same tool: every production file belongs to exactly one band, and every production import must point at a band at or below the importer's. A file in no band fails the check, so a newsrc/root cannot silently acquire whatever band its neighbours have; a production layer inversion is a failed build, not a review catch. (The sibling<name>_tests.zigfiles intentionally form benign 2-cycles with their sources through the forwarding-stanza pattern — see Build And Verification — which is why any cycle check over this tree must be test-aware, asarch-checkis.)
Tensor And Storage Model¶
The raw tensor layer is intentionally small:
- Differentiable math is
f32. Raw storage/view helpers are generic over comptime dtype, and the executor has typed data movement/indexing kernels plus forward-only float kernels for non-f32tensors. - Maximum rank is
8(tensor.max_rank). - No zero-size or zero-rank raw tensors: every dimension is >= 1
(
Shape.initrejects 0) and scalars store as rank-1{1}. This is a deliberate torch divergence, not a gap: emptiness fails loud at the construction boundary instead of surfacing as torch's empty-reduction contract (mean-> NaN,min/max-> runtime error) deep in a graph; data-dependent cardinality lives host-side, where Zig represents and guards it natively (slices of[]usizeindices, optionals — the one data-dependent op pair,maskedSelect/maskedScatter, signals no-match with the recoverableerror.EmptySelection); and op/backend contracts are defined and parity-pinned only over non-degenerate shapes, which keeps that surface small for every current and future backend. - Shape and stride metadata are inline arrays; no heap allocation for metadata.
- Buffers are reference-counted through
storage.BufferOf(dtype). Borrowed storage (mmap'd GGUF tensors, device-resident bytes) uses the borrow constructors;fromBorrowedSliceWithReleaseattaches a release hook that runs when the last reference drops (the Metal weight path uses this to free device bytes and evict the shim wrap-cache slot). cloneView()retains the same buffer and preserves shape/stride/offset.reshape()is a retained view and requires contiguity.viewWithStrides()creates checked retained views.broadcastTo()uses zero strides for broadcast dimensions.data()anddataConst()require contiguity and panic on arbitrary views; recoverable callers usedataChecked()/dataConstChecked().canTakeInPlace()is an ownership optimization, not a synchronization primitive: valid only when the caller has exclusive access to the handle.
Raw tensors are internal to the public API; Tensor.asRawTensor() exposes a
read-only pointer for inspection and interop, and fucina.internal.RawTensor
is the canonical in-tree name for allowed-raw zones.
Execution Runtime¶
ExecContext is the eager runtime boundary. The struct, declared in
src/exec.zig, embeds the substrate as rt: Runtime (declared in
src/exec/runtime.zig), which owns:
- a thread-safe allocator wrapper (
ctx.allocator()is the public accessor), - the published worker team (
rt.parallel_pool, an atomic pointer thatpc()snapshots into theParallelConfigevery kernel call receives), - a reusable
BufferPool(src/exec/buffer_pool.zig), - a lazily initialized
thread.Pool(spin-then-park hot worker team), - the exec-scope stack (
openExecScope/closeExecScope: implicit ownership of training intermediates; seeMEMORY-MODEL.mdandTRAINING.md),
and declares the model/session execution state beside it, including the
grow-only decode scratch arena (decode_scratch, carved by the MoE band).
The substrate functions that operate on these fields (lifecycle, scopes,
worker team, tensor allocation primitives) are free functions in
src/exec/runtime.zig; the struct body aliases them, so ctx.empty(.f32, ...)
and exec_runtime.empty(ctx, ...) are the same call.
The execution context is responsible for allocating outputs, reusing buffers,
classifying layouts, materializing non-contiguous inputs when required
(large strided materializations are chunked across the worker team over the
raw tensor's run-based copyRangeTo), and calling backend kernels only
after validation. Each op is an alias line in the ExecContext struct
body that resolves to its domain module; domain modules take *ExecContext
plus validated arguments.
Important execution paths:
- Elementwise ops support dynamic-rank dispatch and fixed-rank APIs; fast contiguous paths call unchecked backend kernels after shape validation; tail-broadcast paths avoid materializing simple broadcast views.
take*APIs reuse unique contiguous inputs in place when safe.- Reductions, narrow/concat/gather/scatter-add, set-slice/set-rows,
argmax/topK, softmax/
softmaxExt(score scaling, additive masks, sink mass, ALiBi-stylemax_bias; broadcast masks read by stride), RMSNorm (+ fused mul/add and backward variants), LayerNorm/statistics, RoPE (with precomputedRopeTable), convolutions, and cross-entropy execute as first-class eager tensor operations, each specialized by comptime rank/axis. matmul2D/matmulTransA2D/matmulTransB2Dvalidate shape, prepare contiguous inputs, allocate output, and call backend GEMM variants;bmmvariants support broadcasted leading batch dimensions without materializing expanded tensors.- Attention (
src/exec/attention.zig) is a tiled flash-style grouped kernel with windowed and f16/quantized-KV variants. - Quantized/fused matmul (
src/exec/quant_matmul.zig) owns the quantized-RHS dispatch, the fused K-quant FFN paths, and the GPU offload seams.RhsLifetimedistinguishestransientRHS bytes fromstable_processones (process-lifetime mmap or registered device-resident storage) — only the latter may be cached address-keyed by a backend. - MoE (
src/moe/expert_ffn.zig+chain.zig, the moe band above exec) executes batched expert FFNs as a phase chain over the hot team; the route plan is a counting sort shared across families.
The runtime is local and eager. It does not fuse operations, build a planner, or preallocate an entire model execution schedule.
Backend Model¶
Backend selection is build-time (-Dbackend=native|scalar; native is
the default, scalar the reference). The scalar leg is not a second
provider: it sets backend/isa.zig's reference flag, and every kernel
entry in the one provider then selects its scalar reference arm at comptime
(the scalar namespaces in backend/vector/, the .scalar tier in
backend/quant/) while the SIMD/BLAS/GPU/lane-pack arms are never
analyzed. Dispatch is compiled away; adding a variant forces edits through
exhaustive switches.
The kernel set declares its own interface: the declaration list of
native.zig's pub const kernels is the inventory (a comptime-checked
namespace, not a struct of function pointers, because many kernels are
generic over a comptime dtype or op), and backend.zig runs
conformKernels on it at comptime: every declaration is a kernel function,
and ParallelConfig appears first or not at all (a kernel whose first
parameter is anything else is pool-free).
backend.kernels is that set. The signature rule: a kernel that
needs the worker pool takes pc: ParallelConfig as its first parameter;
one that does not use the pool does not take it; the scalar reference arms
ignore pc wherever the native arms thread on it. Exec
calls kernels.X(ctx.pc(), ...), where ExecContext.pc() snapshots the
published worker team.
The native backend uses portable Zig @Vector kernels for elementwise ops,
reductions, dot, and fallback GEMM; optional CBLAS for large GEMM
(-Dblas=none|accelerate|openblas|mkl|blis|nvpl|blas; Accelerate is the macOS
default, -Dblas=none selects the pure-Zig blocked packed GEMM); and
arch-gated int8 dot kernels (NEON sdot/smmla, AVX2/AVX-VNNI/AVX512-VNNI) for
the quantized paths. On -Dgpu=metal builds, f32/f16 GEMM gates in
native.zig and the quantized/MoE entries in the exec layer offload
above-threshold work to src/backend/metal.zig.
Dense f32, f16, and stable-weight quantized GPU calls (Q4_K/Q6_K/Q8_0 on
Metal; those plus Q5_K on CUDA) are eagerly
submitted but are not synchronously joined at every op return. Output storage
carries a completion token: another GPU GEMM stays queue-ordered (CUDA can
consume the producer device pointer), while the first CPU data access waits
for host visibility. F16 kernels write the public f32 output directly; dense
quantized linears bind exec input/output storage instead of copying through
the grouped-MoE panels. Final release waits before recycling but skips an
unused D2H. Metal caches a page wrapper per storage allocation; CUDA pools
eight in-flight typed device/tile slots behind persistent
upload/compute/download streams and uses storage-lifetime page registration so
DMA lands directly in exec-owned tensors. CUDA quantized prefill selects
adaptive N32/N64 f16-input/f32-accumulate tensor-core tiles on capable devices;
underfilled dense grids split K into a grow-only per-slot partial buffer and
queue their fixed-order reduction on the same persistent stream. This fills
idle SMs without adding a host fence, graph node, or steady-state allocation.
The same eager tile-table ABI and scalar-FFMA fallback remain. Reusable events
and one cuBLAS handle order the lanes. Grouped MoE still fences at its CPU
gather/GeGLU/scatter data dependencies, but CUDA transfers/kernel/download are
event-chained before the one required host fence. This is completion tracking
for commands that already
exist, not deferred execution or a graph; see docs/GPU-OFFLOAD.md for the
ordering proof and measurements.
The allocation contract, precisely scoped:
- Output buffers are always supplied by
ExecContext; no backend allocates tensor outputs. The vector/quant compute leaves (backend/vector/*, the dot kernels inbackend/quant/*) are allocation-free. - The quantized-RHS dispatch tier (
matmul2DQuantizedRhsinnative.zig, reference arms invector/matmul_quant.zig) deliberately takes an allocator for per-call LHS quantization scratch (f32 activations → Q8_0/Q8_1/Q8_K blocks); the Q8_0 arm has a 512-block stack fast path (q8_0_lhs_stack_blocks). RHS pack preparation (x4/x8 lane packs) allocates at load time, not per matmul. The exec-tier packed-LHS scratch above this seam is pooled (BufferPool.acquireScratchbyte-slab leases; the pool's byte-slab arm also backs all non-f32emptytransients); pooling the backend-tier scratch below the seam remains an open, bench-gated design task. - Direct native vector kernels accept a
ParallelConfigso the execution context controls thread-pool ownership.
Quantized Matmul Boundary¶
Dense .i8 is a scalar tensor dtype, not a quantized format. The public
quantized inference path is Tensor-backed: DType includes the GGML
block-quantized formats (legacy q1_0/q4_0/q4_1/q5_0/q5_1/q8_0/
q8_1, K-quants q2_k..q8_k, the cold table/nonlinear/FP4 formats
iq1_s..iq4_xs, tq1_0/tq2_0, mxfp4/nvfp4, and the Bonsai ternary
q2_0), and
Tensor(.{ .dtype = .q4_k, ... }) stores logical shape plus GGML-compatible
blocks over the last axis. These tensors are constant inference tensors:
loaded-block construction, dequantize to f32, embedding-style getRows, and
f32 matmul RHS when the dtype has a registered RHS dot kernel and the tensor
is stored [free, contract].
DType is the only identity of a storage format. src/dtype.zig defines
every GGML block struct (dtype.BlockQ4_K, ...) and the comptime
block_formats registry, one row per block-quantized dtype (its block
struct and its logical elements per block); blockSize, blockByteSize,
Storage, isBlockQuantized, and supportsQuantizedMatmulRhs derive
from it, and GGUF derives its type mapping from the same rows. No second
enum names these formats: every RHS container carries pub const dtype:
DType (the W8A8 QuantizedMatmulRhsI8 is not a block format and carries
only its group-size policy), AnyQuantizedMatmulRhs's union tags are
DType tags, and the GPU seam's QuantFormat (backend/gpu_provider.zig)
names the offloadable subset once — a provider's kernel-side integer for a
format is private ABI (abiValue(fmt), null = no kernel), not an identity.
Layout choices are comptime functions of DType and the target:
backend.PackedRhsFor(dt) (aliased as fucina.PackedRhs) is the one
dtype-to-packed-container map (dense f32/f16/bf16 share the f32 panel;
q8_0 → x4, q6_k → x4, q5_k → x8, q4_k → x2mmla on aarch64+i8mm else
x8), and facade ops dispatch on the container's dtype plus its type for
the Q4_K ISA split.
The raw tensor dtype layer owns scalar and block storage. ExecContext owns
validation, materialization, allocation, and dispatch. Backends own numeric
kernels. backend/quant.zig owns block helpers, dequantization, loaded-block
row access, the interleaved pack layouts and RHS containers, and the
portable kernels every arm shares; backend dispatch consumes
AnyQuantizedMatmulRhs internally. Each tier is addressed by one
request type. The backend seam is ops.QuantGemm, the request
{ weight, rhs: RhsPack, lhs: LhsForm, order: LoopOrder }: every packed
container states its interleave as pub const pack, each format file
exports one kernels table (request, tile body) over its tile bodies,
quant.gemm dispatches on that table and quant.supported reads it (the
one matrix of existing kernels), and the dispatch tier (the vector
parallel split gemm2D, the AnyQuantizedMatmulRhs union entry,
kernels.matmulPacked over container (dtype, pack),
kernels.matmulPackedSlice for pre-quantized LHS slices) selects
requests instead of names. The exec seam is QuantMatmul, the request
{ prologue: ?FusedActKind, placement, rhs_lifetime, numerics } with the
Lhs operand union: ExecContext.matmulQuant/matmulQuantInto are the
entries (containers come from packMatmulRhs/packMatmulRhsAs,
packDenseMatmulRhs, or the borrowing compactMatmulRhs/
compactMatmulRhsFromBlocks), plus the try* GPU attempts over raw
quantized bytes. K-quants and the IQ*/TQ* formats dot
against Q8_K activation blocks; IQ4_NL, MXFP4, and NVFP4 (like the
legacy formats) use Q8_0/Q8_1 activation blocks. Decode follows GGML
lookup tables, nonlinear codebooks, and E8M0/UE4M3 FP4 scale rules; every
cold decode format is verified bit-exactly against embedded ggml-golden
fixtures (src/backend/quant/cold_tests.zig). Matmul uses direct
integer/table dot kernels at the dtype/backend boundary, so these paths do
not materialize dense f32 RHS blocks in the inner loop. Encoders (f32 →
blocks) exist for the K-quants (Q4_K/Q5_K/Q6_K) and legacy formats
(quantizeRowForDType, surfaced by gguf.encodeF32); the cold formats decode
and matmul but do not encode.
Tagged Tensor Semantics¶
src/tag_ops.zig is the tag-semantics op library. It applies the comptime
axis-tag algebra from src/tags.zig to runtime raw tensors so the public
autograd tensor and the VJPs can delegate tag alignment and named-operation
semantics without duplicating raw view logic. There is intentionally no tagged
tensor type: tags are comptime-only data, the single runtime currency stays
the raw tensor, and the library's functions take comptime tag tuples plus
*const raw tensors and return owned raw tensors.
Library behavior includes alignTensorTo/permuteTensorTo view reordering
with zero-stride singleton injection, broadcastTensorTo,
splitAxisView/mergeAxesView, tag-driven broadcasting pointwise and
gatedPointwise, sumManyTensor/flattenTensor, and taggedEinsum — the
single contraction lowering: the output tag tuple is the whole einsum
equation (shared tags are batch axes when kept, contraction axes when
dropped; operand-private tags are free when kept, pre-summed when dropped),
operands align to an output-derived order as zero-copy views, each side
picks its plain or transposed GEMM/BMM layout at runtime by contiguity, and
the batch group collapses into one bmm axis; taggedDot is its
single-contract-tag special case — plus the shared dtype-generic
shape/validation helpers (pointwiseShape, dotResultShape,
einsumResultShape, ...). The public autograd Tensor (ag/tensor.zig)
implements the named-op surface once and calls into this library; the VJPs
(ag/backward/) call the same functions directly on raw gradients.
Autograd Model¶
The public Tensor in src/ag/tensor.zig owns exactly one raw tensor value
and optionally one gradient state:
value: RawTensor,
grad_state: ?*GradState = null,
Constants have no gradient state; variables attach a leaf GradState.
Forward execution always happens through the same public tensor operation
path. When no operand requires gradients (or a noGrad scope is active),
operations return a no-grad public tensor without retaining graph state.
When gradients are required, ag/tensor.zig computes the eager forward
value, creates a backward record from ag/backward/, and wraps it in a
GradState from ag/core.zig. ag/core.zig is backward-only: there is no
public Node, no Function.forward, and no separate raw autograd surface (a
guard test in src/ag_tests.zig asserts the legacy declarations stay
removed).
Backward execution (backwardGrad/backwardGradSerial in ag/core.zig):
- Validates every output and pre-allocates the implicit scalar seeds before
prepareBackwardPassinstalls any pending counter, so an error exit during seeding leaves the graph re-runnable instead of stranding counters. - Seeds non-scalar outputs only if a gradient is already present (the facade's
backwardWithGrad, the checkpoint recompute, externalsetGrad); explicitly pre-seeded outputs are respected without an implicit+1on top. Scalar outputs whose gradient appears only mid-pass still accumulate their own seed. - Marks outputs consumed once their pass completes: they keep their
gradients as results, so a later pass reaching a consumed state — as an
output again or as an interior node of a newer graph — would compound it;
the preparation preflights the whole reachable graph and fails with
AgError.BackwardAlreadyRunbefore any gradient moves. Failed passes stay unmarked and re-runnable: every reachable state returns to idle with a zero counter, and no non-leaf keeps a gradient. - Discovers dependencies, unwinds a refused preparation, executes ready
states and tears down released graphs over explicit intrusive worklists
(
GradState.next), never by recursion: graph depth is not a call-stack resource (a 50k-node chain runs and frees in the test suite). Per-state pending-gradient counters cover shared branches; a state is scheduled only when all downstream contributions are present, in the depth-first operand order. - A VJP that consumes its saved state in place calls
core.consumeRecordbefore its first fallible step, so a failure past it leaves the graph consumed (the retry fails at the preflight) rather than replayable over destroyed state. - Refuses to run a VJP over a saved value that was mutated after the
forward (
AgError.SavedValueMutated): every record captures, at creation, the sum of the storage mutation generations of the raw tensors it holds (found by comptime reflection over its fields, nested structs, optionals, arrays and slices), andrecordVTablecompares before the VJP;storage.Buffer.generationadvances on every mutable host access, whichever handle crosses it. - Uses the
ExecContextthread pool for async-capable backward records;backwardGradSerialdisables node-level spawning (required by the checkpoint recompute's threadlocal nesting guard). - Backward records fill only the operand slots that hold a state
(
core.needs), so unnecessary gradients are not computed, and gradients accumulate in-place under a per-state mutex.
Backward coverage spans the pointwise/reduction/view/norm/softmax/RoPE/
cross-entropy surface plus conv1d/convTranspose1d, snake, groupNorm,
quantized/f16 dot, windowed and f16-KV attention, and the tagged
contractions (einsum/dot), whose VJPs exploit closure — the gradient of
a contraction is another contraction, so every contraction backward is
GEMM-lowered (DotBackward and the constant-RHS records delegate to the
einsum records); sampling helpers (argmax/topK selection) are intentionally
no-grad. fucina.checkpoint/checkpointWithContext provide activation
checkpointing (recompute-in-backward); customVjp (ag/custom.zig) admits
user-defined differentiable ops with raw-tensor forward/backward specs; and
gradcheck (ag/gradcheck.zig) is the finite-difference oracle used to
validate both built-in VJPs and custom ops.
LLM Stack¶
src/models.zig is the root of the separate fucina_models module (wired in
build.zig; it consumes only the fucina module).
Where shared model code lives¶
Code shared across families has three homes, and the SUBJECT of the code
decides which. The rule is stated at the top of src/weights.zig; in short:
| Home | Band | Subject |
|---|---|---|
fucina.weights |
core | a weight CONTAINER: building one from GGUF bytes, and multiplying by it (LinearWeight, MoeRhs, linearSeq*, moe*FfnSeq) |
models/model_common.zig |
models | a GGUF FILE's layout: which tensor names a family's layer trio has, how an embed/head/norm set is read |
models/host_ops.zig |
models | raw f32 HOST SLICES, for the host-reference ports that run below the Tensor facade |
models/train/lora_trainer.zig |
models | a LoRA TARGET SELECTION: the per-layer adapter set, its A/B tuple, and the dropout seed stream |
linearSeq* is forward compute inside the model-I/O module on purpose:
LinearWeight.linearSeq is a container dispatch into the per-route arms and
those arms take the container types back, so the container and its multiply
are one mutually-dependent unit. A helper that fits none of these rows wants
a new home with a stated subject, not a fifth un-ruled one.
Code owned by ONE family is not shared and does not get an exec home: a
kernel with a single family consumer lives next to that family and reaches
the runtime through the public ExecContext surface (or the
fucina.internal seams). The gemma fused gate|up engines
(models/gemma/moe_gu.zig), the Kimi KDA recurrence
(models/research/kimi3/delta_attention.zig), and the DeepSeek YaRN
frequency blend (models/ops.zig) live under this rule; the shared MoE
DECODE/BATCH engines (moe/expert_ffn.zig, moe/chain.zig, MoeRhs)
are their own band above exec because every MoE family schedules through
them.
The decoder contract and the architecture registry¶
Every autoregressive text family speaks one comptime-checked surface,
declared in models/decoder.zig and asserted by assertDecoder(Model) at the
top of the generic layers (chat.Conversation,
speculative.SpeculativeDecoder, serving.gguf_chat.GgufChatBackend,
models.text.generate): a Cache type (len()/reset()/deinit(), plus
truncate iff caps.rewind), a caps: decoder.Caps value (rewind,
batch), initCache(self, ctx, capacity), and
forwardStep(self, ctx, cache, tokens, pos0) returning the LAST row's
logits as a caller-owned [1, vocab] tensor (forwardStepAllLogits
returns every row iff caps.rewind; forwardStepBatch decodes N streams
in lockstep iff caps.batch). Conforming: qwen3/qwen3moe, gemma4, SHINE's
AdaptedModel, qwen35, deepseek2, deepseek4, glm4moe, inkling; kimi3
(research tier, cache-less whole-sequence forward) stays outside.
models/registry.zig is the architecture registry: one comptime table from a
GGUF's general.architecture string to the family module (Family decls
in each family's model.zig: Model, the tokenizer module, load,
tokenizer, template_fallback). serving.open dispatches over it with
one inline for; registry.familyFor is the comptime lookup.
models/text/generate.zig is the reference generation loop over the contract
(prefill, sample, emit through a TokenSink until a stop id, the budget,
or capacity); the qwen35 and inkling chat engines and
gemma.Model.generate call it.
Model families live in subdirectories and are exposed as namespaces:
models.qwen3.{model,train}— Qwen3 dense/MoE inference + LoRA fine-tuning;forwardStepBatchis the batch-N lockstep decode entry (one m=N weight pass over N per-stream KV caches).models.qwen35.{model,chat,serving}— Qwen3.5/Qwen3.6 Gated-DeltaNet hybrid plus its ChatML chat/generation engine and serving adapter.models.gemma.{model,train,moe}— Gemma 4 text + MoE; the fused gate|up MoE kernels live with the family (models/gemma/moe_gu.zig),gemma.moeis the tagged family surface over them.models.diffusion_gemma.model— block text-diffusion on the gemma4 backbone.models.deepseek2.model— DeepSeek-V2 family (MLA + MoE).models.deepseek4.{model,serving}— DeepSeek V4 Flash (CSA/HCA attention + streamed experts viaexpert_store) and its serving adapter.models.glm4moe.model— GLM-4.5 family, with native multi-token-prediction speculative decode.models.inkling.{model,mmproj,chat,serving}— Inkling hybrid SWA/global decoder with banded relative-position bias, plus the image/audio mmproj towers, chat glue, and serving adapter.models.parakeet.*— NeMo FastConformer ASR (frontend → subsampling → encoder → CTC/TDT decoder → transcription/streaming).models.text.speculative.{core,mtp,sam_index,recycling,cascade,constrained}— lossless draft-model-free speculative decoding (seeSPECULATIVE.md), including grammar-constrained drafting and native-MTP drafting behind theDraftSourcevtable (mtp.MtpDraftSource, glm4moe).models.research.*— the research tier, one namespace so the facade states it:subq(decode-path attention evaluator, installed through the runner'sAttentionOverrideseam;SUBQUADRATIC-ATTENTION.md),engram(conditional n-gram memory, grafted through the qwen3 trainer'sresidual_hookseam;ENGRAM.md),shine/shine_train(context-to-LoRA adapters, served bymodels.qwen3.shine_serving), andkimi3.model(the Kimi-K3 port).
Shared model machinery and its actual homes (the model-I/O trio lives at
the src/ root in the fucina module; the text runtime lives under
src/models/text/):
src/weights.zig+src/weights/: GGUF weight binding —LinearWeightover resident f32/f16/bf16 and quantized forms;LoadOptions{ .gpu_resident }withloadWithOptions/loadForFusionso pre-fusion parts skip transient device residency; device-resident quant weights are owned via storage release hooks that free device bytes and evict the Metal wrap-cache slot.src/ptqtp_gguf.zig: PTQTP GGUF persistence (docs/PTQTP.md) — decorated models save as one standalone TQ2_0 tensor per trit-plane (<name>.ptqtp0/1/2replaces<name>, everything else byte-verbatim) behind afucina.ptqtp.versionmetadata gate; loader pair-detection (wired in the qwen3 loaders) rebuilds.ptqtparms bitwise, with fused weights row-sliced to source names on save and re-fused throughfuseLinear's ptqtp arm on load.src/gguf_meta.zig: flat loader glue —metaInt/metaFloat(+Opt) readers with an explicitZeroPolicy(families disagree on zero-valued keys on purpose), plus the comptime-genericparallelLoadLayers.models/text/kv_cache.zig: f16-default KV cache (opt-in q8_0 as a capacity option);truncateis the speculative rewind.kv_persist.zig: crash-safe append-only KV-cache sidecar, so conversations reopen warm across process restarts.models/text/cartridge.zig/cartridge_fleet.zig: trainable KV-prefix cartridges and per-document cartridge fleets (CARTRIDGES.md).models/research/engram.zig: conditional n-gram memory — hashed suffix n-gram tables gated into the residual stream of a frozen backbone (ENGRAM.md); exposed asmodels.research.engram.models/text/logit_processor.zig+llguidance.zig: the in-place logit-processing seam and the vendored llguidance grammar/JSON-schema engine behind it (CONSTRAINED-DECODING.md).models/text/tokenizer.zig(byte-level BPE; token-ID-exact pretokenizer chunkers: qwen2, qwen35, glm4, joyai),spm_tokenizer.zig(Gemma SPM),unicode_categories.zig(generated tables),sampler.zig.models/text/data.zig: SFT dataset/dataloader —SftTextJSONL/static pairs,encodePair(template + tokenize + shift + mask), and a deterministicLoaderwhose(seed, epoch) → permutationmapping is a golden-pinned checkpoint contract; the tokenizer parameter is duck-typed so BPE and SPM both fit.models/text/chat.zig:Conversation(comptime Model, comptime Tok)— genuinely generic multi-turn chat over any decoder-contract family withcaps.rewindover the sharedKvCache, paired with a tokenizer module;Templaterenders ChatML/Llama 3/Gemma 1-3/Gemma 4;Optionsincludesextra_stop_ids,stop_sequences, andspeculation(stop_sequencescompose with speculation: the accept gate scans the committed text and the completing token is trimmed, preserving the lossless one-draw contract; the init errors areSpeculationWithBatch/SpeculationWithReuse).sendBatchruns lockstep batch-N decode over N sibling conversations sharing one model viaModel.forwardStepBatch(speculation excluded; ownership contract in §13.8).-
models/text/serving.zig+serving/: the model-shaped half of the serving band.serving/contract.zigis the model-agnostic contract (GenerateRequest/GenerateResult,Caps, the per-familyBackendvtable); the transport is thefucina_servingmodule, one band above:src/serving/http.zig(accept loop, SSE stream pipe, Host guard; needs libc on Linux for thestd.c.recvhang-up probe),src/serving/scheduler.zig(bounded FIFO + single inference worker),src/serving/emitter.zig(per-dialect delta framing),src/serving/openai.zig+src/serving/anthropic.zig(wire dialects, over the sharedsrc/serving/wire_json.zigrequest-parsing head), andsrc/serving/toolcall.zig(hermes tool calling);serving/gguf_chat.zigis the genericGgufChatBackendengine (constraint cache, KV reuse slots + disk tier, RAM guard) for anyConversation-hosted family;serving/open.zigis the load-and-serve entry (serving.open: GGUF in, readyBackendout), dispatched through the architecture registry: each served row'sEntry.Servingnames the family'sserving.zigwiring, with theConversation-hosted set (qwen3, qwen3moe, gemma4) sharing one generic engine box and the engine-hosted set (qwen35, qwen35moe, inkling, deepseek4) driving its own engine adapter over the shared skeleton (serving/adapter_common.zig).apps/lmserveis the CLI front end and keeps only the two non-registry backends (diffusion-gemma, nanochat). -
src/optim.zig(facade) +src/optim/: SGD/AdamW/Muon/APOLLO, grad clipping, LR schedules,OptimizerSetparam groups; positionalFZT1tensor snapshots plus named, dtype-aware safetensors state dicts with name-matched optimizer state. Golden-parity-tested against torch references. src/es.zig: evolution strategies at scale (gradient-free ES-at-scale, arXiv:2509.24372): seed-regenerated gaussian perturbations over registered f32/f16/bf16 parameters (facade tensors or a wholeParamRegistry, frozen entries included — ES needs no GradState), in-place perturb/restore plus member-parallel replica materialization, z-scored update with fp32 accumulation, chunk-parallel kernels bitwise-deterministic for any thread count. Golden-pinned bytools/gen_es_goldens.pyand cross-checked bitwise against the reference implementation bytools/check_es_parity.py(seeTRAINING.md§13).src/param_registry.zig: borrows named f32/f16/bf16 tensors for checkpointing and optimizer registration; a registered name is an on-disk schema path (renames go throughstate_dict.LoadOptions.aliases, never by loosening strictness). The trainers (models/qwen3/train.zig,models/gemma/train.zig) delegate their parameter plumbing here.src/state_dict.zig+src/safetensors.zig: the named checkpoint stream and its safetensors container.src/training_checkpoint.zig: canonical checkpoint directory (model.safetensors/adapters.safetensors, nativeoptimizer.fucina, JSONtrainer_state.jsoncommit sentinel); the state codec is generic over the caller's struct, and the LLM trainers' concrete state isfucina_models.train.trainer_state.TrainerState.src/lora.zig:Adapter(in_tag, out_tag)over frozen weights; named persistence; f32/f16 merge (the fine-tune → merge → quantize → serve loop is documented inTRAINING.md).src/rnn.zig:LstmCell/Lstmover theTensor.lstmsequence op with PyTorch's gate semantics; one op call per layer and block serves the recorded forward (burn-in, truncated BPTT as one call per segment) and the blockStream; stacked-layout import/export views.src/ptqtp.zig: post-training quantization to trit-planes — K ∈ {1,2,3} ternary planes with per-group scales over packed TQ2_0 (PTQTP.md; GGUF persistence insrc/ptqtp_gguf.zig).src/gguf.zig: GGUF parser + writer (byte-verbatim metadata passthrough, llama.cpp-exact offsets;encodeF32is the writer-side quantize seam).
Build And Verification¶
build.zig wires three library modules (fucina from src/fucina.zig,
fucina_models from src/models.zig, fucina_serving from
src/serving.zig) plus the bench_raw (src/bench_raw.zig)
and raw_backend (rooted at src/backend.zig) microbench modules.
build.zig.zon names the package .fucina, so the library modules are
consumable from another project via zig fetch + b.dependency
(§2.5). The full step list and options live in AGENTS.md;
the verification-relevant steps are:
zig build test(+-Dbackend=scalar,-Dblas=none, optimize variants): drives every test root —src/fucina.zig,src/models.zig, and theapps/{lmserve,nam,parakeet,omnivoice,locate_anything,facedetect,voiceagent,nanochat}/main.zigroots (zig build test-fucinaruns the fucina root alone, the routine-Dbackend=scalarleg). Parity suites needing local model/reference assets are env-gated (e.g.OMNIVOICE_PARITY) and skip by default.zig build arch-check: the production import-graph gate (see Layering And Enforcement).zig build x86dot-check: runs the cross-ISA dot parity checker natively (ReleaseSafe, follows-Dtarget) and builds four compile-only legs (x86_64_v3 AVX2, alderlake AVX-VNNI, znver4 AVX512-VNNI, neoverse_v1 smmla) to catch bit-rot of arms no local substrate can execute.zig build doc-check: fails whenAGENTS.md's doc index names a.md(root-level,docs/<name>.md,examples/<name>/README.md, orapps/<name>/README.md) that does not exist, and existence-checks theexamples/<name>/README.mdandapps/<name>/README.mdreferences inRUNNING-MODELS.mdthe same way (tools/check_doc_links.zig).- The model runners (
run,qwen3,qwen35,gemma4,deepseek4,diffusion-gemma,parakeet,omnivoice,locate-anything,facedetect,nam,finetune,export-gguf) double as parity/oracle harnesses;bench*steps are the perf protocol vehicles (BENCHMARK.md).
Test organization: behavioral tests live in sibling <name>_tests.zig files;
each source file keeps only a one-line forwarding stanza
(test { _ = @import("<name>_tests.zig"); }) so the sibling is reachable from
its test root. The one sanctioned exception is tests that must touch non-pub
symbols, which stay inline next to those symbols (e.g. the non-pub dot
kernels in src/backend/quant/cold.zig). The forwarding stanzas are why
production files still contain test blocks; they are not license for inline
behavioral tests, and arch-check ignores the imports inside them.
Known limitations¶
- No stable external API contract: the package manifest and 0.x tags give
consumers a pin (
zig fetch --save git+...#v0.5.1), not a semver stability promise — the public API may change between tags. - The CUDA backend (
-Dgpu=cuda, Linux) covers f32/f16 GEMM + quantized dense/MoE prefill + opt-in decode GEMV; no attention/KV offload and no distributed execution. Mixed-precision training (16-bit params and activations, f32 gradients, optimizer master weights) is CPU-side —TRAINING.md§10. - No graph fusion or compiler layer — deliberate for now (
AGENTS.mdhouse rules); don't add one without a concrete design. - Quantized encoder coverage stops at K-quants + legacy formats; the cold formats (Q2_K/Q3_K/IQ/TQ/FP4) are decode/matmul-only.
- The descriptor runner (
models.qwen3.runner, docs/RUNNER.md) unifies the qwen3 family and the glm4moe trunk, andserving.openis the load-and-serve entry for the Conversation families; the OTHER families still wire their own config/loader/decoder (the shared seams areweights.zig,gguf_meta.zig,chat.zig,host_ops.zig, andmoe.chain), and no descriptor covers recurrent/MLA/hyper-connection vocabularies yet. - No documented thread-safety contract for users sharing tensor handles across threads (the runtime's internal pools are thread-safe; handle sharing is not specified).