scratchy Models Docs

A hyper-specializing
inference stack compiler

Fifty years of software reuse let the compiler delete and specialize, but never restructure. AI knocks that wall down — and once ideas are the unit of reuse, every project gets to be bespoke. Scratchy is what that looks like for inference.

1

A forward DSL

The entire forward pass of an architecture, written as the math.

gemma4-moe
2

A model config

The stock config.json for one instance of that architecture.

gemma-4-26b-a4b-it
3

A quantization

One preset, which replaces the dense emission rather than adding to it.

fp8-dynamic-per-channel
Build it How it works

Blogs

Longer-form writing on why the stack is shaped the way it is.

What “Reuse” Means in the Age of AI

Shared libraries, source modules, bundlers, AoT-compiled crates — every era let the toolchain specialize a little more, and every one of them stopped at the same wall. What happens when the unit of reuse becomes an idea, and what that bought us on IBM Spyre.

Systems Must Evolve to be Compilers

Here is the hot take: AI coding agents will eliminate the software system as a concept. In its place every solution will be bespoke, span the full stack, and will be generated from scratch.

Key numbers

30Mi
AoT binary, Metal

100 Mi for Spyre, 250 Mi for CUDA.

330Mi
Docker image, Spyre

The runtime is the binary. There is no graph executor to ship.

300ms
Warm start, Apple Silicon

Independent of model size. 12 s for 8B on Spyre.

Write the math once

Most inference engines hand-write a Python/CUDA module per architecture, then rely on a runtime graph executor to pick kernels. Scratchy takes the opposite approach: you write the forward declaratively, and the compiler emits the whole pass ahead of time.

Rust's Turing-complete procedural macros run the entire pipeline at expansion — op tape, reroll, layer classes, slot coloring, lifetimes, barriers, arena sizes, kernel selection — and emit static tapes of each target's own types.

Everything is a constant. That turns complex analysis into arithmetic, and hands the optimizer enough constants to cut register pressure in the kernels that matter.

Read the compiler deep dive →
crates/models/arch/dsl/llama.rs.in
#[forward] fn llama() { hidden_states = embed(input_ids, embed_tokens); for layer in 0..num_hidden_layers { normed = rmsnorm(hidden_states, input_layernorm[layer]); q = gemm(normed, self_attn.q_proj[layer]); k = gemm(normed, self_attn.k_proj[layer]); v = gemm(normed, self_attn.v_proj[layer]); (q, k, v) = rope_append(q, k, v, positions, rotary, kv_cache[layer]); attn = attention(q, k, v, kv_cache[layer], block_table); oproj = gemm(attn, self_attn.o_proj[layer]); hidden_states = add(oproj, hidden_states); normed2 = rmsnorm(hidden_states, post_attention_layernorm[layer]); gate = silu(gemm(normed2, mlp.gate_proj[layer])); up = gemm(normed2, mlp.up_proj[layer]); // ... } } See all 25 architectures, side by side →

The hard case: IBM Spyre

Scratchy's first target is not CUDA. It is the IBM Spyre AIU — and that is the real test, because novel silicon is where the library era has nothing to offer you.

Each Spyre core has a 2 MiB scratchpad, of which 1,677,721 bytes are yours. Every tile of every operation in the forward pass must fit. Overflow it and you do not get an error message: you get DtException 1535 on the card, minutes later, about a tile you can no longer inspect. The dominant cost of new hardware isn't writing kernels — it's the feedback loop from a nameless on-card fault back to the line of math that caused it.

Because scratchy knows every tile size at compile time, that fault becomes a cargo build error on your laptop. A general-purpose runtime structurally cannot do this: it doesn't know the shapes until it is already running, on the card, where the only channel back to you is an integer.

Supporting silicon with no ecosystem behind it took about as much Rust as supporting CUDA, and it cost zero lines of model code. The same 25 architectures, the same 1,580 lines of DSL, the same 22-line LLaMA. An 8B model boots in 12 s from a 330 Mi image.

Reuse-by-code quietly puts a population threshold on what hardware is allowed to exist — “support” means a vendor maintaining a general backend inside someone else's general framework until the market justifies the headcount. Reuse-by-idea drops that threshold to one team with a compiler.

Targets

Pick one backend per build; they are mutually exclusive.

-FspyreIBM Spyre AIU · requires the Spyre build toolkit
-FcudaNVIDIA · requires the CUDA build toolkit
-FmetalApple Silicon · requires macOS

Getting started

Install a recent Rust toolchain, then name your target, your model and your quant. Every build names its own scope — naming zero models is a build-time panic, not a silent empty binary.

Apple Silicon · Llama 3.2 3B · MLX 4-bit
# compile a server specialized for exactly this triple cargo build -F metal,model/llama-3.2-3b,quant/mlx --release # run it ./target/release/scr chat mlx-community/Llama-3.2-3B-Instruct-4bit \ -q "why is the sky blue?"
model/<stem>One checked-in config. Widen with model/<arch> or model/all.
quant/<preset>Compiles that quantization instead of dense/bf16.
-Fserve -FbenchBeyond the default scr chat CLI.
Full feature-scoping mechanics →

Documentation

Building

Feature scoping, model and quant selection, air-gapped builds.

The compiler

The whole-forward DSL, and what the proc macro actually emits.

Model architectures

All 25 supported models in the DSL, diffable two at a time.

Adding an architecture

The recipe for teaching scratchy a model it has never seen.

Spyre on OpenShift

Building and shipping an image for the AIU.

Contributing

PR workflow, CI gates, and the architecture invariants.

Source

The repository, issues, and pull requests.