Blogs
Longer-form writing on why the stack is shaped the way it is.
What “Reuse” Means in the Age of AI
Shared libraries, source modules, bundlers, AoT-compiled crates — every era let the toolchain specialize a little more, and every one of them stopped at the same wall. What happens when the unit of reuse becomes an idea, and what that bought us on IBM Spyre.
Systems Must Evolve to be Compilers
Here is the hot take: AI coding agents will eliminate the software system as a concept. In its place every solution will be bespoke, span the full stack, and will be generated from scratch.
Key numbers
100 Mi for Spyre, 250 Mi for CUDA.
The runtime is the binary. There is no graph executor to ship.
Independent of model size. 12 s for 8B on Spyre.
Write the math once
Most inference engines hand-write a Python/CUDA module per architecture, then rely on a runtime graph executor to pick kernels. Scratchy takes the opposite approach: you write the forward declaratively, and the compiler emits the whole pass ahead of time.
Rust's Turing-complete procedural macros run the entire pipeline at expansion — op tape, reroll, layer classes, slot coloring, lifetimes, barriers, arena sizes, kernel selection — and emit static tapes of each target's own types.
Everything is a constant. That turns complex analysis into arithmetic, and hands the optimizer enough constants to cut register pressure in the kernels that matter.
Read the compiler deep dive →The hard case: IBM Spyre
Scratchy's first target is not CUDA. It is the IBM Spyre AIU — and that is the real test, because novel silicon is where the library era has nothing to offer you.
Each Spyre core has a 2 MiB scratchpad, of which 1,677,721 bytes are
yours. Every tile of every operation in the forward pass must fit.
Overflow it and you do not get an error message: you get
DtException 1535 on the card, minutes later, about a tile you
can no longer inspect. The dominant cost of new hardware isn't writing
kernels — it's the feedback loop from a nameless on-card fault back to the
line of math that caused it.
Because scratchy knows every tile size at compile time, that fault becomes
a cargo build error on your laptop. A
general-purpose runtime structurally cannot do this: it doesn't know the
shapes until it is already running, on the card, where the only channel
back to you is an integer.
Supporting silicon with no ecosystem behind it took about as much Rust as supporting CUDA, and it cost zero lines of model code. The same 25 architectures, the same 1,580 lines of DSL, the same 22-line LLaMA. An 8B model boots in 12 s from a 330 Mi image.
Reuse-by-code quietly puts a population threshold on what hardware is allowed to exist — “support” means a vendor maintaining a general backend inside someone else's general framework until the market justifies the headcount. Reuse-by-idea drops that threshold to one team with a compiler.
Targets
Pick one backend per build; they are mutually exclusive.
-FspyreIBM Spyre AIU · requires the Spyre build toolkit-FcudaNVIDIA · requires the CUDA build toolkit-FmetalApple Silicon · requires macOSGetting started
Install a recent Rust toolchain, then name your target, your model and your quant. Every build names its own scope — naming zero models is a build-time panic, not a silent empty binary.
model/<stem>One checked-in config. Widen with model/<arch> or model/all.quant/<preset>Compiles that quantization instead of dense/bf16.-Fserve -FbenchBeyond the default scr chat CLI.Documentation
Building
Feature scoping, model and quant selection, air-gapped builds.
The compiler
The whole-forward DSL, and what the proc macro actually emits.
Model architectures
All 25 supported models in the DSL, diffable two at a time.
Adding an architecture
The recipe for teaching scratchy a model it has never seen.
Spyre on OpenShift
Building and shipping an image for the AIU.
Contributing
PR workflow, CI gates, and the architecture invariants.
Source
The repository, issues, and pull requests.