Skip to content

2026

Holding on to the Wild Elephant

Introduction

Imagine you are the lead developer on a storage system that has been in production for three years. A bug report arrives: under high concurrency, writes are occasionally silently dropped. You open the codebase — 80,000 lines of Rust, hundreds of components, a dozen contributors. You know this system. You wrote parts of it. But scrolling through the concurrency logic, you feel it: the niggling feeling of not quite trusting your own understanding. Rebuilding the mental model takes days[1]. The bug, it turns out, was introduced six months ago in a refactor nobody fully remembers designing.

AI-Native Systems, Six Months In: What We've Learned Building the Loop

Tamar Eilam, Michael Factor, Shila Ofek-Koifman, Fabio Oliveira

In April, we introduced AI-Native Systems as a bet: that a system's evolution, the whole loop from observing a problem to deploying a fix, could be driven primarily by AI, continuously, at machine speed, rather than mediated step-by-step by humans. We described that loop in terms of a Reasoner that observes and hypothesizes, and a Changer that plans and implements, operating over a System Under Control.

Since then we've built and run pieces of that loop against real systems: a distributed inference platform, a hyperspecialized storage engine, and compute kernels for accelerators. Based on that experience, we now have a better understanding of the principles behind building an AI-native system and what's actually required in practice to build such a system. This post is an update, not a reboot: we'll walk through a sharper, more concrete architecture, show with a concrete example how its pieces fit together, and share what we've learned over the past several months of work, including the parts that didn't work the first time.

If you read the April post, this picks up where it left off. If you didn't, this should stand on its own.

What We've Learned About Modeling LLM Latency

The last post in the BLIS series covered how BLIS models the engine, data plane, and control plane to predict full-pipeline latency without touching a GPU. What it didn't answer is whether the predictions can be trusted.

That comes down to one number. Everything BLIS reports at the cluster level rests on its estimate of how long a single forward pass takes, so if that's off, nothing above it can be right. This post is about how we estimate it, how well it holds up, and where it falls short.

The headline, up front: fit once on H100, BLIS predicts held-out configurations (six models across three GPU types) at 6.7% median end-to-end error, roughly 200× faster than running them for real, and those predictions have already steered serving policies we later confirmed on a physical cluster: a better admission controller and soft reflective flow control for llm-d. The rest of this post is how we get there, and where it still falls short.

AI-Native Systems for Portable Kernels Across Hardware Architectures

Modern AI kernels are overwhelmingly written for CUDA, making it difficult to bring state-of-the-art optimizations to emerging hardware backends. In this post, we show how IBM Research collaborated with the K-Search team to automatically translate CUDA kernels to MLX, demonstrating that evolutionary kernel optimization can dramatically reduce the engineering effort required to port high-performance kernels while maintaining competitive performance.

Discovery Is Easy, Composition Is Hard

Why LLM Optimizers Find the Right Answers But Can't Assemble Them

Case Study I: The Certus Cold-Read Path

This is the first in a series of case studies applying AI-driven optimization to Certus, a hyper-specialized storage system for LLM inference and KV-cache workloads. We gave six LLM-driven code optimization frameworks and one coding agent the same task: speed up Certus's cold-read path, the performance-critical path for retrieving cached KV blocks from NVMe storage to GPU memory.

From Simulation to Production Part II: Soft-Reflective Flow Control for llm-d

Part 2 of the sim2real series. A case study in closing the AI-native loop across a second problem class: dispatch prioritization.

Introduction

The first article in this series traced a complete traversal of the AI-native loop for llm-d's admission control layer: an AI agent framework (Nous) used a high-fidelity simulator (BLIS) to discover a probabilistic, proactive admission control policy that cuts critical-traffic TTFT by up to 97% at near-saturation workloads. That article restates the core methodology: simulation lets the loop run at machine speed, reserving expensive GPU time for validating the few candidates that actually merit real-hardware testing.

This article applies the same loop to a different problem: flow control. The distinction matters. Admission control acts at the gate: it decides whether an incoming request enters the system at all, and may reject it. Flow control acts at the dispatch queue: it decides which already-admitted request gets sent to a backend server next. The two mechanisms protect different points in the pipeline and neither replaces the other.

The concrete outcome is the soft-reflective ceiling policy: a new flow control plugin for llm-d-router that reduces critical-class TTFT p99 by up to 98% at near-capacity load without rejecting a single request. It is now merged into llm-d-router, and users can enable it today with a single YAML change.

The Scientific Method on Code: How a Hypothesis-Driven AI Learned What Evolution Couldn't

Part 2 of a series on AI-driven algorithm discovery. Part 1 is here.

The first post ended on a cliffhanger.

I had spent weeks running an evolutionary AI framework — OpenEvolve — against a classic graph theory problem, watching it rediscover a 1979 algorithm called DSatur and then spin its wheels. It kept proposing variations on the same idea. It couldn't reason about why things worked or didn't. When I explicitly asked it to implement Kempe chains — an elegant post-processing technique — it produced buggy code that never ran a single useful operation.

The root cause, I argued, was structural: mutation-based evolution finds better code without understanding it. There's no mechanism to ask why something works, no way to rule out dead ends, no compounding of knowledge across iterations. The AI was playing a slot machine, not doing science.

So I decided to try the other philosophy. Instead of evolution, I used a framework built around the scientific method: the Nous open-source project. Same benchmark. Same problem. Same starting algorithm. Completely different approach.

What followed was one of the more instructive experiments I've run — not because the AI succeeded spectacularly, but because of what the structure of failure and success revealed about where this technology actually is.

Certus: The End of One-Size-Fits-All Storage

What if a storage system tailored to optimize performance based on your exact workload and hardware characteristics could be generated in days rather than hand-built over years? Certus explores that question—using AI Native Systems techniques to synthesize hyper-specialized storage engines from requirements specifications.

Can an Agentic Harness Rediscover the Insights in a Research Paper?

A case study characterizing a streaming GNN inference engine with Nous.

Try Nous yourself: github.com/AI-native-Systems-Research/agentic-strategy-evolution

Can an agentic research harness independently arrive at the same conclusions as a human researcher on a real systems problem? To test this, we pointed Nous, an agentic research harness, at the streaming GNN inference pipeline we had built and characterized in a recent paper. We gave it a research question and access to the target system, then let it run. The goal was twofold: to test whether the harness could reproduce the findings we had published, and to see whether it would surface anything we had missed.

BLIS: Evolving llm-d at Simulation Speed

Deploying llm-d is not just a question of choosing a model server and adding GPUs. In a production inference deployment, operators have to choose routing policies, admission behavior, batching settings, KV-cache reuse strategies, prefill/decode placement, and autoscaling rules under concrete TTFT, ITL, throughput, and cost constraints.

These choices are coupled. A routing change that improves cache locality can concentrate load. A prefill/decode threshold that helps one workload can hurt another. An admission policy that protects critical traffic can reduce total served volume. A change in any one policy can shift TTFT, inter-token latency, throughput, SLO compliance, and accelerator cost in ways that are difficult to predict analytically.

The only reliable way to confirm those tradeoffs is to measure them in a GPU-backed llm-d cluster. But using cluster runs as the first step in every policy or capacity-planning experiment is too slow and expensive. BLIS provides a faster inner loop: a calibrated discrete-event simulator for distributed inference systems like llm-d. Developers can evaluate candidate policies and deployment configurations locally, then reserve cluster validation for the candidates most likely to matter.

We use cookieless Google Analytics to count how many readers each post gets — no cookies, no tracking across sites. Your page URL (without query parameters), browser, and approximate location may be processed. Read what's collected →