Skip to content
duckiec
Go back

Why I Built SREK3S: An AI Incident Responder That Cannot Write to Your Cluster

duckiec/SREK3S: A fail-closed, AI-powered Site Reliability Engineer for your Kubernetes cluster.

Zero-trust, read-only AI incident response.


Most “AI SRE agents” are catastrophic control failures dressed as features. By binding service accounts to cluster-admin and mounting raw credentials, they turn container stdout—which routinely leaks AWS keys and Postgres passwords—into an exfiltration pipeline to third-party LLMs. Worse, granting the model write authority turns a log-line prompt injection into a full cluster mutation requiring zero vulnerabilities in the model itself.

I built SREK3S to take the opposite approach: remove the capability entirely. SREK3S is a zero-trust, read-only AI incident responder. Because it holds no cluster write authority anywhere in its architecture, the question of whether the model would do something dangerous never arises—it physically can’t.

What SREK3S Fixes

Instead of relying on prompt hardening, SREK3S enforces strict, structural security boundaries:

The Rigorous Engineering (Not Just Another Vibecoded Wrapper)

Instead of blindly piping model hallucinations to kubectl apply, SREK3S treats LLM output with extreme suspicion.

We Threw Everything At It: The Verification Gauntlet

We didn’t just write happy-path tests; we tried to break this system in every way imaginable. The codebase passes every gate we could throw at it (go vet, gofmt, -race, black, flake8, mypy, pytest), currently sitting at 1087 passing tests (179 in Go).

SREK3S extracts the genuinely useful capabilities of LLMs—reading crash evidence and forming hypotheses—without handing over the keys to the cluster.

References & Resources


Share this post: