case study · 2025
KubeTrace
An AI-driven Kubernetes testing, security, and reliability platform — the entire platform, built solo. It continuously validates clusters for security, reliability, and configuration drift — and pinpoints root cause in seconds.
Hours → seconds
Root-cause time
5 automated suites
Test domains
Read-only, real-time
Cluster access
What needed solving
Kubernetes clusters fail in ways that are hard to trace. Teams running regulated production workloads had no continuous way to validate security posture, catch configuration drift, or find the root cause of an incident without an expert manually walking the cluster — a process that took hours while services stayed degraded.
How I built it
Designed an MCP-based architecture so AI agents could reach multiple cluster tools through a single consistent interface, rather than hard-coding one integration per tool.
Built a read-only real-time cluster connection, so the platform can investigate a live production cluster without any risk of mutating it.
Implemented an AI root-cause investigation loop that correlates events, logs and resource state across namespaces and clusters to narrow down a failure.
Built automated test suites covering security, networking, performance, storage and API surfaces, producing compliance-ready reports.
Owned the full stack — architecture, backend services, real-time data pipelines and the product interface — from prototype to production.
Where it landed
KubeTrace runs in production at kubetrace.net, cutting root-cause investigation from a manual expert-hours exercise to a matter of seconds, with continuous validation running against multi-cluster regulated environments. Every part of the platform — architecture, backend, AI agents and interface — was designed and built by one engineer, solo.
// stack
// services this demonstrates
// next step
Have a project like this?
Book a 15-minute call and walk me through it. I'll tell you honestly what it would take and whether I'm the right fit.