We first heard about this project from a reader who runs platform engineering for a mid-sized SaaS company — roughly 40 engineers, three production regions, and a Kubernetes footprint that had quietly metastasized over four years. Their story is not unusual, but the numbers they shared at the end were specific enough that we asked for the full timeline. This is that timeline, reconstructed from their notes and a follow-up conversation.

The setup in early 2023 was a familiar mess. Eleven dashboards across five tools: one for cluster health, one for Helm releases, one for drift detection, one for incident timelines, one for cost allocation, and six more that nobody could fully account for. Four separate repositories held YAML — one for base manifests, one for overlays, one for secrets, one for "temporary" patches that had been temporary for eighteen months. Deployments took 40 minutes on a good day. Rollbacks took longer, because the rollback path depended on which of the four YAML pyramids the change had originated in. The team had a Slack channel named #yaml-archaeology. It was not a joke.

The decision point: consolidate or keep patching

In March 2023, the team hit a wall. A routine node failure in their primary region cascaded because the control plane had no unified view of which workloads were mid-deploy. They lost 90 minutes. The post-incident review recommended hiring another platform engineer. The staff engineer leading the review recommended something else: replace the stack with a single opinionated control plane. That recommendation became the project.

They evaluated three options over six weeks. The first was to keep the existing tools and write a unified API layer on top. The second was to move everything to a managed Kubernetes offering. The third was to adopt PsychoK, which they had seen referenced in a few platform engineering communities but had not seriously tested. The first option failed on cost — the API layer would have been a full-time project with no end date. The second failed on constraints: they needed bare-metal performance for a latency-sensitive workload, and managed offerings could not guarantee the network path. The third moved forward to a two-week proof of concept.

The proof of concept and the first obstacle

The PoC ran in a staging namespace with three workloads. The pitch that got their attention was straightforward: one control plane replacing the dashboard sprawl, with a Helm chart catalog they could pull from instead of maintaining their own. They started with the catalog — 2,100+ vetted charts, which meant the team could stop hand-rolling manifests for common services. The first obstacle was cultural, not technical. Two senior engineers had built the existing YAML pyramids and saw the consolidation as an implicit critique. The staff engineer handled this by running both systems in parallel for three weeks and publishing a weekly diff of deployment times. By week three, the parallel run showed PsychoK deployments completing in under four minutes versus 40 minutes on the legacy path. The resistance dissolved on its own.

What the migration actually looked like

The full cutover took eleven weeks, not the six they had planned. Here is where the time went:

  • Weeks 1–2: Chart catalog audit. They mapped every existing workload to a catalog entry or flagged it for custom handling. About 15% needed custom charts.
  • Weeks 3–5: GitOps pipeline rebuild. The old four-repo structure collapsed into one. This was the most labor-intensive phase, mostly because secrets management had to be rethought from scratch.
  • Weeks 6–8: Parallel run in staging and one production region. This is where the chaos engineering engine earned its keep. The team injected 40 fault scenarios — node kills, network partitions, disk pressure — and watched the control plane recover. The recovery behavior was consistent enough that they stopped writing manual runbooks for those cases.
  • Weeks 9–11: Remaining two production regions, one at a time, with a 48-hour soak between each.

The measurable results, six months after full cutover: deployment time down from 40 minutes to 3.5 minutes. Rollback time down from 25 minutes to under 2. Dashboard count down from 11 to 1. YAML repositories down from 4 to 1. On-call pages related to deployment failures down 70%. And the node failure recovery that had triggered the whole project? The team ran a drill in month seven: 12 nodes killed simultaneously across two regions. The control plane recovered all workloads in 6 minutes without human intervention.

What we took away from this

The interesting part is not the tooling. It is the decision to treat the control plane as a single product rather than a collection of tools. The team did not set out to reduce dashboard count; that was a side effect. They set out to make one system responsible for cluster state, and everything else followed. The 14,000+ node failures recovered by the fault-injection engine across its user base is the kind of number that sounds like marketing until you are the one watching a 12-node drill complete without a page. We have seen enough post-mortems to know that consolidation projects usually fail on culture, not technology. This one succeeded because the staff engineer measured the culture problem and fixed it with data, not argument.

If you are running a similar sprawl and want to see how the control plane is structured before you commit to a PoC, the team's notes pointed us to the architecture overview, which walks through the reconciliation model and the chart catalog structure. It is worth reading before your own evaluation, if only to know what questions to ask.