Backstage
over Red Hat Developer Hub
The portal everything else plugs into.
I build on upstream Backstage and add plugins one at a time, so I know exactly what's in it and nothing needs a licence.
Read the decision →The stack · Every tool is a decision
Tool lists are cheap. What I find interesting is the trade-off behind each one, so here is every tool in Hangar as a decision, layer by layer. The 11 cards with an amber edge are the ones where I picked one thing over another, and each card links to where I wrote the decision down.
People should look in one place, and every change they make there should become a pull request.
over Red Hat Developer Hub
The portal everything else plugs into.
I build on upstream Backstage and add plugins one at a time, so I know exactly what's in it and nothing needs a licence.
Read the decision →Releases, promotions, canaries, SLOs and fleet views in one plugin.
A pane of glass that holds no cluster credentials. Every config change it offers becomes a GitOps pull request.
Read the decision →Docs next to the service they describe.
The same mkdocs files feed Backstage and this website, so there is only one copy to keep right.
Read the decision →An agent is just another workload, with less authority by default and every call on the record.
Policy, audit and session gateway for every agent tool call.
One choke point means one place to decide, spend budget and write the audit record. The core is built and tested; the real backends are still ahead.
Read the decision →The tool surface agents talk to.
A standard protocol, so any agent runtime can use Hangar without a custom client.
Read the decision →Fifteen deny rules, each with an id and a fix hint.
Small, fast, side-effect free, and it fails closed. The same language Kyverno uses, so there's one policy idiom.
Read the decision →Deterministic evaluations for every agent definition.
An agent definition is code, so it gets tests, including a check that each test can actually fail.
Read the decision →An ephemeral, scoped run as a Crossplane claim.
Durable changes stay commits; a run is short-lived, so it's a claim with a TTL, not a pull request.
Read the decision →AI-assisted triage when a rollout goes wrong.
Dispatched from an Airframe function, so diagnosis rides the same control plane as everything else and only ever proposes a fix.
Read the decision →One small file for developers; a fixed, platform-owned path from a merged commit to a verified image.
over Vendor-specific hosted CI
The pipeline engine.
Runs on plain Kubernetes with no vendor lock-in, and developers never have to write its YAML.
Read the decision →Everything git-triggered: push, PR, tag, ChatOps.
Webhook signatures and PR status checks are easy to get wrong by hand. Its GitHub App does them properly.
Read the decision →over One monolithic pipeline
Chains stages as events instead of one giant pipeline.
Each stage fails, retries and is observed on its own, and callers are authenticated with TokenReview rather than a key I'd have to guard.
Read the decision →over Docker-in-Docker, privileged buildah
Builds images.
Rootless under Pod Security 'restricted', so the shared build identity never needs privilege on any cluster.
Read the decision →Static analysis in the pipeline.
Fast, rule-based and readable, and its result is attested, so a release can prove SAST actually ran.
Read the decision →Image scanning and the CycloneDX SBOM.
One tool for both jobs, and both results land in the provenance a release gate checks.
Read the decision →Runs each app's test workflows.
Tests as Kubernetes resources, fenced by a Kyverno policy so one tenant's tests can't read another's secrets.
Read the decision →Signatures say who. Provenance says what the build actually did. A release gate needs both.
Signs images and writes SLSA provenance.
Provenance comes from the pipeline controller itself, not from a step a pipeline author could skip.
Read the decision →over Long-lived signing keys
Keyless certificates for in-cluster build identities.
Public Fulcio only trusts a fixed list of CI issuers. My builds are Kubernetes service accounts, so they get their own root.
Read the decision →Transparency log for signatures.
A signature you can't look up later is a claim, not evidence.
Read the decision →Keyless commit signing for people.
Human identities are exactly what public Sigstore already trusts, so running my own for this would add cost for nothing.
Read the decision →over A bare cosign verify
Checks what the build actually did before release.
It validates the provenance's content (did SAST, scan and SBOM really run?), not just that a signature exists.
Read the decision →ArgoCD is the only thing that writes to a cluster, and each cluster has its own.
over A central hub ArgoCD
The only writer to any cluster.
A hub ArgoCD holding every cluster's credentials is a path from dev into prod. Per-cluster instances keep the blast radius to one.
Read the decision →Onboarding as a generator, not a ticket.
A new service appears because a commit landed in the right folder, never because someone ran a command.
Read the decision →over Kustomize remote bases
One application chart for every tier.
Secure defaults live in one versioned chart that every service and every cluster pins.
Read the decision →Canaries with automated analysis.
A canary should feel like a conversation: pause, look, promote or roll back, all visible in Tower.
Read the decision →The platform is an API. A claim says what you want; reconciliation keeps it true afterwards.
The platform's API and reconciler.
People, portals, pipelines and agents all drive the same declarative API, and it keeps the result true afterwards.
Read the decision →What a compliant service is: apps, data services, SLOs.
Node, Spring Boot, Go and Python apps, PostgreSQL, Redis, RabbitMQ and SLOs, each a small, schema-checked claim.
Read the decision →Turn a claim into real resources, and watch them.
go-templating for the plain cases; my own functions for the interesting ones, like watching a rollout.
Read the decision →Repos, teams and branch protection as claims.
A new service's repository is reconciled like anything else, so drift gets noticed and fixed.
Read the decision →PostgreSQL behind the database claim.
A real operator for failover and backups, hidden behind a claim small enough to ask for in one line.
Read the decision →Every cluster starts from the same template and the same bootstrap order, with its own trust roots.
The execution substrate.
Plain upstream Kubernetes and nothing distribution-specific, so everything above it is portable.
Read the decision →Networking and network policy.
The template installs it before anything else, because a NetworkPolicy is only as real as the CNI enforcing it.
Read the decision →Ingress.
Gateway API is where ingress is going, and it separates the platform's listeners from each team's routes.
Read the decision →TLS everywhere, automatically.
Certificates nobody has to remember to renew.
Read the decision →over Hand-copied cluster repos
One cluster.yaml, one script, a known bootstrap order.
Copying the last cluster's repo and hoping nothing diverged is how the first two clusters drifted apart.
Read the decision →No standing credentials where a cluster-issued identity will do, and no raw Secrets in git.
over HashiCorp Vault, a cloud secrets manager
The secrets store.
Open source and self-hosted, with a proper API, so secrets are managed rather than hand-applied.
Read the decision →over Hand-applied Secrets
Delivers secrets into every chart.
Every chart consumes an ExternalSecret; nobody ever applies a raw Secret by hand.
Read the decision →Infisical projects and identities, declared.
The secrets system is configured through the same control plane as everything else, in git.
Read the decision →Workload identity from the cluster itself.
A pod proves who it is with its own audience-bound token. No minting server, no key to leak.
Read the decision →Admission policy where RBAC can't reach.
CEL-native ValidatingPolicy, used only where RBAC genuinely can't express the rule.
Read the decision →A release isn't done when it deploys. It's done when you can see whether it worked.
Traces every change from commit to release.
One trace per flow means a slow release has an answer, not a guess.
Read the decision →Metrics, kept long enough to matter.
The standard, with Thanos so DORA and SLO history survive past local retention.
Read the decision →Logs and traces.
Same query model and same Grafana as the metrics, so one place to look.
Read the decision →Dashboards, linked from Tower.
Deep dives live here; Tower links straight to the right panel.
Read the decision →SLOs as code.
An SLO is a claim in git that generates its own recording and alerting rules.
Read the decision →Delivery metrics from CDEvents.
Just another listener on the event stream, so the numbers come from what really happened.
Read the decision →Archived pipeline history.
Pruned PipelineRuns used to vanish after a day. Now you can look back at what a build really did.
Read the decision →Want to see the same tools in the order a change meets them? The homepage follows one change from intent to production, and the platform page shows how the pieces fit together.