Architecture · 9 min read
What Hangar costs, and where Crossplane stops
Forty-seven tools is a bill, not a feature. What it costs to run, what I'd cut for a smaller team, and the line I draw between things that should be reconciled and things that should just happen once.
When I asked two AI models to review this site, they agreed on a lot, and the sharpest thing either of them said was a warning: operational complexity will become Hangar's limit long before conceptual complexity does. They're right, and I'd rather say it here myself than have you notice it on the stack page and wonder whether I had.
So this is the bill. What the platform costs to run, where I think the cost is worth it, what I'd cut if I were starting a team tomorrow, and the one architectural line I've learned to draw the hard way.
The bill
Count the tools on the stack page and you get forty-seven. Each cluster runs two ArgoCD instances, Crossplane with several providers, Tekton and Pipelines-as-Code, an event broker, External Secrets, Kyverno, Argo Rollouts, a gateway, a full observability stack and an AI triage function. Every app gets a source repo, a GitOps repo and an entry in each cluster's tenants repo. And there's custom code that needs an owner: the broker and its interceptor, the DORA exporter, the Tekton Results relay, the rollout watcher, a Crossplane provider for Infisical, Tower itself and Clearance.
None of that shows up on a developer's screen, which is the point. But it shows up somewhere. The week I wrote this, I took two brand new services from nothing to production, on camera, and the path found four real bugs. Three came from a single, careful refactor that every individual check had passed. That's what seams cost: each piece was right, and the gaps were between them.
A platform isn't finished when it works. It's finished when the team that owns it can afford to keep it working.
Where Crossplane stops
I use Crossplane as Hangar's API, and I'd do it again. A claim says what you want, and reconciliation keeps it true afterwards: a database that should exist, a secret store that should be wired to the right project, an environment that should have its namespace and identity. When something drifts, it gets put back. That's exactly what you want for things that are supposed to stay a certain way.
It's the wrong tool for things that are supposed to happen once. My clearest example is also my most expensive one. Creating a new app writes starter files into its GitHub repository, and I built that as Crossplane resources, so those files are reconciled forever. That design has produced an annotation fight burning about 280 API calls a minute, a retry loop burning about 16,000 an hour against a budget of 5,000 shared with Backstage, a lost external name that turned into repeated create attempts, and a patched fork of the GitHub provider. Every one of those is fixed now. None of them needed to exist.
The rule I use now is short enough to remember in a design review:
If you'd be upset that it changed back, it's a claim. If you'd be upset that it ran twice, it isn't.
A database is a claim. A repository scaffold is a one-time action. A pipeline is a sequence, so it lives in
Tekton. A release is a decision, so it's a pull request. An interactive agent session needs to react in
milliseconds, and Crossplane on my installation re-renders every sixty seconds, so Autopilot's draft
AgentRun claim only describes the sandbox and its scope. The hard stop at expiry comes from the
sandbox controller, not from a reconcile loop.
What I'd keep, and what I'd cut
If you're a small platform team and you like the shape of Hangar, this is the version I'd actually recommend you start with.
Keep from day one
- One ArgoCD per cluster, pulling from git
- It's the security boundary, and it costs almost nothing extra. It's the last thing I'd give up.
- One application chart and one cicd.yaml
- This is the contract developers actually touch. Small, stable, and schema-checked, it pays for itself on the first app.
- External Secrets and a real secrets store
- Hand-applied Secrets are how teams get breached. Any store with a proper API will do; it doesn't have to be the one I chose.
- Tekton, but with fewer stages
- Chaining stages through an event broker is lovely, and it is also a service of its own. One pipeline with build, test and release will carry a small team a long way.
Defer until it hurts
- Self-hosted Fulcio and Rekor
- I run my own because my build identities are Kubernetes service accounts. Until you need that, sign with cosign and the public instance, and spend the time elsewhere.
- Thanos, Tempo and a full tracing story
- Prometheus and Loki answer most of the questions a team of five will ask. Long-term metrics and traces earn their keep later.
- Testkube and the DORA exporter
- Useful, but each is another controller to upgrade. Run tests in the pipeline and count deployments from git until someone asks for more.
- Autopilot
- The agent design is the most interesting thing in Hangar and the least necessary. It belongs after the paved road is boring, not before.
And one thing I'd replace outright: scaffold new repositories with a one-shot action, from Backstage or an onboarding task, and let Crossplane own only what should stay true afterwards.
What I'm changing in Hangar
Hangar is a reference platform, so I'll keep more of it than I'd recommend to you. But the review and this week's runs changed my mind on a few things, and those are now on the list below with honest labels. The biggest one is the repository scaffold. The second is a shared baseline chart, so a new cluster repo is a thin diff instead of a full copy of the template, which is how the first two clusters drifted apart.
The evidence, and what's next
- Built Architecture review, 27 September: "Operating surface is large for the team behind it"
- Built Crossplane writing files to GitHub: the rate-limit incidents and the patched provider fork
- Built Four bugs from one refactor, found only by running the whole path end to end
- Proposal Scaffold repository files once, from an onboarding task, instead of reconciling them forever
- Proposal A shared baseline chart so each cluster repo is a thin diff instead of a full copy
What "the platform disappears" means
People say a good platform disappears. I think that's half true. It should disappear for the people using it: a developer should see one file and one place to look. It should never disappear for the people running it. The bill above is exactly what a platform team should be able to recite, defend and shrink. If I can't explain why a component is there, it's the first thing I'll take out.
James