Hangar

Architecture · Building Airframe, part 1 · 11 min read

Building Airframe: the start of my Crossplane journey

I'd wanted to try Crossplane on one small secret store. Instead I built a whole service catalog on it. What six weeks of Airframe taught me about readiness, API budgets, and what a composition means when it renders nothing.

The proof of concept I never got to run

In my previous role I had a Crossplane proof of concept all picked out. It was a secret store: one claim that would create an Azure Key Vault for an app, plus an External Secrets SecretStore pointing at it. It was small enough to finish, real enough to matter, and a nice way to show a team what a platform API feels like. It never happened.

So when I started building Hangar in August, I skipped the gentle introduction. Airframe, Hangar's service catalog, is Crossplane from the ground up, and it's meant to be used by AI agents as well as by people. That's diving in head first, and I knew it when I started.

Why Crossplane

I like where Upbound is taking Crossplane with intelligent control planes: the control plane as an API that people, pipelines and agents can all reason about and act through, not just a loop that applies YAML. I believe an agent should be able to ask for a database the same way a developer does, and get an honest answer about whether it's ready. If that's the goal, I want the API to be declarative, Kubernetes-shaped and able to describe itself. It's the third of my principles, and choosing Crossplane for it was easy.

Six weeks in, Airframe has fifteen APIs, from app stacks and environments to PostgreSQL, Redis, RabbitMQ, MongoDB and Dex, plus three Composition Functions. And the secret store I'd wanted to prototype was one of the first things I built. It went live on 17 August, on Infisical rather than Azure Key Vault.

Crossplane · 01 of 04 · Where it started The SecretStore I planned, and the one I built Open full page ↗

This is the first in a series. I'll keep writing these as Airframe grows, and I'm starting with what went wrong, because that's where I learned the most. Each story links to the commit that fixed it.

Crossplane · 04 of 04 · The journey so far Six weeks building Airframe, one lesson at a time Open full page ↗

Lesson one: "ready" has to be earned

My first app XRD taught me this. I wanted a new Node.js app to say "everything worked except CI/CD onboarding, which isn't possible yet", so I tried to set Ready to false with a reason. You can't. function-go-templating reserves Ready, Healthy and Synced, and it errors if a Composition sets them. Custom conditions like CicdOnboarded are the supported way to say "done, except this".

The components made it harder. PostgreSQL, running on CloudNativePG, has a Ready condition. A RabbitMQ cluster has ClusterAvailable and no Ready at all. Redis is a Helm release with a state string. Crossplane's own Ready tells you the objects exist and the provider is happy. It doesn't tell you whether you can connect. So each component now sets ComponentReady, worked out from the real resource's own status, with a closed list of reasons like PostgreSQLDegraded or RedisReleaseFailed. I kept the name away from Ready on purpose, so it can't race the auto-ready step that owns Ready. And I kept the reasons closed because something is going to switch on them, and increasingly that something is an agent. When I tested it live, a new PostgreSQL went from Provisioning to Degraded to Ready. That was a real transient state, not a simulated one.

Then the opposite trap. The new Dex component had its pod running and ComponentReady true, and the XR still said "Creating" and stayed there. The auto-ready function decides by reading each composed resource's conditions, and a Service, a PVC or a ConfigMap doesn't have any. The fix was to have the function mark those resources ready itself.

And one readiness bug wasn't mine at all. A GitHub RepositoryFile interrupted mid-create, say by a provider restart, stayed Ready: False forever. The fix belonged in the provider: a deterministic external name, which I patched into my fork of the GitHub provider along with a script that reapplies it every time the code is regenerated.

Green is a promise to whoever reads it next. More and more often, that reader will act on it without asking.

Crossplane · 02 of 04 · Readiness Three layers, three ideas of "ready" Open full page ↗

Lesson two: every reconcile spends someone's budget

Airframe writes to GitHub a lot. Creating an app commits a source repo, its boilerplate, a GitOps repo and entries in each cluster's tenants repo. GitHub gives an account 5,000 API calls an hour, and my account shares that with Backstage. I emptied it three different ways.

Polling. The GitHub provider checks every managed resource every ten minutes by default, forever, whether or not anything changed. Multiply that by every repository and it was a steady drain, and it helped break Backstage plugin testing. Every new cluster from the Apron template now starts with --poll=1h and --max-reconcile-rate=50, so nobody has to find this out again.

A fight. The provider rewrites a new file's external name to its own id, repo:file:main. My templates rendered repo:file:. The two overwrote each other about every two seconds, and every flip triggered an immediate reconcile. Five objects burned roughly 280 calls a minute and emptied the hour's budget in about fifteen. The fix was to render the live object's name when it exists, and compute one only for new objects. When the provider owns a field, render back exactly what it wrote.

A retry loop. I renamed the scaffolded Dockerfile to Containerfile, and assumed that without an Update policy, apps already onboarded wouldn't be touched. They were. Changing a template changes every object it already composed. The file path is a force-replace field, so the provider could neither update nor settle, and it retried at a sixty-second backoff that ignores the poll interval entirely. Seven objects, about 16,000 calls an hour. Now a scaffold file's path and name are pinned to what the provider reports it's managing, so changing a default only affects files that don't exist yet.

The worst one came from the provider release I was running. It read a rate-limit 403 as "this file has been deleted" and helpfully re-created it, overwriting real values.yaml content more than once. Someone upstream had already fixed it, but the fix hadn't been released, so I built and published my fork's main branch and moved on.

Crossplane · 03 of 04 · Pollers and loops Three ways to empty a 5,000-call bucket Open full page ↗

Lesson three: rendering nothing is a delete

This is the one that cost real data, and it came from the secret store, the resource I'd once picked as the safe, gentle place to start.

On 16 September a gate in one environment's template briefly evaluated to false. For that one reconcile, the template didn't render the file that asks for an app's secret store. To Crossplane, a resource that's no longer rendered is one you want gone, so it deleted the file from git. ArgoCD saw the file disappear and pruned the SecretStore it described, and the SecretStore took its Infisical project and secrets with it. Two apps lost their production secrets. In a home lab that's a lesson. At a company it's an incident.

The fixes were plain once I saw it. Every file a tenant owns is now managed without Delete, including files with a single owner, because "only one thing uses it" protected nothing here. The SecretStore never deletes an Infisical project or its environments, though identities and credentials still clean up. And I learned the provider's RepositoryFile had no deletionPolicy: Orphan at all; management policies are the only lever.

A week later I met the mirror image. I tried to "retire" scaffold files by rendering nothing once they'd been created. But a missing object is indistinguishable from a new one, so the next reconcile created it again, then garbage-collected it, then created it again. Absence can't mean "delete" and "leave it alone" at the same time. In a composition, it always means delete.

The rules I keep now

  1. Never call your own condition Ready. Give it a name and a closed list of reasons.
  2. If a resource can't report its own health, say so explicitly, or the XR never goes green.
  3. Know what one reconcile costs, times the number of objects, times how often it runs.
  4. When the provider owns a field, render back exactly what it wrote.
  5. A missing object is a delete. Anything a person owns gets no Delete policy.
  6. Changing a template for objects that already exist is a migration, not a default.

The small print

Smaller things, each found the hard way, each worth a line in someone else's notes:

A namespaced XR can't compose a cluster-scoped resource
Crossplane v2 refuses it. I use the providers' namespaced API groups where they exist, and wrap anything cluster-wide, like a ClusterSecretStore, in a provider-kubernetes Object.
stringData breaks provider-kubernetes
The API server turns stringData into data and never stores it, so the provider can't find its own field and fails with "unable to convert managed fields". Write data, base64-encoded.
One package, one Function object
A second Function pointing at the same package didn't just fail on its own. It broke the dependency lock for every other function, and auto-ready quietly lost its Deployment. I inline templates now instead.
Template delimiters inside YAML comments still count
The templating engine scans for its delimiters before YAML ever sees the comment, so a comment that merely mentions one breaks the whole render.
Crossplane needs RBAC for native kinds
Compose a PrometheusServiceLevel or any other non-provider kind directly and Crossplane's own service account needs permission for it. Nothing tells you until it fails.
Look for the primitive before you build one
I nearly wrote my own deletion-ordering check. Crossplane already ships Usage, which blocks deleting an app while an environment still needs it, enforced by a webhook that was already installed.

Building for agents, not only for people

The part I'm proudest of is the least dramatic. Every XRD now carries a one-line summary written for an agent. A generated contract bundle describes every chart field, every API and every component's outputs in one file, and CI fails if it drifts from the source. There's an llms.txt at the root of the repo, and a small CLI that answers questions like "how do I read the Redis password without guessing a Secret name?" from the contract alone. A cold-start test asks it ten real questions and checks every answer against the source files. All ten pass.

That's what an intelligent control plane means to me in practice. It isn't a chatbot bolted onto Kubernetes. It's an API that explains itself, with conditions and reasons a machine can trust. I track this on the Airframe scorecard, which has gone from 27 out of 100 at its baseline to 72 today, with half of its fourteen checks passing and plenty still to do.

What's next

I'm moving every API into the catalog.hangar.io group. The twins are in and the migration is under way, one app and cluster at a time. The next change I'd most like to make is taking repository scaffolding out of Crossplane altogether: a one-time action from onboarding, not a resource that's reconciled forever. I argued for that in the last essay, and the storms above are most of the reason why. I'll write about both when they're done.

Would I choose Crossplane again? Yes, without hesitating. Almost every bug here was mine, not Crossplane's, and the fixes were usually primitives it already had, like custom conditions, management policies and Usage, once I knew where to look. The secret store I never got to prototype turned out to be a better teacher than any proof of concept would have been.

James

← All architecture essays

The commits behind each lesson