Platform engineering · James Fillman
Good platforms give people their time back. Hangar is how I build one.
Hi, I'm James. I've spent a long career helping software get from a developer's head to production with less friction and fewer surprises. This site is where I share how I think about that work, and Hangar is where I put it into practice: an internal developer platform I build in the open, for container apps and AI agents alike. It's never finished, and neither are my opinions.
Where the opinions come from
- ~500 apps, 9 clusters
- moved from manual
kubectl applyto GitOps with ArgoCD at Best Buy Canada - Jenkins to Tekton
- a CI/CD rebuild with SLSA provenance and Sigstore signing, driven by PCI supply-chain requirements
- 25 years
- from the Linux fleet behind online banking to Kubernetes-native platforms
What I believe
Seven ideas shape almost every decision in Hangar, roughly in order of how much they matter to me.
The platform is a product
The people using the platform are its customers, and the thing it sells is their time.
02Every change is a reviewed commit
If it changes a cluster, it went through git and someone could have said no.
03An API-driven control plane, ready for AI
The platform should be an API that keeps reconciling, so people, portals and agents can all drive it the same way.
04Paved roads, with security built in
The safe way should also be the easy way, so nobody has to choose.
05Event-driven by default
Components say what happened. Whoever cares, listens.
06Honest by design
A platform earns trust by saying exactly what it can and can't do yet.
07Interfaces people trust
A calm, clear, good-looking interface is how a platform earns the trust of people who never read its docs.
The whole SDLC, from the platform side
Follow a change from the first design conversation to the day you measure whether it worked. Each stage pairs what I believe with the part of Hangar that does it, and what's different when the author of the change is an AI agent.
The stack, in the order a change meets it
Every tool in Hangar is here for a reason. Follow one change from intent to production, and pick any tool to see why I chose it. The dashed line is the rule I care about most: the side that builds never touches a cluster.
When the author is an agent: same path, less authority. An agent's change goes through Clearance and lands as a pull request, through the same gates plus two of its own.
- 01
Declare
What do you want?
Intent starts as a form in Tower or a file in git. Either way it ends as a commit.
- 02
Compose
Make it real
Crossplane turns a small claim into a repo, a pipeline, secrets and a database, then keeps them that way.
- 03
Build
Does it build and pass?
A fixed, platform-owned path. Stages are chained by events, so each one fails and retries on its own.
- 04
Prove
Can we trust it?
Who authorised it, and what the build actually did, are separate questions with separate answers.
- Only a reviewed, merged commit crosses here
- 05
Release
Ship it, safely
A release stage never touches a cluster. It opens a pull request, and each cluster's own ArgoCD pulls.
- 06
Watch
Did it work?
Outcomes come back as events, never as an API call reaching the other way.
Secrets, identity and policy
No standing credentials where a cluster-issued identity will do, and no raw Secrets in git.
Observability
A release isn't done when it deploys. It's done when you can see whether it worked.
Why this tool
Clearance
Policy, audit and session gateway for every agent tool call.
builtOne choke point means one place to decide, spend budget and write the audit record. The core is built and tested; the real backends are still ahead.
Read the decision →Why this tool
MCP
The tool surface agents talk to.
draftA standard protocol, so any agent runtime can use Hangar without a custom client.
Read the decision →Why this tool
CEL policy
Fifteen deny rules, each with an id and a fix hint.
builtSmall, fast, side-effect free, and it fails closed. The same language Kyverno uses, so there's one policy idiom.
Read the decision →Why this tool
Preflight
Deterministic evaluations for every agent definition.
builtAn agent definition is code, so it gets tests, including a check that each test can actually fail.
Read the decision →Why this tool
AgentRun claim
An ephemeral, scoped run as a Crossplane claim.
draftDurable changes stay commits; a run is short-lived, so it's a claim with a TTL, not a pull request.
Read the decision →Why this tool
HolmesGPT
AI-assisted triage when a rollout goes wrong.
builtDispatched from an Airframe function, so diagnosis rides the same control plane as everything else and only ever proposes a fix.
Read the decision →Why this tool
Tower
Releases, promotions, canaries, SLOs and fleet views in one plugin.
builtA pane of glass that holds no cluster credentials. Every config change it offers becomes a GitOps pull request.
Read the decision →Why this tool
Backstage
The portal everything else plugs into.
builtI build on upstream Backstage and add plugins one at a time, so I know exactly what's in it and nothing needs a licence.
Chosen over Red Hat Developer Hub
Read the decision →Why this tool
Airframe XRDs
What a compliant service is: apps, data services, SLOs.
builtNode, Spring Boot, Go and Python apps, PostgreSQL, Redis, RabbitMQ and SLOs, each a small, schema-checked claim.
Read the decision →Why this tool
Crossplane
The platform's API and reconciler.
builtPeople, portals, pipelines and agents all drive the same declarative API, and it keeps the result true afterwards.
Read the decision →Why this tool
Composition Functions
Turn a claim into real resources, and watch them.
builtgo-templating for the plain cases; my own functions for the interesting ones, like watching a rollout.
Read the decision →Why this tool
provider-github
Repos, teams and branch protection as claims.
builtA new service's repository is reconciled like anything else, so drift gets noticed and fixed.
Read the decision →Why this tool
provider-infisical
Infisical projects and identities, declared.
builtThe secrets system is configured through the same control plane as everything else, in git.
Read the decision →Why this tool
CloudNativePG
PostgreSQL behind the database claim.
builtA real operator for failover and backups, hidden behind a claim small enough to ask for in one line.
Read the decision →Why this tool
Pipelines-as-Code
Everything git-triggered: push, PR, tag, ChatOps.
builtWebhook signatures and PR status checks are easy to get wrong by hand. Its GitHub App does them properly.
Read the decision →Why this tool
Tekton
The pipeline engine.
builtRuns on plain Kubernetes with no vendor lock-in, and developers never have to write its YAML.
Chosen over Vendor-specific hosted CI
Read the decision →Why this tool
CDEvents broker
Chains stages as events instead of one giant pipeline.
builtEach stage fails, retries and is observed on its own, and callers are authenticated with TokenReview rather than a key I'd have to guard.
Chosen over One monolithic pipeline
Read the decision →Why this tool
Kaniko
Builds images.
builtRootless under Pod Security 'restricted', so the shared build identity never needs privilege on any cluster.
Chosen over Docker-in-Docker, privileged buildah
Read the decision →Why this tool
Semgrep
Static analysis in the pipeline.
builtFast, rule-based and readable, and its result is attested, so a release can prove SAST actually ran.
Read the decision →Why this tool
Trivy
Image scanning and the CycloneDX SBOM.
builtOne tool for both jobs, and both results land in the provenance a release gate checks.
Read the decision →Why this tool
Testkube
Runs each app's test workflows.
builtTests as Kubernetes resources, fenced by a Kyverno policy so one tenant's tests can't read another's secrets.
Read the decision →Why this tool
gitsign
Keyless commit signing for people.
builtHuman identities are exactly what public Sigstore already trusts, so running my own for this would add cost for nothing.
Read the decision →Why this tool
Tekton Chains
Signs images and writes SLSA provenance.
builtProvenance comes from the pipeline controller itself, not from a step a pipeline author could skip.
Read the decision →Why this tool
Fulcio (self-hosted)
Keyless certificates for in-cluster build identities.
builtPublic Fulcio only trusts a fixed list of CI issuers. My builds are Kubernetes service accounts, so they get their own root.
Chosen over Long-lived signing keys
Read the decision →Why this tool
Rekor
Transparency log for signatures.
builtA signature you can't look up later is a claim, not evidence.
Read the decision →Why this tool
Conforma
Checks what the build actually did before release.
builtIt validates the provenance's content (did SAST, scan and SBOM really run?), not just that a signature exists.
Chosen over A bare cosign verify
Read the decision →Why this tool
Helm
One application chart for every tier.
builtSecure defaults live in one versioned chart that every service and every cluster pins.
Chosen over Kustomize remote bases
Read the decision →Why this tool
ArgoCD, one per cluster
The only writer to any cluster.
builtA hub ArgoCD holding every cluster's credentials is a path from dev into prod. Per-cluster instances keep the blast radius to one.
Chosen over A central hub ArgoCD
Read the decision →Why this tool
ApplicationSets
Onboarding as a generator, not a ticket.
builtA new service appears because a commit landed in the right folder, never because someone ran a command.
Read the decision →Why this tool
Argo Rollouts
Canaries with automated analysis.
builtA canary should feel like a conversation: pause, look, promote or roll back, all visible in Tower.
Read the decision →Why this tool
OpenTelemetry
Traces every change from commit to release.
builtOne trace per flow means a slow release has an answer, not a guess.
Read the decision →Why this tool
Prometheus + Thanos
Metrics, kept long enough to matter.
builtThe standard, with Thanos so DORA and SLO history survive past local retention.
Read the decision →Why this tool
Loki + Tempo
Logs and traces.
builtSame query model and same Grafana as the metrics, so one place to look.
Read the decision →Why this tool
Sloth
SLOs as code.
builtAn SLO is a claim in git that generates its own recording and alerting rules.
Read the decision →Why this tool
DORA exporter
Delivery metrics from CDEvents.
builtJust another listener on the event stream, so the numbers come from what really happened.
Read the decision →Why this tool
Infisical
The secrets store.
builtOpen source and self-hosted, with a proper API, so secrets are managed rather than hand-applied.
Chosen over HashiCorp Vault, a cloud secrets manager
Read the decision →Why this tool
External Secrets
Delivers secrets into every chart.
builtEvery chart consumes an ExternalSecret; nobody ever applies a raw Secret by hand.
Chosen over Hand-applied Secrets
Read the decision →Why this tool
TokenReview
Workload identity from the cluster itself.
builtA pod proves who it is with its own audience-bound token. No minting server, no key to leak.
Read the decision →Why this tool
Kyverno
Admission policy where RBAC can't reach.
builtCEL-native ValidatingPolicy, used only where RBAC genuinely can't express the rule.
Read the decision →Why this tool
Grafana
Dashboards, linked from Tower.
builtDeep dives live here; Tower links straight to the right panel.
Read the decision →Why this tool
Tekton Results
Archived pipeline history.
builtPruned PipelineRuns used to vanish after a day. Now you can look back at what a build really did.
Read the decision →Six products, one Hangar
Home base for building, releasing, and watching every service you run.
The platform as a whole: design, docs and the running status.
The structural core: what a compliant service is, defined once.
Crossplane XRDs, Compositions and Functions, and the one application chart every tier deploys through.
The guarded descent from a merged commit to a verified release.
CI/CD on Tekton and Pipelines-as-Code, chained by CDEvents, released by GitOps.
Release orchestration and operational intelligence for Backstage.
The single pane of glass: releases, deployments, SLOs and fleet views.
Ground infrastructure a cluster needs before a Hangar can run on it.
The template for a new cluster repo: one config file, one script, a known bootstrap order.
Any AI agent workload, bounded and audited.
Clearance (policy, audit, sessions), agent definitions, Preflight evaluations and the AgentRun claim.
Start here
Seventeen practices, one worked example
What I believe about CI/CD, told through Glidepath's architecture decision records.
AI workloadsAgents on a platform, bounded and audited
Autopilot, in nineteen diagrams: how Hangar runs any AI agent workload.
TowerA three-minute walkthrough
The Backstage plugin at the front of Hangar, from the fleet view to a promotion, a release record and a canary in prod.
Hangar, runningFrom nothing to production, for real
Two recorded runs: a new service from its first request to a canary in prod, bugs and all.
The stackEvery tool, as a decision
What I chose, what I chose it over, and why, layer by layer.
ArchitectureHow the platform fits together
The control plane, the delivery path and the pane of glass, for containers and agents.
LogWhat changed, and why
Hangar is never finished. This is its history, one dated entry at a time.
Away from the keyboard
I live in Squamish, BC, with my dog Riley. When I'm not building platforms, I'm usually at the crag, on the trails, or deep in a chess game. If you're building a platform, wrestling with one, or just want to swap notes, I'd love to hear from you.