Docs · Hangar
Service catalog design (Crossplane XRDs)
Status: DRAFT — first pass, for discussion. Goal 8 from the project’s core goals:
the service catalog is what a Backstage-driven Crossplane plugin turns into templates
(per [[project_dream_idp]]’s memory — “the Backstage service catalog is generated
from Crossplane’s CRDs, not hand-authored separately”). This doc also resolves the
one thing docs/gitops-strategy.md deliberately deferred to this design: where
Crossplane actually runs.
Revised same session: dependency-ordering questions (can an env exist without an app, does a ConfigMap need a link back to one) exposed a real gap in the first pass’s Attached-tier mechanism — fixed by routing Attached-tier resources through one shared chart (§3) instead of separately-committed files, which resolves those questions structurally rather than by validation. Components (Redis, OAuth servers, …) and a scoped-down look at a secret vault were added the same round. Crossplane v2 confirmed as the target — Claim/XR terminology from the first pass updated throughout; see Terminology section for what that actually changes, not just renames.
Revised again same session: the rendering mechanism for every list-shaped
values.yaml field (§3) is now explicit — generic range-loop templates, not Helm
sub-charts (a real mechanical limitation, not a style call), fixing an actual gap in
how multiple ConfigMaps would’ve worked. Argo Rollouts confirmed as the default
deployment resource — rollout: replaces image/Deployment in §3’s schema, and
AnalysisTemplate joins the Embedded tier (app-specific) alongside a new
platform-curated ClusterAnalysisTemplate library (cluster config, not the service
catalog) — same curated-default-plus-escape-hatch pattern already used for Components
and idp-cluster-baseline.
Revised a third time same session: confirmed Backstage holds zero Kubernetes
credentials at all — even the one remaining live API call from the previous round
(Bootstrap-tier XR creation) is now a GitOps commit, closing the last exception §0 had
carved out. Database/Queue joined the catalog; secrets settled on a self-hosted
Infisical backend (new SecretStore XRD, item 8); ArgoCD’s Rollout health check
verified for real, with a concrete caveat. See “Open questions” for the full changelog.
Revised once more: SecretStore scope narrowed to one per (app, cluster) —
shared across that app’s envs on the same cluster, never across clusters — which also
caught and reverted a wrong call from the previous round (a namespaced SecretStore
doesn’t survive cross-namespace sharing; back to ClusterSecretStore, namespace-scoped
via spec.conditions). See item 8.
Revised 2026-08-13 (separate session, after Phase 2’s first slice — see
[[idp_session_phase2_holmesgpt]] — was built and live-verified): three additions to
§3’s schema, confirmed before implementation started. rollout.podSpec: {} — a raw
map, deep-merged (Sprig mergeOverwrite) onto the rendered .spec.template.spec after
every curated/generated field — is the actual “any pod-spec field” escape hatch this doc
had only gestured at before. Deliberately excludes containers itself (the curated
fields above it cover the main container; extraContainers: [] covers sidecars) —
mergeOverwrite replaces whole arrays rather than merging by index, so letting this
escape hatch touch containers would silently clobber the chart’s own rendered
container list instead of extending it. networkPolicy joins the Embedded tier,
default enabled: true: deny cross-namespace ingress except from the ingress controller
(only relevant when ingress.enabled), leave egress open by default — an ingress-only
default was a deliberate choice, not an oversight; defaulting egress closed too would
break DNS/external-API calls in a much more confusing way than ingress isolation does,
so egress-tightening is the escape hatch (extraEgressRules), not the default.
volumes (PVC support) joins the Embedded tier, using the exact same list-of-objects
- generic range-loop pattern already established for
configMaps— one PVC + onevolumeMountper entry, no new rendering mechanism needed. Chart implementation was deferred to a later session at the time this paragraph was written; see the revision note below — it’s since been built.
Revised 2026-08-13 (third session of the day): the idp-application chart
itself is now built (idp-service-catalog/charts/idp-application), helm lint/helm template verified against three fixtures (minimal, full-featured
including blueGreen, and an appType: infra standalone-component release
with no workload). §3 below is unchanged as the schema source of truth: the
chart’s own values.yaml documents every field from that schema plus a small
number of fields §3 left genuinely unspecified (an app-facing secrets:
entry’s Infisical path selection, configMaps: mount paths, and similar) -
see charts/idp-application/README.md for the concrete list, not repeated
here since it’s implementation detail, not design. One real naming collision
worth recording here rather than just in the chart README, since it’s a trap
this doc’s own schema shape invites: cluster/env-identity fields needed
for spec.environmentRef (§ Framework) cannot be named env - §3’s schema
already uses the top-level key env for the Embedded-tier container env-var
list (env: [{name: FOO, value: bar}]); a same-named identity field silently
collides with it in one flat values map. Caught live by helm lint during
this build (range .Values.env failed with “can’t iterate over dev”) - named
envName instead. Two things stayed genuinely unresolved, not implementation
gaps but real open questions this pass didn’t answer: the actual Crossplane
API group for components:/slos: XRs (none of those XRDs exist yet - the
chart uses a placeholder, catalog.idp.io/v1alpha1) and the platform’s
default canary step sequence (§ “Still open” item 3, below - the chart renders
a deliberately inert single-step placeholder, not a real default).
Revised 2026-08-13 (fourth session of the day, immediately after the above):
a deliberate v1 resource-coverage pass on the now-built chart, prompted by “what
else should be in v1 so we don’t have to keep changing this chart” - entirely
beyond §3’s schema, not a gap in it, so not folded into §3 itself; recorded here
only as a pointer, full detail in charts/idp-application/README.md. Added: a
dedicated ServiceAccount (the identity every pod now runs as, instead of the
namespace’s implicit default - the one thing on this list that’s genuinely hard
to retrofit once real deployments exist), jobs:/cronJobs: batch tasks sharing
the main workload’s env/secrets/config automatically, a ServiceMonitor
(kube-prometheus-stack is already installed cluster-side), and a raw
extraManifests: escape hatch matching idp-cluster-baseline’s own pattern. One
more real bug caught live, general enough to be worth a line here too: Sprig’s
default function treats an explicit false/0 exactly like “unset” and
silently substitutes the default anyway - hook: false on a jobs: entry still
rendered as hook: true until fixed with an explicit hasKey check instead. Any
future schema field on this chart where the Go zero value is a legitimate,
meaningful setting (not just “not configured”) needs the same treatment, not a
bare | default.
Revised 2026-08-13 (fifth session of the day): code review of the built
chart surfaced three more real gaps in §3’s configMaps:/secrets: schema
itself (not the resource-coverage pass above, and unlike that pass, folded
directly into §3’s own schema block below, since these are genuine additions
to fields §3 already specifies, not new resource kinds outside it) - full
detail in charts/idp-application/README.md. secrets: could only ever become an env
var, never a mounted file; configMaps: could only ever be volume-mounted,
never envFrom’d; and configMaps: could only ever be chart-owned via
data:, with no way to reference one created outside this chart - the last
one a real, concrete need (a Kustomize configMapGenerator’s output). Both
lists gained an as: env | volume | both field; configMaps: gained
existingConfigMap: as a data: alternative. Your call on the Kustomize
case specifically: a fixed name (disableNameSuffixHash: true), not a
hash-suffixed one requiring external sync - the accepted tradeoff is that
ConfigMap loses Kustomize’s own automatic-rollout-on-content-change property.
Closed a related, adjacent gap at the same time: configMaps[].data/secret
edits previously didn’t change the Rollout’s pod template at all (same
name/keys), so Argo Rollouts never started a new revision - fixed with
checksum/configmaps/checksum/secrets pod-template annotations, though the
secrets one only catches a declaration change, not a value rotated in
Infisical without touching values.yaml (a different, harder problem, not
solved here).
Revised 2026-08-13 (sixth session of the day): networkPolicy:’s
extraIngressRules/extraEgressRules escape hatch works but requires knowing
the real K8s NetworkPolicyPeer/NetworkPolicyPort shape - not simple for the
dominant real case, “let this other namespace (optionally narrowed to some
pods) or this CIDR reach me on this port.” allowIngressFrom:/allowEgressTo:
(folded directly into §3’s schema block below, same reasoning as the
configMaps:/secrets: revision above - a genuine addition to a field §3
already specifies) cover that flatly: {namespace, podLabels?, ports?} or
{cidr, ports?}, ports as bare TCP port numbers. namespace: resolves via the
same kubernetes.io/metadata.name-label mechanism already used for
ingressControllerNamespaceSelector, not a new idiom. Both raw escape hatches
stay, unchanged, for UDP/SCTP or multiple ORed peers in one rule - full detail
in charts/idp-application/README.md.
Terminology (Crossplane v2 primitives, for reference)
Confirmed: this design targets Crossplane v2. v2’s real, load-bearing change from v1: an XR can be namespaced directly — a developer creates the XR itself, in a real namespace, with no separate cluster-scoped XR hidden behind a namespaced Claim proxy. “Claim” language from the first draft (written before this was confirmed) is replaced below — flagging here rather than silently, since it’s a genuine simplification, not just a rename: the Attached tier (§ Framework) no longer has a hidden indirection layer to explain — the XR the chart renders is the real resource, sitting directly in the app’s own namespace.
- XRD (
CompositeResourceDefinition) — defines a new custom API type: the schema for the XR (Composite Resource) a developer creates directly — namespaced by default in v2, cluster-scoped only where an XRD deliberately declares it (none of this catalog’s XRDs need cluster scope; see below). - Composition — the implementation: how an XR’s spec becomes real managed resources
(or other XRs). Multiple Compositions can implement one XRD, selected by label — this
is how “same XRD, different backing resource per env tier” works, flagged as a real
future need in §5 of
gitops-strategy.md(a lower-env XR resolving to a cheaper Composition than the same XR kind in an upper env). - Composition Function — a real code pipeline (Go, KCL, or a templating function) that computes a Composition’s output, the modern replacement for pure declarative patch-and-transform. Needed here — see §2.
- Provider / managed resource — a Crossplane-managed integration with an external API (cloud provider, GitHub, a Helm release, another Kubernetes API) and the CRD-shaped resources it manages.
§0. Where Crossplane runs, and how Backstage reaches it — with zero K8s write credentials
Revised this round — Backstage never holds any Kubernetes credential, of any kind, on any cluster. The first pass had one exception (Backstage calling the dev cluster’s K8s API directly to create Bootstrap-tier XRs) — your call: no k8s write creds on Backstage at all, every XR follows the GitOps pattern. This turns out to be a real simplification, not just a constraint satisfied grudgingly — it removes the one special case §0 used to carve out, and makes the answer to “does an Attached-tier XRD have to go through Backstage-calls-K8s-directly, or can it be a template too” (your question 1) simply: everything is a git commit, so yes, absolutely — see below.
Crossplane still runs per-cluster (confirms gitops-strategy.md’s tentative
default). Every XR, Bootstrap or Attached, is created the same fundamental way — commit
a manifest into a GitOps-synced location, let that cluster’s own pull-based ArgoCD and
Crossplane do the rest — the only variable is which location:
- Bootstrap-tier (
NodeJSApplication,SpringBootApplication,ApplicationEnvironment): commit intogitops-cluster-<cluster>-tenants/<app>/— the same directory that already carriesidentity.yamland already drives the tenant-onboardingApplicationSet. A newxr-requests/subdirectory carries the XR manifest itself. This closes the “namespace must pre-exist” wrinkle from the first pass for free, rather than needing Backstage to make a separate namespace-creation call: the SAME per-appApplicationthat already rendersapp-<name>-cicd’s namespace/RBAC/Triggers now also renders whatever’s inxr-requests/, in one sync — a namespace at a lowersync-wavethan the XR that lands inside it, both applied by ArgoCD in the same operation. For a brand-new app, Backstage’s action is exactly one commit (identity.yaml+ the firstxr-requests/file, together) — there’s no longer a live API call anywhere in this path, at any point in an app’s life. - Attached-tier (
SLO,Redis,OAuthServer, …): unchanged from the first pass — a block in that env’s ownvalues.yamlingitops-<app-name>, rendered by theidp-applicationchart. This tier never needed a K8s credential in the first place.
What Backstage actually needs, then: a scoped GitHub write, not a Kubernetes one —
and not even a new mechanism for that. platform-cicd’s token-review-interceptor
already mints per-repo-scoped GitHub installation tokens on demand
(/github-installation-token), TokenReview-authenticated, never persisting the GitHub
App’s own private key outside platform-system. Backstage’s scaffolder actions call
that same endpoint rather than holding a standing GitHub credential of its own —
directly the pattern [[feedback_credential_persistence_preference]] already established
(“never saved anywhere” beats “cached/refreshed”), reused instead of re-solved.
One nuance worth being precise about, so “zero K8s creds” is verifiably true and not just asserted: Backstage still needs something to authenticate to that endpoint — a Kubernetes ServiceAccount token, checked via TokenReview, the same mechanism every other caller of that endpoint already uses. This is not a contradiction of “no k8s write creds”: that SA token carries no RBAC grants to create, patch, or delete any Kubernetes resource — it’s an identity credential for one HTTP call, not a write credential for the K8s API. Worth stating explicitly rather than eliding, since the distinction is exactly the kind of thing worth getting precise rather than hand-waved.
This keeps the guiding constraint from gitops-strategy.md (“no cluster ever holds
credentials for another cluster’s API”) intact, and extends it one layer further than
the first pass did: now it’s not just “no cross-cluster credential,” it’s “no live
write credential of any kind, anywhere, for anything” — every mutation, at every layer
of this whole design, is a reviewed or self-service git commit, synced by the cluster
that owns the result.
Resolved: app-<name>-cicd, uniformly, for all three Bootstrap-tier XRDs — no
separate shared platform-catalog namespace. Unchanged conclusion from the first
pass, reached by a cleaner mechanism now (above) instead of a two-call workaround.
Does -cicd still fit, now that it also hosts provisioning objects, not just pipeline
execution? Recommend keeping the name — the boundary it draws didn’t actually move.
naming-conventions.md already defined -cicd as conceptually distinct from dev/
staging/pr-<n> before any of this: “the Application’s own pipeline-execution
namespace,” explicitly not a deploy target. That distinction was always really
control-plane namespace vs. deploy-target namespace — pipeline execution just
happened to be the only control-plane resident that existed yet. Crossplane’s Bootstrap-
tier XRs are a second resident of the same role (control-plane objects that act on
behalf of the app, never receive live traffic), not a new category the existing
boundary has to stretch to cover. So the name is arguably always been a slight
under-description of its own role, not a description that’s now wrong.
Weighed against an actual rename: -cicd is baked into a live, working,
two-real-tenant system — the envNamespace helper in platform-cicd-app, every
existing namespace, RBAC, broker CEL filters, docs. A rename means a real namespace
migration (Kubernetes can’t rename a namespace in place), not a find-and-replace, for a
label-precision gain that doesn’t change any actual behavior. Not worth it. If a broader
naming pass ever happens for independent reasons, “control” or “platform” would be a
more literal name for the role this namespace has always actually played — worth
remembering then, not a reason to act now.
Where Crossplane runs across a multi-cluster fleet (resolved 2026-08-15)
Closes the doc’s own opening claim (“resolves the one thing gitops-strategy.md
deliberately deferred… where Crossplane actually runs”) for real, once a fleet with
more than one dev cluster and real upper-env clusters is on the table (a second, real
prod cluster now exists for testing this). Reasoning worked through live in
conversation, not asserted — kept here so it isn’t lost:
The two XRD tiers have different locality requirements, and that difference is the
whole answer. Bootstrap-tier (NodeJSApplication, ApplicationEnvironment) composes
only provider-github resources — every mutation is a GitHub API call, never a
Kubernetes API call to any cluster. Attached-tier (SLO today; Redis/OAuthServer/
Database/Queue once built) composes native, in-cluster resources directly (SLO’s
Composition renders a real PrometheusServiceLevel into whichever cluster it’s
reconciled on) — that has no meaning unless Crossplane is actually running there.
- Bootstrap-tier stays centralized on one dev cluster, permanently, regardless of fleet size. There is nothing about creating a GitHub repo or committing a file that benefits from running per-cluster, and running N copies of the same Composition against N clusters would just create N controllers racing to own the same repo.
- Attached-tier (and, see below, AI-triage) must run per-cluster — on every cluster
that hosts real app deployments, dev and upper-env alike. This was always implied by
how
SLOalready works; it just had nothing to contradict it while only one cluster existed.
A cluster registry is the missing piece that makes any of this checkable, once
there’s more than one cluster of either kind. A small, cluster-admin-authored,
PR-reviewed object per cluster — a labeled ConfigMap is enough, no new CRD needed
(matches the low-ceremony, operator-authored-data pattern this doc already uses for
tenants/*/app.yaml):
apiVersion: v1
kind: ConfigMap
metadata:
name: prod # cluster name
namespace: crossplane-system
labels: {platform.io/cluster-registry: "true"}
data:
type: upper # dev | upper
cicdReady: "false" # dev only - flips once that cluster's CICD control plane is live
crossplaneReady: "false" # both types - flips once Crossplane + the Attached-tier
# catalog subset is installed and healthy there
Both readiness flags are manual, PR-reviewed attestations, not automated health
probes — same “CI gates are lint/syntax, human review carries the real weight”
philosophy gitops-strategy.md §8 already applies to cluster config generally, and
consistent with [[feedback_live_verification]]’s standing caution against trusting a
passing check without confirming the thing it’s gating is actually on. An automated
probe could report into this as a second opinion later; it shouldn’t be the sole gate.
Registering a cluster (adding its entry) and bringing it fully online (§4’s bootstrap
sequence, extended to also install the Attached-tier catalog once 10-crds-operators
includes Crossplane) stay the same cluster-admin action — the registry doesn’t invent
new toil, it just makes the fact checkable.
NodeJSApplication gains a required devCluster field, validated via an
extra-resources lookup against this registry (must resolve to type: dev +
cicdReady: "true", or the Composition creates nothing and reports a blocking
condition — same “structural backstop, refuse rather than partially succeed”
instinct already used for WorkloadDeployed/CicdOnboarded). Immutable once set —
a CEL transition rule (self.devCluster == oldSelf.devCluster, rejected after
creation), not a mutable field Crossplane would try to reconcile toward. devCluster
isn’t just “which tenants repo app.yaml lands in” — it’s which cluster’s CICD control
plane owns the app’s pipeline history/secrets/webhooks, and which cluster’s own
lower-env (§10) ApplicationSet live-reads this app’s platform/envs/. None of that is
something a declarative reconcile loop can safely migrate; a spec-field change would at
best orphan everything on the old cluster while partially standing up a new one, not
actually move anything. Moving an app to a different dev cluster is a deliberate
decommission-and-re-onboard (delete NodeJSApplication, which per the fix below already
can’t happen until every ApplicationEnvironment child is gone — then create a new one),
not a field edit. As-built procedure, proven live on boarding-api 2026-09-24:
airframe/docs/user/decommission-app.md (deleting the XR also deletes both GitHub repos).
ApplicationEnvironment’s cluster field stops being a hardcoded Composition
constant and becomes a real, required spec field, gated the same way: extra-resources
lookup against the registry, must resolve to type: upper + crossplaneReady: "true".
Explicitly rejecting type: dev targets is a real, load-bearing enforcement, not just
tidiness — §10 of gitops-strategy.md is explicit that gitops-<app-name> (which
ApplicationEnvironment writes into) carries upper environments only; a dev cluster’s
environments belong to the separately-scoped platform/envs/-live-read mechanism
instead, with its own narrower AppProject. Nothing currently stops ApplicationEnvironment
from targeting a dev cluster — it’s only ever pointed at dev today because that’s
the sole cluster that exists, not because anything enforces the boundary. This closes
that gap once the registry exists to check against.
One more real gap the registry surfaces: the first ApplicationEnvironment for a
given (app, cluster) pair must also seed that cluster’s own app.yaml-equivalent
tenant entry, not just identity.yaml. §6 scopes AppProject per cluster — an app
deployed to three clusters gets three independently-generated AppProjects, each built
by that cluster’s own tenant-appprojects from that cluster’s own app.yaml.
NodeJSApplication only ever writes the devCluster’s copy; every other cluster an app
gets deployed to needs its own, and ApplicationEnvironment is the only thing that ever
learns about a new cluster for an app, so it’s the natural (idempotent,
overwriteOnCreate: true, same as everywhere else) place to seed it.
AppProject/Application ownership stays with ArgoCD’s ApplicationSets
(cluster-admin-templated), not moved into direct Crossplane composition — considered
and rejected, for two compounding reasons. First, it’s structurally incompatible with
Bootstrap-tier staying centralized: ApplicationEnvironment targeting an upper-env
cluster runs on the dev cluster’s Crossplane, which has no credential to that upper-env
cluster’s API and structurally shouldn’t get one — direct composition would require
exactly the cross-cluster credential the guiding constraint forbids, or force
Bootstrap-tier to un-centralize after all. Second, AppProject.sourceRepos/
destinations is the actual security enforcement boundary (§6) — its shape is
deliberately authored in gitops-cluster-<name>, a repo gitops-strategy.md keeps
“close to read-only… rare, high-stakes… real review weight,” owned by cluster
admins. Letting a Composition generate that shape directly would move the security
boundary’s definition into idp-service-catalog’s own release cadence instead — a
different, less rigorous bar than the one gitops-strategy.md deliberately wants for
anything that shapes an AppProject.
The same reasoning extends to a related idea considered and set aside: having
ApplicationEnvironment itself trigger installation of the Attached-tier catalog onto a
new cluster (e.g. by committing into that cluster’s gitops-cluster-<name>) the first
time an app targets it. Appealing (removes a manual step), but it’s the same
ownership-boundary cross one level further down the stack — an app-owner-facing XRD
would be expanding what’s installed on a cluster, which gitops-strategy.md’s own
terminology section assigns to cluster admins exclusively (“app owner… never touches
cluster config”). crossplaneReady in the registry is the alternative that gets most of
the practical value (a single checkable fact ApplicationEnvironment gates on) without
crossing it — if the real goal is reducing cluster-admin toil rather than shifting who
controls cluster infrastructure, the lever is a more turnkey admin-run bootstrap script
(extending the existing hack/bootstrap-upper-cluster.sh precedent), not moving the
trigger to the app side.
Fixes the real, twice-confirmed AppProject-deletion-ordering bug — built and
live-verified 2026-08-15 (found live building ApplicationEnvironment, see
idp_session_applicationenvironment_xrd — the two ApplicationSets prune
independently with no ordering between them, so deleting an AppProject before its
dependent Application finishes its own finalizer cleanup permanently stuck that
Application) — at the Crossplane layer, not the ArgoCD one, since ArgoCD’s
ApplicationSet doesn’t expose an ordering primitive for this and teaching it one
isn’t obviously possible. Originally planned as a homegrown extra-resources lookup on
NodeJSApplication’s own Composition (query for remaining ApplicationEnvironment
XRs, refuse deletion if any exist) — superseded before building it once kubectl api-resources on dev confirmed Crossplane itself already ships a real
primitive for exactly this: protection.crossplane.io/v1beta1 Usage (“defines a
deletion blocking relationship between two resources”), enforced by a live
crossplane-no-usages admission webhook already installed with this cluster’s
Crossplane — nothing new to deploy. ApplicationEnvironment’s own Composition now
composes one, unconditionally (not gated by the $clusterOk cluster-registry check
elsewhere in the same template — the app/env relationship holds regardless of
deployment-gate status): spec.of = the parent NodeJSApplication (by
spec.appName), spec.by = the ApplicationEnvironment XR itself. NodeJSApplication’s own Composition needed zero changes — a real simplification
versus the original design, since the webhook and Usage controller do all the
blocking purely by watching Usage objects, regardless of what composed them.
Live-verified end-to-end on dev: a real NodeJSApplication + referencing
ApplicationEnvironment, confirmed the Usage object and the crossplane.io/in-use
label it drives, confirmed kubectl delete on the app is cleanly rejected at
admission time (not a finalizer hang) while the env exists, confirmed the Usage is
garbage-collected the moment the env is deleted, confirmed app deletion then succeeds
— see idp-service-catalog’s own README (v0.3.2) for the full pass.
The xr-requests/ mechanism itself (point 1 above) — built and live-verified
2026-08-15. gitops-cluster-dev/02-argocd-apps/xr-requests/ adds a dedicated
ApplicationSet (a git directories generator on tenants/*, not a files generator
on app.yaml — gating XR creation on a file the XR itself produces would be circular)
and a narrowly-scoped idp-onboarding AppProject (not default, not the per-app one
— both are circular too for a brand-new app; scoped to just
NodeJSApplication/ApplicationEnvironment from gitops-cluster-dev-tenants only,
same shape as platform-cicd’s own platform-onboarding AppProject). Live-verified
end-to-end, twice, with a real throwaway app (xr-onboarding-verify): a git commit
into tenants/<app>/xr-requests/nodejsapplication.yaml created the app-<app>-cicd
namespace and the XR, which provisioned real GitHub repos and committed app.yaml
back, which the pre-existing tenant-appprojects ApplicationSet then turned into a
real per-app AppProject — closing the loop with zero manual kubectl apply anywhere.
A second commit (applicationenvironment.yaml, targeting prod) deployed a real
env on the second cluster the same way. The idp-onboarding AppProject boundary was
attack-tested, not just asserted: a committed Secret was rejected with resource :Secret is not permitted in project idp-onboarding.
Two real bugs found live during this build:
- Fixed: a directory-type
Applicationsource pointed straight atxr-requests/errors manifest generation entirely (app path does not exist) the moment its last file is removed, since git doesn’t track empty directories — exactly the moment a real deletion needs a clean diff to zero resources instead. Fixed by pointing the source at the always-presenttenants/<app>directory (guaranteed to exist — it’s what the generator just matched) withdirectory: {recurse: true, include: "xr-requests/*.yaml"}instead of the subfolder directly. - Suspected resolved, not proven — downgraded 2026-08-15 after 3 clean live
reproductions: deleting an
ApplicationEnvironmentxr-request through thisApplicationwas confirmed twice (2026-08-15 morning session) to deadlock against theUsageit composes — theUsage’s own controller refuses to release its finalizer until theApplicationEnvironmentis actually gone (WaitingUsingDeleted), while theApplicationEnvironmentitself won’t finish going away until that sameUsage(an owned,blockOwnerDeletion: truedependent) is gone first — circular.PrunePropagationPolicy=backgroundwas tried as a fix and did not resolve it on a same-day retest. A later same-day pass (afternoon) live-reproduced the same deletion path 3 times, including one attempt matching the original failure’s timing and target cluster almost exactly (prod, ~13 minutes dwell before deletion, same git-commit-removal path) — all 3 tore down cleanly with no manual finalizer-clearing. No code change was made to theUsage/finalizer mechanism itself between the confirmed failures and the clean runs. The one relevant thing that did change: a real, separate ArgoCD Redis-cache staleness bug (seegitops-cluster-dev’s01-argocd/README.md) was found and fixed immediately before the clean runs, and the original failures happened during a session doing many rapid onboard/ teardown cycles — exactly the load that would build up stale ArgoCD cache state. This is a plausible link, not a proven one; forcing a stale-cache condition deliberately and re-testing would be needed to confirm causation. Recovery, if this ever recurs: manually clear theUsageobject’s own finalizer (kubectl patch usage <name> -n <ns> --type=merge -p '{"metadata":{"finalizers":[]}}'). Treat env deletion throughxr-requests/as usable but monitor rather than fully routine until this has more soak time.
AI-triage (function-rollout-watcher/diagnosis-holmes-dispatch) redesign — DONE,
live-verified 2026-08-18/19 on prod (built directly on this section’s own
design below, written the session before). Confirmed live, not just by reading code,
that the old mechanism — watching req.observed.resources["rollout"], a Rollout
composed by step 1 of the function’s own pipeline (the ai-rollout-derived
Application XRD, which rendered the Rollout directly) — had nothing to attach to in
the real idp-application/Helm/ArgoCD deployment model. Fixed by switching to a real
Crossplane extra-resources requirement (response.require_resources/
request.get_required_resource — the modern SDK call, not the deprecated
extra_resources field or function-go-templating’s YAML-only ExtraResources
meta-resource, which doesn’t apply to a native Python function): a small, always-on
Attached-tier XR (RolloutWatch, new XRD) idp-application renders unconditionally
alongside any release with rollout: set (same treatment as ServiceMonitor), same
name/namespace as the Rollout it watches. Its Composition (single-step, just
function-rollout-watcher — no function-go-templating involved at all) matches the
live Rollout by name and reads its observed status from there; the diagnosis Job is
still the only thing this Composition composes.
GitOps/source repo coordinates turned out simpler than planned: NodeJSApplication has
no live appRepoUrl field to read (it was only ever a template-computed string, never
persisted) and can’t be cross-cluster-looked-up anyway (Bootstrap-tier centralizes on
dev; RolloutWatch runs on whichever cluster the app is actually deployed to).
Derives them deterministically instead — gitops-<appName> / <appName>, same fixed
platform owner, <cluster>/<env>/values.yaml — mirroring exactly what
ApplicationEnvironment’s own Composition already computes for the same app. No per-XR
annotations needed at all (the old gitops.example.org/*/src.example.org/* scheme is
gone, it was never wired to the real system).
Holmes locality, resolved: runs per cluster, not shared — same reasoning that
already kept AppProject/SecretStore per-cluster (it needs live in-cluster access to
actually investigate a Degraded Rollout). Installed via its own GitOps Application
(gitops-cluster-<name>/30-ai-triage/holmesgpt/), robusta/holmes v0.38.0, per-cluster
ANTHROPIC_API_KEY/GitHub-PAT secrets (never templated into git). One real gotcha hit
live worth remembering: an ANTHROPIC_API_KEY env var alone does not register a
usable model — Holmes’ own /api/model reported only the built-in "Robusta" hosted
model until an explicit modelList entry (envRef:ANTHROPIC_API_KEY sugar) was added.
Live-verified end-to-end on prod’s real (not manufactured) checkout-api
ImagePullBackOff — see [[idp_session_ai_triage_extra_resources]] for the full account.
Upper-env half built and live-verified 2026-08-15: the registry,
ApplicationEnvironment.spec.cluster becoming real, the type: upper/crossplaneReady
gating (both the rejection and success paths), and the first-time app.yaml seeding
are all real and proven end-to-end against prod — see Item 3’s own “Built and
live-verified for real” note below for the detail, including one real bug found and
fixed (managementPolicies, not deletionPolicy). Still not buildable: a second
dev cluster (NodeJSApplication.spec.devCluster and its own registry gate) — no
second dev cluster exists yet, and that field was explicitly out of scope for this
pass. AI-triage’s own redesign (this section, above) is now built and live-verified
(2026-08-18/19), no longer open.
§1. provider-github is the mechanism behind every “create/commit to a repo” step
Category-1 XRDs (§0) need to create a GitHub repo and commit files into it, without a
publish/package step and without a full clone-commit-push flow. The real package
(confirmed live building NodeJSApplication, not just this doc’s original shorthand) is
crossplane-contrib/provider-upjet-github — provider-github was a guess at the name,
the actual xpkg reference is xpkg.upbound.io/crossplane-contrib/provider-upjet-github.
It wraps the Terraform GitHub provider as plain managed resources: Repository (create
the repo), RepositoryFile (create/update a single file via GitHub’s Contents API — this
is what writes boilerplate files, cicd.yaml, a tenant identity.yaml-equivalent, or an
env’s values.yaml), BranchProtection (required checks/reviewers, matching what
platform-cicd already enforces on gitops-<app-name> today — not built in
NodeJSApplication’s first pass, since it would need to reference status-check names
from a CICD pipeline that doesn’t exist yet on this cluster, see Item 1/2’s status note).
Credential: not the GitHub App after all — corrected live, this was the doc’s biggest
wrong assumption. The original plan (“one GitHub App credential, same shape as
token-review-interceptor already mints from — no new credential class”) turned out to
be structurally impossible for this catalog’s actual GitHub account: jfillman is a
personal User account, not an Organization, and GitHub Apps are unconditionally
blocked from POST /user/repos (403 Resource not accessible by integration) — a
documented platform restriction, not a permissions/scope gap on the App. GitHub Apps can
only create repositories inside an Organization they’re installed on. Confirmed live
building NodeJSApplication: the App credential 403’d on every Repository create,
with zero App-permission configuration able to fix it. Live fix: a classic PAT
(repo + delete_repo scopes), stored the same way (source: Secret, JSON
{"owner": "jfillman", "token": "..."} — the provider’s githubConfig.Token field,
plain PAT auth alongside its app_auth field, not instead of it as a provider
limitation — this catalog just doesn’t use app_auth now). This is a new credential
class, contradicting the original plan — tracked here as a known, deliberate deviation,
not silently reconciled. Revisit if jfillman ever becomes/moves under an Organization,
which would reopen the GitHub-App path. owner stays a ProviderConfig-wide setting
either way (required, confirmed live against the provider’s own credential-parsing
source) — not a per-Repository field, so no XRD in this catalog takes an owner as a
spec field regardless of which credential type backs it.
Managed-resource family: the namespaced one (repo.github.m.upbound.io), not the
Cluster-scoped one — also corrected live, a second wrong first guess. The provider
ships both a Cluster-scoped family (repo.github.upbound.io) and a namespaced one
(repo.github.m.upbound.io); the Cluster-scoped one looked like the safer choice at
first because it has real, complete examples in the provider’s own repo (the namespaced
example tree uses a stale, never-filled-in placeholder apiVersion,
template.m.crossplane.io/v1beta1, that reads like unfinished scaffolding). That
first guess broke immediately on a real cluster: Crossplane v2 hard-rejects a namespaced
XR composing a Cluster-scoped managed resource at all (cannot apply cluster scoped composed resource ... for a namespaced composite resource) — not a permissions issue,
a structural one. The namespaced family’s CRDs are genuinely installed and working
despite its misleading examples; NodeJSApplication composes those instead, via a
ClusterProviderConfig (github.m.upbound.io/v1beta1, not the legacy Cluster-scoped
ProviderConfig) so the one credential stays referenceable from every namespace without
duplicating the Secret per app. Lesson for the next Bootstrap-tier XRD
(ApplicationEnvironment) or any future provider adoption: check the actual installed
CRD scope (kubectl api-resources) before trusting which variant a provider’s example
directory happens to document best.
§2. Composition authoring: function-go-templating, not patch-and-transform + KCL
Revised after actually building NodeJSApplication. This section originally
recommended a pipeline of function-patch-and-transform for field-mapping plus a
dedicated KCL-or-Go-templating function for file-rendering, shared across
NodeJSApplication and SpringBootApplication. What got built instead, matching the
real convention the SLO Composition already established: pure
function-go-templating, source: Inline, generated from templates/*.yaml via a
build-composition.sh script (idp-service-catalog/compositions/slo/, now also
compositions/nodejsapplication/) — no function-patch-and-transform step, no second
Function registration. A second Function object pointing at a package reference
already installed (e.g. a dedicated file-rendering function alongside the shared
function-go-templating) corrupted Crossplane’s shared dependency-lock graph
cluster-wide, a real bug hit live building the SLO Composition (see that Composition’s
build-composition.sh header) — reusing the one already-installed function-go-templating
Function via Inline templates avoids the problem entirely rather than working around it.
Revisited now that SpringBootApplication is built (2026-08-24), decision: still not
worth it. Four of five render-github-resources/*.yaml files
(00-devcluster-gate.yaml, cicd-identity-yaml.yaml, gitops-repo.yaml,
secretstore-xr.yaml) plus cicd-onboarding-status/status.yaml are now genuinely
byte-for-byte duplicated between the two Compositions, with only cosmetic
NodeJSApplication/SpringBootApplication naming swapped in comments — real
duplication, not a hypothetical. Only src-repo.yaml (boilerplate/Dockerfile) is
actually stack-specific. A “shared function, parameterized by stack” refactor would
remove that duplication, but was deliberately not built this pass: it would touch the
one pipeline mechanism both live-verified XRDs depend on, for a payoff that’s
copy-paste-drift risk (a real but low-severity cost — these five files change rarely)
against a nontrivial redesign of an already-working, already-verified mechanism. Worth
revisiting if a third Bootstrap-tier language XRD is ever built (three copies is a
different call than two), or if one of these five files needs a real bugfix and the
other Composition’s copy would otherwise silently drift — not before then.
Confirmed: Crossplane v2 target. See the Terminology section above for what that changes — namespaced XRs directly, no separate Claim type. Composition Functions (this section) are orthogonal to the v1/v2 split (a 1.14+ change, still how v2 Compositions work) — nothing here needed revising because of the version confirmation, only the Claim/XR terminology used throughout the rest of this doc did.
Framework: identity/bootstrap resources vs. attached resources vs. embedded fields
Three tiers, not a blanket policy — this is the direct answer to your “should we
separate these out” questions (items 5, 6). Revised from the first pass after your
dependency-ordering questions below — Attached-tier resources no longer commit their
own separate file; they’re values blocks inside the same values.yaml the Deployment
already comes from, rendered by one shared chart (§3). This isn’t just simpler, it’s
what makes “can this exist without an app/env” unaskable rather than merely validated —
see “Dependency ordering” below.
| Tier | Examples | Mechanism | Lifecycle |
|---|---|---|---|
| Bootstrap | NodeJSApplication, SpringBootApplication, ApplicationEnvironment |
A commit into gitops-cluster-<cluster>-tenants/<app>/xr-requests/, synced by the app’s own onboarding Application, reconciled via provider-github (§0) — no live API call, no k8s credential |
Creates a new addressable git location that didn’t exist before — a repo, or a new <cluster>/<env>/values.yaml |
| Attached | SLO, Redis, OAuthServer (§ Components) |
A block inside that env’s values.yaml (slos:, components:); the idp-application chart (§3) renders it into an XR, directly in the app’s own namespace, auto-stamped with which app/env it came from |
Independent provisioning lifecycle, but only expressible inside an existing env’s file |
| Embedded | config, secrets, HPA, PodDisruptionBudget, AnalysisTemplate, resource limits |
Plain fields on the same values.yaml, rendered directly (no XR at all) |
1:1 with the single workload (Argo Rollout — §3) the Application already owns |
Linking mechanism, revised: because Attached-tier blocks live inside the same
values.yaml file the chart already renders the Deployment from, the chart already
knows appName/cluster/env from its own release context — it stamps
spec.environmentRef: {name: <app>-<cluster>-<env>} and platform.io/app: <app> onto
every XR/resource it renders from a components:/slos: block automatically. A
developer adding Redis to their app never types an app reference by hand; it’s implicit
in which file they’re editing. environmentRef names the specific ApplicationEnvironment
XR this resource depends on (deterministic name <app>-<cluster>-<env>, not three
separate fields) — see “Dependency ordering” for what that reference is actually for.
On the Backstage side, the Crossplane plugin still needs to translate these refs into
dependsOn/dependencyOf catalog-info.yaml relations — unchanged from the first draft,
still not built.
Dependency ordering — what can and can’t exist without what
Direct answers to your questions, in order:
“Can an environment be added without specifying an application?” No, structurally.
ApplicationEnvironment.spec.environmentRef’s app-name component is a required XRD
schema field — the Kubernetes API server itself rejects an XR missing it, before any
Composition runs. Semantically it couldn’t work anyway: the Composition’s only job is to
commit into gitops-<app-name>, so without an app name there’s no repo to commit into.
“Does it need to be required to link a ConfigMap to an application? Does an env need
to exist? Where would it deploy with no env defined?” These questions dissolve under
the revised mechanism rather than needing a runtime check to catch them: a config field
(Embedded tier) isn’t an independently-Claimable resource at all, it’s a key inside a
specific env’s values.yaml. There is no “orphan ConfigMap” state to guard against,
because there’s no path to creating one that isn’t “edit a file that, by definition,
already belongs to one app and one env.” Same answer for Attached-tier blocks
(components:, slos:) — they’re keys in that same file.
The one place a real dependency-ordering gap remains: environmentRef proves an
app name was supplied, not that the named ApplicationEnvironment XR actually
exists yet — OpenAPI schema validation can’t express cross-resource existence checks.
Two layers, not one:
- Structural backstop: the chart only ever renders Attached-tier blocks as part of
a Helm release that’s already scoped to one specific, already-provisioned env’s own
Application/AppProject(§0.2) — so in practice this can’t actually be reached through the normal self-service path at all. It only becomes reachable if something hand-constructs the XR directly, bypassing the chart. - For that edge case: the Attached XRD’s Composition Function does an “extra
resources” lookup of the referenced
ApplicationEnvironmentXR and reportsReady: Falsewith a clear reason if it’s not found yet, rather than erroring opaquely — Crossplane’s native way to express “this waits on that,” visible viakubectl describeand, once built, on the resource’s own Backstage catalog page.
§3. The idp-application Helm chart — the actual center of this design
Everything above funnels into one chart (idp-service-catalog/charts/idp-application),
because every env is exactly one Helm release: one values.yaml, one Application,
one namespace. Worth designing this concretely rather than leaving “plain Helm values”
abstract, since it’s what items 5/6 and the new Components/SLO mechanism above actually
resolve to.
# gitops-<app-name>/<cluster>/<env>/values.yaml — one file, one release, one namespace
appName: my-app
appType: app # app | infra (naming-conventions.md's existing distinction —
# see §"Components" below for why infra-type reuses this, not a new concept)
rollout: # Embedded — Argo Rollouts is the default workload, see § below.
image: {repository: ..., tag: ...} # omit `rollout` entirely for an infra-type release with no custom workload
replicas: 2
resources: {...}
strategy: canary # canary | blueGreen — platform supplies default steps unless overridden
# Curated, safe-default escape hatches for the main container/pod — common enough to
# deserve real fields rather than forcing every app through the raw podSpec escape
# hatch below. Each is toYaml'd straight in when set; the chart supplies a sane
# default (e.g. an httpGet probe on the first port) when omitted.
command: []
args: []
ports: [{name: http, containerPort: 8080}] # first entry doubles as the Service/ingress target
livenessProbe: {}
readinessProbe: {}
podSecurityContext: {}
containerSecurityContext: {}
extraContainers: [] # full container specs, appended as-is (sidecars) — kept
# separate from podSpec.containers, see below
podSpec: {} # the actual "any pod-spec field" escape hatch — a raw map,
# deep-merged (Sprig mergeOverwrite) onto the rendered
# .spec.template.spec AFTER every curated/generated field
# (main container, volumes from configMaps/secrets/volumes
# below, extraContainers). Anything not already covered by a
# curated field goes here: tolerations, affinity,
# nodeSelector, topologySpreadConstraints, dnsPolicy,
# hostAliases, terminationGracePeriodSeconds, etc.
# Deliberately NOT for containers[0] (use the curated fields
# above) or additional containers (use extraContainers) —
# mergeOverwrite replaces whole arrays rather than merging by
# index, so mixing container edits into this escape hatch
# would silently clobber the chart-rendered container list.
analysisTemplates: # Embedded — app-specific custom AnalysisTemplates, see § below.
- name: checkout-conversion-rate
metrics: [...]
env: [{name: FOO, value: bar}] # Embedded
configMaps: # Embedded — revised this round, see "Rendering mechanism" below.
# `as: env|volume|both` (default volume) and
# `existingConfigMap:` (a data: alternative, for one this
# chart doesn't own — e.g. a Kustomize configMapGenerator's
# output) added 2026-08-13, see revision note above.
- name: app-settings
data: {app-config.yaml: "...", logging.yaml: "..."}
- name: feature-flags
data: {flags.json: "..."}
secrets: [{name: db-password, key: DB_PASSWORD}] # Embedded — same appSecretStores/ESO
# mechanism platform-cicd already built,
# not reinvented here — unlike ConfigMap,
# NOT revised to a multi-object shape, see below.
# `as: env|volume|both` (default env)
# added 2026-08-13, see revision note above —
# volume mode mounts one key at an exact
# file path, not a directory.
volumes: # Embedded — PVC support, added 2026-08-13. Same list-of-objects
# + generic range-loop pattern as configMaps: one PVC + one
# volumeMount per entry.
- name: uploads
size: 10Gi
storageClassName: standard # omit for cluster default
accessModes: [ReadWriteOnce] # default if omitted
mountPath: /data/uploads
autoscaling: {enabled: false, min: 2, max: 10, targetCPUPercent: 70} # Embedded — scaleTargetRef.kind: Rollout
podDisruptionBudget: {enabled: false} # Embedded
ingress: {enabled: true, host: my-app.example.com} # Embedded
networkPolicy: # Embedded — added 2026-08-13. Default reflects "isolate
# traffic to its own namespace": deny cross-namespace ingress
# except from the ingress controller (only relevant when
# ingress.enabled) and other pods in this same namespace;
# egress left open by default — see revision note above for why.
enabled: true
allowIngressFromIngressController: true
allowIngressFrom: # simplified peers, added 2026-08-13 — {namespace, podLabels?,
# ports?} or {cidr, ports?}; namespace resolves via the
# auto-populated kubernetes.io/metadata.name label, same
# mechanism as ingressControllerNamespaceSelector. See
# revision note above and idp-application's own README for
# why extraIngressRules alone wasn't simple enough.
- namespace: app-payments-prod
ports: [8080]
allowEgressTo: [] # same shape as allowIngressFrom, for egress
extraIngressRules: [] # raw NetworkPolicyIngressRule list — escape hatch for
# whatever allowIngressFrom can't express (UDP/SCTP,
# multiple ORed peers in one rule)
extraEgressRules: []
components: # Attached — rendered as Component XRs, see below
- type: redis
name: cache
spec: {size: small}
slos: # Attached — rendered as SLO XRs
- name: availability
objective: 99.9
indicator: {...}
Rendering mechanism: one generic pattern, not sub-charts, applied to every list field
Not Helm sub-charts — this is a mechanical Helm limitation, not a preference. A
sub-chart is a statically-declared, 0-or-1 unit in Chart.yaml — Helm’s dependency
model has no way to say “instantiate this sub-chart once per entry in a values list,
with different values each time.” Every field above that needs N instances
(components, slos, configMaps, analysisTemplates) genuinely can’t be expressed
as a sub-chart at all, regardless of style preference — a sub-chart could give you one
optional Redis, never “however many components a developer lists.”
The actual mechanism: one range loop per field, in the main chart’s own
templates — the standard Helm idiom for “N instances driven by a values list.”
Concretely, for components:/slos: (Attached tier), the loop is generic, not
type-specific: each entry carries a type (redis, oauth-server, slo) mapped to a
kind via a small lookup table, and its spec: block is passed straight through to the
rendered XR untouched — the chart doesn’t validate or branch on component-specific
fields at all, Crossplane’s own XRD schema does that when the XR lands. This is what
makes the catalog extensible without editing this chart: adding a new Component XRD
next month (Kafka, say) needs one new row in the lookup table, never new template logic.
Backstage’s own form curation is still the primary UX guardrail (§ Dependency ordering) —
the generic passthrough is what happens underneath a curated form, not a replacement for
one.
ConfigMap, revised to answer “what if a user needs several”: the first draft’s flat
configFiles: [...] never actually said whether multiple entries meant multiple keys in
one ConfigMap or multiple ConfigMap objects — a real gap. Fixed with a two-level shape:
configMaps: is a list of objects (name, data: {key: content, ...}), rendered by
the same generic range-loop. This answers both cases at once — multiple keys in one
logical ConfigMap (nest more entries in one data: block) and multiple separate
ConfigMap objects (add another list entry) — without needing to pick one shape over the
other. Real reasons an app might want the split (separate mount paths, separate
restart-on-change semantics via a checksum annotation scoped to just one object, the
1MiB per-object size limit) all fall out naturally once it’s list-of-objects rather than
list-of-files.
Secrets stays as it was, deliberately not given the same two-level treatment:
platform-cicd’s existing ESO pattern is already one ExternalSecret per app with
multiple keys — a working, live-verified mechanism, not something this doc is revising.
Only ConfigMap had the gap; secrets never did.
Argo Rollouts as the default deployment resource
New decision this round, replacing plain Deployment as idp-application’s core
workload — the rollout: block above renders an Argo Rollout, not a Deployment.
Downstream effects worth being explicit about:
autoscaling’s HPA now targetsscaleTargetRef: {kind: Rollout, ...}instead ofDeployment— Argo Rollouts supports this natively, no different mechanism needed, just a differentkindin the same field.AnalysisTemplate: same curated-default-plus-self-service-escape-hatch pattern already used for Components (§ item 7) andidp-cluster-baseline(§8 ofgitops-strategy.md), applied a third time — worth naming as a real, repeating pattern in this design, not a coincidence:- Platform-curated, shared, reusable across every app:
ClusterAnalysisTemplateresources (Argo Rollouts’ own cluster-scoped variant, referenceable from any namespace) — a small golden-path library (error-rate-check,success-rate-check) living inidp-cluster-baseline, installed alongside the Argo Rollouts controller itself in cluster config, not a service-catalog XRD — developers reference these by name, they never Claim or create one. - App-specific, custom: the
analysisTemplates:Embedded-tier block above, rendered as namespacedAnalysisTemplateresources scoped to that app’s own namespace — for genuinely app-specific metrics a shared library wouldn’t cover (e.g. a business KPI unique to this app). Argo Rollouts lets a singleRollout’s canaryanalysis.templates[]mix references to both cluster-scoped and namespace-scoped templates in the same step, so an app can use the platform defaults and its own custom one together without any special wiring.
- Platform-curated, shared, reusable across every app:
- A real integration dependency worth flagging, not verified yet: the multi-cluster
release-outcome mechanism from
gitops-strategy.md(PostSync/SyncFailhooks) only fires correctly once ArgoCD’s ownstatus.health.statusreachesHealthy— which means ArgoCD has to understand aRollout’s health correctly, not just “resource applied.” ArgoCD does ship a built-in health check forargoproj.io/Rolloutin reasonably current versions, but this doc hasn’t confirmed it against the actual ArgoCD version in use — worth a real live check (same instinct as the hook-timing test inplatform_cicd_session_multicluster_argocd) before relying on it, not an assumption.
What’s deliberately NOT in this chart: anything cluster-scoped, or anything a
per-app AppProject (§6 of gitops-strategy.md) shouldn’t be allowed to create outside
its own namespace — ClusterSecretStore, the Application/AppProject resource
itself (can’t render its own container), platform-shared infra (ingress controller,
Vault server if one exists — see the Component/Vault discussion below). Those live in
cluster config, referenced by name, never templated here.
Resolved 2026-08-13, when the chart was built: rollout: is genuinely optional
(set to null/omitted) rather than needing a separate, lighter sibling chart — the
“optional in the same chart” leaning below, confirmed. Implemented and fixture-tested
(an appType: infra release with only a components: block renders cleanly, no
Rollout/Service/HPA/PDB/AnalysisTemplate, just the NetworkPolicy and the Component XR).
Item 1/2: NodeJSApplication / SpringBootApplication
Separate XRDs, not one generic Application XRD with a language field. A discriminated-
union schema would technically work, but goal 8’s own framing — one XRD becomes one
Backstage template — argues for it directly: a developer picking “New Service” wants two
distinct, clearly-labeled template cards, not a generic form with a language dropdown
buried inside. Matches the existing memory note on this (favor XRD designs that “read
cleanly as a Backstage template input” over internally-convenient ones). The repo-
creation/CICD-onboarding logic that’s ~90% identical across languages is real,
confirmed duplication now that both are built — see §2’s revision note for the
decision (still two separate, unshared Compositions; not worth a shared-Function
refactor yet).
Scope, deliberately narrow: src repo + boilerplate + an empty, scaffolded
gitops-<app-name> repo + a tenants/<app-name>/app.yaml commit into
gitops-cluster-dev-tenants (corrected from this section’s original identity.yaml-
equivalent guess — app.yaml is the real, already-built file this catalog’s tenant
ApplicationSets read for the app-level AppProject, see
gitops-cluster-dev-tenants/README.md). Not upper-env provisioning — that’s item 3.
This split maps directly onto the lower/upper security boundary already designed:
everything these two XRDs do is inherently dev-cluster, self-service, no-review-gate-needed
territory; promoting to a real environment is a deliberately separate, higher-trust action.
Status: NodeJSApplication built 2026-08-13, live-verified on dev
(idp-service-catalog/xrds/nodejsapplication.yaml, compositions/nodejsapplication/).
At build time, the “CICD onboarding” half of this scope genuinely couldn’t complete
yet — platform-cicd’s control plane wasn’t running on dev — so the Composition
surfaced the gap as an explicit custom condition (CicdOnboarded: False, reason
CicdOnboardingPending) rather than silently succeeding or blocking. Real as of
2026-08-15: platform-cicd’s control plane now runs as a second, independent instance
on dev (see platform-cicd/docs/bootstrap.md’s own note), and the Composition
gained a real step committing tenants/<app-name>/identity.yaml into
platform-cicd-dev-tenants — platform-cicd’s own tenant-onboarding
ApplicationSet picks it up and stands up the app’s actual CICD pipeline. Redirected
2026-08-16 (idp-service-catalog v0.3.5): that dedicated repo was eliminated once
live history showed it only ever held throwaway apps and dev’s platform-cicd
instance was confirmed idp-exclusive - the same commit now lands in
gitops-cluster-dev-tenants instead, alongside app.yaml. See that repo’s own README
and cicd-identity-yaml.yaml’s own header comment for the full reasoning.
CicdOnboarded now reflects the real observed status of that commit (True once it’s
Ready), not a hardcoded False — live-verified end-to-end with a throwaway app,
including a real signed build (genuine .att OCI attestation in the registry, Fulcio
cert chained to dev’s own independently-generated root CA — not just the
chains.tekton.dev/signed: "true" annotation, which has lied on this platform before).
Custom condition, not an override of the standard Ready condition, which
function-go-templating reserves and errors on if a Composition tries to set it
directly (confirmed live; the framework’s own custom-condition mechanism, target
CompositeAndClaim, is the supported way to surface this). BranchProtection was also
deliberately left out of this pass — it would need to reference status-check names from
a CICD pipeline, and while one now exists, wiring BranchProtection itself to it is
still separate, unstarted work.
Live verification (real kubectl apply, throwaway nodejsapp-verify-test) produced two
real corrections, not assumed in the original design — see §1 for the full detail: the
GitHub App credential can’t create repos under jfillman’s personal account at all (a
PAT backs the ProviderConfig instead, for now), and the Composition composes
repo.github.m.upbound.io (namespaced), not repo.github.upbound.io (Cluster-scoped) —
Crossplane v2 rejects the latter for a namespaced XR outright. Both real repos, all five
boilerplate files, and the tenants/nodejsapp-verify-test/app.yaml commit were confirmed
via the GitHub API before teardown. One more live footnote worth recording: the PAT
needs delete_repo alongside repo — without it, Repository deletion 403s and the
composed resource gets stuck Terminating (hit live during cleanup; not a
NodeJSApplication bug, but relevant to anyone deprovisioning through this provider).
Status: SpringBootApplication built 2026-08-24, live-verified on dev
(idp-service-catalog/xrds/springbootapplication.yaml,
compositions/springbootapplication/) — a structural port of NodeJSApplication onto
the Java/Spring Boot stack, same devCluster-gated onboarding mechanism, same
CicdOnboarded-not-Ready status convention, same six-field identity.yaml. Only the
Java/Spring-specific pieces differ: spec.javaVersion ("17"/"21", replacing
nodeVersion), spec.buildTool (maven/gradle, replacing packageManager, branching
the Composition’s boilerplate and multi-stage Dockerfile the same way), and
spec.groupId (no NodeJSApplication equivalent — the Maven groupId / Gradle group,
and deliberately this app’s only Java package, not groupId+artifactId nested,
because $xrName’s kebab-case metadata.name isn’t a legal Java package-name segment;
the scaffolded Application class lives directly in the groupId package instead of
attempting a kebab-to-camelCase derivation). springBootVersion is a template literal
(3.3.4), not a spec field, matching the reasoning against a spec.githubOwner field on
NodeJSApplication — the platform, not each app, picks and centrally upgrades the
framework version. The Dockerfile is genuinely multi-stage here (JDK build stage, bare
JRE runtime stage) rather than NodeJSApplication’s single-stage script — a real
functional difference, not just a style choice, since a compiled JVM app has no runtime
need for the build toolchain at all (unlike Node’s runtime interpreter, which is the
same binary either way).
Live verification (real kubectl apply, throwaway springbootapp-verify-test on
dev, both maven- and gradle-buildTool branches exercised via
crossplane render first, then a real end-to-end kubectl apply for the maven
branch) needed no corrections beyond what NodeJSApplication already discovered and
fixed in this same provider/pipeline plumbing — confirms this catalog’s “port the
already-verified mechanism, vary only the stack-specific content” approach holds up a
second time. Real pom.xml, identity.yaml, and the xr-requests/secretstore.yaml
commit were confirmed via the GitHub API; DevClusterReady, CicdOnboarded, and the
XR’s own Ready condition all reached True live before teardown.
Why gitops-<app-name> gets created here, empty, rather than by item 3: its
lifecycle is app-level (create once), not env-level (create per cluster×env) — creating
it alongside the src repo, at the one moment both are being bootstrapped together, avoids
a “which of possibly-several env Claims owns the shared repo” ownership question with no
clean answer under Crossplane’s per-Claim resource ownership model.
Still not built: a required, creation-immutable devCluster field (gated against
the new cluster registry — type: dev + cicdReady), designed 2026-08-15 — see
“Where Crossplane runs across a multi-cluster fleet” in §0 for the full reasoning.
The AppProject-deletion-ordering bug fix, also designed 2026-08-15, is built and
live-verified — but it doesn’t touch this XRD/Composition at all, contrary to the
original plan recorded here. It’s a protection.crossplane.io Usage composed by
ApplicationEnvironment’s own Composition instead (§0 has the full mechanism and
live-verification detail) — NodeJSApplication needed zero changes.
Item 3: ApplicationEnvironment (renamed from UpperEnv)
Naming and mechanism, both addressed:
Granularity — one XR per (app, cluster, env), not one XR per app covering all its
envs. Matches gitops-<app-name>’s own directory structure 1:1 (<cluster>/<env>/ values.yaml), and avoids the failure mode of a list-shaped spec where adding a new env
means editing (and risking corrupting) an existing resource rather than creating a new,
independent one.
Mechanism — pure provider-github, no direct access to the target cluster at all.
The Composition commits two things: the env’s <cluster>/<env>/values.yaml into
gitops-<app-name>, and an onboarding entry into gitops-cluster-<cluster>-tenants
(creating that entry if this is the app’s first env on that cluster). The target
cluster’s own argocd-apps ApplicationSet — already watching its own tenants repo, per
gitops-strategy.md §1 — picks up the new entry entirely on its own and creates the
namespace/Application/AppProject as part of its normal sync. This is why a remote
provider-kubernetes credential is never needed here, and why this XRD is Bootstrap-tier
(§0.1), not Attached-tier: at XR-creation time, nothing on the target cluster exists
yet for it to commit into.
Naming: ApplicationEnvironment over the placeholder UpperEnv — names what it
does (deploys one environment for one application) rather than where it sits in a
lower/upper taxonomy that’s really a property of which repo the request lands in
(gitops-/platform/envs/), not of the XRD itself. Easy to bikeshed
further, low cost to rename later.
Built and live-verified 2026-08-15 (idp-service-catalog/xrds/ applicationenvironment.yaml, compositions/applicationenvironment/). Two design
calls resolved concretely, both confirmed against already-live code before deciding,
not guessed:
Superseded 2026-08-15, now that a second, real upper-env cluster (clusterstays a fixed Composition constant ("dev"), not a spec field — confirmed the already-builttenant-onboardingApplicationSet (gitops-cluster-dev/02-argocd-apps/tenant-onboarding/applicationset.yaml) already hardcodes the same literal in two places (valuesObject.clusterand itsvalueFilespath); there’s no real multi-cluster wiring anywhere downstream yet to make a spec field meaningful. MatchesNodeJSApplication’s own precedent (platformOwner/tenantsRepoas fixed constants).prod) exists to design and test against:clusterbecomes a required spec field, gated via the new cluster registry (type: upper+crossplaneReady) — explicitly rejectingtype: devtargets, since §10 scopesgitops-<app-name>to upper environments only. See “Where Crossplane runs across a multi-cluster fleet” in §0 for the full reasoning (also covers the first-time-on-a-clusterapp.yaml-seeding gap this surfaces).- Initial
values.yamlis an identity-only stub,rollout: null— no real image exists to deploy at XR-creation time regardless of CICD control-plane availability (platform-cicdnow runs ondevas of 2026-08-15, but a real image only exists once a developer’s own push actually completes a real release through it). Rather than seed a placeholder image that would sit in permanentImagePullBackOff, the Composition reports a siblingWorkloadDeployed: Falsecustom condition — same mechanismNodeJSApplication’s ownCicdOnboardedcondition used before that one went real (see Item 1/2’s own status note), same accepted limitation (doesn’t auto-clear once a developer’s own follow-up PR sets a realrollout.image).
env opened up from a closed enum to team-chosen names, 2026-08-15. Was
enum: ["dev", "staging", "prod"] since the field’s introduction — traced every real
consumer (this Composition’s own template, the tenant-onboarding ApplicationSet,
the AppProject‘s own destinations wildcard app-<appName>-*) and confirmed
nothing branches on the specific value; it was always pure path/name interpolation
(the k8s namespace app-<appName>-<env>, git paths <cluster>/<env>/values.yaml and
tenants/<appName>/<env>/identity.yaml), never encoded business logic. Replaced the
enum with a pattern matching Kubernetes’ own DNS-1123 namespace-label rule plus
maxLength: 20, so a value Kubernetes would reject still fails at XR admission with a
clear message rather than downstream as an ArgoCD sync failure. Live-verified on
dev: a custom name (perf-test, previously impossible) reconciles end-to-end
for real; an invalid one (Staging!) is rejected at admission with the expected
pattern-mismatch error.
Built and live-verified for real 2026-08-15 (same day as the design above,
different session): cluster is now a real required field, the cluster registry
exists (gitops-cluster-dev/00-bootstrap/cluster-registry/), and prod was
bootstrapped as a real second cluster (gitops-cluster-prod, reusing its
pre-existing ArgoCD instance) specifically to prove both the rejection path
(crossplaneReady: "false" → ClusterReady: False, zero resources created) and the
success path (real commits, prod’s own ArgoCD picking up the new tenant on
its own, a real namespace/ServiceAccount/NetworkPolicy) end-to-end, not just in
design. One real bug found live and fixed: the cluster’s shared app.yaml can’t use
spec.deletionPolicy: Orphan as originally planned — provider-upjet-github
v0.19.1’s RepositoryFile CRD has no such field, confirmed via a real
ReconcileError (.spec.deletionPolicy: field not declared in schema) plus
kubectl explain. Fixed with managementPolicies excluding "Delete" instead, same
intent, correct field for the actually-installed CRD schema. See
idp-service-catalog’s own README and [[idp_session_applicationenvironment_xrd]]
follow-on memory for the full detail.
Also confirmed before building: SLO’s own Composition doesn’t actually implement
the “extra resources lookup” dependency-ordering mechanism described below — it’s
aspirational text, not built anywhere in this catalog yet. ApplicationEnvironment
doesn’t need to expose anything for that lookup as a result; it’s purely the write
side of the environmentRef contract SLO’s XRD and the chart’s own
environmentRef helper already fix as <appName>-<cluster>-<envName>.
Item 4: SLO (narrowed from SLA/SLO/SLI)
Start with one XRD, SLO, not three. SLI (the measured indicator — e.g. p99
latency) is naturally just a field within an SLO spec (which query defines success),
not an independent concept worth its own XRD. SLA (an external, often contractual
commitment, sometimes aggregating multiple SLOs with consequences attached) is a
reporting/business layer that can be built later, on top of SLOs that already exist —
nothing about building SLO first forecloses adding an SLAReport XRD afterward.
Superseded again, same day: switched to wrapping Sloth after all. The hand-rolled
revision immediately below was built and live-verified first (own PromQL/burn-rate
math, no Sloth dependency) - then, discussing it further, the user asked for a
pros/cons comparison and was persuaded back to wrapping Sloth, for two concrete
reasons: the hand-rolled version only implemented a 2-tier simplification of the SRE
workbook’s real 4-window pattern (page + ticket tiers each have a fast AND a slow
alert; the hand-rolled version only had one per severity), and “wrap an existing tool”
is this project’s convention everywhere else (Argo Rollouts, ESO, component charts) -
the hand-rolled version was the outlier, not the house style. Also built and
live-verified on dev, on a second pass after this: the XRD got SIMPLER, not
more complex, switching to Sloth - spec.window and spec.alerting.burnRates are
both gone, since Sloth computes the compliance period (a controller-wide default, not
per-SLO - confirmed against Sloth’s own CRD schema) and the full canonical burn-rate
pattern automatically from just objective. The Composition generates a Sloth
sloth.slok.dev/v1 PrometheusServiceLevel, not a PrometheusRule directly - Sloth’s
own controller (gitops-cluster-dev/10-crds-operators/sloth/) does that translation.
See idp-service-catalog/README.md’s Status section for the real bugs hit switching
(most notably: a second function-go-templating Function registration, added to keep
the SLO Composition’s templates isolated, corrupted Crossplane’s package-manager lock
graph for every OTHER Function on the cluster - fixed by using source: Inline
instead, which also meant no separate Function/mount was needed at all).
Superseded 2026-08-13: built hand-rolled, not wrapping Sloth (first pass, later superseded again above). The reasoning below (“don’t reinvent SLO-to-alerting-rule translation”) was the original lean, but the user explicitly chose the hand-rolled approach instead (inspired by a kube-slo-style article proposing exactly that), trading Sloth’s battle-tested math for one dependency-free artifact with no extra in-cluster controller. Original text kept below for the record:
Don’t reinvent SLO-to-alerting-rule translation — wrap an existing tool.
Sloth (or OpenSLO/Pyrra, same space) already solves “SLO spec →
multiwindow-multi-burn-rate Prometheus recording/alerting rules” correctly; the SLO
XRD’s Composition should generate whatever CR that tool consumes (or generate
PrometheusRule directly if going dependency-free), not hand-roll the burn-rate math in
a Composition Function.
Attached-tier, per the revised mechanism (§ Framework): a slos: block in that
env’s own values.yaml (an app can have different SLOs for staging vs. prod, since each
env has its own file) — the idp-application chart renders one SLO Claim per entry,
auto-stamped with environmentRef. Same namespace, same AppProject, no new credential,
no separately-committed file. Structurally upper-env-only in practice (nothing stops a
lower-env slos: entry from existing, but it’s a meaningless concept on an ephemeral
test namespace) — worth a Kyverno policy in the cluster-config policy layer rejecting it
there, rather than teaching the XRD itself about environment tiers.
Item 5: config/secrets — recommend no standalone XRD
Static, developer-authored config: plain Helm values on the Application chart, not a
separate XR. A ConfigMap has no independent provisioning lifecycle — it’s data
1:1-coupled to the Deployment that reads it, and it churns with ordinary app iteration
far more often than infrastructure topology does. Routing every config edit through a
full Crossplane reconcile, a separate Backstage catalog entry, and the linking
machinery is ceremony with no matching payoff. Mechanism: the existing per-env
values.yaml (upper) / platform/envs/*.yaml (lower) carries config fields directly,
rendered into a ConfigMap by the same chart that renders the Deployment.
Secrets: reuse the ESO/ExternalSecret consumption pattern platform-cicd already
built, with Infisical replacing the backend — confirmed this round; see the revised
Item 8 below for the backend/SecretStore design. The developer-facing half is
unchanged: the secrets: Embedded-tier list in §3, one ExternalSecret per app,
consumed by the workload exactly as before. Only what’s behind the SecretStore
object changes.
Dynamic/derived config (e.g. a database’s connection string) is not a ConfigMap
concern at all — it’s an output of whatever XRD provisions the underlying dynamic
resource (a future Database XRD’s own Composition writes its own connection-info
Secret as one of its managed resources). No separate ConfigMap XRD is needed even for
this case.
Item 6: HPA — mostly embedded, but names the real dividing line
Recommend: an optional field block on the Application XR’s spec (spec.autoscaling: {enabled, min, max, targetCPUPercent}), same mechanism as ConfigMap — not a separate
XRD. Reasoning: HPA is still 1:1 with the single Deployment the Application XR
already owns; there’s no independent lifecycle or cross-app sharing that a separate
XR would buy. The “genuinely optional, added later, real operational stakes” concern
you raised is fully addressed by it being an optional field with a safe default
(disabled) that a developer turns on via a normal PR to their own values file — same
review bar as any other config change in that repo, no new XR needed. Concrete
schema: autoscaling/podDisruptionBudget in §3.
The actual dividing line this surfaces, worth stating explicitly since it’ll keep coming up as the catalog grows: “1:1-coupled, in-namespace, no independent provisioning lifecycle” → embed as a field (ConfigMap, HPA, PodDisruptionBudget, resource limits). “Independent provisioning lifecycle, real external footprint” → separate Attached-tier XRD (SLO; Components — Redis, OAuth server — immediately below).
Item 7: Components — Redis, an OAuth server, and others (new, from this round)
Real infrastructure with its own provisioning lifecycle, not embeddable — squarely
Attached-tier, using the mechanism above: a components: block entry in some env’s
values.yaml, rendered by the idp-application chart into an XR (Redis,
OAuthServer, …), reconciled locally by that cluster’s Crossplane.
Yes to your question 1 — these are real, independent XRDs, so they’re independently
selectable Backstage templates, not just a hand-edited values.yaml block. Nothing about
being Attached-tier prevents that; it only determines the mechanism behind the
template action. “Add Redis” in Backstage means the same thing “New NodeJS Application”
now means (§0): the action commits — here, appending an entry to the target env’s
components: list in gitops-<app-name>, via the same scoped GitHub token mechanism —
rather than a live API call either way. Bootstrap and Attached tiers were never
different in whether Backstage can offer them as templates, only in which git
location the resulting commit lands in.
Each wraps a real upstream Helm chart via provider-helm’s Release resource,
confirmed as the right mechanism — not hand-assembled Deployment/Service/PVC
manifests in a Composition. Recommend the platform maintain a thin wrapper chart per
component type (idp-service-catalog/charts/component-redis, wrapping Bitnami’s Redis
chart or similar), not a bare pass-through to an arbitrary upstream chart — same
curated-golden-path reasoning that already won for #1/#2 over a generic
discriminated-union XRD: a Redis XRD with a real schema (size: small|medium|large,
persistence: bool) gives a legible Backstage form and platform-controlled defaults; a
generic HelmComponent XRD (chartRepo/chartName/arbitrary values: map) is more
flexible but gives up both the guardrails and the self-service simplicity goal 1/6 both
ask for. Worth having the generic escape hatch too, later, for the case the curated
catalog doesn’t cover yet — not designed now, same pattern as idp-cluster-baseline’s
extraManifests: escape hatch in gitops-strategy.md §8.
Standalone deployment (“its own namespace”) reuses appType: infra — not a new
concept. platform-cicd’s docs/naming-conventions.md already defines infra as
“a shared/platform-adjacent service onboarded with its own pipeline” alongside app —
a standalone Redis or OAuth server is exactly that, unchanged. Mechanically: it goes
through the same ApplicationEnvironment Bootstrap flow as any real app (its own
gitops-infra-<name> repo, its own env directory), but that env’s values.yaml
contains only a components: block with itself as the sole entry — no rollout:
block, no custom workload. This is the open question flagged in §3: whether the chart
needs to tolerate an absent rollout: block. If yes (recommended, for mechanism reuse), “attached to
an app” and “standalone” are the same XRD, the same chart, the same Bootstrap
flow — the only variable is which env’s values.yaml the components: entry ends up
in, an app’s own or a dedicated infra one.
Item 8: SecretStore — Infisical-backed, self-hosted (resolved this round)
Status (2026-09-18): fully built and live. The “mechanism not yet confirmed” question two paragraphs down is resolved —
github.com/jfillman/provider-infisical(upjet-generated from the officialInfisical/infisicalTerraform provider) is real and live. Every app-tierSecretStoreon the dev cluster (the kubernetes-auth branch: the cluster that hosts Infisical itself) andplatform-cicd-control-plane’s own project are cut over to it; the hand-rolledinfisical-secretstore-operatorthis section’s “Mechanism not yet confirmed” paragraph anticipated as a possible fallback was decommissioned entirely on the dev cluster the same day (CRDs, RBAC, Deployment all removed — zeroInfisicalProject/InfisicalEnvironmentCRs remain). Universal-auth clusters (prod, mgmt — every cluster that ISN’T the Infisical host) deliberately still run their own separate copy of the old operator: a real, unfixed-upstream bug interraform-provider-infisical’s Go SDK (IdentityUniversalAuthClientSecret’s response unmarshal) blocks that auth path on the new provider. See the Round 2026-09-18 section below for the full build/cutover record and what it means for Item 7’s own provider-per-backend plan.
Same platform-infra-vs-per-app-capability split the Vault discussion in the first
pass already reasoned through — Infisical just answers which concrete backend.
Infisical itself (the community edition, self-hosted on-prem) is shared platform
infrastructure, not a service-catalog item — installed once, cluster-admin-owned, in
idp-cluster-baseline’s 10-crds-operators/ group alongside ESO’s own controller, the
same way the earlier hypothetical Vault install would have been. A SecretStore XRD
provisions per-app isolation within that already-running shared instance — it does
not stand up Infisical itself.
Scope: one SecretStore per (app, cluster) — not per (app, cluster, env) — sharing
secrets across envs on the same cluster, never across clusters. Your addition this
round, and it’s a real, deliberate narrowing, not just an implementation detail: the
store itself is what draws the sharing boundary, and that boundary can’t cross a
cluster line for a structural reason, not a policy one — a SecretStore-family object
is always local to one cluster’s API, so “shared across envs, scoped to one cluster” was
always the widest this could go.
Correction to the first pass, caught by this narrowing: I’d recommended a
namespaced SecretStore last round, reasoning that Infisical’s own project isolation
made the cluster-scoped ClusterSecretStore (platform-cicd’s current pattern)
unnecessary. That reasoning silently assumed one store per namespace — it breaks under
cross-env sharing: ESO’s namespaced SecretStore can only be referenced by an
ExternalSecret in the same namespace; sharing across app-<name>-dev/
-staging/-prod (different namespaces, same cluster) needs ClusterSecretStore,
full stop, regardless of backend isolation. Reverted: ClusterSecretStore, matching
what platform-cicd already does — restricted to exactly this app’s own namespaces on
this cluster via ESO’s spec.conditions namespace selector, so “cluster-scoped” doesn’t
mean “visible to every app,” just “referenceable from more than one of this app’s own
namespaces.” Worth flagging plainly: this is me catching my own earlier recommendation
being wrong once new information changed the tradeoff, the same thing this platform’s
own history already has a habit of doing openly rather than quietly — see
platform_cicd_session_argocd_onboarding’s tracked-copy → live-read reversal for the
precedent.
Infisical-side structure that operationalizes “shared, scoped to a cluster”: one
Infisical project per (app, cluster), with sub-paths inside it — /dev/*,
/staging/*, /prod/* for env-specific secrets, plus a /shared/* path for the
ones meant to be reused. The ClusterSecretStore connects to the project; which
path a given env actually reads is controlled by that env’s own ExternalSecret
(remoteRef), not by the store — so “shared” is opt-in per secret, not automatic
exposure of everything to every env.
Ownership: the first ApplicationEnvironment for a given (app, cluster) pair creates
it; later ones for the same pair reference the existing one. This is the same
“multiple Claims can’t cleanly co-own one shared resource” problem already reasoned
through for who creates gitops-<app-name> (§ Item 1/2) — resolved the same way: an
“extra resources” existence lookup (the identical mechanism §“Dependency ordering”
already uses for environmentRef waits) checks whether a ClusterSecretStore named
<app>-<cluster> already exists before deciding whether this env’s
ApplicationEnvironment Composition also needs to create one.
The XRD’s remaining jobs, matching what you described:
- Configure the Infisical backend — create the per-(app,cluster) project.
Mechanism not yet confirmed: a native Crossplane
provider-infisicalmay or may not exist with the maturity this needs — worth verifying before committing, rather than assuming. If it doesn’t,provider-terraformwrapping Infisical’s own (real, actively maintained) Terraform provider is the fallback, same “wrap a real tool, don’t reinvent it” instinct already applied to Sloth for SLOs. - Create the
ClusterSecretStorepointing at that project,spec.conditionsrestricted to this app’s own namespaces on this cluster. ExternalSecretcreation stays with the existing Embedded-tiersecrets:mechanism (§3), not this XRD — unchanged reasoning from last round: which keys an env actually pulls (and from which path) is ordinary, fast-churning config, not this XRD’s concern.
Linking field deliberately different from every other Attached-tier resource: not
environmentRef (which names one specific env) — spec.appRef: {name: <app>} +
spec.cluster: <cluster-name>, since this resource explicitly must not pin to one
env. A second, now-explicit exception to the general Attached-tier pattern, alongside
“auto-provisioned rather than developer-selected” from last round — both worth keeping
visible as named exceptions rather than quietly special-cased.
Proposed v1 catalog
| XRD | Tier | XR scope | Creates |
|---|---|---|---|
NodeJSApplication |
Bootstrap | one per app | src repo, boilerplate, empty gitops-<app-name>, dev-cluster CICD onboarding |
SpringBootApplication |
Bootstrap | one per app | same, Java/Spring-specific |
ApplicationEnvironment |
Bootstrap | one per (app, cluster, env) | env’s values.yaml in gitops-<app-name>, onboarding entry in gitops-cluster-<cluster>-tenants |
SLO |
Attached | one slos: entry per (app, cluster, env) |
Sloth/PrometheusRule, rendered by the chart from that env’s values.yaml |
Redis |
Attached | one components: entry, in an app’s own env or a dedicated infra-type one |
provider-helm Release of a platform-wrapped Redis chart |
OAuthServer |
Attached | same as Redis |
provider-helm Release of a platform-wrapped OAuth/identity chart |
Database |
Attached | same as Redis |
provider-helm Release of a platform-wrapped Postgres (or similar) chart |
Queue |
Attached | same as Redis |
provider-helm Release of a platform-wrapped queue (e.g. RabbitMQ) chart |
SecretStore |
Attached, auto-provisioned | one per (app, cluster) — shared across that app’s envs on the same cluster — created by the first ApplicationEnvironment for that pair, never developer-selected |
Infisical project + a ClusterSecretStore scoped to this app’s namespaces — see item 8 |
ConfigMap and HPA are deliberately not on this list — they’re values.yaml fields,
§3 has the schema.
Open questions for discussion
Resolved this round:
- Crossplane v2 confirmed as the target (Terminology section) — namespaced XRs directly, no separate Claim type.
- Zero K8s write credentials for Backstage, anywhere (§0, revised) — every XR,
Bootstrap or Attached, is a git commit; Bootstrap-tier lands in
gitops-cluster-<cluster>-tenants/<app>/xr-requests/, synced intoapp-<name>-cicdby the same Application that already renders that namespace (sync-wave ordering, not a live namespace-creation call). Backstage authenticates only totoken-review-interceptor’s existing/github-installation-tokenendpoint via a K8s ServiceAccount token that carries no resource-write RBAC at all. - Attached-tier XRDs (Redis,
OAuthServer,SLO) are independently selectable Backstage templates, same as Bootstrap-tier — confirmed, see item 7. Database/Queueadded to the v1 catalog, same Component pattern asRedis.- Secret vault question resolved to Infisical, self-hosted — see the revised item 8.
- ArgoCD’s
Rollouthealth check verified (below) — real, but not guaranteed bundled; needs explicitargocd-cmconfiguration, not an assumption. ClusterAnalysisTemplateconfirmed as the primary path — custom app-specificanalysisTemplates:entries should be the rare exception, not a co-equal option.
ArgoCD Rollout health check — verified, with a real caveat: ArgoCD’s own repo
ships a Lua health script for argoproj.io/Rollout
(resource_customizations/argoproj.io/Rollout/health.lua) that correctly distinguishes
Healthy/Progressing/Degraded/Suspended — confirmed by fetching it directly.
But this is not necessarily compiled into every ArgoCD version’s binary the way
Deployment/StatefulSet/DaemonSet health checks are — Argo Rollouts is a separate
project from Argo CD itself, and current guidance is to explicitly add this script to
argocd-cm under resource.customizations.health.argoproj.io_Rollout rather than
assume it’s already active. Concrete action, not just a caveat: idp-cluster-baseline
should carry this argocd-cm entry explicitly, alongside installing the Argo Rollouts
controller itself — belongs in cluster config, not something to leave to chance per
cluster.
Sources: Argo CD Resource Health docs, How to Configure Health Checks for Argo Rollouts in ArgoCD
Resolved 2026-08-13: idp-application’s rollout: block is genuinely optional
(§3, chart built) — a standalone infra-type component (§ Components) reuses the exact
same chart, no second sibling chart needed.
Resolved 2026-08-17: Item 8 built and live-verified end-to-end
(idp-service-catalog’s xrds/secretstore.yaml +
compositions/secretstore/ + operators/infisical-secretstore-operator/).
Neither a native provider-infisical nor provider-terraform was used in the
end — no mature native provider exists (only an open feature request,
Infisical/infisical#3240), and a small purpose-built kopf/Python controller
against Infisical’s real REST API turned out to be the better fit than
Terraform’s state/HCL overhead, matching this project’s existing convention of
small first-party Python controllers over heavier general-purpose tools (see the
operator’s own README for the full “Q1” reasoning). Real chain proven live: a
SecretStore XR → a real Infisical project + environment + machine identity +
Universal Auth credentials → a Ready: True ClusterSecretStore (wrapped in a
provider-kubernetes Object - Crossplane v2 rejects composing a cluster-scoped
resource directly from this namespaced XR, same fix NodeJSApplication’s
provider-github already needed) → a real secret pulled through into a real K8s
Secret. Still not built: wiring into ApplicationEnvironment’s
auto-provisioning (create-on-first-env-for-a-cluster, reference on later ones) -
this XRD is standalone-creatable only for now, same as SLO before any
Attached-tier auto-provisioning existed.
Multi-cluster revision, resolved and built 2026-08-17 (separate session, same
day). Two real, connected pieces of follow-on work, prompted by three questions
about how this should actually operate: does onboarding auto-provision the store,
does a second cluster get real per-environment isolation (not just a
secretsPath convention), and how should secrets actually be organized inside
Infisical.
Auth method, per cluster, not universal. Kubernetes Auth (above) only works
because Infisical and the workload calling it are the same cluster - Infisical
only ever runs on the dev cluster. A ClusterSecretStore on any OTHER cluster (prod
today) would mean Infisical calling TokenReview against a DIFFERENT cluster’s API,
which Infisical CE can only do via Gateway mode - Enterprise-only. Two real options
weighed with you: a second Infisical instance per cluster (keeps Kubernetes Auth
everywhere, zero persisted credentials anywhere, real infra cost per cluster) vs.
Universal Auth on non-dev clusters only (one shared instance, cheaper, but a
real persisted clientId/clientSecret on upper-env clusters specifically - the
higher-stakes environments). You chose Universal Auth for upper clusters -
InfisicalProject gained spec.authMethod (kubernetes | universal), set by
the SecretStore Composition from a plain string comparison on spec.cluster
("dev" → kubernetes, else → universal) - a static topology fact, not an
ExtraResources cluster-registry lookup like ApplicationEnvironment’s own gate.
Secrets organization: environmentSlug gained a real second mode, not a new
field - "shared" (the default, original behavior, one project/identity/store
per (app, cluster)) or any other value, which creates ONE additional Infisical
environment inside the ALREADY-existing project plus a SEPARATE
ClusterSecretStore narrowed to exactly one namespace
(^app-<appName>-<environmentSlug>$). Two options short of this were considered
and rejected: keeping the single-shared-environment secretsPath convention
(what item 8 originally shipped) doesn’t give real isolation - nothing stops an
ExternalSecret from pointing its own remoteRef at another env’s path, since
ESO’s Infisical provider pins auth/scope at the STORE level, not per secret; a
separate Infisical PROJECT per (app, cluster, env) (instead of environments
within one project) was rejected as needless multiplication - it buys no more
isolation than the per-environment-store design already delivers (both ultimately
reduce to “which environmentSlug can this store query”), while losing “shared”
as a natural concept. Project slug is now computed from spec fields
(<appRef.name>-<cluster>), not derived from either XR’s own metadata.name -
needed so a per-environment XR (a different object, different name) can compute
the identical slug its “shared” sibling already used, with zero lookups.
Auto-provisioning mechanism - NOT ApplicationEnvironment’s own Composition
directly, a real correction of the original plan. Both NodeJSApplication and
ApplicationEnvironment are provider-github-only (§0/§1) - neither can create a
native Kubernetes resource on ANY cluster, including the dev cluster’s own. The real
mechanism is the same one every other Attached-tier resource already uses:
idp-application’s own chart (templates/attached/secretstore.yaml) renders the
SecretStore XR - unconditionally, "shared" mode, into a dedicated
app-<appName>-secrets namespace (redundant-but-harmless idempotent writes from
every env release, same convention cluster-app-yaml.yaml already established -
needs a namespace that doesn’t vary per env for this to actually be one shared
object, not N different ones), plus, on any cluster other than the dev cluster, this
env’s own non-"shared"-mode XR too. SecretStore itself (and its own
InfisicalProject) therefore needed installing on the prod cluster for the first
time, alongside a new, small InfisicalEnvironment CRD/reconcile loop in the
same operator (ensures one environment exists in an already-existing project;
never creates a project itself).
Genuinely cross-cluster, proven live, not simulated: a real secret written
into Infisical on the dev cluster, pulled by a real ExternalSecret on the prod cluster, over
a live-verified NodePort path from the prod cluster’s own node to the dev cluster’s
(infisical-nodeport.yaml) - both clusters happen to share one L3 network in
the home lab, confirmed live (a real 403 from the dev cluster’s own API server proved
TCP reachability before ever exposing Infisical), a real but explicitly
lab-specific stand-in for what a routable endpoint between genuinely
separate clusters would be. Isolation proven live too: a correctly-namespaced
ExternalSecret pulls the right value; a wrong-namespace one hard-fails
(could not get secret data from provider), not just an authz warning. Five real
bugs found live along the way (beyond the four in the Kubernetes Auth swap,
already recorded above): ensure_project_membership genuinely wasn’t idempotent
(a live 400 on retry, the first time that code path ever ran against real data);
_session’s own default Content-Type: application/json broke every empty-body
DELETE (Fastify rejects it as a 500, not the 400 it actually was); the chart’s
own {{appName}}-{{cluster}} naming pattern renders as a bare - under helm lint’s all-empty defaults, which YAML parses as an ambiguous block-sequence
indicator, not a plain scalar - fixed by quoting; Crossplane’s own core
controller had no RBAC for the new infisicalenvironments CRD, on either
cluster; and a real off-by-one in this operator’s own INFISICAL_API_URL for
the prod cluster’s copy (/api included when it shouldn’t be - main.py’s own client
already prefixes every path with /api/v1/...), caught live as a real
/api/api/v1/... 404.
Still open: the org-level INFISICAL_ADMIN_TOKEN now has real blast radius on
two clusters, not one - an accepted, flagged cost of one shared instance, not
solved further here. A genuinely separate, real upper-env cluster (not sharing
the lab’s network with the dev cluster) would need a real routable endpoint in
place of the NodePort stand-in - not designed here, flagged as the thing to revisit
first if this pattern is ever extended beyond this local sandbox.
Third revision, same day, real correction of the above: the trigger moved from
idp-application’s chart to the Bootstrap XRs themselves. Your objection, and a
correct one: “the first deployed Helm chart sets things up” meant secrets
infrastructure was an emergent side effect of someone shipping a real release, not
a real provisioning step - not discoverable from the Bootstrap XRD list at all,
and not created until long after an app might actually need it. Considered
composing the SecretStore XR directly from NodeJSApplication/
ApplicationEnvironment’s own Compositions first - real nuance worth recording:
NodeJSApplication structurally could (its Composition already runs on
the dev cluster, exactly where the dev-cluster store needs to live too), but
ApplicationEnvironment structurally can’t for an upper-cluster store (its
Composition also runs on the dev cluster - Bootstrap-tier’s centralization - but the
resource has to land on the target cluster, and provider-kubernetes remote
credentials would violate “no cluster ever holds credentials for another
cluster’s API,” the same constraint that already kept AppProject/Application
ownership out of direct Crossplane composition, §0). Real fix: both XRs commit a
SecretStore XR manifest via xr-requests instead - the exact mechanism every
other Bootstrap XR already goes through, just aimed at the target cluster’s own
tenants repo for ApplicationEnvironment’s copy. idp-application’s
attached/secretstore.yaml is deleted entirely, not just made conditional.
One real, necessary piece of new infrastructure this exposed: xr-requests was
dev-only before (Bootstrap-tier never ran anywhere else). The prod cluster now has
its own copy (gitops-cluster-prod/02-argocd-apps/xr-requests/), narrowly
scoped to just SecretStore in its AppProject - NodeJSApplication/
ApplicationEnvironment still correctly never run there.
Live-verified against real, already-onboarded apps (not fresh test fixtures) on
both clusters - updating the live Composition objects triggered
compositionUpdatePolicy: Automatic re-reconciliation of existing
NodeJSApplication/ApplicationEnvironment XRs for real, which committed the new
manifests, which xr-requests picked up and applied, entirely independent of
whether that app had ever shipped a release. Two real bugs found live in the
process, both from the previous revision’s chart-triggered mechanism having
already run for real apps before this fix landed (tag/pin bump + sync happened
in-between the two revisions) - not bugs in this design itself, but real
contamination it had to clean up: (1) the dev cluster’s own idp-onboarding AppProject
(xr-requests’ scoping project) didn’t whitelist SecretStore yet - “resource
catalog.idp.io:SecretStore is not permitted in project idp-onboarding”, same real
gap independently caught and fixed on the prod cluster’s copy already, missed on
the dev cluster’s because that file wasn’t touched building the first pass; (2) the
old chart-rendered SecretStore XRs (a duplicate, same deterministic project
slug, different Kubernetes namespace) had already created real Infisical
projects for checkout-api on both clusters - deleting them out from under the
new xr-requests-created XRs (same slug) left the new ones pointing at
now-deleted project/credential state, since Crossplane’s own per-namespace
Object resources don’t know about each other’s same-named remote target
colliding. Recovered live by deleting the stale namespaces/duplicate Objects
and forcing an operator resume (pod restart) - real cleanup, not a design flaw,
but a genuine hazard worth naming for next time two provisioning mechanisms
briefly coexist during a mid-flight redesign like this one.
Also fixed the same day, a separate but related real bug: idp-application’s own
external-secret.yaml never got rewired when the per-environment stores were
built in the second revision above - every secret, shared: or not, still
resolved through the one wide-open shared store via a path-prefix convention.
The per-environment stores existed and were correctly isolated (proven by
hand-written ExternalSecrets during that build) but nothing developer-facing
ever used them. Fixed via ESO’s real, confirmed (not guessed)
ExternalSecretData.sourceRef.storeRef per-entry override: a non-shared secret
on any cluster but dev now targets its own per-environment store directly
(no path prefix needed - the store itself is already scoped to that one
environment); the dev cluster, which has no per-environment store at all by design,
keeps the original path-prefix convention unchanged.
One more real naming bug, caught by user review, fixed the same day: the prod cluster’s
new xr-requests (above) originally reused the dev cluster’s app-<appName>-cicd
destination namespace, copy-pasted for mechanical consistency without weighing
what the name actually claims. On the dev cluster that name is accurate (the real CICD
control plane also lives there); prod never runs anything CICD-related, so
reusing it there falsely implied it did. Renamed to app-<appName>-xrs -
describes what’s actually there and generalizes to any future XR kind this
mechanism ever carries on an upper cluster, not just SecretStore. Migrated the
one real app that had live state under the old name (checkout-api) the same
live-delete-and-let-the-operator-recreate way as the mid-migration cleanup above.
Real bug found and fixed 2026-08-18: InfisicalProject/InfisicalEnvironment
never reported readiness Crossplane could see, so every SecretStore (and any
Bootstrap XR waiting on one, e.g. NodeJSApplication’s own src-repo-adjacent
readiness gate) stayed stuck Creating forever even once actually provisioned.
Root cause: the operator (main.py) only ever wrote its own status.phase: Ready
convention - the SecretStore Composition’s function-auto-ready pipeline step
(same mechanism every other Composition in this catalog uses) only ever checks
the standard status.conditions[type=Ready].status=="True" shape, which neither
CRD’s schema even had a field for. Fixed by adding status.conditions to both
CRDs and having reconcile()/reconcile_environment() write a real Ready: True
condition on success (ready_condition() helper) - live-verified on both
the dev cluster (checkout-api-dev SecretStore, Kubernetes Auth) and prod
(all three checkout-api-prod* SecretStores, Universal Auth) by rebuilding
the operator image, applying the updated CRDs, and restarting the Deployment on
both clusters; checkout-api-xr-requests/nodejs-demo-app-xr-requests ArgoCD
Applications went Healthy immediately after, no other change needed.
A separate, unrelated bug surfaced investigating the same live symptom
(checkout-api-xr-requests stuck Progressing): three provider-github
Repository managed resources (checkout-api-src, nodejs-demo-app-src,
nodejs-demo-app-gitops) had lost their crossplane.io/external-name annotation
- the provider pod’s creation timestamp lined up almost exactly with each
resource’s
external-create-succeededtimestamp, consistent with a provider restart racing the post-create annotation write. Each kept retryingPOST /user/reposon every reconcile and hitting a real422 name already exists on this account, since the GitHub repo genuinely already existed (confirmed live viagh repo view) but Crossplane no longer knew that. Recovered by manually restoring each resource’sexternal-nameannotation to match itsspec.forProvider.name- not a code bug, an operational data-loss incident this catalog has no automated recovery for yet, flagged here rather than silently patched away. Also newly visible on the same three repos once this unblocked their next reconcile, and fixed the same day once diagnosed: a live422 Secret scanning is not available for this repositoryon everyPATCH. Root cause confirmed live via the GitHub API directly (gh api), not guessed:security_and_analysisisn’t set anywhere in this catalog’s own Compositions -terraform-provider- github’s own schema defaultssecret_scanning/secret_scanning_push_protectionto"enabled"when the block is left entirely unset, and upjet late-inits that default intospec.forProvideras real desired state on first reconcile. This account’s private repos have no GHAS entitlement, so GitHub silently creates the repo with it already “on” but rejects an explicit PATCH trying to (re-)assert"enabled"- only an explicit"disabled"PATCH is accepted (confirmed by hand against the real API before touching any code). Fixed by renderingsecurityAndAnalysisexplicitly (disabled) on bothsrc-repo.yamlandgitops-repo.yaml(compositions/nodejsapplication/templates/render-github- resources/), taggedv0.3.24, rolled out via the existing git-tag-pinned Application on both clusters (forced a real ArgoCD app-of-apps refresh + manual sync, ending up withCompositionrevision 15) - live-verified all fourRepositoryresources across both apps (checkout-api,nodejs-demo-app)Synced: Trueagain.
Still open:
- What the platform’s default canary step sequence should be (§ Argo Rollouts) —
idp-applicationshould ship a sensible default (weights, pauses, whichClusterAnalysisTemplate(s) attach automatically) so most apps never need to specifyrollout.stepsat all, matching the “golden path, fully automated” goal, now sharpened by your confirmation that customanalysisTemplates:should be rare — not designed here. The chart (built 2026-08-13) ships a deliberately inert single-step placeholder in the meantime — seecharts/idp-application/README.md. - The chart-architecture question raised this round (one
idp-applicationchart vs. splitting pieces like ConfigMap out) — addressed in chat, not yet folded into a doc revision pending your read. (Built as one chart, per the original leaning — worth confirming this is still the intended resolution now that it’s real code, not just the leaning.) - The actual Crossplane API group for
components:/slos:XRs — not fixed anywhere yet (none of those XRDs exist). The chart uses a placeholder,catalog.idp.io/v1alpha1, one value to update once this is decided. - The real namespace/labels that identify the ingress controller (§3’s
networkPolicy.allowIngressFromIngressController) — no ingress controller has been installed or named anywhere in idp’s docs yet. The chart defaults to theingress-nginxproject’s conventional namespace label as a placeholder.
Round 2026-09-17: Item 7 goes from design to build plan — Postgres/Redis/OAuth/RabbitMQ/MongoDB/nginx, and a correction to how Attached-tier stateful services get provisioned
Prompted by a real ask: add six new Component types (postgresql, redis,
oauth-server, rabbitmq, mongodb, http-proxy/nginx), each installable either as
components: entry inside an existing app’s own env (its own namespace, alongside that
app) or as a standalone shared service (its own dedicated namespace, managed the same
way in Tower). Both modes are already named in Item 7 (appType: infra reusing the same
XRD/chart/flow) — this round is mostly “build the thing already designed,” but two real
gaps surfaced doing it for real, one structural (stateful multi-tenant sharing) and one
almost-repeated mistake (reaching for a custom controller out of habit rather than
re-checking the provider landscape for each new backend).
Status check: how much of Item 7 actually exists
Re-verified against the live repos rather than assumed from the doc:
- Zero Component XRDs exist.
airframe/xrds/hasapplicationenvironment,goapplication,nodejsapplication,pythonapplication,rolloutwatch,secretstore,slo,springbootapplication,tektoncicd— noredis,database,oauth-server, orqueue. The chart’s generic renderer (airframe-application/templates/attached/components.yaml) and itscomponentKindslookup table (redis/oauth-server/database/queue) are real and already built — they’ve just never had a real XRD to point at. provider-helmis not installed anywhere.apron’s Crossplane install (10-crds-operators/crossplane/) hasprovider-kubernetesandprovider-github(each with its ownDeploymentRuntimeConfig—provider-github-runtime.yaml— a pattern worth repeating, see below) but noprovider-helm. Item 7’s “wrap an upstream chart viaprovider-helm’sRelease” plan is unbuilt infrastructure, not just missing XRDs.appType: infrahas never been used for real. It’s fixture-tested in the chart’s own test suite (aninfra-type release with only acomponents:block renders cleanly, no Rollout/Service/HPA/PDB) but nogitops-infra-*repo has ever been created through the real onboarding flow — “standalone shared” is a proven chart capability, not a proven operational path.- Tower has no UI for any of this.
ConfigTab.tsx/useConfigData.tscover ConfigMaps, Secrets, Rollout, Scaling, Resources, Service ports, health checks, cron jobs — nocomponents:section exists on either the frontend or (unconfirmed, needs checking before scoping) the backendconfig/schemaroute.
What the outstanding provider-github bug actually teaches for this build
The stuck-Ready/403-misread bug is
provider-upjet-github-specific and only touches Bootstrap-tier git-commit resources —
it doesn’t block Component XRDs directly. But the live-ops lessons from surviving it
generalize directly, at higher stakes, because these six services carry real persistent
data where a RepositoryFile mostly didn’t matter beyond its own content:
- A provider’s “does this resource still exist” logic can be wrong, and you don’t get to assume it isn’t — that misclassification is the entire root cause of the github bug’s data loss. Before trusting any new provider with a stateful resource, check how it distinguishes a transient API error from “really gone,” don’t assume correctness by lineage.
- Shared
DeploymentRuntimeConfigblast radius is real (a--max-reconcile-rateflag on the fleet’sdefaultconfig once crash-looped every Composition Function cluster-wide for ~20 minutes).provider-githubalready gets its own dedicated runtime config — every new provider this round (provider-sql,provider-rabbitmq,provider-keycloak,provider-helm) needs the same treatment from day one, never inheritingdefault. - A provider restart forces every resource it manages through a simultaneous
cold-start re-Observe — for
RepositoryFilethat turned a 10-resource problem into 46 and is what actually exposed the data-loss bug. Aprovider-helm/provider-sqlrestart doing the same to every tenant’s Postgres/Mongo/RabbitMQ Release simultaneously is a materially scarier event if any of them ever resolves drift via delete+recreate instead of upgrade-in-place. - Delete+recreate as a live “fix” is not durable and must never be reached for on a
stateful component — already proven not to hold even for a stateless resource.
Concrete guardrail: every PVC-backed component (Postgres, Mongo, RabbitMQ; Redis
too if persistence is enabled) renders with
managementPoliciesexcludingDelete, the same fix already applied tovalues-yaml.yaml(7769ad9) for the identical reason — and an explicitRetain-equivalent reclaim policy, since no StorageClass inapronsets one explicitly today. - A published git tag is not an installable package — verify against the actual
registry (
xpkg.upbound.ioor wherever a given provider is distributed) before pinning any of the four new providers below, same as theprovider-upjet-github v0.20.0trap.
Correction: not every Attached-tier stateful service should be “one Helm Release per XR”
Item 7’s original design treats every Component the same way — one provider-helm
Release per XR. That’s right for Redis-as-a-cache and for a solo, per-app
Postgres/Mongo/RabbitMQ. It’s wrong for the “standalone shared” case this round
explicitly asks for: five apps sharing one Postgres shouldn’t mean five Helm releases
pretending to be one shared thing — it means one physical instance, N logical
tenants (database/vhost/realm + credentials per consuming app). This project has
already built and live-verified exactly that shape once, for SecretStore/Infisical —
worth generalizing rather than re-deriving per service:
- Dedicated mode (its own instance, in an app’s own namespace or a solo infra
namespace):
provider-helmRelease, one per XR. Fine for Redis, fine for a standalone Postgres/Mongo/RabbitMQ one app doesn’t want to share. - Shared-tenant mode (attach to an already-running shared instance): provisions a
logical database/vhost/realm + credentials inside it,
writeConnectionSecretToRefhanding the consuming app a Secret — no new instance rendered. Acomponents:entry needs amode: create | attach(or a distincttype, e.g.postgres-tenantvs.postgres-standalone) referencing the shared instance by name.
Correction 2: the reflex to write a custom controller for shared-tenant mode was almost repeated without re-checking — don’t do that
The first draft of this round’s recommendation proposed a small kopf/Python operator
for shared-tenant Postgres/Mongo/RabbitMQ/OAuth, pattern-matching directly off the
SecretStore/Infisical build. That was the wrong move, caught before building
anything: the Infisical operator exists because no mature native Crossplane provider
covered Infisical at all — that was the justified exception, not a template to
reapply to every new backend without checking again. The Infisical operator’s own bug
history is direct, first-party evidence of what a hand-rolled controller costs relative
to a native provider: it never reported a real Ready condition Crossplane could see
(it only wrote its own status.phase: Ready convention, which function-auto-ready’s
standard status.conditions[type=Ready] check couldn’t read), and it needed RBAC
hand-wired for its own CRD that Crossplane’s core controller didn’t know about by
default. Both are exactly the class of thing crossplane-runtime’s generated
reconciler — the machinery every native provider is built on — gets right for free:
standard Ready/Synced conditions, retry/backoff, scheduled drift detection, and
native writeConnectionSecretToRef for exactly the “hand the consumer app a Secret”
job Shared-tenant mode needs.
The corrected principle: check per-backend whether a native Crossplane provider
already covers “manage a logical resource inside an already-running shared instance”
before writing anything custom; when nothing exists, prefer composing an existing,
actively-maintained upstream operator (the same “wrap a real tool, don’t reinvent it”
instinct already applied to Sloth for SLO) over writing new reconciliation logic from
scratch. A hand-rolled controller is the last resort, justified explicitly per service,
not a default reached for because it worked once before.
Checked against the real current landscape (not assumed) for each of the six services:
| Service | Native provider for shared-tenant mode? | Recommendation |
|---|---|---|
| PostgreSQL | Yes — crossplane-contrib/provider-sql (PostgresqlDatabase/PostgresqlRole/PostgresqlGrant; connects to an existing server via a connection secret, doesn’t provision the server itself — exactly this shape) |
Use it, no controller. Caveat, not hypothetical: pre-1.0 (v0.9.0), 39 open issues, with live open issues specifically about managementPolicies not being supported and grant-existence detection being unreliable (#206, #240) — verify both against the actual version pinned before trusting it with real tenant data, same discipline the provider-github incident demands generally. |
| RabbitMQ | Yes — pnowy/provider-rabbitmq (VHost/User/Permissions/Exchange/Queue, v2.0+, namespaced-resource support, listed on Upbound Marketplace) |
Use it, no controller. Single-maintainer community project, not a crossplane-contrib-org provider — lower trust bar than the others here, a deliberate call to make explicitly rather than a default. |
| OAuth (Keycloak) | Yes — crossplane-contrib/provider-keycloak (Realm/Client, actively maintained, v2.24 recent) |
Use it, no controller. Assumes Keycloak as the backend — if a different IdP is chosen, re-check this row, don’t assume it carries over. |
| MongoDB (self-hosted) | No. crossplane-contrib/provider-mongodbatlas only covers MongoDB Atlas (the managed cloud service) — nothing native covers a self-hosted mongod. Real exception, same shape as Infisical was. |
Still not a from-scratch controller: compose the upstream MongoDB Community Kubernetes Operator’s own CRDs via the already-installed provider-kubernetes, the same “wrap a real tool” pattern already used for Sloth/SLO. That operator manages users/databases natively against a self-hosted replica set; Crossplane’s job is orchestrating an existing, actively-maintained upstream operator, not reimplementing Mongo admin commands. |
| Redis | N/A — dedicated-cache use case, not a multi-tenant-within-one-instance sharing question the way the other four are | Unchanged from Item 7: provider-helm Release, one per XR. |
| nginx / http-proxy | N/A — same reasoning as Redis | provider-helm Release wrapping a thin platform chart. Scope question, not yet resolved: Item 7 already scopes the cluster ingress controller out of this chart entirely (“platform-shared infra… those live in cluster config, referenced by name, never templated here”). Confirm this Component is meant as a per-app/per-team reverse proxy in front of a couple of backend Services, not a second ingress layer or API gateway — if the latter, it’s a cluster-scoped concern and doesn’t belong in this catalog as a namespaced Component at all. |
Net new providers to install, each with its own DeploymentRuntimeConfig per the
lesson above, each version-pin verified against its actual registry before use:
provider-helm, provider-sql, provider-rabbitmq, provider-keycloak. MongoDB adds
no new provider — it adds a Composition that renders the MongoDB Community Operator’s
CRDs through the provider-kubernetes already installed.
A missing Bootstrap XRD: standalone shared infra has no source code to scaffold
Every Bootstrap-tier XRD today (NodeJSApplication, SpringBootApplication,
GoApplication, PythonApplication) exists to scaffold a src repo + boilerplate —
that’s their whole job. A standalone shared service (a shared Postgres, a shared
Keycloak) has no application code. Onboarding one today would mean creating a pointless
src repo just to get the gitops-infra-<name> repo and an ApplicationEnvironment to
hang a components: block off of. Proposing a fifth Bootstrap XRD, InfraService,
doing only what NodeJSApplication does minus the src-repo/boilerplate step: creates
the empty gitops-infra-<name> repo and the tenants/<name>/app.yaml commit.
Everything downstream — ApplicationEnvironment, the components: block, Tower — is
unchanged.
Consumer access: mostly already solved, just not wired up automatically
Connecting an app to a shared component needs two things: credentials (solved by
Shared-tenant mode’s writeConnectionSecretToRef above) and network access. The
NetworkPolicy mechanism for the latter already exists and is already live —
networkPolicy.allowIngressFrom (a {namespace, podLabels?, ports?} peer, §3) — so no
new field is needed, just automation: when a Shared-tenant component XR is created for
a consuming app, its Composition should also commit the reciprocal
allowIngressFrom entry into the shared instance’s own env’s values.yaml, rather
than leaving that as a manual two-sided edit two different teams have to remember to
keep in sync.
Tower/Backstage gaps to close alongside the XRDs
“Configurable from Tower” is a real, current gap on both new axes this round introduces:
- A
components:section inConfigTab.tsx/useConfigData.ts— a curated form per component type (size/persistence/version fields, matching the curated-golden-path principle already used forrollout:/autoscaling:/etc.), for both Dedicated and Shared-tenant (attach) modes. - For Shared-tenant instances specifically: a page showing which apps currently
consume a given shared component — the
environmentRef/sharedInstanceRef→ catalog-relation translation flagged as “still not built” since the very first draft of this doc, now with a real consumer this round to build it for. - Confirm (not yet checked) whether the backend
config/schemaroute already round-trips acomponents:field at all before scoping the frontend work.
Sequencing
- Install
provider-helmwith its ownDeploymentRuntimeConfig, version verified against its actual registry. - Build
InfraService— unblocks standalone shared infra for real. - Ship
redisfirst (Dedicated only,provider-helm) — lowest risk, proves the wrapped-chart pattern end to end and gives Tower’scomponents:UI a real target to build against. - Install
provider-sql(ownDeploymentRuntimeConfig) and shippostgresqlin Shared-tenant mode second — highest-value, highest-risk item (real data, real multi-tenant credential isolation), reusing theSecretStoreoperator’s pattern of proof (live-verified against real already-onboarded apps, not just fixtures) before it’s trusted beyond a throwaway app. Explicitly test themanagementPoliciesand grant-existence-detection gaps flagged above before relying on either. mongodb(via the Community Operator composition) andrabbitmq(viaprovider-rabbitmq) follow the same Shared-tenant shape once Postgres proves it.oauth-server(viaprovider-keycloak) andnginxlast — both need an explicit scope decision first (how often does a single app legitimately need its own IdP instance vs. always being Shared-tenant against one Keycloak; whethernginxrisks duplicating the ingress controller’s job) rather than being built on an assumption.
Open questions for this round
- Confirm the Dedicated-vs-Shared-tenant split above, especially for Postgres/Mongo/RabbitMQ — a real deviation from Item 7’s “every Component is a Helm Release” assumption.
- What
nginx/http-proxyis actually for — per-app reverse proxy (fits this catalog) vs. something closer to a gateway (doesn’t, per Item 7’s own cluster-shared-infra exclusion). - Is Keycloak the confirmed OAuth backend, or still open?
provider-keycloak’s applicability depends on it. provider-rabbitmq’s single-maintainer status - acceptable for this project’s risk tolerance, or worth the extra scrutiny (fork-and-vendor, or fall back to aprovider-helmRelease without shared-tenant support) before depending on it for anything beyond a lab?
Round 2026-09-18: SecretStore/Infisical fully built, live-cut-over, and the old operator decommissioned - the real-world proof for Item 7’s whole provider-first argument, and a live-ops playbook worth reusing verbatim
Everything the “Correction 2” section above argued for in the abstract (native
provider over custom controller, checked per-backend rather than assumed) is now a
completed, live-verified build, not a design position. Recorded here because Item 7’s
next phase (provider-sql/provider-rabbitmq/provider-keycloak, all real upjet
providers of exactly the same shape) will hit the same category of live-ops problems
this build already hit and solved - worth reusing the playbook, not rediscovering it.
What actually got built and shipped
github.com/jfillman/provider-infisical- upjet-generated from the officialInfisical/infisicalTerraform provider, 7 namespaced resources (Project,ProjectEnvironment,Identity,IdentityKubernetesAuth,IdentityUniversalAuth+IdentityUniversalAuthClientSecret,ProjectIdentity). Built, live-verified standalone, packaged, published toghcr.io/jfillman/provider-infisical.- All 7 real app-tier
SecretStoreXRs on the dev cluster (the kubernetes-auth branch - the cluster that hosts Infisical itself) cut over from the hand-rolledinfisical-secretstore-operator’sInfisicalProjectCR to this provider’s native resource chain, one app at a time, each verifiedReady: Truebefore moving to the next. platform-cicd-control-plane’s own Infisical project (a Helm-rendered resource, not one of the catalog’s own XRs) also migrated - by rendering aSecretStoreXR from inside the Helm chart rather than reinventing the resource chain in plain Helm templates (see “The plain-Helm-chart problem” below for why).- Every secret in every migrated project restored from a pre-migration backup
(both a local file export and live
-backupsibling Infisical projects, both verified secret-count-for-secret-count beforehand) and verified key-for-key, path-for-path against that backup after cutover - zero data loss across all 8 projects (7 apps + platform-cicd), despite the cutover being a real destroy-and-recreate of each Infisical project along the way (see “Crossplane does NOT auto-garbage-collect” below for why that was unavoidable). - The old operator fully decommissioned on the dev cluster: Deployment, RBAC, both CRDs
(
InfisicalProject/InfisicalEnvironment) deleted. Zero CRs of either kind remain anywhere on the cluster. The two genuinely shared objects it used to also carry (the Kubernetes token-reviewer ServiceAccount/Secret everyIdentityKubernetesAuthresource reads from, and the cross-cluster Infisical NodePort the prod cluster’s own separate operator instance still needs) were split into their own directory/Application first and adopted there viaServerSideApply- zero disruption, confirmed via unchanged object creation timestamps. - The prod and mgmt clusters each keep running their own separate copy of the old operator, unaffected and out of scope - their universal-auth branch deliberately still uses the old CR kind, blocked on a real, unfixed upstream bug (see below).
Live-ops problems hit and fixed - reusable playbook for the next provider
- A parse error in a Composition’s own explanatory comment took down every
already-migrated
SecretStoreXR the moment it synced (comment ends before closing delimiter) - a space between*/and the closing>>broke Go’stext/templatelexer, which requires the right delimiter to immediately follow a comment’s close with zero characters between. Root-caused fast by extracting the rendered template and parsing it with a tiny standalone Go program using the realtext/templatepackage (custom delimiters, stubFuncMap) - reproduced the exact error at the exact line offline, in seconds, instead of iterating against the live cluster. Reusable for any future Composition bug in this family: this catalog’sfunction-go-templatingCompositions are ordinary Go templates outside Crossplane’s own machinery; Go’s own parser is a free, accurate offline reproduction tool for anything shaped like “invalid function input: cannot parse the provided templates.” stringDatadoesn’t work for aprovider-kubernetesObjectmanaging aSecret. The API server convertsstringDatatodataserver-side and never persistsstringDataas a real field, soprovider-kubernetes’ managedFields-based Observe (which extracts only the fields it owns by walking the desired manifest’s own field paths) finds nothing at thestringDatapath and errors converting a nil result into a map. Fix: always usedatawith an explicitb64enc, neverstringData, for anyprovider-kubernetes-managedSecretgoing forward - applies identically to any future service composing a synthetic credentials Secret this way.- Crossplane does NOT auto-garbage-collect a composed resource just because it
fell out of a Composition’s new render set, contrary to this project’s own
working assumption going into the cutover. Confirmed live: a composed resource’s
reference disappeared from the XR’s own
spec.resourceRefsimmediately on the new Composition’s first render, but the actual object sat untouched for 8+ minutes across 20+ reconciles - the new resource kept failing on an “already exists” slug collision precisely because the old one hadn’t been removed. The real mechanism for this class of cutover is manual, one resource at a time: delete the old CR yourself (verified backup first), which frees whatever uniqueness constraint the new resource is colliding on, then let the new chain create cleanly on retry. Budget for this explicitly in any future provider-family cutover with a global uniqueness constraint (a slug, a name, a vhost) - it will not resolve itself. - A
Provider’s reconcile backoff can genuinely stretch past 10 minutes after repeated errors on one resource - waiting is not always the fastest path. Annotating the stuck resource with any new value (kubectl annotate ... touch=... --overwrite) forces an immediate requeue, bypassing the backoff timer entirely - a fast, safe, reusable trick for unsticking any single Crossplane managed resource without restarting the whole provider pod (which would force every other resource it manages through a simultaneous cold-start re-Observe, the exact fleet-wide-outage shape already documented in theprovider-githubincident above). - A default upjet-generated provider publishes an EMPTY
writeConnectionSecretToRefSecret unless a resource’sconfig.goexplicitly setsr.Sensitive.AdditionalConnectionDetailsFn- confirmed live (anIdentityresource’s connection Secret had zero keys). Needed once a consumer had no Crossplane Composition (a plain Helm chart, see below) and so no.observed.resourcesaccess to read a provider-assigned id back out any other way. Fixed with a ~10-lineconfig.gochange (attr["id"]→ a named key), rebuilt, republished, live-verified before use. Worth checking proactively forprovider-sql/provider-rabbitmq/provider-keycloak: any resource whose provider-assigned id or credential a plain (non-Composition) consumer will need should get this wired in from the start, not discovered as a blocker mid-migration the way it was here.
The plain-Helm-chart problem, and why it matters for InfraService (Item 7’s missing Bootstrap XRD)
platform-cicd-control-plane’s own Infisical project is rendered by a plain Helm
chart, not a Crossplane Composition - and that turned out to matter structurally,
not just cosmetically. A Composition Function gets live .observed.resources access
on every reconcile, so SecretStore’s own Composition can read a just-created
Identity’s provider-assigned id back out and stitch it into a synthetic Secret. A
plain Helm chart has no equivalent: helm template (what ArgoCD actually runs to
render a Helm source) has no live cluster access, so neither a lookup call nor any
hand-rolled equivalent can read a resource’s post-creation state at render time - and
baking a real credential into values.yaml/git was rejected outright, matching this
platform’s standing “never persist credentials” posture.
The fix here was two-pronged and worth generalizing: (a) the AdditionalConnectionDetailsFn
change above, so Crossplane’s own native connection-secret mechanism (evaluated by the
managed resource’s own controller at runtime, not by Helm at template time) carries
the value instead; and (b) for the one piece that mechanism still couldn’t reach
(IdentityKubernetesAuth’s literal JWT/CA-cert fields, with no SecretRef
alternative on that CRD), reusing the existing SecretStore XRD/Composition
wholesale from inside the Helm chart (rendering a SecretStore XR instead of
hand-rendering the resource chain) rather than inventing new plain-Helm machinery to
solve a problem Crossplane’s own Composition layer already solves.
This is the exact shape Item 7’s proposed InfraService Bootstrap XRD exists to
avoid needing per-service: any standalone shared infra component that isn’t a real
onboarded “app” still needs the full Composition-based machinery, not a
hand-rendered plain-Helm shortcut - platform-cicd-control-plane hit this problem
because it predates InfraService and is deliberately not modeled as an app (see
Item 8’s own header comment on why). Once InfraService exists, this exact class of
problem shouldn’t recur for postgresql/redis/oauth-server/rabbitmq/mongodb/
nginx’s own standalone-shared instances - they go through the real XRD/Composition
path from day one instead of a bespoke Helm-chart special case.
Handoff: Item 7’s build is now unblocked, next phase is unstarted
Nothing in Item 7’s own Sequencing (above) has been built yet - this round’s work
was entirely SecretStore/Infisical (Item 8, already scoped as “resolved” before this
round; this round just finished actually building and cutting it over live). Item 7’s
own six services (postgresql/redis/oauth-server/rabbitmq/mongodb/nginx)
remain exactly where the Round 2026-09-17 section left them: zero Component XRDs
exist, provider-helm is not installed, InfraService doesn’t exist, Tower has no UI
for any of it. The next session picking this up should start at Sequencing step 1
(provider-helm + its own DeploymentRuntimeConfig) and step 2 (InfraService),
then step 3 (redis, Dedicated-only, lowest risk) - applying the five live-ops
lessons and the plain-Helm-chart lesson above from day one, not rediscovering them
partway through the way this round’s own build did. The “Open questions for this
round” list above (Postgres/Mongo/RabbitMQ Dedicated-vs-Shared-tenant split,
nginx’s real scope, Keycloak-as-confirmed-backend, provider-rabbitmq’s
single-maintainer risk) are all still genuinely open and should be resolved before,
not during, their respective build steps.
Round 2026-09-24: Postgres backend is CloudNativePG, and every component must label what it creates
Two decisions from live testing on the dev cluster. Both supersede the provider-sql-first plan in
the 2026-09-17 round above for Postgres.
Postgres on Kubernetes: CloudNativePG, not Bitnami and not a hand-rolled StatefulSet
CloudNativePG (operator 1.30.1, chart 0.29.1) is installed on the dev cluster
(gitops-cluster-dev/10-crds-operators/cloudnative-pg/). Operator and default Postgres
images are multi-arch (amd64 + arm64). Why it, and what was proven on a throwaway Cluster:
- Tenant isolation. By default any Postgres role can
CONNECTto any database and list its table names. RevokingCONNECTontemplate1does not help (CREATE DATABASE does not copy the ACL). The rule that works is a custompg_hbaplus naming each tenant’s database after its role:
CNPG’shost all postgres all scram-sha-256 # admin (provider-sql / operator) host samerole all all scram-sha-256 # a role may reach only the database named for it host all all all rejectspec.postgresql.pg_hbaputs user rules after its fixed rules and before its defaulthost all all all scram-sha-256, so ours win (first match). Tested over TLS: a tenant reaches its own database and is rejected on another tenant’s, onpostgres, ontemplate1and on CNPG’s ownappdatabase, before authentication, and the rules survive a pod restart with data intact. The XRD must therefore set database name == role name. - Native per-resource CRDs. CNPG ships
DatabaseandDatabaseRole(one resource each,cluster:reference), withdatabaseReclaimPolicy/databaseRoleReclaimPolicy: retain. Tested: deleting both CRs leaves the database and role in place, with none of the ownership deadlockprovider-sqlhad (a Delete-protected database blocks its owner role’s drop,2BP01). This makesprovider-sqlunnecessary for CNPG-hosted Postgres. It was installed on the dev cluster for the early tests and has been removed (nothing used it). - Gotchas found. CNPG’s
pg-superusersecret carrieshost: <cluster>-rw, a short name that a controller in another namespace cannot resolve; the Composition must use<cluster>-rw.<ns>.svc.DatabaseRole.passwordSecretneeds akubernetes.io/basic-authSecret the Composition creates. - Rejected: Bitnami’s chart (its versioned images were removed and the chart re-published as a different appVersion, see the Redis note in the compositions), and the official image with our own StatefulSet (we would own HA, backups and upgrades).
Every attached-tier component must propagate hangar.io/* labels to what it creates
The application chart stamps hangar.io/app, env, cluster and component-type on the
component XR, but nothing composed below the XR inherited them (checked live: a Redis cache’s
StatefulSet, Service, Secret and Pod carried only the chart’s own labels), so
kubectl get all -l hangar.io/app=<app> missed them. Requirement: a component Composition copies
the XR’s hangar.io/* labels onto its composed object and onto everything the backend creates,
using whatever the backend offers:
| Backend | Mechanism | Verified |
|---|---|---|
| Bitnami Redis chart | commonLabels value; labels every object and the pod template, not the StatefulSet selector or volumeClaimTemplates (so PVCs are not labelled) |
live, chart 19.0.2 |
| CloudNativePG | spec.inheritedMetadata.labels on the Cluster; labels Pod, PVC, Services, Secrets, ServiceAccount (the Cluster CR itself is labelled by the Composition) |
live, 1.30.1 |
Only hangar.io/* is copied. The XR’s app.kubernetes.io/* labels describe the app and would
collide with the backend’s own.
Status (later 2026-09-24): built, and verified on both clusters
The dedicated PostgreSQL component is built (airframe xrds/postgresql.yaml,
compositions/postgresql/) and installed on the dev and prod clusters, each with the CNPG operator and
Crossplane RBAC for postgresql.cnpg.io and networking.k8s.io.
- The NetworkPolicy is necessary, proven on the prod cluster (Calico). The same
Clusterin a namespace with the app baseline policy but without the component’s operator-allow policy never became healthy (Instance Status Extraction Error: HTTP communication issue); with it, Ready in about a minute. The dev cluster’s CNI does not enforce policy, so only prod could show this. Cross-namespace access to the database was blocked. - Kubernetes 1.37 on the prod cluster. CNPG 1.30 lists 1.34-1.36 as supported and 1.37 as “tested, but not supported”. Works here; not covered by upstream support.
- The component’s credentials are consumed with
envvalueFrom(airframe v0.3.88), not copied into Infisical: the chart used to dropvalueFromfromenventries. Redis’s password can be read the same way. - Tower gap, fixed (backstage
366ea8c): the App Configuration Environment variables section is a name/value form; editing it used to rewrite the wholeenvlist as name/value and replace avalueFromentry with an empty value. It now showsvalueFromrows read-only and keeps them. - First consumer: Skyport’s
flight-api, seeairframe/docs/user/quickstart-flight-api.md.
Round 2026-09-25: RabbitMQ is a shared broker, attached per app with generated permissions
Decisions (with the user) and what live testing on the dev cluster showed. The component is
airframe/xrds/rabbitmq.yaml + compositions/rabbitmq/.
Shape
- Shared, not dedicated. flight-api publishes; boarding-api and baggage-api consume. Nothing
crosses a vhost without federation/shovel, so they must share a broker and a vhost. Unlike
PostgreSQL/Redis this is one broker per cluster and environment, owned by an
InfraService(skyport-broker). Backend: RabbitMQ Cluster Operator 2.23.0 + Messaging Topology Operator 1.20.3 (both multi-arch), vendored under10-crds-operators/rabbitmq-operatorsin each cluster’s gitops repo. Both need cert-manager. Rejected:provider-rabbitmq(single maintainer), Bitnami’s chart. - One vhost per domain (
flights), not per app (breaks fan-out) and not one for the whole demo (no isolation between domains). Each app gets its own user with least-privilege permissions on it. - Two modes in one XRD.
mode: brokercomposes theRabbitmqCluster, itsVhosts and a NetworkPolicy.mode: attachcomposes aUser+Permission+ a connection ConfigMap in the consumer’s own namespace, referencing the broker across namespaces. This avoids the cross-namespace limit that ruled out shared PostgreSQL. - Permissions are generated from names, never raw regexes:
queuePrefix(queues the app may create/bind/consume),publish(exchanges it may declare and publish to),consume(exchanges it may bind to). Binding a queue needswriteon the queue andreadon the exchange. - Credentials are the operator’s
<xr>-user-credentialsSecret (username, password) in the consumer’s namespace, read withenvvalueFrom; host/port/vhost are in<xr>-connection.
Verified live (dev, arm64), over real AMQP
A producer declared the exchange and published; a consumer declared its own queue, bound it, and received the message. Refused: the consumer publishing to the exchange, the consumer declaring a queue outside its prefix, the producer reading the consumer’s queue.
Gotchas found
- Trust boundary: a User/Permission from another namespace is refused unless the broker’s
rabbitmq.com/topology-allowed-namespacesannotation lists that namespace (allowedNamespaces). Anyone who can write a namespace’s manifests can request any permission there, so the broker owner controls who may attach, not what they ask for. Acceptable for the demo; a real multi-team setup needs an admission policy. - Use RabbitMQ 4.2, not 4.1: the operator’s startup probe calls
/api/health/checks/reached-target-cluster-size, which 4.1.8 answers 404; the pod never becomes ready. Pinned torabbitmq:4.2.9-management. - Ownership conflict: the Topology Operator makes a
Permission’suserReferencetarget its controlling owner, which fails when Crossplane already controls the Permission. The composition refers to the user by name instead, read from the observedUserstatus (the operator generates the username), so the Permission renders one reconcile after the User. - No
Readycondition: aRabbitmqClusterreportsAllReplicasReady, so function-auto-ready never completes; readiness is derived from that condition. - PVCs are not labelled
hangar.io/*(pods, Services, StatefulSet and Secrets are), the same gap as Redis.
Not yet verified
- prod (amd64, Calico). The NetworkPolicy (operator on 15672/5672, allowed namespaces on 5672) is written but only dev has run it, and its CNI does not enforce policy.
- prod capacity. The prod cluster is resource-constrained (observability is scaled to 0). RabbitMQ (~1Gi) plus two operators is a real addition to it.
Walked live (2026-09-25): the first real InfraService, and Phase 2’s first consumers
skyport-broker was created through the real path on the dev cluster: an InfraService XR
(tenants/skyport-broker/xr-requests/), the gitops-infra-skyport-broker repo it scaffolds, and a
platform/envs/dev.yaml with a rabbitmq component. This retires the earlier “appType: infra has
never been used for real” caveat for the dev cluster. Guide: airframe/docs/user/quickstart-broker.md.
- Gap found and fixed: the
idp-onboardingAppProject did not whitelistInfraService(not permitted in project idp-onboarding), so the XR never applied. Added in gitops-cluster-dev and apron. The prod cluster’s copy is deliberately narrower (Bootstrap kinds are dev-only), so anInfraService’s prod environment is anApplicationEnvironmentcreated on dev. - How an InfraService environment is defined:
platform/envs/<env>.yamlingitops-infra-<name>, because the tenant’sappRepoUrlis that same repo and the lower-env ApplicationSet readsplatform/envs/*.yamlfrom it. Namespace:app-<name>-<env>. - End to end, live: flight-api (publisher) and boarding-api (consumer, evicts its Redis cache)
attached through
mode: attachcomponents and exchanged real events; a gate change reached the board withcached: falseat once. Both apps read their broker credentials withenvvalueFrom(secretKeyRefandconfigMapKeyRef, no chart change needed). - Image scan is a real gate: flight-api on Spring Boot 3.3.4 failed the pipeline’s Trivy scan (39 HIGH/CRITICAL, all in the base: Tomcat, Spring, Jackson). Boot 3.5.16 plus pom overrides for Tomcat, the Postgres driver, the RabbitMQ client and Netty cleared it. Part 2’s flight-api would fail the same scan today.
- Provenance gate: boarding-api’s release PR fails
provenanceon an unsigned commit; the user’s own earlier release PR (#11) did too and was merged, so signing is not enforced in practice. - Still open: the prod flight environments for the broker and both apps (not walked; the
host was memory-short),
baggage-api, and the PVC labelling gap.