Airframe · Scorecard
Airframe, measured
Autopilot only works if an agent can change an app safely without guessing. So I score Airframe, the contract every app is built on, against what an agent needs: ten dimensions, fourteen acceptance checks, all automated. A+ means 97 or better and all fourteen checks. The number starts low on purpose, because it measures the finished contract, and it moves only when the work is real.
Progress over time
Each filled point is a committed run of the tool. The hollow points are the roadmap's milestone targets still ahead, not measurements.
- Baseline · 2026-09-26 Before any of the A+ work. One acceptance check passed: every live values file rendered.
- M0 · 2026-09-26 The chart guard, chart tests and CI, AGENTS.md and a first airframe validate. M0 closed the next day against a target of 35.
- M1 and M2 · 2026-09-29 A strict schema, component outputs, a contract bundle, the release-file split, an ownership gate, field-level agent scope and replayable walkthroughs.
Ten dimensions, measured 2026-09-29
The bar is the latest score; the thin mark on it is where the baseline stood, and the amber tick is the A+ line.
The fourteen acceptance checks
A+ needs every one. A high average cannot hide a missing check.
| Check | Baseline2026-09-26 | M02026-09-26 | M1 and M22026-09-29 |
|---|---|---|---|
| typo acceptance is 0% | · | · | ✓ |
| schema description coverage >= 95% | · | · | · |
| every non-passthrough object rejects unknown keys | · | · | · |
| contract bundle published | · | · | ✓ |
| AGENTS.md at root and on every component | · | · | · |
| all component XRDs declare outputs and verify checks | · | · | · |
| no derived-name literals in live env entries | · | · | · |
| airframe validate exists and is a required check | · | ✓ | ✓ |
| no live file mixes human and machine-owned keys | · | · | ✓ |
| shared base layer in the ApplicationSets | · | · | ✓ |
| airframe.* tools exist | · | · | · |
| AppSpec schema exists | · | · | · |
| executable walkthroughs replace prose quickstarts | · | · | ✓ |
| every live values file renders cleanly | ✓ | ✓ | ✓ |
Principle honored
Measured, not felt. Every check is automated, so A+ is something the tool prints, not something I claim. Every run that moves the number is committed, so the history on this page is the tool's own output.
A correction to the tool
The first run on 2026-09-29 said 58. The tool was looking for component contracts and for Clearance where they were first planned to live, not where M1 and M2 put them, and two checks were hard-coded to zero. I fixed the paths and turned those two checks into real ones; the fix is in the same commit as the score. Snapshots before it were not re-scored.
Limits
The scorecard can't tell whether an agent would actually succeed. The parachute sentence, run as a Preflight case, does. I use both.
Re-run it yourself: python3 tools/airframe-scorecard/scorecard.py --json out.json.
The tool, every committed run, the rubric
and the roadmap are on GitHub. More on Autopilot on the AI workloads page.