Hangar

Airframe · Scorecard

Airframe, measured

Autopilot only works if an agent can change an app safely without guessing. So I score Airframe, the contract every app is built on, against what an agent needs: ten dimensions, fourteen acceptance checks, all automated. A+ means 97 or better and all fourteen checks. The number starts low on purpose, because it measures the finished contract, and it moves only when the work is real.

Score now72.4/1002026-09-29 · grade C-
Acceptance checks7/14from 1 at the baseline
Since the baseline+45.12026-09-26 to 2026-09-29
A+ bar97 and 14/14not yet

Progress over time

Each filled point is a committed run of the tool. The hollow points are the roadmap's milestone targets still ahead, not measurements.

Airframe scorecard over time Baseline 2026-09-26: 27.3 out of 100, 1 of 14 checks; M0 2026-09-26: 41.3 out of 100, 2 of 14 checks; M1 and M2 2026-09-29: 72.4 out of 100, 7 of 14 checks. Targets ahead: M2 75, M3 88, M4 93, M5 97. 0 25 50 75 100 A+ at 97 27.3 Baseline 2026-09-26 · 1/14 41.3 M0 2026-09-26 · 2/14 72.4 M1 and M2 2026-09-29 · 7/14 75 M2 target Safe write 88 M3 target Autopilot core 93 M4 target Skyport AI 97 M5 target Evidence
  1. Baseline · 2026-09-26 Before any of the A+ work. One acceptance check passed: every live values file rendered.
  2. M0 · 2026-09-26 The chart guard, chart tests and CI, AGENTS.md and a first airframe validate. M0 closed the next day against a target of 35.
  3. M1 and M2 · 2026-09-29 A strict schema, component outputs, a contract bundle, the release-file split, an ownership gate, field-level agent scope and replayable walkthroughs.

Ten dimensions, measured 2026-09-29

The bar is the latest score; the thin mark on it is where the baseline stood, and the amber tick is the A+ line.

DimensionScoreHistory What the measurement foundWhat A+ still needs
1 Discoverability 84 35 → 50 → 84 AGENTS.md, llms.txt and a generated contract bundle exist; 6 of 15 XRDs are in the catalog AGENTS.md on every component
2 Schema precision 76 45 → 45 → 76 86% of schema nodes described, 46 of 76 objects strict; typos now fail 95% described, every object strict, examples
3 Component contracts 68 4 → 3 → 68 5 components declare outputs and verify checks; 33 of 45 env entries still hard-code a name fromComponent in every live env entry, SecretStore outputs
4 Pre-merge validation 94 22 → 76 → 94 airframe validate is a required check, with 7 dead-end rules more chart guards with actionable messages
5 Write safety 67 13 → 25 → 67 no live file mixes owners; base layer and release file are real an owner and a risk class on every field
6 Observe and verify 57 21 → 21 → 57 5 components declare a closed set of reason codes and verify checks reason codes on every XRD, a describe tool
7 Docs for agents 100 39 → 54 → 100 all three quickstarts replay as live walkthroughs, green held; every new component ships its walkthrough
8 Interaction surface 27 27 → 27 → 27 still Tower and Git only; no AppSpec schema, no airframe.* tools AppSpec, the planner and airframe.* tools (M3)
9 Safety integration 50 20 → 20 → 50 field-level agent scope is built into Clearance; no risk classes yet risk-classed fields, escape hatches denied by default
10 Hygiene and determinism 100 46 → 92 → 100 helm lint, 22 chart tests and chart CI; every live file renders held; keep the fleet rendering clean
Overall 72.4 27.3 → 41.3 → 72.4 7 of 14 acceptance checks · A+ needs 97 and 14 of 14

The fourteen acceptance checks

A+ needs every one. A high average cannot hide a missing check.

CheckBaseline2026-09-26M02026-09-26M1 and M22026-09-29
typo acceptance is 0% ··✓
schema description coverage >= 95% ···
every non-passthrough object rejects unknown keys ···
contract bundle published ··✓
AGENTS.md at root and on every component ···
all component XRDs declare outputs and verify checks ···
no derived-name literals in live env entries ···
airframe validate exists and is a required check ·✓✓
no live file mixes human and machine-owned keys ··✓
shared base layer in the ApplicationSets ··✓
airframe.* tools exist ···
AppSpec schema exists ···
executable walkthroughs replace prose quickstarts ··✓
every live values file renders cleanly ✓✓✓

Principle honored

Measured, not felt. Every check is automated, so A+ is something the tool prints, not something I claim. Every run that moves the number is committed, so the history on this page is the tool's own output.

A correction to the tool

The first run on 2026-09-29 said 58. The tool was looking for component contracts and for Clearance where they were first planned to live, not where M1 and M2 put them, and two checks were hard-coded to zero. I fixed the paths and turned those two checks into real ones; the fix is in the same commit as the score. Snapshots before it were not re-scored.

Limits

The scorecard can't tell whether an agent would actually succeed. The parachute sentence, run as a Preflight case, does. I use both.

Re-run it yourself: python3 tools/airframe-scorecard/scorecard.py --json out.json. The tool, every committed run, the rubric and the roadmap are on GitHub. More on Autopilot on the AI workloads page.