A fleet that had quietly drifted into a dozen firmware and configuration combinations, spread across sites, with no practical way to recall units.
Sector: Commercial fleet operator
The constraint
Years of field service, incremental fixes and site-level workarounds had left the operator unable to answer a basic question: what is each unit actually running? Incidents could not be reproduced reliably, because the aircraft involved might not match the one on the bench.
Units were geographically distributed and recall was not viable. Any update mechanism had to assume interruption — power loss, link loss, an operator closing a laptop mid-write.
What we did
We designed the failure cases first. Updates are atomic: a unit either completes the transition or rolls back to the version it was running, with no partial state that leaves it unbootable. Recovery assumes interruption at the worst possible moment, because eventually it happens.
Rollout is staged. A release reaches a small cohort, dwell time and health metrics are checked, and only then does it widen — so a bad release costs a handful of units of downtime rather than a fleet.
Underneath it is a fleet record that is verified rather than declared: each unit reports what it is actually running and that is reconciled against what it should be, with divergence raised as an exception.