Updating a distributed fleet without bricking a single unit

A fleet that had quietly drifted into a dozen firmware and configuration combinations, spread across sites, with no practical way to recall units.

Sector: Commercial fleet operator

A fleet shown as connected nodes, coloured by release state, update progress and units needing attention.

The constraint

Years of field service, incremental fixes and site-level workarounds had left the operator unable to answer a basic question: what is each unit actually running? Incidents could not be reproduced reliably, because the aircraft involved might not match the one on the bench.

Units were geographically distributed and recall was not viable. Any update mechanism had to assume interruption — power loss, link loss, an operator closing a laptop mid-write.

Stylised map of a distributed fleet, each unit shown at its last reported position.

What we did

We designed the failure cases first. Updates are atomic: a unit either completes the transition or rolls back to the version it was running, with no partial state that leaves it unbootable. Recovery assumes interruption at the worst possible moment, because eventually it happens.

Rollout is staged. A release reaches a small cohort, dwell time and health metrics are checked, and only then does it widen — so a bad release costs a handful of units of downtime rather than a fleet.

Underneath it is a fleet record that is verified rather than declared: each unit reports what it is actually running and that is reconciled against what it should be, with divergence raised as an exception.

Result

  • Fleet-wide alignment achieved with zero units lost to failed updates
  • Verified inventory of running versions, replacing an assumed one
  • Incidents reproducible, because the configuration is known
  • Staged rollout now standard for every release

Knowing exactly what every unit is running is not administration. It is a precondition for investigating anything.