The problems these choices solve

Everything here is on the list because we have shipped a product on it and then had to maintain that product for years afterwards. That is a different test from whether a tool demos well.

You need a part supported, and nobody supports it

A sensor arrives with a driver for one kernel version nobody ships any more. A programme loses a quarter to it.

The other version of this is quieter: the image that works is a snowflake nobody can rebuild. Two years on, the engineer who made it has gone, the aircraft is still flying, and no one dares touch it.

Open, reproducible from source, carrying nothing it does not need, small enough to update over a field link — embedded Linux built with Yocto. Not a preference. It is what makes year three survivable.

LinuxYocto Project

It worked in the demo

Then it was split across two compute units and the timing came apart. Then a fault appeared once in forty flights and never on the bench. Neither problem existed in the demo, and both were designed in before it.

The control loop does not belong in middleware. Everything above it does — perception, mission logic, recording — and ROS 2 carries it well: module boundaries that survive being distributed, and a recording that turns that once-in-forty fault into a bench test.

None of which is free. What ships with it is a foundation, not a system — the reliability came from the nodes we wrote around it: supervision, liveness and health checks, bounded queues, and defined behaviour for every source that goes quiet. The stock tooling gets you something that works on a good day.

ROS 2

It benchmarked fine on the desk

The desk had airflow.

In the aircraft the module sits in a sealed bay, in still air, at altitude. It throttles, and the perception budget you sized the whole mission around is gone — at the point in a programme where redesigning costs the most.

So the first question about compute is what the airframe can cool, not what the part can do. NVIDIA where the work is perception and video; Qualcomm where connectivity and power budget decide it.

NVIDIAQualcomm

The autopilot does almost what you need

Almost leaves two roads, and both cost more than they look:

  • Bolt a workaround onto the outside of a stack you do not control — and own that seam for the life of the product.
  • Fork it — and strand the programme on a version that can never take an upstream security fix.

There is a third road, open only if the stack is and you are willing to work inside it: change it properly, and carry the changes as patches against upstream. That is how we work on ArduPilot, and it is why a vehicle we built years ago can still be updated today.

ArduPilotCubePilot

The position looked right

It looked right until the RF environment changed, and nothing said so. How would it? A receiver that quietly absorbs interference reports the same confident position either way.

The second version is worse, because it looks like success: sensors that agree on where and disagree on when. Fusion does not fail. It returns a confident wrong answer, which is far harder to catch than an obvious break.

So: multi-band receivers that report interference rather than swallow it, and one clock that everything is stamped against. Depth, lidar, radar and vision get added where they earn their mass, and nowhere else.

u-bloxVzense

The link that worked on the test range

The test range is flat, empty and legal. The job is in a valley, in a band somebody else is also using, under rules written for a different country. And unlike every other subsystem, you cannot walk over and inspect the link — you only ever see what arrived.

Which is why we carry more than one datalink supplier. SIYI and Taisync answer different versions of the question, RFDesign answers the one where range matters more than bandwidth, and no single radio wins everywhere. Where there is coverage, a Quectel modem adds a second way home. Where latency or integrity is the actual problem, we write the protocol.

SIYITaisyncRFDesignQuectel

Catalogue figures are measured somewhere else

Read the conditions printed under a thrust figure and you will usually find:

  • sea level, around twenty degrees
  • a bench, not an airframe
  • a test lasting minutes

Your aircraft works cold, high, and for hours. The margin you think you have is the margin somebody else measured, in conditions you will never fly in.

Propulsion therefore gets matched at the real operating point — T-MOTOR, sized to the airframe rather than to the number. Power gets measured everywhere it is spent and regulated at the point of load, so one subsystem browning out cannot take the flight controller with it.

T-MOTOR

The update reaches four hundred units

One loses power halfway through and does not come back. Another turns out to be running firmware nobody can identify. A third has been degrading quietly for a month and will fail next week — you will find out when it fails.

Three problems, three requirements. Updates that are atomic and roll themselves back, so an interrupted write is only a failed one. A history in which every unit traces to the exact commit it runs. And telemetry kept long enough to show a trend, because without months of it you cannot tell a part that is failing from one that always read like that.

Mender, Git and Grafana, in that order.

MenderGitGrafana

Recognising it early is the whole job

None of these problems is exotic. What makes them expensive is meeting them late — during integration, during qualification, or in the field — when the schedule has no slack and the answer costs a redesign instead of a decision.

We are frequently paid to do nothing but see them coming: to say which path a programme should take before anyone commits to hardware. Getting that wrong never costs the licence fee or the module price. It costs the year you spend finding out.

How we approach system architecture, or tell us what you are building.