I’ve been looking at several PX4 incident logs recently, and one pattern keeps coming up: the most visible failsafe or termination event is often not the event that actually initiated the failure.
For example, in one recent analysis the vehicle eventually exceeded ~80° attitude and entered failsafe. Looking only at the final events could easily make the failsafe or tilt threshold look like the cause.
But reconstructing the ULog chronologically showed a different sequence:
- commanded attitude remained small,
- measured attitude began diverging from the setpoint,
- actuator outputs subsequently reached extremes,
- the vehicle continued losing attitude control,
- failure detection asserted later,
- and flight termination/failsafe occurred after that.
That made me think about a more general question for PX4 log analysis:
What is the most reliable way to identify the first causal failure in a ULog, rather than simply identifying the final detected failure?
The workflow I’ve been using is roughly:
command/setpoint → measured response → estimator state → controller/rate response → actuator outputs → vehicle state transitions → failure detection → failsafe
Then I work backwards only after finding the earliest meaningful divergence between expected and observed behaviour.
For attitude-control incidents, for example, I find the distinction between these two cases particularly useful:
A. Setpoint diverges first
Investigate where that command originated and which control path generated it.
B. Setpoint remains reasonable but measured attitude diverges
Move downstream toward rate tracking, control authority, actuator saturation and the physical motor/ESC/power path.
I’m interested in how PX4 developers and experienced log analysts approach this.
Do you generally have a preferred set of ULog topics/events for establishing the first divergence, especially when estimator, controller, actuator and failsafe events occur within only a few seconds of each other?
Also, are there cases where working from the first observable divergence can be misleading because the actual initiating condition is only detectable later in the logging chain?
I’d be interested in comparing methodologies.