On the morning of September 8, a flaw dormant within the code of the United Kingdom's air traffic control system awoke for a fraction of a second — and in that fraction, set in motion nine and a half hours of cascading failure. It is a story as old as complex systems themselves: the danger that hides not in what is broken, but in what appears to have healed. For the third time since summer 2023, hundreds of thousands of travellers paid the price for infrastructure that was trusted beyond its own resilience, and a nation is now asking whether the institutions charged with keeping the skies safe
Millisecond software glitch caused UK air traffic chaos, but questions remain on detection
A single fault cascaded into nationwide disruption
So the software error happened at 10am, but nothing serious happened until 12:30pm. Why the delay?
The system actually recovered on its own at 10:06am. Engineers saw it drop, investigated, and it came back up. They kept monitoring, but there was no indication of danger. Then the same defect triggered again two and a half hours later, and this time it didn't recover.
But that raises the real question: if they knew something had gone wrong at 10am, even if it seemed fixed, shouldn't that have triggered a deeper investigation? Why wasn't the system taken offline for a full diagnostic?
That's exactly what the government is asking now. The preliminary report says it was a legacy defect in combination with very specific events. But ministers are saying the report doesn't explain why this wasn't caught before it caused chaos.
And this is the third major failure since summer 2023?
Yes. Which is why airlines are now demanding not just compensation but systemic change. Ryanair is calling for the CEO to resign.
The report does say the defect was previously unknown. So it's not like Nats ignored a known risk. But that's almost worse—it suggests their testing and monitoring aren't catching these things until they blow up.
What happens now?
The Civil Aviation Authority will do an independent review within six months. They'll assess whether Nats is even fit for purpose and demand investment in resilience.
Six months is a long time when you're a passenger who just lost your flight. And we don't know yet if the review will actually change anything structural, or if it'll just be another report with recommendations.
Il Polso
- A millisecond-long software glitch corrupted flight data at 10am, but engineers declared recovery within minutes — a false calm that delayed any major incident response for over two hours.
- When the same defect struck again at 12:32pm under slightly different conditions, the system could not self-correct, forcing controllers into manual procedures and ultimately a full reboot that took nearly ninety minutes.
- More than 2,000 flights were cancelled and hundreds of thousands of passengers stranded, with the ripple of disruption taking over two days to clear across the network.
- Airlines are no longer accepting apologies — executives from Ryanair, EasyJet, and Airlines UK are demanding resignations, mandatory investment, and structural reform to ensure no single fault can ever again ground the nation's airspace.
- The Civil Aviation Authority has been tasked with an independent review of Nats' fitness for purpose, its investment plans, and its regulatory accountability, with findings due within six months.
On the morning of September 8, a flaw dormant within the code of the United Kingdom's air traffic control system awoke for a fraction of a second — and in that fraction, set in motion nine and a half hours of cascading failure. It is a story as old as complex systems themselves: the danger that hides not in what is broken, but in what appears to have healed. For the third time since summer 2023, hundreds of thousands of travellers paid the price for infrastructure that was trusted beyond its own resilience, and a nation is now asking whether the institutions charged with keeping the skies safe are truly fit for the age they inhabit.
On the morning of September 8, a defect buried in the UK's air traffic control software activated in the space of a millisecond. A routine request to assign an identification code to an aircraft was interrupted by a higher-priority task; when the system resumed, it did not resume correctly. Flight data downstream became unreliable.
Engineers at National Air Traffic Services detected the problem at 10:02am and, by 10:06am, reported that the system had recovered. No major incident was declared. The morning continued. Then, at 12:32pm, the link between the national airspace system and London's control centre dropped again — this time without self-correcting. Controllers shifted to manual procedures. By 1:32pm, the system had failed completely. A full reboot ran from 2:50pm to 4:09pm. Full operations did not resume until 7:30pm, nine and a half hours after the initial error.
The toll was severe: over 2,000 flights cancelled, hundreds of thousands of passengers displaced, and a backlog that took more than two days to clear. It was the third major UK air traffic control failure since summer 2023. Nats chief executive Martin Rolfe apologised and described the defect as a legacy issue — previously unknown, only triggered by a precise and rare sequence of simultaneous events. A permanent fix, he said, was being safety-tested.
Transport Secretary Heidi Alexander found the explanation insufficient and tasked the Civil Aviation Authority with an independent review of Nats' systems, investment plans, and regulatory accountability, to be published within six months. The airline industry was less measured: executives called for mandatory resilience investment, and Ryanair's representative demanded Rolfe's resignation outright.
What the timeline ultimately reveals is a system capable of detecting its own failure but not of recognising its own danger. The defect was seen at 10:06am and appeared resolved — because, in that moment, it genuinely was. The deeper failure came when the same flaw returned under slightly altered conditions, and the window for prevention had already closed. The question now before the review is whether that window can be structurally widened — whether a system can be built not merely to recover from errors, but to stop them before recovery becomes the only option left.
On the morning of September 8, a defect buried in the UK's air traffic control software activated without warning. The problem unfolded in the space of a millisecond—so fast that no human operator could have caught it in real time. A request to assign an identification code to an aircraft was interrupted by a higher-priority task. When the system resumed processing the original request, the software did not resume correctly. The output corrupted. Flight data downstream became unreliable.
Engineers at National Air Traffic Services noticed the problem at 10:02am, just two minutes after it occurred. By 10:06am, they reported that the system had recovered and appeared to be working normally. No operational impact, they said. They continued investigating, but there was no sense of urgency. No major incident was declared. The morning proceeded.
Then, at 12:32pm—two and a half hours later—the link between the national airspace management system and London's control center dropped again. This time, the failure did not self-correct. Controllers had to shift to manual procedures for some tasks. By 12:45pm, Nats began restricting incoming flights as the system continued to malfunction. At 1:32pm, the system failed completely. Airlines were notified of restrictions at 1:40pm. Between 2:50pm and 4:09pm, engineers rebooted the entire system. Full operations did not resume until 7:30pm—nine and a half hours after the initial error.
The damage was immense. Over 2,000 flights were cancelled. Hundreds of thousands of passengers found their travel plans erased or delayed. Aircraft were displaced across the network. The backlog of disruption took more than two days to clear. This was the third major failure of the UK's air traffic control system since the summer of 2023, and it raised a question that government ministers and airline executives could not ignore: why was a known software defect never found and fixed before it brought the system down?
Martin Rolfe, the chief executive of Nats, issued an apology and insisted that safety had never been compromised. The defect was a legacy issue, he said—something inherited from earlier versions of the software, previously unknown, that only manifested when a very specific sequence of events occurred simultaneously. A permanent fix was being safety-tested. Mitigation was in place. But his words offered little comfort to an industry already fractured by repeated failures.
Heidi Alexander, the transport secretary, said the Nats report left critical questions unanswered. She tasked the Civil Aviation Authority with conducting an independent review to examine not only Nats' findings but also the company's investment plans and regulatory accountability. The review would be published within six months. The message was clear: Nats' leadership and systems were now under formal scrutiny.
Airlines demanded more than apologies. Tim Alderslade, chief executive of Airlines UK, called for mandatory investment in system resilience so that a single fault could never again cascade into nationwide disruption. EasyJet's Kenton Jarvis said passengers deserved action, not promises. Ryanair's Neil McMahon went further, calling for Rolfe's resignation and rejecting what he characterized as an excuse masquerading as an explanation. The industry was signaling that the tolerance for failure had expired.
What emerges from the timeline is a system that detected its own failure but did not recognize the danger. Engineers saw the problem, confirmed recovery, and moved on. The system appeared stable. No alarm bells rang at 10:06am because the system, at that moment, genuinely was stable. The real failure came later, when the same defect triggered again under slightly different conditions. By then, the window for prevention had closed. The question now is whether the review will reveal how such a window can be kept open—how a system can be designed not just to recover from errors, but to prevent them from ever reaching the point where recovery is the only option left.
Citazioni salienti
The disruption we saw last week was completely unacceptable... we need to urgently understand why this issue was not discovered and fixed before it caused chaos.— Transport Secretary Heidi Alexander
Passengers deserve more than just another promise that lessons will be learned. Firm actions must be taken.— EasyJet Chief Executive Kenton Jarvis