The F-35 program did not fail to test its software. It tested constantly, for over two decades, at a scale few programs ever will. Its problem was sequencing: production ran ahead of the verification evidence that was supposed to justify it — a gap known in the program by its own official name, "concurrency."
Building the aircraft before the evidence existed
The U.S. Government Accountability Office has published dozens of reports on the F-35 Joint Strike Fighter program since the mid-2000s, and a consistent thread runs through them: aircraft were manufactured and delivered while the software and systems verification meant to confirm their capabilities was still in progress, or had not yet started. GAO's own term for this is concurrency — production overlapping with development and test, rather than following it.
The consequence was structural, not incidental. When flight testing later surfaced a deficiency, the fix had to be retrofitted into aircraft that were already built, or already in service — instead of being resolved before those aircraft left the production line. GAO reporting through the 2010s repeatedly flagged retrofit and concurrency costs running into the billions of dollars, directly attributable to capability being fielded ahead of the evidence that it worked as intended.
The pattern continued into the program's most recent chapter. In 2023, the Pentagon paused acceptance of newly built F-35s after unresolved issues surfaced in Technology Refresh 3 (TR-3), the hardware and software package underpinning the jet's next combat-capable software load, Block 4. Deliveries resumed only once enough verification evidence existed to justify accepting a scaled-back interim version — a smaller version of the same story: capability arriving before its proof was ready.
The pattern, in GAO's own recurring findingAcross successive reports, the same root cause reappears: schedule and production commitments were set before verification evidence existed to support them — so the evidence was always being assembled after the fact, on aircraft already built.
Evidence that trails the milestone is evidence you pay for twice
This is F-05's failure mode at program scale. Validation evidence that should have accumulated alongside development instead arrived after key decisions — production go-ahead, delivery, acceptance — had already been made. Every retrofit is the cost of that gap made physical: rework on hardware that verification, done in time, would have caught before it was built.
Audit-ready is a standing state, or it's a debt
The lesson scales down as cleanly as it scales up. Whether it's an aircraft program or a single software release, evidence assembled after a milestone is evidence that arrives too late to inform the decision it was meant to support — it can only confirm or contradict a decision already made. Evidence that accumulates continuously, tied to the requirement and design it verifies as the work happens, is available when the decision is actually being made, not after.
That is the difference PES is built around: validation and audit readiness maintained as a standing state on the same thread as the requirement and the design — not a document set assembled once development outruns the proof.