whitepaper
Inherited Blind Spots Part-3

Inherited Blind Spots Part-3

September 3, 2026Nicholas Edwards
Alchemist AI Pro™TradewindsDefense

Generation capacity has risen dramatically, while human review capacity is still where Fagan measured it in 1976. An AI auditor narrows the aperture cheaply, but only if the unflagged items can be trusted. Otherwise the program buys the compute and then asks the reviewer for more hours anyway.

Inherited Blind Spots (Part 3)

Paying Twice for One Review

BLUF: Large language models (LLMs) systematically overlook errors in their own output while catching it reliably from external sources. Programs are beginning to rely on LLMs to narrow the aperture to get a clear picture of what is filtered from the firehose of output generated by LLMs. This series has explored the risk in letting a model audit its own output, and left unhandled it feeds a well-known cost driver: rework that has been raising DoD development costs for decades. How do you reduce the chance of paying twice for one review?

This part can be read without Part 1, which covers the technical, or Part 2, which covers the human cost.

The Review Line Item Nobody Priced

A single user with a frontier model can produce more requirements as output in an afternoon than a review board can get through in a week.

Reviewers have a limit, and that has been established since Fagan (1976) measured them. Yes, technology has expanded the range of items that can be reviewed or deemed appropriate, but that "aperture" has remained the same. Fifty years have not proven IBM's initial findings wrong. Once a threshold is reached in human review, the reviewer's detection effectiveness collapses (Fagan, 1976). The more pages, the harder it is for someone to be effective.

That single human reviewer can only handle up to that same threshold written down in 1976 (Fagan, 1976). No program manager or developer has insight into "unreviewed requirements"; only failed tests well after development begins, when the costs of those undiscovered problems are raised.

Figure1 review capacity gap
Figure 1 - A representation of the advent of generative AI against the capacity of a human reviewer working two two-hour sessions a day necessitating AI assisted review (Fagan, 1976)

The Problem

Parts 1 and 2 covered the technical and human cost, but this paper is dedicated to the financial cost.

What drives the cost?

In 2023, Congress requested that auditors review how the military buys software. The GAO recommends that the military adopt better requirements engineering, stronger oversight, and improved tooling to manage program requirements (U.S. Government Accountability Office [GAO], 2023).

To compound the issue, Triezenberg et al. (2025) conducted a study on productivity losses in the DoD workforce and found that, conservatively, $2.5 billion was lost in fiscal year 2023 due to underperforming software and IT systems. This paints an interesting picture of 2023 as a year when the federal government began to realize the scope of these problems. RAND's figure covers lost workforce productivity from underperforming IT and software across DoD programs as a conservative estimate to the speculated true cost.

These are not just evidence, but conclusions drawn from years of identification. To understand the mechanisms behind why software is a risky investment for the government, you can look at the Defense Science Board, which is the Pentagon's senior technical advisory board. In 2018, they studied how the DoD buys and builds software. It offered 7 recommendations, the last of which was that machine learning systems need independent verification and validation (Defense Science Board [DSB], 2018). Software drives program risk for around 60% of new-system acquisitions programs (DSB, 2018, p. 4). Requirements were highlighted as surfacing during or after development, causing substantial additional effort (DSB, 2018). In government, substantial additional effort equates to higher spending on contracting.

So far it has been established that software is costly for the government. Requirements are part of that calculation with multipliers the longer they take to surface after development begins. Reviewers are the people in the seats that can cost your program the most money if they neglect to account for something. They are expensive to underserve with the right tooling and oversight.

Looping back to the review line nobody priced: generation capacity has risen dramatically; review capacity for a human-in-the-loop is still where it was in 1976; and now you can see how cost climbs as the standard development lifecycle (SDLC) advances towards completion. Whatever helps narrow the aperture for the reviewer to meet Fagan's threshold needs to be trustworthy enough to not trigger substantial additional effort (Fagan, 1976).

The Remedy That Looks Cheap

One of the main conceded points of this series is that putting an LLM in the audit seat to filter what the reviewer sees is reasonable. It catches real problems, and it can do it at a capacity that a human reviewer can then process. It provides the right aperture. It is also cheap. There is no doubt about any of these things. In fact, this series encourages this course of action, although with a word of caution.

If you have read the first two parts, then you will know this part well. Skip ahead. For the rest, read on.

There have been many studies on self-preference, self-identification, the amplification AI brings, and the problems these present. In a study comparing 14 open-source non-reasoning models, Tsui (2025) found that injected errors were corrected reliably when they were created external to the model's own output. However, 64.5% of identical errors were missed by models reviewing their own output (Tsui, 2025). This series has dubbed that a same-lineage problem. When models from the same family share the same corpus, weights, and system guardrails, a risk arises. Panickssery et al. (2024) found that models can distinguish their own output as uniquely their own and don't share that preference for another model. Not only can it identify its analogous output, but it has a self-preference for its own output that rises in kind with that recognition. This is regardless of how humans might perceive the output. A dangerous proposition for a system which might be auditing its own work.

Figure2 independent vs self review
Figure 2 - Same error but one side catches it in independent review of an external source, and another misses the error due to self-recognition.

There is a crack in our reasonable assumption that LLMs should be allowed to undergo first-pass audits of their own output. In a 2025 DORA report published on Google's own blog, Harvey and DeBellis (2025) found that AIs tend to amplify what is already present in the delivery system they are added to. An LLM added to a thin verification process will amplify the problems of a system with very little oversight.

All of this was explored in detail in parts 1 and 2. Part 1 was about eliminating common-mode failures through technical means (Edwards, 2026a), and part 2 was about closing the aperture, which was a thorough dive into trust and the human cost of relying on same-lineage models (Edwards, 2026b). Both go into these research projects and studies in more detail.

Taking all these studies, then pairing them with the findings of RAND and the GAO (Triezenberg et al., 2025; GAO, 2023), we start to see an interesting picture of what the cost might be, but the DSB's 7th recommendation, which was titled Independent Verification and Validation for Machine Learning, underlines what the risk is specifically when it comes to AI. The DSB's recommendation directs DARPA, the SEI FFRDC, and DOD laboratories to build programs around the practical use of machine learning with efficient testing, independent verification and validation, with a helping of cyber resiliency as a focus (DSB, 2018, p. 28). While this predates GPT-3 by two years, the concept is still relevant; independent verification and validation translate into this series' core thesis: that a single model lineage can't be trusted to audit its own work as a simple, conservative method for achieving trust in the process. If a program purchases tooling that shows self-preference for its output and ignores problems with what it was fed, then that is a risk. This series presents a pragmatic approach towards resolving that risk.

Paying Twice

So here is the main point of this paper. You are using a tool with the latest frontier model that does everything for you. You need to verify the mountain of output text it has produced, so you give it an agent that acts as an auditor. This auditor carries a degree of self-preference that increases the more it recognizes its own output and amplifies any dysfunction in its process, with a chance of not identifying it as a problem due to that self-preference.

Let's cover a program's options and the cost of each:

  1. Review all the pages manually. A reviewer spends hours, days, or weeks reviewing each item that the model outputs. Every new item is overhead. This is honest work, but it is the most expensive option for a program and carries the greatest risk of schedule slippage.
  2. Accept the clean report. Accept that the technology is presenting a clean report at face value. This accepts the risk that whatever gets missed will only be caught downstream. Someone will eventually run into the error and need to test it when caught, but that could be in test or even production before it surfaces. Every additional step into the SDLC increases the cost of the fix. Catching it in review is advantageous (Consortium for Information and Software Quality [CISQ], 2022).
  3. Staff up the review. Extra reviewers hit Fagan's rate limits, so each one gives back less than the last, and no program office can hire fast enough to keep pace with the output of an LLM's generation capacity anyway (Fagan, 1976). Which was the argument for the AI auditor to begin with.

The argument isn't that human auditors don't have their purpose nor that a same-lineage model can't catch errors in an audit. They can catch a lot and deliver substantial savings. What it can't do, though, is provide an unbiased look at the errors it created, and if it gives errors a clean bill of health to a human reviewer, then is that reviewer going to look at the unflagged items to verify them as an independent reviewer or accept the pages upon pages of unflagged content as correct and ready for development? Most would accept the report as clean, like step 2. It is worth the little bit of effort to employ a second model to provide an independent audit, validate and verify the audit, and raise the marginal increase in flagged items that would otherwise go unfound.

The Failure That Never Gets Named

All three options assume somebody eventually finds out. Usually nobody does.

A smoke detector with a battery that died when no one was home to hear its beeps for attention behaves exactly like a house that is not on fire. Both are quiet. Most don't bother to check the battery until something makes it a serious concern. Hopefully, not a fire.

What if that important issue is finding the defect deep into an integration, in a test environment, in sustainment, or even worse: in production. Weeks or months have gone by. What if the person who signed the review moved on?

The requirement handed to a human is unclear, or, if passed to an agent, it assumes the outcome based on context that should never have passed review.

The finding is right. That ambiguity is set into the record, and the corrective action that follows is the appropriate one for what gets documented by whoever found the problem. It now requires a meeting, redeployment, and/or a schedule slip. It costs the program money. The well-oiled machine screeches to a halt.

This is a bug, an issue, or a ticket. It isn't a bad requirement because it was agreed upon. No one flags these as an issue with the review stage, as a clean audit typically leaves no artifacts. Traceability can't be established, so it goes to a tracker as an extended scope of the program's schedule.

If this happens too often, then the program does the responsible things. Go back and review the requirements and the language in the specification, improve them through elicitation, and provide the engineers who write requirements with more training in the tool. Nobody touches the audit configuration, because nothing in the record pointed there. The next audit goes to the same seat and the same lineage and comes back just as quiet.

The GAO told the DoD to fix their requirements processes, oversight policies, and engineering tools (GAO, 2023). No one checks the verification layer to see whether it might look like a trusted one at all. Most accept the tools' output and authorize it.

The most common solution for a program is to use project management tools to run a full correction loop in good faith, without acknowledging that the broken process might sit outside the loop.

The Question for the Next Program Review

Part 1 asked whether the review was truly independent. Part 2 asked how many pages a person read before approval. This paper asks a different question:

How many hours of human review did the AI audit remove, and what makes that number reliable?

Most teams can answer the first half of that: how much it removes. The second half most can't answer, but it is what separates a reliable architecture from one that risks costly rework.

How Alchemist AI Pro™ Separates the Seats

Alchemist AI Pro™ collaborates with teams using Socratic elicitation and elaboration. The output is a specification-first approach to producing a human-readable but AI-recognized package of documents to support development. Built with the thesis of this series in mind. The requirements-generating model and the model that audits the outputs come from different vendors and are in different model lineages. A human holds complete disposition authority over every finding under a narrowed aperture.

Google models generate the requirements, and OpenAI's models audit them. This is the commercial offering, but the system itself is vendor-agnostic, and different models can serve as generator and auditor roles. Part 1 covers the architecture, the audit stage, and the defect taxonomy that people write and settle before a run begins (Edwards, 2026a).

More at https://alchemistaipro.com.

748 to 21, Priced

Alchemist AI Pro™ can output a package recognized by most agentic development platforms. ACC3 International uses its own agentic lifecycle framework called Alchemy SDLC™. As evidence of the advantages of the requirements approach and fully traceable development lifecycle, ACC3 developed an application several employees have worked with in the past, a CRM, and published what came out of it (ACC3 International, 2026a):

  • 125 of 125 specifications (100%) referenced directly in shipped code
  • 105 of 105 use cases (100%) verified and functional
  • 228 of 229 test cases (99.6%) functionally covered
  • 748 tasks executed by the automated Alchemy Crew, with 21 carried by the human Away Team

The deterministic estimate for the build was 2,978 story points, or 11,912 hours of manual development effort, which is 74.5 team-weeks. The delivered system consumed 61.8 tracked human hours in its final mile (ACC3 International, 2026a).

That last ratio is the labor line. Roughly 97% of task execution ran automated, with human effort reserved for the cases that required judgment. The split only holds up if the engineer can trust whatever sorted the cases. When the engineer cannot, the same pages still need to be read, and only the label on them has changed.

The Product Campaign Manager numbers are the ones a cost buyer will weigh. 10 features, 395 use cases, 100 percent feature completeness, three to four people where the plan called for nine, and roughly $250,000 against a $2.27 million plan (ACC3 International, 2026b).

ACC3 ran both builds on its own work, inside the vendor. They show the pattern. No other program should expect to reproduce those numbers, and results elsewhere will depend on the complexity of the requirements, the maturity of the defect taxonomy, how well the approach meshes with proven workflows, and the reviewers' experience. Programs should calibrate against their own environment. Receipts are published at https://alchemistaipro.com/proof.

Tradewinds Awardable Positioning

Alchemist AI Pro™ has been assessed Awardable through the Tradewinds Solutions Marketplace, the Department of War's post-competition repository for AI, data, and analytics solutions (Chief Digital and Artificial Intelligence Office, n.d.). Awardable is a procurement-readiness signal following independent assessment. It is not an award, endorsement, or contract.

Conclusion

Currently, the generation of textual output is not a bottleneck anymore, but it has been known since 1976 that the constraint lies with the person doing the reading. AI changed how much gets delivered and how cheaply it can be done, but how does it narrow the aperture or number of things it asks a human to review and sign off on?

It looks easy and cheap, but this paper suggests that it remains so only if the unflagged items can be trusted. Otherwise, the program buys the compute and then returns to ask the reviewer for more hours anyway.

Get in touch with the author below for an executive brief focused on one named software initiative: what it costs to review that baseline today and where lineage separation moves that number. Not a demo. Try or find out more about Alchemist AI Pro™ and Alchemy SDLC™ at https://alchemistaipro.com.

References

ACC3 International. (2026a, July). Alchemy Pro CRM: Replacing our own CRM with the Alchemy SDLC™. https://alchemistaipro.com/library/replacing-our-crm-using-alchemy-sdlc

ACC3 International. (2026b). Product Campaign Manager [Build receipt]. https://alchemistaipro.com/proof/pcm

Chief Digital and Artificial Intelligence Office. (n.d.). Tradewinds. U.S. Department of War. Retrieved August 11, 2026, from https://www.ai.mil/Industry/Tradewinds/

Consortium for Information and Software Quality. (2022). The cost of poor software quality in the US: A 2022 report. https://www.it-cisq.org/the-cost-of-poor-quality-software-in-the-us-a-2022-report/

Defense Science Board. (2018). Design and acquisition of software for defense systems. U.S. Department of Defense. https://www.cto.mil/wp-content/uploads/2023/07/DSB-SWA-Report-2018.pdf

Edwards, N. (2026a). Inherited blind spots (Part 1): Eliminating common-mode failure through dual-model audits. ACC3 International. https://alchemistaipro.com/library/inherited-blind-spots-part-1

Edwards, N. (2026b). Inherited blind spots (Part 2): Closing the aperture. ACC3 International. https://acc3int.com/whitepapers/inherited-blind-spots-part-2

Fagan, M. E. (1976). Design and code inspections to reduce errors in program development. IBM Systems Journal, 15(3), 182-211. https://doi.org/10.1147/sj.153.0182

Harvey, N., & DeBellis, D. (2025, September 23). Announcing the 2025 DORA report: State of AI-assisted software development. Google Cloud Blog. https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report

Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems, 37, 68772-68802. https://proceedings.neurips.cc/paper_files/paper/2024/hash/7f1f0218e45f5414c79c0679633e47bc-Abstract-Conference.html

Triezenberg, B. L., Zabel, S., Steratore, R., Salas, A., Lepetic, I., Wilson, K. A., Henriquez Sanchez, N., Fan, J., Levedahl, A., & Denton, S. W. (2025). Underperforming software and information technology in the Department of Defense (RR-A2927-1). RAND Corporation. https://www.rand.org/pubs/research_reports/RR-A2927-1.html

Tsui, K. (2025). Self-correction bench: Uncovering and addressing the self-correction blind spot in large language models. arXiv. https://doi.org/10.48550/arXiv.2507.02778

U.S. Government Accountability Office. (2023). Defense software acquisitions: Changes to requirements, oversight, and tools needed for weapon programs (GAO-23-105867). https://www.gao.gov/products/gao-23-105867

© 2026 AI Pro Holdings, Inc. All rights reserved.