Proof of Build
THE RECEIPTTenkiller Receipts
Scope11 features · 106 use cases · 765 planned tasks
Estimated the traditional way2,892 story points × 4 hours = 11,568 hours ≈ 72 team-weeks
Build one — human hours38, one non-developer directing the crew
Build two — human hours433 logged, across five senior engineers
CalendarBuild one: 21 days · Build two: 11 weeks, part-time
CommitsBuild one: 6,813 · Build two: 162 — separate repos, no shared authors
Scope carriedBuild one: all 106 · Build two: 58, phase one of a deliberate two-phase split
The 2-hour testBuild one: ≈ 21 min per use case against a 120-minute budget
Requirements, both buildsunder 12 hours, once, with no standing meeting after it

Everything on this page comes from the project’s own records: two git repositories, two time logs, and the estimate written before the first commit.

The yardstick

Before either build started, the package was estimated the traditional way — story points, times hours per point.

11 features · 106 use cases (52 with a user interface) · 765 planned tasks
2,892 story points  ×  4 hours  =  11,568 estimated development hours
≈ 72 team-weeks. Call it seventeen months.

The arithmetic is on the page on purpose. Same inputs, same number, every time — that is what makes it a yardstick rather than a claim.

Why this counts as an experiment

Two repositories. Not one person committed to both.

BUILD ONEautomated crew 6,788engineer 13product owner 10recovery + backstop 26,813 commits
BUILD TWOdeveloper 74developer 80supporting 4manager + supporting 4162 commits

Separate repositories, separate authors, no overlap. We are not asking you to accept that the builds were independent. These are the commit logs.

Measure
BUILD ONEA product owner with no development background, directing the automated crew
BUILD TWOFive people, all onshore, senior — using a coding assistant
Human hours
38Reconstructed from commit timestamps, day by day.
433Logged, across five contributors.
Calendar
21 days
11 weeks, part-time around client work
How it was built
The crew executed 714 tasks in three build waves, then twelve final-mile episodes hardened it. Four deployment attempts failed before one worked. We print that, because that is what a receipt is.
Story by story: specify, code, review, merge, deploy — seventeen stories, worked in the gaps around client delivery.
Scope carried
All 106 use cases stayed in scope.The crew does not triage — it executes the package it is given.
58 of 106 use cases — phase one of a deliberate two-phase split.The team could not timeline the full package in the hours available, so they made the professional call: phase it. The other 48 use cases wait for phase two — with its own hours, calendar, and budget.
What shipped
11 features · 516 API routes · 70 database tables · customer portal and owner admin. About 85% of the estimated scope, with the remainder backlogged and listed: payments, email, auth wiring, rate limiting, the security checklist.
Production. This is the build the business chose to run — carrying the reduced scope above.
Traceability
Task per use case, by construction.
The team went its own way from the specs, so an item-for-item trace is no longer possible. Alchemist takes this system back in the way it takes in any brownfield system — as a legacy application — read from the code, with the original documents as inputs. The round trip closed the day the specs were left behind.

The number nobody prints

Build two’s five senior engineers could not timeline the whole package in the hours they had. So they made the professional call: split it into two phases, and carry 58 of the 106 use cases first.

That is not a criticism of the team — phasing is exactly what disciplined delivery does when human bandwidth cannot carry full scope on the calendar. But notice what the deliberate call actually is: the schedule is where human-hour delivery pays for scope. Phase two is real work still ahead — its own hours, its own calendar, its own coordination — and none of it appears in the build’s headline numbers. It happens on delivery floors every day, and it is almost never printed.

Build one — a single non-developer directing the automated crew — never needed a phase two. All 106 use cases stayed in one build — because the crew’s capacity is not priced in human hours, so the scope never had to be split to fit them.

BUILD ONE — one non-developer, directing the crew  →  38 hours ÷ all 106 use cases carried  ≈  21 minutes per use case
BUILD TWO — five senior engineers, coding assistant  →  433 hours ÷ 58 phase-one use cases  ≈  7.5 human hours per use case

The requirements didn’t change. The delivery method turned them into a schedule.

The standard we hold ourselves to

We budget two hours of human time per use case for the final mile. No build has spent it. Build one came in at about 21 minutes per use case against a 120-minute budget — 38 hours across 106 use cases.

What the estimate leaves out

The 11,568 hours cover setup, design, use cases, integration, QA and final tasks. There is no discovery line. That estimate begins the moment good requirements already exist — and producing them is its own project.

In a traditional delivery that means workshops, stakeholder interviews, process mapping, a written specification, review cycles and sign-off. Then, if any of it is built at distance, a standing review rhythm: twice a day at the start, once a day from the third sprint onward, for the life of the build. Those hours are real, they are senior, and they are almost never in the estimate.

Alchemist produced this package in under 12 hours of human time. Once. With no standing meeting after it.

The review rhythm is worth a second look. Twice a day early, once a day later — because ambiguity is highest before anything is built, and those meetings exist to resolve it. The cadence is a readout of how unclear the requirements were. Every delivery organisation already budgets for that. Few of them price it.

We are not going to tell you what your developers cost

You know your blended rate. We do not, and any number we invented for it would be the weakest thing on this page. So here are the hours. Bring your own rate.

Estimated the traditional way  →  11,568 hours  ×  your rate  (full 106-use-case package)
Built by five senior engineers  →  433 hours  ×  your rate  (phase one: 58 of 106 use cases, deliberately split)
Built by one non-developer  →  38 hours  ×  your rate  (all 106 in scope)
Requirements, both times  →  under 12 hours  ×  your rate

Whatever number you multiply by, it is the same number in all four rows. That is the point.

None of this is a surprise to anyone who has measured it. From the Consortium for Information & Software Quality’s 2022 report on the cost of poor software quality in the United States:

“The cost of finding and fixing deficiencies is the largest single expense element in the software development lifecycle. Over a 25-year life expectancy of a large software system, almost fifty cents out of every dollar will go to finding and fixing bugs.

The earlier in the development lifecycle deficiencies are found, the more economical the overall delivery will be.

Krasner, H. (2022). The Cost of Poor Software Quality in the US: A 2022 Report. Consortium for Information & Software Quality. The report puts the total at $2.41 trillion, with accumulated technical debt near $1.52 trillion.

What we are not claiming

That these two builds finished in the same place. They did not — in either direction. Build one carried all 106 use cases to about 85% depth and stopped where we chose to stop it; finishing it would take more hours. Build two went to production and runs the business — carrying phase one, 58 of the 106 use cases, after a deliberate decision to split the package into two phases because the hours available could not timeline all of it. One build is broader, the other is harder. We print both facts rather than flatter either build.

That the estimate was wrong. It was computed properly from the package in front of it. The question this page asks is not whether the arithmetic held — it is what the package was worth before anyone started multiplying.

That your team would hit 433 hours. Ours did, and none of the five had used the tool before. That is one result on one system, and we would rather show you one honest number than a range we cannot support.

That software delivery stops needing judgement. Three seats still belong to people: deciding what to build, finishing the last mile, and accepting the result. Everything on this page is what happens when those three are done well and the rest is not done by hand.

Build one — Tenkiller cabin listings, demonstrable.
Build one: demonstrable.
Build two — the live Tenkiller Hideouts booking page, in production.
Build two: in production - Visit the site.

If your own estimate says months and your instinct says it should not, bring us the scope and we will run the requirements against it.

jamie@acc3int.com

Alchemist AI Pro™ is free to use — alchemistaipro.com

Figures drawn from the project’s git history, contributor time logs, and the estimate written before the first commit. Tenkiller Hideouts is a real customer system; the build is published with their knowledge.

© 2026 AI Pro Holdings, Inc. All builds verified. All receipts public.