By SigmaForge, October 2026
At a glance
Executive summary
Most improvement programs do not fail for lack of training. They fail in the gap between a certificate on the wall and a result that finance will sign. AI can narrow that gap or widen it. Used carelessly, it produces fluent charts and confident savings figures that nobody can trace. Used with discipline, it takes the slow work off a team while every number stays checkable and every decision stays with a named person.
This paper sets out three tests for AI in improvement work, shows how SigmaForge is built to meet each one, and ends with a checklist you can put to any vendor, including us.
- The numbers are computed by code, and the model only explains them.
- Tollgates and charters are decided by people, and nobody approves their own work.
- Savings are projected, claimed and then validated by someone who did not lead the project.
01
The problem: training that does not become results
Every operations leader has seen the pattern. A cohort is trained and certified. Projects are chartered with energy. Six months later a few have closed, more have stalled somewhere in Measure, and the savings total presented to the leadership team is the sum of the estimates written on day one.
Call it slideware ROI. A figure is typed into a charter at kickoff, copied into every review deck, and never reconciled against an invoice or a ledger. Nobody lied. Nobody checked either, because the person presenting the number was usually the person who produced it, and often the person approving the next phase as well.
The symptoms are familiar: projects that stop once the data gets hard; reviews in which the presenter and the approver are the same person; and a program total that finance quietly discounts. Over time the program loses the one thing it cannot buy back, which is belief in its numbers.
AI adds a new version of the old risk. A language model will happily describe a control chart it never computed, quote a p-value it never calculated, and estimate a saving from a sentence. The prose is fluent, so the error is harder to see. If a program was already struggling to tell an estimate from a result, a tool that generates plausible numbers makes the gap wider, not narrower.
02
Computed statistics, not generated ones
Language models are good at explaining and poor at arithmetic they cannot show. So the first test is simple: for every number in a deliverable, could you say which line of code produced it?
In SigmaForge, 28 of the 32 DataForge analysescompute every figure in a deterministic statistics engine: capability indices, ANOVA, regression, control limits, Gage R&R, hypothesis tests and the rest. The model is handed the finished result and asked to do the part it is good at, which is saying what the result means and what to do next. The same data in gives the same numbers out, every time.
That matters most at review. A reviewer can rerun an analysis and get the identical figure. A Cpk in a tollgate pack can be traced back to its data. If the wording of an interpretation is off on a given day, the numbers underneath it are not.
You do not have to take this on trust. The free capability calculator on this site runs the same engine in your browser, with no account and no server call. Give it the same measurements twice and compare.
03
Decisions stay human
The second test is about who decides. AI can do a great deal of preparation, and it should. What it should not do is approve work, and neither should the person who did the work.
Every SigmaForge project runs through 5 DMAIC tollgates. The agents prepare for each one: phase coaching names the one thing blocking the phase, which evidence is attached and which is missing, and what the numbers say today. When the team submits a gate, the reviewer receives it already read. Then a person decides. Only a manager, an admin or the owner can review a gate, and the product refuses an approval from the person who submitted it. A gate sent back carries the reviewer’s reason.
The same rule runs earlier in the life of a project. Ideas are scored on impact, effort, confidence and fit by arithmetic, not by who asked loudest, and the person who submitted an idea cannot charter it into a project. In a university course the pattern is the same: an AI pre-read of every capstone submission against the rubric, and the professor decides every grade.
Segregation of duties only works if it leaves evidence, so every submission, approval, return, comment and deliverable is recorded with who did it and when. A whole project exports as one dossier that a sponsor, an auditor or a successor can read without asking anyone.
04
Savings you can defend
The third test is the one finance cares about. SigmaForge keeps three savings figures apart.
- Projectedthe estimate written at charter, before anything has been tried.
- Claimedthe figure the team reports when the project closes.
- Validatedthe figure signed off by an owner or admin who did not lead the project, with a note on what it was checked against. The product refuses a validation from the project lead.
The executive results view shows all three side by side, so the distance between them is visible rather than discovered at year end. That distance is useful information. A claim that validates close to its figure says the team measured well. A large gap says the estimating needs work, and says it early.
The product’s own demo companies show how this reads. In the Corvane Pumps demo, three closed projects claimed $770,000 and were validated at $717,000, about 93% of the claim, each signed off below the team’s figure by a manager who led none of them. In the Acme Manufacturing demo, a forklift project claimed $92,000 and was validated at $88,000, because $4,000 of the claim was a one-time saving counted as annual. Both are demo data, not customer results; the illustrative case studies tell them in full.
05
The AI that shows its assumptions
Every model fills gaps. The question is whether it tells you. A control plan drafted from a short description has to assume a measurement frequency, a reaction plan, an owner. A savings estimate has to assume a volume and a cost. Those assumptions are where a reviewer should start reading, so they should never be hidden in the prose.
DocuForge and ImpactMatrix return the inferences they made as a separate list of assumptions, shown with the result and carried into the exported PDF. Where a fact matters and was not given, such as a baseline, an element time or a person’s name, the draft writes [TBD] in its place instead of inventing one. A team table names real people only where the input named them. The phase coach shows a missing number as [TBD] too, because its absence is itself worth seeing.
The tools also read from the project record: the charter and the deliverables already saved to the project are passed in as facts, not as suggestions. A control plan is written from the same record the team keeps, so it cannot quietly contradict the baseline the team measured in Measure.
For a reviewer this changes the job. Instead of asking whether a document sounds right, ask which of its assumptions are wrong. That is a faster and far more honest conversation.
06
From training to results
Training is the input, not the output. SigmaForge teaches Lean Six Sigma across 4 belts, White to Black, with narrated lessons, a quiz at the end of every module, and an exam with a 70% pass mark. A certificate carries an ID and a QR code that open a public record, so anyone can check it without contacting the holder or us.
But a certificate is not a saving. The point of training is a project, so the projects live in the same place as the lessons. In an enterprise workspace, the people who learn are the people who run DMAIC projects with the agents, and a leader sees workforce readiness by belt beside the project results. In a university course, every student who joins gets their own DMAIC capstone, reviewed at 5 gates by the professor, and the course certificate names who supervised it.
For individual learners the same holds. The Green and Black Belt certificates also require a capstone project: a real DMAIC project in the learner's own workspace, reviewed phase by phase at5 tollgates, so the certificate is issued only once the work is proven.
07
How to evaluate a platform
These questions work on any vendor, including us. Ask them in a demo, on your own data if you can, and judge the answer you see on screen rather than the one on a slide.
- 01
Which numbers does code compute, and which does the model write?
A good answer: A named list of analyses whose figures come from a statistics engine, and a plain statement of which outputs are written by a model.
- 02
If I run the same data twice, do I get the same figures?
A good answer: Yes, to the last decimal. A capability index or a p-value that changes between runs was generated, not computed.
- 03
Who can approve a tollgate, and can the person who submitted it approve it?
A good answer: Named roles review gates, and the product itself refuses a self-approval. A policy document is not the same thing.
- 04
How are savings recorded at each stage?
A good answer: Projected, claimed and validated figures kept apart, a validator who did not lead the project, and a note saying what the figure was checked against.
- 05
What does the AI do when it lacks a fact?
A good answer: It marks the gap where the fact belongs and lists every inference it made, instead of filling the space with something plausible.
- 06
Is there a trail I could hand to an auditor?
A good answer: Every submission, decision, comment and deliverable recorded with who and when, and a project that exports as one document.
- 07
Does the training connect to the project work?
A good answer: The same people learn and run projects in the same place, so a manager can see readiness beside results rather than in a separate system.
- 08
Can a third party check a certificate without contacting you?
A good answer: A public record behind an ID or QR code on the certificate, open to anyone, with no sign-in.
- 09
What does the platform not do?
A good answer: A short, specific list. A vendor that cannot name its gaps has not looked for them.
- 10
Can I see it on my own data before I commit?
A good answer: A session on a spreadsheet of yours, with the analysis saved into a project you can inspect afterwards.
08
About SigmaForge
SigmaForge brings Lean Six Sigma training, project work and proof into one platform: 4 belts, 13 AI agents with one job each, 56 tools including 32 DataForge analyses, and 5 DMAIC tollgates that people decide. It is used by organizations running improvement programs, by universities teaching quality engineering, and by individual learners working through the belts.
The fastest way to judge it is to see it on a spreadsheet of yours. Book thirty minutes and we will run the analyses on your data, save them into a project, and show you where every number came from.
