Anyone who works in manufacturing ends up making decisions on data that is never perfectly clean. Is this measurement good enough to act on? Did the change actually make a difference? How does this parameter drive the result? Can we accept this delivery? Six Sigma gives you the statistics to answer those questions. Most of us just do not use them often enough to remember which test applies, or how to set it up.

This toolkit started during my Six Sigma Black Belt. I wanted one place where the theory, the choice of test and the calculation sit side by side, and where every result shows its working. It grew into something useful for anyone who has to deal with Six Sigma and statistics, so I am sharing it.

A use case for building with AI

I built the toolkit with AI as a development partner. A tool of this scope used to be far too big for a side project: dozens of statistical methods, each with its own formulas, edge cases and explanation. With AI doing much of the heavy lifting, the effort goes where it should: deciding what the tool needs to do, checking that the statistics are right and shaping it around how people actually work.

That is the broader point for manufacturing. AI does not replace the domain knowledge; it lets you turn that knowledge into working solutions faster, and iterate on them until they fit. Faster building does not make checking optional, though. A built-in self-test compares every calculation against reference values from scipy, the standard scientific library for Python.

What is in it

The toolkit is a single HTML file that runs entirely in the browser. Nothing is sent anywhere, and once downloaded it works offline. Each module starts with the theory and then moves to the calculation. Every result shows the hypotheses, the formula, the formula with your numbers filled in, the result and the equivalent Excel formula. You can paste data straight from Excel.

It covers more than the four tools below: distributions, process capability (Cp, Cpk, Pp, Ppk, DPMO), SPC control charts, ANOVA, 2^k designed experiments, non-parametric tests and a formula sheet. One caveat: the interface and explanations are currently in Dutch. The formulas and Excel functions travel well regardless.

Hypothesis testing

A hypothesis test is a decision rule under uncertainty, not a proof. The null hypothesis is the boring default: nothing changed, there is no difference. You keep it until the data become too unlikely under it. There are two ways to be wrong. Alpha is the false alarm you accept in advance; beta is the real effect you miss. Think of a smoke detector that goes off without a fire, or stays silent when there is one. Failing to reject the null means there is not enough evidence, not that the two things are equal.

What people tend to forget is not the p-value but the choice of test. Is the question about a mean, a spread, a proportion or a relationship? One sample or two? Paired or independent? The toolkit starts with a test selector for exactly that reason. The other thing worth deciding up front is sample size, from alpha, beta, the spread and the smallest effect you care about. Too small and you will miss real effects. Very large and trivial differences become significant.

Use it when you need to know whether the new setting really reduced cycle time, whether one supplier’s material varies more than another’s, or whether the scrap rate moved after a fix.

Gauge R&R

Every measurement carries error. The variation you observe is the process variation plus the measurement variation. If the measurement part is large, you reject good parts and pass bad ones, the process looks less capable than it is, and real improvements disappear in the noise. So check the measurement system before you judge the process.

A Gauge R&R study splits measurement variation into repeatability, the spread when the same operator measures the same part several times, and reproducibility, the difference between operators measuring the same part. A typical study uses ten parts that cover the process range, two or three operators and two or three repeats each, in random order and without the operator knowing which part is which. As a rule of thumb, measurement variation below 10% of the total is good, 10 to 30% may be acceptable depending on the application, and above 30% is not. The number of distinct categories should be at least five. The toolkit covers both the average and range method and the ANOVA method, plus bias and linearity. Precision and accuracy are different things: a gauge can be precise and still consistently wrong, which calibration fixes and a Gauge R&R will not show.

Use it before a capability study, before trusting data in an improvement project, when a new gauge or a new team starts measuring, and whenever the line and the lab disagree.

Regression

Regression describes how a response changes, on average, with one or more inputs. The least squares line gives you a slope, how much the output moves per unit of input. A t-test on that slope tells you whether the relationship is real, and R² tells you how much of the variation it explains.

The distinction worth remembering is between the two intervals. The confidence interval around the line is about the average response at a given setting. The prediction interval is about a single new part, and it is always wider. Set a process window from the confidence interval and individual parts will surprise you. Check the residuals, do not extrapolate beyond the range you measured, and remember that a relationship found in historical data is not proof of cause. That is what designed experiments are for.

Use it to relate a process parameter to a quality characteristic, to set an operating window, or to build a calibration curve.

Acceptance sampling

A lot of 10,000 parts arrives. Inspecting every one is expensive, slow, sometimes destructive and not even error-free. So you take a sample of n parts and accept the lot if the number of defects is at most c. That is a hypothesis test on the lot.

The binomial distribution gives the probability of accepting a lot for any true fraction defective. Plot it and you get the operating characteristic curve, which tells you what a plan actually does. Two points anchor it. The acceptable quality level is the quality you want to accept almost always; the producer’s risk is the chance of rejecting such a lot anyway. The limiting quality level is the quality you want to reject; the consumer’s risk is the chance of accepting it anyway.

A plan of 100 parts with an acceptance number of 4 sounds strict. It accepts a lot with 2% defects 95% of the time, which is what you want. It also accepts a lot with 5% defects 44% of the time, which you may not. Sampling does not improve quality; it only helps you decide. The toolkit designs single and double plans as well as plans for measured characteristics, and draws the curve for each.

Use it for incoming goods inspection, for agreeing a sampling plan with a supplier, and for releasing batches where testing destroys the part.

Try it

The toolkit is free to use at sixsigma.genabyte.be. For offline use, download the single file from GitHub and run the self-test first: everything should be green.

It is a living tool. If you need a test that is missing, a version in English or a module shaped around a specific problem on your floor, get in touch and I am happy to improve it.