- What it is
- A one-stop reference for the formulas in the dojo
- Covers
- Descriptive statistics through DOE
- Includes
- Control chart constants, critical values, and tail areas
- Each formula links to
- The page with the worked example
- Use the browser Print command
- Excel
- Function names are listed beside the formulas
- Minitab
- Menu paths are on each topic page
- Why it matters
- Know which formula answers which question
How to Use This Sheet
Every formula used in Stat Dojo is here, grouped by topic, with a link to the page that explains it and works an example. Symbols: x̄ sample mean, s sample standard deviation, μ and σ population mean and standard deviation, n sample size, p̂ sample proportion, α significance level, df degrees of freedom.
Describing Data
| Quantity | Formula | Notes | Page |
|---|---|---|---|
| Mean | x̄ = Σx / n | Excel AVERAGE | Descriptive |
| Median | Middle value of the sorted data | Average of the two middle values when n is even | Descriptive |
| Sample variance | s² = Σ(x − x̄)² / (n − 1) | Excel VAR.S | Descriptive |
| Sample standard deviation | s = √s² | Excel STDEV.S | Descriptive |
| Range, IQR | R = max − min; IQR = Q3 − Q1 | Excel QUARTILE.EXC matches Minitab | Descriptive |
| Coefficient of variation | CV = s / x̄ × 100% | Ratio-scale data only | Descriptive |
| z score | z = (x − μ) / σ | Standard deviations from the mean | Distributions |
| Standard error of the mean | SE = s / √n | Shrinks with the square root of n | Sampling and CLT |
| Histogram bins (Sturges) | k = 1 + log2(n) | A starting point | Graphical Analysis |
| Outlier fences (box plot) | Q1 − 1.5 IQR and Q3 + 1.5 IQR | Graphical Analysis |
Distributions
| Distribution | Probability | Mean and standard deviation | Excel |
|---|---|---|---|
| Normal | z = (x − μ) / σ; look up the area | μ, σ | NORM.DIST, NORM.INV |
| Binomial | P(k) = C(n, k) pk (1 − p)n−k | np; √(np(1 − p)) | BINOM.DIST |
| Poisson | P(k) = e−λ λk / k! | λ; √λ | POISSON.DIST |
| Weibull | R(t) = exp[−(t/η)β] | Mean = η Γ(1 + 1/β) | WEIBULL.DIST |
| Yield from defects per unit | Yield = e−DPU | EXP |
Pages: Distributions, Reliability and Weibull.
Confidence Intervals and Sample Size
| Interval | Formula | Page |
|---|---|---|
| Mean (unknown σ) | x̄ ± tα/2, n−1 s / √n | Confidence Intervals |
| Proportion (Wald) | p̂ ± zα/2 √(p̂(1 − p̂)/n); prefer the Wilson or exact interval for small counts | Confidence Intervals |
| Standard deviation | s √((n − 1)/χ²upper) to s √((n − 1)/χ²lower) | Confidence Intervals |
| Difference of two means (Welch) | (x̄1 − x̄2) ± t √(s1²/n1 + s2²/n2) | t-Tests |
| Sample size for a mean | n = (zα/2 σ / E)² for margin E | Sample Size and Power |
| Sample size for a proportion | n = z² p(1 − p) / E²; use p = 0.5 if unknown | Sample Size and Power |
| Sample size for a two-sample t-test | n per group ≈ 2 (zα/2 + zβ)² σ² / δ² | Sample Size and Power |
Hypothesis Tests
| Test | Statistic | Distribution | Page |
|---|---|---|---|
| One-sample t | t = (x̄ − μ0) / (s / √n) | t, n − 1 df | t-Tests |
| Two-sample t (Welch) | t = (x̄1 − x̄2) / √(s1²/n1 + s2²/n2) | t, Welch df | t-Tests |
| Paired t | t = d̄ / (sd / √n) on the differences | t, n − 1 df | t-Tests |
| One proportion | z = (p̂ − p0) / √(p0(1 − p0)/n) | Normal; use the exact test for small counts | Tests for Proportions |
| Two proportions | z = (p̂1 − p̂2) / √(p̂(1 − p̂)(1/n1 + 1/n2)), pooled p̂ | Normal; Fisher exact for small counts | Tests for Proportions |
| Two variances (F) | F = s1² / s2² | F, n1 − 1 and n2 − 1 df | Tests for Variances |
| One variance | χ² = (n − 1) s² / σ0² | Chi-square, n − 1 df | Tests for Variances |
| Levene (Brown-Forsythe) | ANOVA on |x − group median| | F | Tests for Variances |
| Chi-square association | χ² = Σ(O − E)²/E; E = row total × column total / grand total | Chi-square, (r − 1)(c − 1) df | Chi-Square Tests |
| Goodness of fit | χ² = Σ(O − E)²/E | Chi-square, categories − 1 − fitted parameters | Chi-Square Tests |
| Mann-Whitney, Wilcoxon, Kruskal-Wallis | Based on ranks | Normal approximation or exact tables | Nonparametric Tests |
| Anderson-Darling normality | A² measures distance between the data and the normal cumulative curve | Compare to the critical value or p-value | Normality Tests |
| Decision | Rule | Page |
|---|---|---|
| p-value | Reject H0 if p < α; p is the chance of a result this extreme if H0 is true | P-Values and Error Types |
| Type I error | Reject a true H0; probability = α | P-Values and Error Types |
| Type II error | Fail to reject a false H0; probability = β; power = 1 − β | P-Values and Error Types |
ANOVA
| Quantity | Formula | Page |
|---|---|---|
| Between-group sum of squares | SSB = Σni(x̄i − x̄̄)², df = k − 1 | One-Way ANOVA |
| Within-group sum of squares | SSW = Σ(x − x̄i)², df = N − k | One-Way ANOVA |
| F statistic | F = MSB / MSW = (SSB/(k − 1)) / (SSW/(N − k)) | One-Way ANOVA |
| R² | SSB / SStotal | One-Way ANOVA |
| Tukey comparison | |x̄i − x̄j| > (q / √2) √(MSW(1/ni + 1/nj)) | Post-Hoc Comparisons |
| Two-way ANOVA | Separate sums of squares for A, B, A × B, and error | Two-Way ANOVA |
Correlation and Regression
| Quantity | Formula | Page |
|---|---|---|
| Correlation | r = Σ(x − x̄)(y − ȳ) / √(Σ(x − x̄)² Σ(y − ȳ)²) | Correlation |
| Test of r | t = r √(n − 2) / √(1 − r²), n − 2 df | Correlation |
| Slope | b1 = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)² | Simple Regression |
| Intercept | b0 = ȳ − b1x̄ | Simple Regression |
| R² | 1 − SSerror / SStotal | Simple Regression |
| Adjusted R² | 1 − [SSerror/(n − p − 1)] / [SStotal/(n − 1)] | Multiple Regression |
| Residual standard error | s = √(SSerror / (n − p − 1)) | Simple Regression |
| Variance inflation factor | VIF = 1 / (1 − Rj²) | Multiple Regression |
| Confidence interval for the mean response | ŷ ± t s √(1/n + (x0 − x̄)²/Sxx) | Simple Regression |
| Prediction interval for one value | ŷ ± t s √(1 + 1/n + (x0 − x̄)²/Sxx) | Simple Regression |
Control Charts, Capability, Reliability, and DOE
| Quantity | Formula | Page |
|---|---|---|
| Individuals chart | x̄ ± 2.660 M̄R; MR upper limit 3.267 M̄R | Control Chart Theory |
| X-bar chart | x̄̄ ± A2 R̄ | Control Chart Theory |
| R chart | D3 R̄ to D4 R̄ | Control Chart Theory |
| Sigma from ranges | σ̂ = R̄ / d2; from moving ranges, M̄R / 1.128 | Control Chart Theory |
| Average run length (in control) | ARL0 = 1 / (2 P(Z > 3)) = 370 | Control Chart Theory |
| Cp, Cpk | Cp = (USL − LSL)/(6σwithin); Cpk = min(CPU, CPL), CPU = (USL − μ)/(3σ) | Capability Statistics |
| Pp, Ppk | The same with the overall standard deviation | Capability Statistics |
| Z bench | Z = Φ−1(1 − total proportion out of spec) | Capability Statistics |
| Weibull B10 life | η (−ln 0.9)1/β | Reliability and Weibull |
| Factorial effect | Mean at +1 − mean at −1; coded coefficient = effect / 2 | Analyzing Designed Experiments |
| Factorial sum of squares | N × effect² / 4 | Analyzing Designed Experiments |
| Standard error of an effect | 2s / √N | Analyzing Designed Experiments |
Constants and Critical Values
Control chart constants (computed from the distribution of the range of normal samples):
| Subgroup n | d2 | A2 | D3 | D4 |
|---|---|---|---|---|
| 2 | 1.128 | 1.880 | 0.000 | 3.266 |
| 3 | 1.693 | 1.023 | 0.000 | 2.575 |
| 4 | 2.059 | 0.729 | 0.000 | 2.282 |
| 5 | 2.326 | 0.577 | 0.000 | 2.114 |
| 6 | 2.534 | 0.483 | 0.000 | 2.004 |
| 7 | 2.704 | 0.419 | 0.076 | 1.924 |
| 8 | 2.847 | 0.373 | 0.136 | 1.864 |
| 9 | 2.970 | 0.337 | 0.184 | 1.816 |
| 10 | 3.078 | 0.308 | 0.223 | 1.777 |
Normal and t critical values for two-sided intervals:
| Confidence | z | t (df = 5) | t (df = 10) | t (df = 20) | t (df = 30) |
|---|---|---|---|---|---|
| 90% | 1.645 | 2.015 | 1.812 | 1.725 | 1.697 |
| 95% | 1.960 | 2.571 | 2.228 | 2.086 | 2.042 |
| 99% | 2.576 | 4.032 | 3.169 | 2.845 | 2.750 |
| 99.9% | 3.291 | 6.869 | 4.587 | 3.850 | 3.646 |
Tail areas beyond k standard deviations (normal):
| k | One-sided | Two-sided | Two-sided per million |
|---|---|---|---|
| 1 | 0.158655 | 0.317311 | 317,310.5 |
| 1.645 | 0.049985 | 0.099970 | 99,969.8 |
| 1.96 | 0.024998 | 0.049996 | 49,995.8 |
| 2 | 0.022750 | 0.045500 | 45,500.3 |
| 2.576 | 0.004998 | 0.009995 | 9,995.1 |
| 3 | 0.001350 | 0.002700 | 2,699.8 |
| 4 | 0.000032 | 0.000063 | 63.3 |
| 4.5 | 0.000003 | 0.000007 | 6.8 |
| 5 | 0.000000 | 0.000001 | 0.6 |
| 6 | 0.000000 | 0.000000 | 0.0 |
Sigma levels and defects per million (long-term, with the conventional 1.5 sigma shift): 3 sigma = 66,807 ppm; 4 sigma = 6,210 ppm; 5 sigma = 233 ppm; 6 sigma = 3.4 ppm.
Statistics Formula Sheet: Frequently Asked Questions
Is this formula sheet printable?
Yes. Use your browser’s Print command. The tables are laid out to print on a few pages.
Do I need to memorize these?
No. Software does the arithmetic. What matters is knowing which formula answers which question and what each part means. Use this sheet to look things up.
Where do the control chart constants come from?
They are derived from the distribution of the range of samples from a normal distribution. The Control Chart Theory page shows the derivation and the table is computed numerically.
Which formulas does Minitab use?
For the methods covered here Minitab uses the same standard formulas. Where options differ, for example quartile definitions or capability intervals, the topic page says so.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods (itl.nist.gov/div898/handbook).
- Douglas C. Montgomery, Introduction to Statistical Quality Control and Design and Analysis of Experiments, Wiley.
- David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
- Minitab Support, “Methods and formulas” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.