1 · Choose the design

Three factors, each at a low level (−1) and a high level (+1). The choice is how many of the eight possible combinations you are willing to pay for.

Experimental design
Defining relation:

2 · The coded matrix and your responses

Every column below is a contrast: multiply the letters together, run by run. Type a measured response for each run, or paste the whole column at once.

Coded design matrix with one editable response per run

Tab moves from one response to the next in run order. The estimates below update as you type.

3 · What each contrast actually estimates

In a fraction, two terms can share a column exactly. When that happens the experiment cannot tell them apart — not with more care, not with more precision, not ever. Only more runs can.

4 · Estimated effects

An effect is the mean response at the high level minus the mean at the low level. The regression coefficient is half of it, because the coded levels are two units apart.

Fill in a response for every run to see the estimates.

5 · What to run next

Screening is a first move, not a verdict. These notes change with the design and with your numbers.

6 · Assumptions this lab is making

  1. Two levels, coded −1 and +1, and nothing in between. Every estimate here is a difference between two averages. It says nothing about what happens between the levels. If the response bends — if the best setting is in the middle — a two-level design will report a small effect and be badly wrong about why. Centre points are the standard check, and this lab does not include them.
  2. The response is measured once per run. No replication, so no independent error estimate, so no p-values. This is the single most consequential assumption on the page.
  3. Runs are performed in random order, even though they are listed in standard order. The table is sorted so the sign patterns are readable. Executing the runs in that order confounds every effect with time. Randomise before you run; the analysis does not care about the order, but the physics does.
  4. Effects add up. The model behind the arithmetic is linear in the coded factors plus their products. Multiplicative behaviour, saturation and thresholds all violate it, and the usual repair is to analyse a transformed response rather than the raw one.
  5. Errors have constant spread across the design region. If the response is noisier at high settings than at low ones, the effects are still unbiased but the visual comparison of their sizes is misleading.
  6. In the fraction, the alias chains are exact, not approximate. The chains shown in section 3 are algebraic identities of the design, not statistical statements. Interpreting ℓ(C) as C alone is an assumption about the physics — namely that AB is negligible — and the design gives you no evidence for it.
  7. Lenth’s pseudo standard error assumes effect sparsity. It only works because most contrasts are expected to be inert. In a design where nearly everything is active, the trimmed median is inflated by real effects and the margin of error becomes too wide to reject anything.
  8. The foldover blocks are two separate batches. The lab treats them that way and puts the batch difference on ABC, which is where the algebra puts it. If both blocks genuinely ran under identical conditions, that contrast is a clean ABC estimate — but that is a claim about your lab, not about the design.

7 · A worked example, start to finish

The numbers loaded into the table are from this example. Change the design above and watch the same underlying truth get reported differently.

The setting. A distribution centre wants to raise picks per labour hour on one line. Three changes are on the table: A, doubling the batch size a picker carries; B, re-slotting the fastest movers to the golden zone; C, swapping the cart for a powered one. Each is either in place (+1) or not (−1).

The truth, which the experiment does not know. Suppose the real response is

y = 50 + 5A + 3B + 1C + 1·AB

so the true effects — twice the coefficients — are A = 10, B = 6, C = 2 and AB = 2, with AC, BC and ABC all exactly zero. The eight runs then read 42, 50, 46, 58, 44, 52, 48 and 60 picks per hour, in standard order.

  1. The full 2³ finds all of it. A’s effect is the average of the four high-A runs minus the average of the four low-A runs: (50+58+52+60)/4 − (42+46+44+48)/4 = 55 − 45 = 10. Working the other six columns the same way returns B = 6, C = 2, AB = 2 and zero for the rest. Eight runs, seven contrasts, no confounding, and — note — nothing left over to estimate error with.
  2. The half fraction gets three of them cheaply, and one of them wrong. Setting C = AB keeps runs c, a, b and abc, giving responses 44, 50, 46 and 60. Now
    ℓ(A) = (50+60)/2 − (46+44)/2 = 55 − 45 = 10,
    ℓ(B) = (46+60)/2 − (50+44)/2 = 53 − 47 = 6,
    ℓ(C) = (44+60)/2 − (50+46)/2 = 52 − 48 = 4.
    The first two happen to be right. The third is not: ℓ(C) estimates C + AB = 2 + 2 = 4. Four runs bought the same three main effects for half the money, and the price was that the cart effect now reads double its true size because the batch-and-slotting interaction is sitting on the same column.
  3. The foldover pulls them apart. Running the mirror image — the same four runs with every sign reversed — gives a second estimate of that column, and this time the generator is C = −AB, so it estimates C − AB = 2 − 2 = 0. Two numbers, 4 and 0, for two unknowns:
    C = (4 + 0) / 2 = 2, and AB = (4 − 0) / 2 = 2.
    Both recovered exactly. The eight runs together are the full 2³, which is the point of the mirror foldover: it does not add information from nowhere, it completes the design.
  4. What the foldover costs. The two halves were run as two batches, and the contrast that separates batch 1 from batch 2 is precisely the ABC column. If the second batch happened on a different shift, that shift difference lands on ABC and cannot be told apart from a genuine three-factor interaction. Load the foldover design in the table and add, say, 6 to each of the four block-2 responses: every main effect and every two-factor interaction stays exactly where it was, and ABC alone moves by −6.
  5. Where this leaves the decision. Batch size is worth about 10 picks per hour and slotting about 6, and they reinforce each other by another 2. The powered cart is worth about 2 on its own. With one observation per run none of that carries a standard error, so the honest next step is to replicate the two or three settings that matter rather than to quote a significance level for all seven.

References and further reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, section 5.3.3.4.3, “Fractional factorial design specifications and design resolution” — the source for the generator-and-defining-relation machinery and the resolution numbering used above.
    https://www.itl.nist.gov/div898/handbook/pri/section3/pri3343.htm
  • NIST/SEMATECH, e-Handbook of Statistical Methods, section 5.3.3.8.1, “Mirror-image foldover designs” — the source for what a foldover does to resolution and for the single-factor foldover mentioned in section 5.
    https://www.itl.nist.gov/div898/handbook/pri/section3/pri3381.htm
  • Lenth, R. V. (1989), “Quick and easy analysis of unreplicated factorials”, Technometrics 31(4), 469–473 — the pseudo standard error and the two margins of error reported in section 4.
  • Daniel, C. (1959), “Use of half-normal plots in interpreting factorial two-level experiments”, Technometrics 1(4), 311–341 — the plot in section 4.

These references explain the methods used in the lab. For acknowledgements across the site, see References and acknowledgements.