故事 · Fisher 与「不必一次只问一个问题」
Origin Story · Fisher and "Why Ask Just One Question at a Time?"
1920 年代,农学家 Ronald A. Fisher 在英国 Rothamsted 试验站研究怎样提高作物产量。
当时的传统是「一次只变一个变量」 —— 这一季只试化肥种类、下一季只试灌溉量。Fisher 看穿了它的两大浪费:
既慢,又看不见交互(某种肥料只有在充分灌溉时才显威力)。他提出革命性的因子设计:
把多个因子的所有水平组合同时安排在一批试验里,让每一次收成都同时携带多个因子的信息。
于是 2² 设计只需 4 块地,就能同时解出 A 的主效应、B 的主效应,以及它们的交互。
这套「全因子」思想,从此成为现代实验设计(DOE)的基石 —— 用同样的试验数,榨出多得多的因果信息。
In the 1920s, agronomist Ronald A. Fisher at England's Rothamsted Experimental Station was wrestling with crop-yield improvement.
The convention then was "change one variable at a time" — fertilizer this season, irrigation next. Fisher saw two wastes baked into that habit:
it was slow, and it was blind to interactions (a fertilizer that only pays off under generous irrigation looks worthless on its own).
His revolutionary alternative was factorial design: lay out every combination of factor levels at once within a single batch of trials,
so every harvest carries information about multiple factors simultaneously.
A 2² design on just four plots delivers the main effect of A, the main effect of B, and their interaction all at the same time.
Full-factorial thinking became the cornerstone of modern Design of Experiments (DOE) — same run count, vastly more causal information.
1 2² 设计方块:四个角,三个效应一次解出 The 2² Square: Four Corners, Three Effects in One Shot
拖滑块改响应Drag sliders2 交互图:两条线平行 = 无交互,越交叉 = 越强 Interaction Plot: Parallel = None, More Crossed = Stronger
横轴是因子 A 的低/高,两条线分别代表B 低与B 高。 若两线平行,A 的效应不随 B 变 —— 没有交互;若不平行甚至交叉,A 的最佳值取决于 B —— 这就是交互,OFAT 完全看不见它。 The horizontal axis runs A from low to high; the two lines stand for B low and B high. Parallel lines mean A's effect doesn't depend on B — no interaction. Non-parallel or crossing lines mean A's optimum depends on B — that's interaction, and OFAT can't see it at all.
3 现实里的全因子设计 Full-Factorial Design in the Real World
实验设计:注塑、焊接、化工反应的多因子优化,2^k 用最少试验解出全部主效应与交互。
Experimental design: For multi-factor optimization in injection molding, welding, and chemical reactions, 2^k extracts every main effect and interaction with the fewest possible runs.
主效应:每个因子单独平均能把响应抬高多少 —— 排出「谁最重要」的因子重要度。
Main effects: How much each factor lifts the response on its own, averaged across the others — it ranks the factors by importance.
交互作用:催化剂只在高温下起效、保压只在高压下有用 —— 全因子是唯一能稳稳抓到它的工具。
Interactions: A catalyst that only works at high temperature, a hold-pressure that only matters at high pressure — full-factorial is the only tool that catches them reliably.
Fisher 遗产:从田间试验到半导体良率,因子设计是现代 DOE 与六西格玛 Analyze/Improve 阶段的核心。
Fisher's legacy: From field trials to semiconductor yield, factorial design is the backbone of modern DOE and the Analyze/Improve phases of Six Sigma.
一句话In One Line
全因子的算法极其优雅:把每个响应贴上 A、B 的+/−符号,主效应就是「+组平均 − −组平均」,
交互效应就是把 A、B 符号相乘再做同样的差。同一批 4 个数据,被「正交」地榨出了三条独立信息 ——
这就是为什么 2^k 这么省。而真正的杀手锏是交互项:它回答的不是「A 有没有用」,
而是「A 的用处会不会因 B 而变」。OFAT 把因子拆开一个个问,结构上就丢掉了这个问题。
现实世界的最优,往往就藏在「A 和 B 一起调到某个组合」的斜坡上 —— 看得见交互,才看得见真正的山顶。
The math behind a 2^k is beautifully simple: tag every response with A's and B's +/− signs,
a main effect is just "average of + group minus average of − group," and an interaction effect multiplies the A and B signs first, then does the same difference.
The same four numbers, squeezed orthogonally for three independent pieces of information — that's why 2^k is so efficient.
But the real killer feature is the interaction term: it doesn't ask "is A useful?" but "does A's usefulness change because of B?"
OFAT splits the factors apart and asks them one by one, structurally throwing that question away.
The true optimum often hides on a ridge where A and B are dialed in together — without seeing the interaction, you'll never see the real summit.
常见误用Common Mistakes
只报主效应,无视交互。交互显著时,单看主效应会得出错误的最优组合 —— 必须连交互一起读。
Reporting only main effects and ignoring interactions. When an interaction is significant, main effects alone point to the wrong optimum — always read main effects alongside interactions.
因子一多就硬上全因子。7 因子全因子要 128 次 —— 试验数随 k 指数爆炸,该用部分析因先筛选。
Forcing full-factorial as the factor count grows. Seven factors fully crossed already costs 128 runs — run counts explode exponentially in k. Use fractional factorials to screen first.
不重复就当结论可靠。无重复无法估误差、判显著,关键结论需中心点或重复来验证。
Treating un-replicated runs as conclusive. Without replicates you can't estimate error or judge significance — back critical conclusions with center points or replication.