故事 · 心理学家 Cohen 揪出了"运气"
Origin · Psychologist Cohen Calls Out the Luck
1960 年,统计学家 Jacob Cohen 注意到一个被长期忽视的陷阱:当两位评估员都做"是 / 否"判断时,哪怕完全靠瞎猜,也会有相当比例凑巧判得一样。
如果只报告"一致率 90%",听起来很厉害,可如果待判的件里九成本来就该合格,那两人闭眼都点"合格"也能轻松到 90% —— 这一致几乎全是运气。
于是他提出 Cohen's Kappa:用观测一致率减去偶然一致率,再除以"理论上还能改进的空间"(1 − Pe),得到一个扣掉运气后的净一致系数。
后来 Fleiss 把它推广到任意多位评估员。在属性 MSA(质检员一致性研究)里,Kappa 成了判断"判官们到底齐不齐心"的标准答案 —— κ > 0.75 才算可靠。
In 1960, psychologist Jacob Cohen spotted a long-overlooked trap: when two raters make a yes/no call, even pure guessing produces a sizeable share of agreement.
Reporting "90% agreement" sounds impressive, but if 90% of the parts truly are good to begin with, two raters could close their eyes, stamp "pass" on everything, and still hit 90% — almost all of it luck.
Cohen's fix: take observed agreement, subtract the chance-expected rate, then divide by the room that was actually left to improve (1 − Pe). The result is a net agreement coefficient with luck stripped out.
Fleiss later extended the idea to any number of raters. In attribute MSA — the formal "are our inspectors aligned?" study — Kappa is the standard answer, and κ > 0.75 is the bar a measurement system has to clear.
1 2×2 一致性矩阵:对角线是一致,反对角是分歧 The 2×2 Agreement Matrix: Diagonal = Agree, Off-Diagonal = Disagree
κ 0.002 原始一致率 vs Kappa:运气被扣掉多少? Raw Agreement vs Kappa: How Much Was Luck?
Po(看着很高)里有 Pe 那一截是运气。Kappa 把分母也换成"扣掉运气后还剩的空间",于是 κ 往往比 Po 低不少 —— 那道落差,就是被运气虚高的水分。虚线 0.75 是合格门槛。 Po (looks impressive) hides a Pe slice that's pure luck. Kappa also swaps the denominator for "room left after luck", so κ usually lands well below Po — that gap is the inflation luck added. The dashed line at 0.75 marks the acceptance bar.
3 现实里的 Kappa Kappa in the Real World
属性 MSA:质检员判合格/不合格、外观分级、缺陷分类 —— 计数数据用 Kappa 而非 Gauge R&R。
Attribute MSA: pass/fail, cosmetic grade, defect class — count data goes to Kappa, not Gauge R&R.
κ > 0.75:常用判读 —— >0.75 好、0.4~0.75 一般、<0.4 差;质检场景越严越好。
Landis & Koch bands: κ > 0.75 substantial, 0.4–0.75 moderate, < 0.4 poor. Critical-quality inspections demand the high end.
Cohen vs Fleiss:两位评估员用 Cohen κ,三位及以上用 Fleiss κ;还能算与基准(标准答案)的一致。
Cohen vs Fleiss: Cohen's κ for two raters, Fleiss's κ for three or more — and either can also be scored against a known reference standard.
Kappa 悖论:当某一类极度占多数时,κ 可能偏低,需结合 Po 与边缘分布一起看。
The Kappa paradox: when one class dominates, κ can read low even with high Po. Always read it together with the marginal distribution.
一句话In One Line
Kappa 的全部精神就一句话:把瞎猜也能拿到的分数先扣掉。
分子 Po − Pe 是"真本事净赚的一致",分母 1 − Pe 是"理论上还能净赚的上限"。
两人完美一致时 κ = 1;只达到瞎猜水平时 κ = 0;比瞎猜还差(系统性唱反调)时 κ 甚至为负。
这正是为什么不能只报"一致率 90%" —— 当一类样本占绝大多数,Pe 本就很高,90% 可能等于 κ 只有 0.2。
属性数据没有方差可拆,Kappa 就是它的 %GRR:它衡量的不是产品,而是判官们的判断系统可不可信。
The whole spirit of Kappa is one sentence: first subtract the score blind guessing would already earn.
The numerator Po − Pe is the agreement skill actually contributed; the denominator 1 − Pe is the room left to contribute in.
Perfect agreement gives κ = 1. Pure guessing gives κ = 0. Worse than guessing — systematic disagreement — drives κ negative.
That's why "90% agreement" alone is dangerous: when one class dominates, Pe is already huge and 90% can correspond to κ as low as 0.2.
Attribute data has no variance to decompose — Kappa is its %GRR equivalent. It doesn't grade the product; it grades whether the raters' judgment system can be trusted.
常见误用Common Mistakes
只报原始一致率就下结论。一致率会被偶然虚高,必须用扣除偶然的 Kappa 才公平。
Concluding from raw Po alone. Observed agreement is inflated by chance — only the chance-corrected Kappa is a fair score.
给属性数据算方差 / Gauge R&R。计数(合格/不合格)是分类数据,没有方差,要用 Kappa 做一致性。
Running Gauge R&R on attribute data. Pass/fail counts are categorical — variance is undefined. Use Kappa-based attribute agreement instead.
无视 Kappa 悖论。类别极不平衡时 κ 偏低,要结合 Po 与边缘分布解读,别只盯一个数。
Ignoring the Kappa paradox. With severely unbalanced classes, κ can read low despite high Po. Read κ alongside Po and the marginal totals — never as a single number.