起源故事 · Cohen 与"碰巧一致" Origin Story · Cohen and "Chance Agreement"
1960 年,心理统计学家 Jacob Cohen 在研究两位临床医生诊断是否一致时发现一个漏洞: 直接数"两人判得一样"的比例(观察一致率),会被运气严重高估 —— 如果大家都偏向某一类,光靠瞎猜就能凑出很高的一致。 他提出把"碰巧能达到的一致"先减掉,再看实际超出了多少。这个扣除碰巧的系数就叫 Kappa, 今天它是属性 MSA、医学诊断、评分者信度的标准量尺。 In 1960, the psychometrician Jacob Cohen was studying whether two clinicians agreed on a diagnosis and spotted a loophole: simply counting "the share of cases they call the same way" (the observed agreement) is heavily inflated by luck — if both lean toward one category, blind guessing alone can rack up a high rate. He proposed subtracting "the agreement chance alone would deliver" first, then asking how much further the raters actually went. That chance-corrected coefficient became kappa, now the standard yardstick for attribute MSA, medical diagnosis, and inter-rater reliability.

1 2×2 配对方格:两人到底有多一致? 2×2 Confusion Grid: How Closely Do the Two Inspectors Really Agree?

高一致

2 把一致率拆开:哪些是真本事,哪些是运气 Split the Agreement: How Much Is Skill, How Much Is Luck

Kappa 的几何意义:观察一致率 Pₒ 高出碰巧一致率 Pₑ 的部分,占"满分到 Pₑ 这段可改进空间"的比例。运气越大(Pₑ 越高),同样的 Pₒ 换来的 Kappa 越低。 The geometric meaning of kappa: the slice by which observed agreement Pₒ exceeds chance Pₑ, expressed as a fraction of "the room left between Pₑ and a perfect 100%". The more luck contributes (higher Pₑ), the lower the kappa you earn from the same Pₒ.

3 现实里的 Kappa Kappa in the Real World

外观检验:两名质检员判"划痕合不合格",属性 MSA 用 Kappa 评判读一致性,<0.7 就得重新培训标准。 Visual inspection: two QC inspectors judge whether a scratch is acceptable. Attribute MSA uses kappa to grade their agreement — below 0.7 and the acceptance standard must be retrained.
医学诊断:两位放射科医生看同一张片子判"有无病灶",Kappa 是诊断一致性的金标准。 Medical diagnosis: two radiologists read the same scan for "lesion present or not". Kappa is the gold standard for diagnostic agreement.
评分者信度:论文评审、作文打分、面试评级,多名评委的评分一致性用 Kappa 衡量。 Inter-rater reliability: paper reviewing, essay scoring, interview rating — when multiple judges grade the same item, kappa quantifies their consistency.
数据标注:机器学习训练集的人工标注质量,Kappa 量化标注员之间的可靠程度。 Data labeling: for machine-learning training sets, kappa quantifies how reliable the human annotators are relative to one another.
一句话 In One Line
一致率回答"两人判得一样的比例",Kappa 回答"扣掉运气后还剩多少真本事"。 量具是连续数据用 Gauge R&R(上一页),但很多检验是"合格/不合格"这种属性判断,没有刻度可读,只能比一致性 —— 这时 Kappa 就是属性 MSA 的核心指标。 记住:批次越偏(合格率越极端),同样的一致率对应的 Kappa 越低,因为运气贡献被算得明明白白。 Raw agreement answers "how often the two inspectors call it the same way"; kappa answers "how much real skill remains once luck is stripped out". Continuous gauges use Gauge R&R (previous page), but many inspections are pass/fail attribute judgments with no scale to read — there you can only compare agreement, and kappa is the core attribute-MSA metric. Remember: the more skewed the batch (the more extreme the pass rate), the lower the kappa you earn for the same observed agreement, because luck's contribution is being accounted for explicitly.
常见误用 Common Mistakes
只看一致率就下结论必须算 Kappa:高一致率可能全是碰巧凑的。 Calling it a day after looking at raw agreement. Always compute kappa — a high agreement rate can be entirely the work of chance.
样本严重偏向一类还硬解读 Kappa极端偏度下 Kappa 会被压低,需配合样本结构一起看。 Reading kappa in isolation when the sample is heavily skewed. Extreme prevalence drags kappa down; interpret it alongside the marginal distribution, not on its own.
把 Kappa 当准确率Kappa 衡量"一致",不代表"对";判对要另比标准答案。 Treating kappa as accuracy. Kappa measures "agreement", not "correctness" — to judge correctness you must compare against a known reference answer.

Kappa 属性一致性