S4-17 · 基础 · 概念实验 S4-17 · Foundation · Concept Lab

幻觉 · 像真的,不等于有根据 Hallucination · Plausible Is Not the Same as Supported

evidence removed → fluent sentence remains → anchor breaks → warning

幻觉(Hallucination)是模型生成了流畅但缺乏事实依据或与事实冲突的内容。要识别它,必须把“像真”与“有证据”拆成两条独立判断。 A hallucination is fluent model output that lacks factual support or conflicts with facts. Detecting it requires separating plausibility from evidence.

01 · 先从日常问题开始 01 · Start with an everyday question

同事讲得特别顺,却拿不出关键文件——你会因为语气自信就签字吗? A colleague sounds completely confident but cannot produce the key document. Would you approve on tone alone?

02 移走关键证据,盯住不断线的流畅度 Remove Key Evidence and Watch Fluency Stay Intact

拖动保留证据数 drag evidence count
语言流畅度 Fluency
style score
锚定主张 Anchored claims
of 3 claims
无锚点警报 No-anchor alerts
claim-level

03 四个零件解释“为什么仍像真的” Four Parts Explain Why It Still Sounds Plausible

把刚才的变化拆开 Map the change to parts
01 · EVIDENCE

证据片段 Evidence snippets

三片来源各自支持一个可核查主张;移走后,来源本身不会被模型自动补回。 Three source snippets each support one checkable claim. Removing one does not make the source magically return.

02 · CLAIM

可核查主张 Checkable claims

回答里的日期、数量和归因需要逐项检查,而不是给整段文字一个绿色勾。 Dates, quantities, and attributions need claim-level checks, not one green tick for an entire paragraph.

03 · FLUENCY

语言流畅度 Fluency

句子是否顺滑与是否有依据是两条轴;流畅度可以在证据断开后保持很高。 Smooth language and factual support are separate axes. Fluency may stay high after evidence disappears.

04 · ALARM

无锚点警报 No-anchor alarm

只要主张没有来源连接,就明确降级,不用语气自信替代证据。 Any claim without a source link is downgraded explicitly; confident tone never substitutes for evidence.

04 最小机制:流畅度不参与证据计数 Minimal Mechanism: Fluency Is Not Evidence

anchor rate = linked / 3
anchor rate = linked claims / 3
本页固定三个主张。每移走一片关键证据,就断开一条主张—来源连接;流畅度分数故意保持不变。 This lab fixes three claims. Each removed source breaks one claim-to-source link while the fluency score deliberately stays unchanged.
生成句子 Generate sentence 语言模式保持顺滑 language pattern stays smooth
核对主张 Check each claim 逐项寻找来源 look for a source per claim
断线就报警 Warn on broken link 不把语气当证据 never treat tone as evidence

05 边界实验:证据清空后仍给整段绿灯 Boundary Test: Keep the Paragraph Green with No Evidence

主动制造失败 Create a failure

制造一个“流畅但零锚点”的回答 Create a fluent answer with zero anchors

一键移走三片证据。若系统仍只看语言质量,这段回答会显得完全正常;逐主张检查则会亮出三个警报。 Remove all three sources. A style-only check still sees a normal answer, while claim-level checking raises three warnings.

尚未执行。先在上方改变主变量。 Not run yet. Change the main variable above first.
写得顺,所以大概率是真的。 It reads smoothly, so it is probably true. → ✓ 流畅度与证据支持分开测。 Measure fluency and support separately.
整段有一个引用就算有依据。 One citation supports the whole paragraph. → ✓ 每个可核查主张都要有对应来源。 Every checkable claim needs its own source.
删掉证据后继续显示绿色。 Keep the claim green after its source is removed. → ✓ 连接断开就降为“未证实”。 Downgrade to unverified when the link breaks.
06 · 一句话带走 06 · One line to keep

幻觉之所以危险,是因为关键证据被移走后回答仍可能很流畅;可靠系统必须逐主张检查锚点,断线就报警。 Hallucinations are dangerous because an answer may stay fluent after key evidence is removed. Reliable systems check anchors claim by claim and warn on broken links.