v0.2.3 · Verification Prompt Ablation
初始实验 60/60 Tool / Task 固定,仅替换 Verification Prompt这一版做什么
image-only · native claude -p · no harness hard gate这版不复制 VDiff 的 PDF binding、固定 Step 1–5 或其它“大操作”。它只把可迁移的部分抽象为三段 Prompt:通用外部 MLLM 工具、简单 image-to-code 任务、不同密度的验证策略,然后观察模型是否主动形成过程验证与完成前的全面验证闭环。
Viewer 调整
12 datasets · 60 trajectories · rubric-only score contractv0.2.3 Unified Visual Trajectory Viewer真实 target / render / HTML、运行事件与 VDiff Gold+discovery 离线 rubric 分数
打开 v0.2.3 Viewer →Viewer 边界本实验没有 harness hard gate,也没有 epoch。Viewer 使用 rubric-only 评分合同,不把离线 rubric 结果伪装成在线 geometry / visual gate。
结果、时间与验证行为
5 cases × 3 replicates per condition质量看 VDiff Gold+discovery rubric 分数;成本看平均时间、主模型成本和 MLLM 调用。验证行为再拆成过程中的 Focused Validate 与候选页面级的 Gate-like Checkpoint。
| 条件 | Macro | Pass@4 | 相对 C0 | 平均时间 | 平均成本 | MLLM 调用 | 过程验证 | Gate-like | 平均修复循环 |
|---|---|---|---|---|---|---|---|---|---|
| C0Tool + Task 无 Verification |
3.675micro 3.644 | 5/15 | baseline | 2m 10s | $1.50 | 4other 4 | 0 | 0 | 0.00 |
| C1+ V0 Mentor Minimal Mentor 极简版 |
3.927micro 3.919 | 6/15 | +0.251CI [0.052, 0.425] | 4m 04s | $2.87 | 5other 5 | 0 | 0 | 0.00 |
| C2+ V1 Initial Full 初始完整版 |
3.830micro 3.770 | 7/15 | +0.155CI [-0.133, 0.387] | 12m 29s | $11.00 | 81other 10 | 19 | 5213P / 38F / 1? | 3.07 |
| C3+ V2 Overall Compact 整体压缩版 |
3.875micro 3.829 | 8/15 | +0.200CI [0.017, 0.345] | 5m 50s | $5.24 | 36other 7 | 10 | 1919? | 0.93 |
行为口径过程验证:局部裁图、局部主张或中间产物检查。Gate-like:同时使用完整 reference 与完整 candidate、面向整页/多维度结论的调用。它是轨迹中的自发行为,不是 harness 强制 Gate;
? 表示调用范围像 Gate,但回答没有明确 PASS / FAIL / INSUFFICIENT。当前 V3 后续:验证行为
12/12 generation complete · offline quality evaluation pending这一表只描述已经发生的外部验证行为,不并入上面的正式质量结论。C3 共 3 次过程验证、13 次 Gate 且全部 FAIL;C4 共 18 次过程验证、81 次 Gate(61F / 20P),严格双 Checkpoint closure 仅 3/6 成立。
展开 12 条轨迹Focused / Gate / Other / closure
| 轨迹 | 过程验证 | Gate | Other | 最终行为 |
|---|---|---|---|---|
| C3 · case25-r1 | 0 | 22F / 0P | 0 | 最终 Gate FAIL 后完成 |
| C3 · case25-r2 | 0 | 22F / 0P | 0 | 最终 Gate FAIL 后完成 |
| C3 · case25-r3 | 1 | 22F / 0P | 0 | 最终 Gate FAIL 后完成 |
| C3 · case5-r1 | 0 | 22F / 0P | 0 | 最终 Gate FAIL 后完成 |
| C3 · case5-r2 | 0 | 22F / 0P | 0 | 最终 Gate FAIL 后完成 |
| C3 · case5-r3 | 2 | 33F / 0P | 0 | 最终 Gate FAIL 后完成 |
| C4 · case25-r1 | 11 | 1815F / 3P | 2 | Closure 违反 |
| C4 · case25-r2 | 3 | 117F / 4P | 2 | Closure 成立 |
| C4 · case25-r3 | 1 | 95F / 4P | 0 | 成立;最终 PASS 有保留 |
| C4 · case5-r1 | 0 | 139F / 4P | 1 | Closure 违反 |
| C4 · case5-r2 | 0 | 1712F / 5P | 0 | Closure 成立 |
| C4 · case5-r3 | 3 | 1313F / 0P | 1 | Closure 违反 |
当前读法V3 显著增加了 Gate 调用,但“调用更多”不等于严格闭环成立:3/6 C4 轨迹仍违反同一最新 artifact 上连续两次全面验证的 closure。离线 rubric 分数完成前,不判断 V3 的最终质量收益。
逐个 case 对比
case mean + 60 条单轨迹跳转先看每个 case 的三次平均分,再展开到每条轨迹。四条件对比链接始终固定同一个 case 和同一个 replicate。
| case | C0 | C1 | C2 | C3 | 最佳 | Viewer |
|---|---|---|---|---|---|---|
| case5 | 3.753 | 4.017 | 3.420 | 3.617 | C1 | r1 四条件对比r2 四条件对比r3 四条件对比 |
| case9 | 4.180 | 4.060 | 4.180 | 4.300 | C3 | r1 四条件对比r2 四条件对比r3 四条件对比 |
| case22 | 3.450 | 3.807 | 3.750 | 3.767 | C1 | r1 四条件对比r2 四条件对比r3 四条件对比 |
| case23 | 3.217 | 3.750 | 3.493 | 3.610 | C1 | r1 四条件对比r2 四条件对比r3 四条件对比 |
| case25 | 3.777 | 4.000 | 4.307 | 4.083 | C2 | r1 四条件对比r2 四条件对比r3 四条件对比 |
case512 trajectories · best mean C1 4.017
| 条件 | 分数 | 时间 | MLLM | 过程验证 | Gate-like · final | 修复循环 | 查看 |
|---|---|---|---|---|---|---|---|
| C0 · r1 | 3.68 | 2m 57s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r2 | 4.11 | 1m 38s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r3 | 3.47 | 2m 02s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r1 | 4.05 | 2m 49s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r2 | 4.05 | 5m 17s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r3 | 3.95 | 4m 18s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C2 · r1 | 3.21 | 32m 17s | 9 | 1 | 8 · pass | 6 | 轨迹 · 回放 |
| C2 · r2 | 3.68 | 15m 40s | 7 | 1 | 6 · fail | 5 | 轨迹 · 回放 |
| C2 · r3 | 3.37 | 12m 56s | 6 | 1 | 5 · fail | 4 | 轨迹 · 回放 |
| C3 · r1 | 4.00 | 4m 20s | 1 | 0 | 1 · unknown | 0 | 轨迹 · 回放 |
| C3 · r2 | 3.74 | 5m 54s | 1 | 0 | 1 · unknown | 0 | 轨迹 · 回放 |
| C3 · r3 | 3.11 | 23m 29s | 11 | 6 | 2 · unknown | 7 | 轨迹 · 回放 |
case912 trajectories · best mean C3 4.300
| 条件 | 分数 | 时间 | MLLM | 过程验证 | Gate-like · final | 修复循环 | 查看 |
|---|---|---|---|---|---|---|---|
| C0 · r1 | 4.18 | 0m 45s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r2 | 4.27 | 0m 41s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r3 | 4.09 | 0m 38s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r1 | 4.45 | 2m 24s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r2 | 3.82 | 2m 17s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r3 | 3.91 | 1m 36s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C2 · r1 | 4.09 | 5m 46s | 3 | 0 | 2 · pass | 1 | 轨迹 · 回放 |
| C2 · r2 | 4.18 | 7m 29s | 3 | 3 | 0 | 2 | 轨迹 · 回放 |
| C2 · r3 | 4.27 | 12m 51s | 12 | 5 | 5 · fail | 7 | 轨迹 · 回放 |
| C3 · r1 | 4.36 | 4m 40s | 2 | 2 | 0 | 1 | 轨迹 · 回放 |
| C3 · r2 | 4.09 | 3m 10s | 2 | 0 | 1 · unknown | 0 | 轨迹 · 回放 |
| C3 · r3 | 4.45 | 3m 42s | 2 | 0 | 2 · unknown | 1 | 轨迹 · 回放 |
case2212 trajectories · best mean C1 3.807
| 条件 | 分数 | 时间 | MLLM | 过程验证 | Gate-like · final | 修复循环 | 查看 |
|---|---|---|---|---|---|---|---|
| C0 · r1 | 3.30 | 2m 03s | 1 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r2 | 3.47 | 2m 20s | 1 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r3 | 3.58 | 4m 04s | 1 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r1 | 4.03 | 4m 45s | 1 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r2 | 3.58 | 2m 12s | 2 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r3 | 3.81 | 3m 08s | 1 | 0 | 0 | 0 | 轨迹 · 回放 |
| C2 · r1 | 4.26 | 33m 40s | 7 | 2 | 3 · pass | 3 | 轨迹 · 回放 |
| C2 · r2 | 3.52 | 6m 53s | 4 | 0 | 4 · pass | 3 | 轨迹 · 回放 |
| C2 · r3 | 3.47 | 7m 21s | 5 | 0 | 4 · pass | 2 | 轨迹 · 回放 |
| C3 · r1 | 4.03 | 5m 20s | 3 | 0 | 2 · unknown | 1 | 轨迹 · 回放 |
| C3 · r2 | 4.09 | 2m 37s | 1 | 0 | 1 · unknown | 0 | 轨迹 · 回放 |
| C3 · r3 | 3.18 | 3m 57s | 2 | 1 | 0 | 0 | 轨迹 · 回放 |
case2312 trajectories · best mean C1 3.750
| 条件 | 分数 | 时间 | MLLM | 过程验证 | Gate-like · final | 修复循环 | 查看 |
|---|---|---|---|---|---|---|---|
| C0 · r1 | 2.57 | 1m 00s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r2 | 3.75 | 1m 06s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r3 | 3.33 | 1m 29s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r1 | 4.79 | 4m 28s | 1 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r2 | 3.61 | 13m 59s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r3 | 2.85 | 2m 36s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C2 · r1 | 3.40 | 10m 31s | 4 | 0 | 4 · fail | 3 | 轨迹 · 回放 |
| C2 · r2 | 3.47 | 8m 02s | 7 | 0 | 5 · fail | 3 | 轨迹 · 回放 |
| C2 · r3 | 3.61 | 16m 17s | 6 | 6 | 0 | 4 | 轨迹 · 回放 |
| C3 · r1 | 3.33 | 4m 14s | 1 | 1 | 0 | 0 | 轨迹 · 回放 |
| C3 · r2 | 3.68 | 6m 30s | 2 | 0 | 2 · unknown | 1 | 轨迹 · 回放 |
| C3 · r3 | 3.82 | 5m 45s | 3 | 0 | 3 · unknown | 2 | 轨迹 · 回放 |
case2512 trajectories · best mean C2 4.307
| 条件 | 分数 | 时间 | MLLM | 过程验证 | Gate-like · final | 修复循环 | 查看 |
|---|---|---|---|---|---|---|---|
| C0 · r1 | 3.33 | 2m 09s | 1 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r2 | 4.08 | 1m 28s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C0 · r3 | 3.92 | 8m 14s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r1 | 3.83 | 2m 13s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r2 | 3.50 | 3m 34s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C1 · r3 | 4.67 | 5m 19s | 0 | 0 | 0 | 0 | 轨迹 · 回放 |
| C2 · r1 | 4.17 | 3m 01s | 2 | 0 | 1 · pass | 0 | 轨迹 · 回放 |
| C2 · r2 | 4.50 | 4m 08s | 3 | 0 | 2 · pass | 1 | 轨迹 · 回放 |
| C2 · r3 | 4.25 | 10m 24s | 3 | 0 | 3 · pass | 2 | 轨迹 · 回放 |
| C3 · r1 | 4.67 | 7m 46s | 2 | 0 | 2 · unknown | 1 | 轨迹 · 回放 |
| C3 · r2 | 3.33 | 2m 57s | 2 | 0 | 1 · unknown | 0 | 轨迹 · 回放 |
| C3 · r3 | 4.25 | 3m 12s | 1 | 0 | 1 · unknown | 0 | 轨迹 · 回放 |
三部分 Prompt
分别查看 · 分别切换中英文页面按 Tool → Task → Verification 展示,不再按完整实验条件重复三套内容。每一部分独立切换中英文;Verification 额外切换清单中的多个版本。
Tool Prompt
如何把外部 MLLM 暴露给主模型# Tool V1 · Initial MCP-Aligned / 初始版 · 对齐当前 MCP:中文审阅
状态:待审候选
对应 Module ID:`tool-V1-initial-mcp-aligned`
用途:人工审阅;不会注入运行时
```text
<multimodal_model_tool>
你可以使用 `mcp__mllm__call`,下文称为 `CallMLLM`。每次调用会把一条自包含的
`instruction` 和带有明确 label 的 `inputs` 发送给一个全新、独立的多模态模型
上下文。Inputs 可以包含内联文字、工作区内的 UTF-8 文本文件或图片。
被委托模型可能具有比主模型更强或互补的视觉能力,并提供独立的上下文窗口。在需要视觉
证据的工作中,应主动并根据需要多次使用它。当多个 model profile 可用时,可以对不同
profile 进行多次聚焦调用,以获得互补观察或独立意见。
CallMLLM 可以用于视觉读取、定位、布局理解、图片比较、图片与代码或历史产物之间的
对应,以及可见差异诊断。这些只是能力示例,不构成固定 workflow。
每次调用彼此独立,只能看到本次提供的 instruction 和 inputs。提供必要上下文,并清楚
标注多个输入。只在 instruction 中写出路径并不等于附加文件。已有代码或产物文件使用
`kind: "text_file"` 附加,图片使用 `kind: "image"` 附加。
由你决定询问什么、提供哪些输入、选择哪个可用 model profile,以及需要什么返回格式。
`capability_tags` 和 `output` 可以省略;默认值分别为 `["Z1"]` 和文本输出。
返回内容是当前工作的证据;如何使用仍由你负责。
可用 profiles:
{{MODEL_PROFILE_CATALOG}}
</multimodal_model_tool>
```
# Tool V1 · Initial MCP-Aligned / 初始版 · 对齐当前 MCP
Status: frozen 2026-07-19 as verification-line control base (SHA in manifest.json/VERSIONING.md); do not edit — introduce a new version instead
Module ID: `tool-V1-initial-mcp-aligned`
Language: English, proposed runtime text
Runtime placeholder: `{{MODEL_PROFILE_CATALOG}}`
```text
<multimodal_model_tool>
You have access to `mcp__mllm__call`, referred to as `CallMLLM`. Each call
sends a self-contained `instruction` and explicitly labeled `inputs` to a
fresh, independent multimodal model context. Inputs may contain inline text,
workspace-relative UTF-8 text files, or images.
The delegated model may have stronger or complementary visual capabilities than
you and provides a separate context window. Use it proactively and as often as
useful for visually grounded work. Multiple focused calls, including calls to
different model profiles when more than one is available, are encouraged when
they provide complementary observations or an independent second opinion.
CallMLLM can help with visual reading, localization, layout interpretation,
image comparison, correspondence between images and code or prior artifacts,
and diagnosis of visible differences. These are examples, not a fixed workflow.
Each call is independent and sees only the instruction and inputs supplied in
that call. Include the necessary context and label multiple inputs clearly. A
path mentioned only inside the instruction is not an attachment. Attach existing
code or artifact files with `kind: "text_file"` and images with
`kind: "image"`.
You decide what question to ask, which inputs to provide, which available model
profile to use, and what response format is useful. `capability_tags` and
`output` may be omitted; they default to `["Z1"]` and text output. The
returned response is evidence for your work; you remain responsible for
deciding how to use it.
Available profiles:
{{MODEL_PROFILE_CATALOG}}
</multimodal_model_tool>
```
Task Prompt
要完成的 image-to-code 任务# Task V1 · Initial / 初始版:中文审阅
状态:待审候选
对应 Module ID:`task-V1-initial`
用途:帮助人工审阅英文运行时原文;本文件不会注入 Claude
运行时占位符:`{{REFERENCE_IMAGE_PATH}}`
本模块只定义图像转代码任务。产物入口发现和依赖策略属于 Runtime Contract;验证、外部工具、
Checkpoint、Gate 和评分均不属于 B2。英文文件是唯一规范运行时来源。
```text
<image_to_code_task>
参考图片:`{{REFERENCE_IMAGE_PATH}}`
创建一个忠实、live 且可编辑的 HTML 页面,复现参考图片所展示的可见状态。主要目标是在参考图片原始像素尺寸下渲染时保持视觉忠实。
- 可读文字和界面内容必须保持为真实、可见的 DOM 元素。不得使用完整参考图片或包含可读文字的栅格图片替代真正的页面实现。
- 每个可见元素都应尽可能匹配参考图片中对应区域的位置和 bbox。
- 对于可编辑文本流,应使用普通文档流、Flexbox 或 Grid。不得使用绝对定位、固定布局表格、定位文本框或手工 `<br>` 模拟自然文字换行。
</image_to_code_task>
```
# Task V1 · Initial / 初始版
Status: frozen 2026-07-19 as verification-line control base (SHA in manifest.json/VERSIONING.md); do not edit — introduce a new version instead
Module ID: `task-V1-initial`
Language: English, proposed runtime text
Runtime placeholder: `{{REFERENCE_IMAGE_PATH}}`
This module defines only the image-to-code task. Artifact discovery and dependency policy belong to
the runtime contract; verification, external-tool use, checkpoints, gates, and scoring remain outside B2.
```text
<image_to_code_task>
Reference image: `{{REFERENCE_IMAGE_PATH}}`
Create a faithful, live, and editable HTML reproduction of the visible state shown in the reference image. The primary target is fidelity when rendered at the reference image's native pixel dimensions.
- Keep readable text and interface content as real, visible DOM elements. Do not use the complete reference image, or rasterized readable content, as a substitute for the implementation.
- Match each visible element as closely as possible to the position and bounding box of its corresponding region in the reference image.
- For editable text flow, use normal document flow, flexbox, or grid. Do not use absolute positioning, fixed-layout tables, positioned text boxes, or manual `<br>` elements to imitate natural text wrapping.
</image_to_code_task>
```
Verification Prompt
验证与 Checkpoint 行为# Verification V0 · Mentor Minimal / Mentor 极简版:中文审阅
状态:保留的基线比较版本;尚未冻结,也没有用于实验
对应 Module ID:`verification-V0-mentor-minimal`
来源:Mentor HTML 验证条款第 3–5 条;有意不包含任务条款第 1–2 条
```text
迭代调整 HTML 页面,直到空间差异处于容差范围内(页面坐标中小于 5px),并且没有文字裁切或重叠。
空间差异解决后(漂移小于 5px),执行视觉验证。在宣布构建完成之前修复视觉失败。如果视觉修复造成漂移回归,重新校准间距,并重新验证空间与视觉两方面。
对持续存在的视觉差异,判断哪些差异对人眼而言显著。视觉上显著的差异必须修复——不要在没有改变方法的情况下宣布它们无法修复。
```
# Verification V0 · Mentor Minimal / Mentor 极简版
Status: used as C1 in initial v0.2.3 ablation; editable and not SHA-frozen
Module ID: `verification-V0-mentor-minimal`
Language: English, proposed runtime text
Source: Mentor HTML verification clauses 3–5; task clauses 1–2 intentionally excluded
```text
Iteratively refine the HTML page until spatial discrepancy is within tolerance
(< 5px in page coordinates), and no text clipping/overlapping.
After spatial discrepancy is resolved (<5px drift), run visual verification.
Fix visual failures before declaring the build complete. If visual fixes cause
drift regression, recalibrate gaps and re-verify both.
For persistent visual discrepancy, distinguish what is significantly noticeable
to human eyes. Visually significant differences must be fixed — do not declare
them unfixable without changing the approach.
```
# Verification V1 · Initial Full / 初始完整版:中文审阅
状态:待审候选
对应 Module ID:`verification-V1-initial-full`
用途:人工审阅;不会注入运行时
```text
<verification_policy>
# Meta Verification and Checkpoint Policy
验证是任务执行过程的一部分,而不只是结束前的一次检查。
你必须使用 `CallMLLM` 对重要的中间结论和最新候选产物执行外部多模态验证;其他外部
工具可以提供补充证据。不得仅根据自己的分析、源码阅读、实现意图、记忆或未经外部验证的
视觉判断,宣布某项要求已经满足。
验证分为两种:
1. `Focused Validate`:针对中间过程中的具体产物或具体主张进行局部验证。
2. `Checkpoint Verification`:在形成有意义的候选版本时,对任务目标进行全面验证。
## 1. Verification Inputs
根据当前验证目标,可以使用以下一种或多种输入:
- 原始参考材料,例如参考图片;
- 最新 HTML 页面渲染截图;
- 当前 HTML、CSS、JavaScript 或其他代码;
- 当前 DOM 结构和计算样式;
- 浏览器测量结果;
- 文字、图片、区域或元素的绑定关系;
- OCR、转录、裁图和素材提取结果;
- 之前阶段产生的 JSON、截图、日志或测量文件;
- 上一个 checkpoint 的代码、截图和验证结果;
- 当前 checkpoint 与上一个 checkpoint 之间的代码或视觉差异。
验证输入必须对应当前最新产物。不得使用旧截图验证已经修改过的代码。
当已有绑定、元素对应或测量产物时,应优先复用这些证据,而不是要求外部工具从零重新
猜测。
输入外部工具时,只提供支持当前验证主张所必需的相关材料,并说明每个输入的身份、来源和
版本。
## 2. Focused Validate:中间过程验证
### WHEN
在以下情况下执行 Focused Validate:
- 一个中间产物将被后续步骤依赖;
- 从参考图片中提取了文字、图片、颜色、bbox 或结构;
- 建立或修改了参考图片区域与 HTML 元素之间的绑定;
- 一个非验证工具返回了影响后续实现的重要结果,或者现有工具结果尚未包含足以支持下一步
决策的验证证据;
- 当前观察存在歧义或可能影响大量后续工作;
- 修改可能造成文字裁切、布局漂移、元素缺失或结构错误;
- 发现多个证据之间存在冲突。
不要机械验证每个小动作。只验证错误会向后传播、影响最终结果或改变后续决策的重要中间
结论。
### WHAT
每次 Focused Validate 必须先定义一个具体、可证伪的验证主张,例如:
- “这段文字与参考图片中的标题完全一致。”
- “这个裁图对应参考图片中的主要照片。”
- “DOM 元素 `.hero-title` 对应参考图片中的主标题区域。”
- “这个修改没有造成正文裁切。”
- “当前渲染中三个卡片保持等间距。”
- “最新代码已经消除了上一次验证发现的重叠。”
禁止使用过于宽泛的主张,例如:
- “页面是正确的。”
- “工具输出没问题。”
- “布局看起来不错。”
### HOW
1. 明确写出待验证的具体主张。
2. 选择能够推翻该主张的外部证据。视觉或空间主张必须包含 CallMLLM 验证;浏览器、
DOM、测量、测试或静态分析工具可以提供补充证据。
3. 提供相关参考图片、候选截图、代码、绑定、测量结果或此前产物。
4. 要求外部验证器返回:
- `PASS`;
- `FAIL`;
- `INSUFFICIENT`。
5. 要求外部验证器给出支持结果的具体证据。
6. 如果结果为 `FAIL`,证据必须影响后续实现。
7. 如果结果为 `INSUFFICIENT`,不得将其视为通过;应补充证据、缩小问题或改变验证
方法。
8. 修复重要失败后,重新验证受影响的主张。
Focused Validate 不要求每次都执行全面视觉比较。它只检查当前最重要的中间主张。
## 3. Checkpoint:候选版本与全面验证触发点
Checkpoint 表示当前代码已经形成一个有意义、可运行、值得进行全面评估的候选版本。
Checkpoint 是由主模型维护的逻辑候选状态,不要求固定文件名或固定 JSON schema;但其
代码版本、截图、输入、已知问题和验证结果必须能够从工作产物或执行轨迹中识别。
Checkpoint 不能只是任意时间点的文件备份。每次提交 checkpoint,都必须同时触发一次
`Checkpoint Verification`。
### WHEN
在以下情况下提交 checkpoint:
- 首次产生完整、可渲染的候选页面;
- 完成一批相关布局或视觉修复;
- 解决了一组重要验证失败;
- 准备判断是否继续迭代;
- 准备宣布任务完成。
不要在每次微小编辑后提交 checkpoint。应在一个连贯的修复阶段结束后提交。
### CHECKPOINT CONTENT
每个 checkpoint 至少应关联:
- 当前代码或代码版本;
- 最新候选渲染截图;
- 当前使用的参考材料;
- 关键中间产物或绑定信息;
- 从上一个 checkpoint 以来的主要修改;
- 当前已知问题;
- 本次全面验证结果。
Checkpoint 必须标记为以下状态之一:
- `CANDIDATE`:等待全面验证;
- `ACCEPTED`:全面验证通过;
- `REJECTED`:存在必须修复的失败;
- `INSUFFICIENT`:外部证据不足,不能判断。
只有 `ACCEPTED` checkpoint 才能作为完成依据。如果 `INSUFFICIENT` 涉及完成条件、
视觉上显著的差异或重要空间主张,则 checkpoint 不能成为 `ACCEPTED`。与完成无关的
非重要不确定性应被记录,但不自动阻止完成。
## 4. Checkpoint Verification:全面验证
每次 checkpoint 都必须对最新候选执行全面验证。
全面验证必须使用 CallMLLM 进行外部多模态验证,不得由主模型自己独立判断通过。
验证输入应根据需要包括:
- 原始参考图片;
- 最新候选渲染截图;
- 当前 HTML、CSS 或 JavaScript;
- 可用的元素绑定、DOM bbox 或测量结果;
- 上一个 checkpoint 的截图和失败项;
- 当前 checkpoint 的修改摘要。
全面验证依次执行两个阶段。
### A. Spatial Verification
检查:
- 页面尺寸与边界;
- 显著元素的位置和尺寸;
- 左、上、右、下边界;
- 宽度和高度;
- 对齐、居中、间距和包含关系;
- 文字换行;
- 文字裁切、遮挡和溢出;
- 元素重叠和页面越界。
如果已有可靠的参考区域与 DOM 元素绑定,应使用绑定和坐标证据计算空间差异。
如果没有可靠绑定,只能由外部验证器从图片中估计对应关系,则必须明确标记证据来源和
不确定性。
只有外部证据提供明确的元素对应及坐标证据时,才可以声明漂移小于 5px。如果一个重要
空间主张无法建立可靠对应,标记为 `INSUFFICIENT`,不得视为通过。
空间通过条件:
- 所有可可靠测量的显著边界漂移小于 5px;
- 没有文字裁切、遮挡或意外换行;
- 没有非预期重叠、溢出或越界;
- 没有尚未处理的显著空间差异。
### B. Visual Verification
空间验证通过后,检查:
- 内容完整性;
- 页面结构;
- 视觉层级;
- 字体、字重和文字风格;
- 颜色、背景、边框和阴影;
- 图片裁切、比例和位置;
- 图标、装饰和关键视觉细节;
- 已知验收项;
- 未被既有验收项覆盖的新差异。
要求外部验证器报告:
- 参考材料中的观察;
- 候选产物中的观察;
- `PASS / FAIL / INSUFFICIENT`;
- 差异是否对人眼显著;
- 支持结果的具体证据;
- 可选的修复方向或仍需补充的证据。
视觉上显著的失败必须修复。
## 5. Feedback and Iteration
全面验证结果必须影响 checkpoint 状态和后续行为。
### 如果通过
- 将 checkpoint 标记为 `ACCEPTED`;
- 保留验证证据;
- 如果满足全部任务目标,可以宣布完成。
### 如果失败
- 将 checkpoint 标记为 `REJECTED`;
- 将失败项转成具体修复计划;
- 优先修复影响最大、可能引起连锁问题的差异;
- 修改代码;
- 重新渲染;
- 提交新的 checkpoint;
- 再次执行全面验证。
### 如果证据不足
- 将 checkpoint 标记为 `INSUFFICIENT`;
- 补充代码、截图、绑定、局部裁图或测量输入;
- 或改变外部验证器的问题与方法;
- 如果证据不足涉及完成条件,不得宣布完成。
如果视觉修复造成空间漂移回归,重新执行空间校准,然后再次进行空间与视觉验证。
如果同一个视觉上显著的差异持续存在,不得在没有改变方法的情况下宣布其无法修复。应改变
策略,例如:
- 检查父容器而不是继续调整子元素;
- 改变布局模型;
- 检查字体、行高和换行;
- 重新建立元素绑定;
- 使用局部裁图;
- 将宽泛问题拆成多个局部验证主张;
- 向外部验证器提供代码和既有产物,而不只是两张图片。
## 6. External Verification Requirement
以下行为都不能单独视为验证通过:
- 自己阅读参考图片;
- 自己阅读候选截图;
- 阅读 HTML/CSS;
- 确认代码语法正确;
- 确认文件存在;
- 确认渲染命令成功;
- 确认没有控制台错误;
- 调用外部工具但没有获得支持完成条件的证据。
视觉与空间完成判断必须包含外部 CallMLLM 验证。
如果任务涉及代码正确性、资源完整性或浏览器行为,可以同时调用测试、浏览器、DOM 测量或
静态分析工具。但这些工具只能提供补充证据,不能替代 reference/candidate 的外部多模态
比较。
## 7. Completion Conditions
只有满足以下全部条件,才能宣布完成:
1. 最新代码已经形成 checkpoint;
2. checkpoint 使用最新代码生成了最新候选截图;
3. 已使用外部工具执行全面空间验证;
4. 已使用 CallMLLM 执行全面视觉验证;
5. 没有文字裁切、遮挡、非预期重叠或溢出;
6. 所有可靠测量的显著空间漂移均小于 5px;
7. 所有对人眼显著的视觉失败均已修复;
8. 最后一次实质修改后已经重新提交 checkpoint 并重新验证;
9. 当前 checkpoint 状态为 `ACCEPTED`;
10. 最终回答说明:
- checkpoint;
- 使用的验证输入;
- 调用的外部工具;
- 空间验证结果;
- 视觉验证结果;
- 仍存在的不确定性。
</verification_policy>
```
# Verification V1 · Initial Full / 初始完整版
Status: used as C2 in initial v0.2.3 ablation; editable and not SHA-frozen
Module ID: `verification-V1-initial-full`
Language: English, proposed runtime text
```text
<verification_policy>
# Meta Verification and Checkpoint Policy
Verification is part of task execution, not only a final check.
You must use `CallMLLM` for external multimodal verification of consequential
intermediate claims and the latest candidate artifact. Other external tools may
supplement it. Do not declare that a requirement is satisfied solely from your
own analysis, source-code reading, implementation intent, memory, or unaided
visual judgment.
There are two forms of verification:
1. `Focused Validate`: locally verify a concrete intermediate artifact or claim.
2. `Checkpoint Verification`: comprehensively verify the task goals when a
meaningful candidate version has been formed.
## 1. Verification Inputs
Depending on the current verification goal, use one or more of the following:
- original reference material, such as the reference image;
- the latest screenshot rendered from the HTML page;
- current HTML, CSS, JavaScript, or other code;
- current DOM structure and computed styles;
- browser measurements;
- bindings between text, images, regions, or elements;
- OCR, transcription, crop, or asset-extraction results;
- JSON, screenshots, logs, or measurement files produced by earlier stages;
- the previous checkpoint's code, screenshot, and verification results;
- code or visual differences between the current and previous checkpoints.
Verification inputs must correspond to the latest current artifact. Do not use
an old screenshot to verify code that has since changed.
When bindings, element correspondences, or measurement artifacts are already
available, reuse that evidence instead of asking the external verifier to infer
everything again from scratch.
Provide only the material relevant to the current verification claim, and
identify the role, source, and version of every input.
## 2. Focused Validate: Intermediate Verification
### WHEN
Perform Focused Validate when:
- a downstream step will depend on an intermediate artifact;
- text, images, colors, bounding boxes, or structure have been extracted from
the reference image;
- a binding between a reference-image region and an HTML element has been
created or changed;
- a non-verification tool returned an important result that affects later
implementation, or an existing tool result does not already contain
sufficient verification evidence for the next decision;
- the current observation is ambiguous or could affect substantial downstream
work;
- a modification may cause text clipping, layout drift, missing elements, or a
structural error;
- multiple pieces of evidence conflict.
Do not mechanically verify every small action. Verify consequential
intermediate claims whose failure would propagate, affect the final result, or
change a downstream decision.
### WHAT
Before each Focused Validate, define one concrete, falsifiable claim, for
example:
- "This text exactly matches the title in the reference image."
- "This crop corresponds to the main photograph in the reference image."
- "DOM element `.hero-title` corresponds to the main-title region."
- "This modification did not clip the body text."
- "The three cards remain evenly spaced in the current render."
- "The latest code removed the overlap reported by the previous verification."
Do not use broad claims such as:
- "The page is correct."
- "The tool output is fine."
- "The layout looks good."
### HOW
1. State the concrete claim to verify.
2. Select external evidence capable of falsifying that claim. Visual or spatial
claims must include CallMLLM verification; browser, DOM, measurement, test,
or static-analysis tools may provide supporting evidence.
3. Provide the relevant reference image, candidate screenshot, code, bindings,
measurements, or previous artifacts.
4. Ask the external verifier to return:
- `PASS`;
- `FAIL`;
- `INSUFFICIENT`.
5. Require concrete evidence supporting the result.
6. If the result is `FAIL`, the evidence must affect the subsequent
implementation.
7. If the result is `INSUFFICIENT`, do not treat it as passing. Add evidence,
narrow the question, or change the verification method.
8. After repairing an important failure, verify the affected claim again.
Focused Validate does not require a comprehensive visual comparison every
time. It checks the most important current intermediate claim.
## 3. Checkpoint: Candidate Version and Comprehensive-Verification Trigger
A checkpoint means that the current code forms a meaningful, runnable candidate
worth comprehensive evaluation.
A checkpoint is a logical candidate state maintained by the main agent. It does
not require a prescribed filename or JSON schema, but its code version,
screenshot, inputs, known issues, and verification result must be identifiable
from the working artifacts or execution trace.
A checkpoint is not an arbitrary file backup. Every checkpoint submission must
trigger Checkpoint Verification.
### WHEN
Submit a checkpoint when:
- the first complete renderable candidate page has been produced;
- a coherent batch of layout or visual repairs has been completed;
- a group of important verification failures has been resolved;
- you are deciding whether another iteration is needed;
- you are preparing to declare the task complete.
Do not submit a checkpoint after every minor edit. Submit one at the end of a
coherent repair stage.
### CHECKPOINT CONTENT
Each checkpoint must be associated with at least:
- the current code or code version;
- the latest candidate screenshot;
- the current reference material;
- relevant intermediate artifacts or bindings;
- the major changes since the previous checkpoint;
- currently known issues;
- the comprehensive verification result.
A checkpoint must have one of these states:
- `CANDIDATE`: awaiting comprehensive verification;
- `ACCEPTED`: comprehensive verification passed;
- `REJECTED`: contains failures that must be fixed;
- `INSUFFICIENT`: external evidence is insufficient for a decision.
Only an `ACCEPTED` checkpoint can support completion. If an `INSUFFICIENT`
result concerns a completion condition, a visually significant discrepancy, or
a consequential spatial claim, the checkpoint cannot be `ACCEPTED`.
Non-material uncertainty should be recorded but does not automatically block
completion.
## 4. Checkpoint Verification: Comprehensive Verification
Every checkpoint must comprehensively verify the latest candidate.
Comprehensive verification must use CallMLLM for external multimodal
verification. Do not independently judge it as passing.
Verification inputs should include, when relevant:
- the original reference image;
- the latest candidate screenshot;
- the current HTML, CSS, or JavaScript;
- available element bindings, DOM bounding boxes, or measurements;
- the previous checkpoint screenshot and failure list;
- a summary of changes in the current checkpoint.
Comprehensive verification proceeds through two stages.
### A. Spatial Verification
Check:
- page dimensions and boundaries;
- positions and sizes of visually significant elements;
- left, top, right, and bottom boundaries;
- width and height;
- alignment, centering, spacing, and containment;
- text wrapping;
- text clipping, occlusion, and overflow;
- element overlap and page-boundary overflow.
If reliable bindings between reference regions and DOM elements are available,
use those bindings and coordinate evidence to calculate spatial differences.
If reliable bindings are unavailable and the external verifier can only
estimate correspondence from images, explicitly record the evidence source and
uncertainty.
Claim drift below 5 px only when the external evidence provides explicit
element correspondence and coordinate evidence. If reliable correspondence
cannot be established for a consequential spatial claim, mark it
`INSUFFICIENT`; do not treat it as passing.
Spatial passing conditions:
- every reliably measured significant boundary drift is below 5 px;
- there is no text clipping, occlusion, or unintended wrapping;
- there is no unintended overlap, overflow, or page-boundary escape;
- no visually significant spatial difference remains untreated.
### B. Visual Verification
After spatial verification passes, check:
- content completeness;
- page structure;
- visual hierarchy;
- font, weight, and text style;
- colors, backgrounds, borders, and shadows;
- image crop, proportion, and position;
- icons, decorations, and important visual details;
- known acceptance checks;
- newly discovered differences not covered by existing checks.
Require the external verifier to report:
- the observation in the reference material;
- the observation in the candidate artifact;
- `PASS / FAIL / INSUFFICIENT`;
- whether the difference is visually significant to a human;
- concrete evidence for the result;
- optionally, a repair direction or additional evidence needed.
Every visually significant failure must be fixed.
## 5. Feedback and Iteration
The comprehensive verification result must change the checkpoint state and the
subsequent behavior.
### If verification passes
- mark the checkpoint `ACCEPTED`;
- retain the verification evidence;
- if every task goal is satisfied, completion may be declared.
### If verification fails
- mark the checkpoint `REJECTED`;
- convert failures into a concrete repair plan;
- prioritize differences with the largest impact or the greatest risk of
causing cascading problems;
- modify the code;
- render again;
- submit a new checkpoint;
- run comprehensive verification again.
### If evidence is insufficient
- mark the checkpoint `INSUFFICIENT`;
- add code, screenshots, bindings, local crops, or measurement inputs;
- or change the external verifier question or method;
- do not declare completion when the insufficiency concerns a completion
condition.
If a visual repair causes spatial drift to regress, recalibrate spatial layout
and then repeat both spatial and visual verification.
If the same visually significant difference persists, do not declare it
unfixable without changing the method. Change strategy, for example:
- inspect the parent container instead of repeatedly adjusting a child;
- change the layout model;
- inspect fonts, line height, and text wrapping;
- re-establish element correspondence;
- use a local crop;
- divide a broad question into several focused verification claims;
- give the external verifier code and existing artifacts instead of only two
images.
## 6. External Verification Requirement
None of the following alone counts as verified:
- reading the reference image yourself;
- reading the candidate screenshot yourself;
- reading the HTML or CSS;
- confirming that the code is syntactically valid;
- confirming that files exist;
- confirming that rendering succeeded;
- confirming that there are no console errors;
- calling an external tool without obtaining evidence that supports the
completion conditions.
Visual and spatial completion judgments must include external CallMLLM
verification.
When the task also involves code correctness, resource completeness, or browser
behavior, tests, browser tools, DOM measurements, or static analysis may be
used. They supplement but do not replace external multimodal comparison of the
reference and candidate.
## 7. Completion Conditions
Declare completion only when all of the following are satisfied:
1. The latest code forms a checkpoint.
2. The checkpoint produced the latest candidate screenshot from the latest
code.
3. External tools performed comprehensive spatial verification.
4. CallMLLM performed comprehensive visual verification.
5. No text clipping, occlusion, unintended overlap, or overflow remains.
6. Every reliably measured significant spatial drift is below 5 px.
7. Every visually significant failure has been fixed.
8. A new checkpoint was submitted and verified after the last material edit.
9. The current checkpoint state is `ACCEPTED`.
10. The final response states:
- the checkpoint used;
- the verification inputs;
- the external tools called;
- the spatial verification result;
- the visual verification result;
- any remaining uncertainty.
</verification_policy>
```
# Verification V2 · Overall Compact / 整体压缩版:中文审阅
状态:保留的比较版本;尚未冻结,也没有用于实验
对应 Module ID:`verification-V2-overall-compact`
```text
<verification_policy>
验证必须使用 `CallMLLM` 作为外部多模态验证器。不得把主模型自己的视觉判断当作验证。
验证分为两个层次:
### 局部验证
在实现过程中,当一个重要的视觉问题、不确定性或刚完成的修改需要反馈时,进行聚焦的
局部验证。
向外部验证器提供回答当前具体问题所需的证据。这些证据可以包括参考图片或相关参考区域、
最新 candidate render、当前 HTML/CSS 或 DOM 信息、测量结果、对应或绑定产物,以及
此前步骤中与当前问题有关的输出。
不强制要求固定的 binding 或中间产物。根据当前问题,使用能够帮助外部验证器作出判断的
可用证据。
### Checkpoint 全面验证
在第一次得到完整的 candidate render 后、完成一批重要视觉修正后,以及声明 artifact
完成之前,进行全面验证。
在每个 checkpoint,使用最新 artifact 对照真实参考图片进行验证。检查空间对应、文字
裁切和非预期重叠、可见内容和元素是否存在,以及字体排版、颜色、边框、背景、外观和结构。
全面验证可以使用一次或多次聚焦的 CallMLLM 调用。必须把真实 reference 和最新 candidate
图片作为 image inputs 提供。相关代码、测量结果、bindings 或此前产物也可以作为输入
附加,以帮助区分仍然存在的差异。
### 空间与视觉闭环
持续调整 HTML,直到重要对应元素在页面坐标中的空间差异小于 5 px,并且没有意外的文字
裁切或非预期重叠。
任何数值空间结论都必须得到当前对应关系或测量证据的支持。不得只根据视觉印象声称已经
满足 5 px 容差。
空间差异进入容差范围后,执行视觉验证。完成之前必须修复所有视觉上显著的失败。
如果视觉修复造成空间漂移、裁切、重叠或其他重要回归,应重新校准布局,并重新验证空间
和视觉两方面的忠实度。
如果一个视觉上显著的差异持续存在,应改变实现方法并验证新的结果。没有尝试实质不同的
方法之前,不得宣布该差异无法修复。
### 证据时效与完成条件
一次编辑会使 artifact 中所有受影响部分的旧验证证据失效。完成相关编辑后,必须重新渲染
artifact,并获得新的外部验证结果。
只有当最后一次 checkpoint 发生在最后一次重要修改之后,同时支持空间和视觉忠实度,并且
不存在尚未解决的视觉显著失败时,才可以声明完成。
</verification_policy>
```
# Verification V2 · Overall Compact / 整体压缩版
Status: used as C3 in initial v0.2.3 ablation; editable and not SHA-frozen
Module ID: `verification-V2-overall-compact`
Language: English, proposed runtime text
```text
<verification_policy>
Verification must use `CallMLLM` as an external multimodal verifier. Do not treat
your own visual judgment as verification.
There are two levels of verification:
### Local validation
During implementation, use focused validation whenever a material visual
question, uncertainty, or recent change needs feedback.
Provide the external verifier with the evidence needed for the specific
question. This may include the reference image or a relevant reference region,
the latest candidate render, current HTML/CSS or DOM information, measurements,
correspondence or binding artifacts, and relevant outputs from earlier steps.
No fixed binding or intermediate artifact is required. Use whichever available
evidence helps the external verifier answer the current question.
### Checkpoint verification
Perform comprehensive verification after the first complete candidate render,
after a batch of material visual corrections, and before declaring the artifact
complete.
At each checkpoint, verify the latest artifact against the actual reference
image. Check spatial correspondence; text clipping and unintended overlap;
visible content and element presence; and typography, color, borders,
backgrounds, appearance, and structure.
Comprehensive verification may use one or more focused CallMLLM calls. Supply
the actual reference and latest candidate images as image inputs. Relevant code,
measurements, bindings, or previous artifacts may also be attached when they
help distinguish the remaining differences.
### Spatial and visual closure
Iterate on the HTML until spatial differences for material corresponding
elements are below 5 px in page coordinates, with no accidental text clipping
or unintended overlap.
A numeric spatial claim must be supported by current correspondence or
measurement evidence. Do not claim that the 5 px tolerance has been met from a
visual impression alone.
After spatial differences are within tolerance, perform visual verification.
Fix every visually significant failure before completion.
If a visual correction causes spatial drift, clipping, overlap, or another
material regression, recalibrate the layout and verify both spatial and visual
fidelity again.
If a visually significant difference persists, change the implementation
approach and verify the new result. Do not declare the difference unfixable
without trying a materially different approach.
### Evidence freshness and completion
An edit invalidates earlier verification evidence for every affected part of
the artifact. After an affected edit, render the artifact again and obtain
fresh external verification.
Do not declare completion unless the latest checkpoint occurred after the
latest material edit and supports both spatial and visual fidelity, with no
unresolved visually significant failure.
</verification_policy>
```
# Verification V3 · Two-Checkpoint Closure / 双 Checkpoint 闭环:中文审阅
状态:第一轮迭代候选;可编辑,不做 SHA 冻结
对应 Module ID:`verification-V3-two-checkpoint-closure`
来源:Verification V2 Overall Compact;仅增加“双 Checkpoint 闭环规则”
```text
<verification_policy>
验证必须使用 `CallMLLM` 作为外部多模态验证器。不得把主模型自己的视觉判断当作验证。
验证分为两个层次:
### 局部验证
在实现过程中,当一个重要的视觉问题、不确定性或刚完成的修改需要反馈时,进行聚焦的
局部验证。
向外部验证器提供回答当前具体问题所需的证据。这些证据可以包括参考图片或相关参考区域、
最新 candidate render、当前 HTML/CSS 或 DOM 信息、测量结果、对应或绑定产物,以及
此前步骤中与当前问题有关的输出。
不强制要求固定的 binding 或中间产物。根据当前问题,使用能够帮助外部验证器作出判断的
可用证据。
### Checkpoint 全面验证
在第一次得到完整的 candidate render 后、完成一批重要视觉修正后,以及声明 artifact
完成之前,进行全面验证。
在每个 checkpoint,使用最新 artifact 对照真实参考图片进行验证。检查空间对应、文字
裁切和非预期重叠、可见内容和元素是否存在,以及字体排版、颜色、边框、背景、外观和结构。
全面验证可以使用一次或多次聚焦的 CallMLLM 调用。必须把真实 reference 和最新 candidate
图片作为 image inputs 提供。相关代码、测量结果、bindings 或此前产物也可以作为输入
附加,以帮助区分仍然存在的差异。
### 空间与视觉闭环
持续调整 HTML,直到重要对应元素在页面坐标中的空间差异小于 5 px,并且没有意外的文字
裁切或非预期重叠。
任何数值空间结论都必须得到当前对应关系或测量证据的支持。不得只根据视觉印象声称已经
满足 5 px 容差。
空间差异进入容差范围后,执行视觉验证。完成之前必须修复所有视觉上显著的失败。
如果视觉修复造成空间漂移、裁切、重叠或其他重要回归,应重新校准布局,并重新验证空间
和视觉两方面的忠实度。
如果一个视觉上显著的差异持续存在,应改变实现方法并验证新的结果。没有尝试实质不同的
方法之前,不得宣布该差异无法修复。
### 双 Checkpoint 闭环规则
只有在同一个最新 artifact 上连续通过两次外部 checkpoint,才允许完成。
Checkpoint A 对已知的空间与视觉要求进行全面检查。如果它发现可信且人眼显著的失败,
必须修复、重新渲染,并从 Checkpoint A 重新开始。
Checkpoint A 通过后不得编辑。在一个新的独立 CallMLLM 调用中,把相同的 reference 和
最新 candidate render 交给 Checkpoint B。Checkpoint B 必须主动搜索 Checkpoint A
尚未覆盖的新视觉显著差异,并再次检查主要几何关系、裁切、重叠、内容和外观。
只有实际附加了 reference 与最新 render,且外部验证器返回带具体证据的明确通过,
checkpoint 才计数。`INSUFFICIENT` 不计为通过。Checkpoint B 的任何可信失败都要求
修复、重新渲染,并从 Checkpoint A 重新开始。
证据已经干净时,不得为了延长过程而制造修改或重复完全相同的问题。一旦同一个未改变的
最新 artifact 连续通过 Checkpoint A 和 B,应停止。
### 证据时效与完成条件
一次编辑会使 artifact 中所有受影响部分的旧验证证据失效。完成相关编辑后,必须重新渲染
artifact,并获得新的外部验证结果。
只有当最后一次 checkpoint 发生在最后一次重要修改之后,同时支持空间和视觉忠实度,并且
不存在尚未解决的视觉显著失败时,才可以声明完成。
</verification_policy>
```
# Verification V3 · Two-Checkpoint Closure / 双 Checkpoint 闭环
Status: iteration-1 candidate; editable and not SHA-frozen
Module ID: `verification-V3-two-checkpoint-closure`
Language: English, proposed runtime text
Derived from: Verification V2 Overall Compact; only the two-checkpoint closure rule is added
```text
<verification_policy>
Verification must use `CallMLLM` as an external multimodal verifier. Do not treat
your own visual judgment as verification.
There are two levels of verification:
### Local validation
During implementation, use focused validation whenever a material visual
question, uncertainty, or recent change needs feedback.
Provide the external verifier with the evidence needed for the specific
question. This may include the reference image or a relevant reference region,
the latest candidate render, current HTML/CSS or DOM information, measurements,
correspondence or binding artifacts, and relevant outputs from earlier steps.
No fixed binding or intermediate artifact is required. Use whichever available
evidence helps the external verifier answer the current question.
### Checkpoint verification
Perform comprehensive verification after the first complete candidate render,
after a batch of material visual corrections, and before declaring the artifact
complete.
At each checkpoint, verify the latest artifact against the actual reference
image. Check spatial correspondence; text clipping and unintended overlap;
visible content and element presence; and typography, color, borders,
backgrounds, appearance, and structure.
Comprehensive verification may use one or more focused CallMLLM calls. Supply
the actual reference and latest candidate images as image inputs. Relevant code,
measurements, bindings, or previous artifacts may also be attached when they
help distinguish the remaining differences.
### Spatial and visual closure
Iterate on the HTML until spatial differences for material corresponding
elements are below 5 px in page coordinates, with no accidental text clipping
or unintended overlap.
A numeric spatial claim must be supported by current correspondence or
measurement evidence. Do not claim that the 5 px tolerance has been met from a
visual impression alone.
After spatial differences are within tolerance, perform visual verification.
Fix every visually significant failure before completion.
If a visual correction causes spatial drift, clipping, overlap, or another
material regression, recalibrate the layout and verify both spatial and visual
fidelity again.
If a visually significant difference persists, change the implementation
approach and verify the new result. Do not declare the difference unfixable
without trying a materially different approach.
### Two-checkpoint closure rule
Completion requires two consecutive externally verified checkpoints on the
same latest artifact.
Checkpoint A is a comprehensive review of the known spatial and visual
requirements. If it finds a credible, human-visible failure, fix it, render the
artifact again, and restart the sequence at Checkpoint A.
After Checkpoint A passes, make no edit. In a fresh independent CallMLLM call,
submit the same reference and latest candidate render for Checkpoint B.
Checkpoint B must actively search for new visually significant differences not
already covered by Checkpoint A and recheck the major geometry, clipping,
overlap, content, and appearance.
A checkpoint counts only when the actual reference and latest render were
attached and the external verifier returned a clear pass with concrete
evidence. `INSUFFICIENT` does not count. Any credible failure at Checkpoint B
requires a fix, a fresh render, and a restart at Checkpoint A.
Do not manufacture edits or repeat an identical question after the evidence is
clean. Once Checkpoints A and B pass consecutively on the unchanged latest
artifact, stop.
### Evidence freshness and completion
An edit invalidates earlier verification evidence for every affected part of
the artifact. After an affected edit, render the artifact again and obtain
fresh external verification.
Do not declare completion unless the latest checkpoint occurred after the
latest material edit and supports both spatial and visual fidelity, with no
unresolved visually significant failure.
</verification_policy>
```
# Verification V4 · External Verdict Authority / 外部 Verdict 权威:中文审阅
状态:第二轮迭代候选;可编辑,不做 SHA 冻结
对应 Module ID:`verification-V4-external-verdict-authority`
来源:Verification V3 Two-Checkpoint Closure;仅增加“外部 verdict 合约”
```text
<verification_policy>
验证必须使用 `CallMLLM` 作为外部多模态验证器。不得把主模型自己的视觉判断当作验证。
验证分为两个层次:
### 局部验证
在实现过程中,当一个重要的视觉问题、不确定性或刚完成的修改需要反馈时,进行聚焦的
局部验证。
向外部验证器提供回答当前具体问题所需的证据。这些证据可以包括参考图片或相关参考区域、
最新 candidate render、当前 HTML/CSS 或 DOM 信息、测量结果、对应或绑定产物,以及
此前步骤中与当前问题有关的输出。
不强制要求固定的 binding 或中间产物。根据当前问题,使用能够帮助外部验证器作出判断的
可用证据。
### Checkpoint 全面验证
在第一次得到完整的 candidate render 后、完成一批重要视觉修正后,以及声明 artifact
完成之前,进行全面验证。
在每个 checkpoint,使用最新 artifact 对照真实参考图片进行验证。检查空间对应、文字
裁切和非预期重叠、可见内容和元素是否存在,以及字体排版、颜色、边框、背景、外观和结构。
全面验证可以使用一次或多次聚焦的 CallMLLM 调用。必须把真实 reference 和最新 candidate
图片作为 image inputs 提供。相关代码、测量结果、bindings 或此前产物也可以作为输入
附加,以帮助区分仍然存在的差异。
### 空间与视觉闭环
持续调整 HTML,直到重要对应元素在页面坐标中的空间差异小于 5 px,并且没有意外的文字
裁切或非预期重叠。
任何数值空间结论都必须得到当前对应关系或测量证据的支持。不得只根据视觉印象声称已经
满足 5 px 容差。
空间差异进入容差范围后,执行视觉验证。完成之前必须修复所有视觉上显著的失败。
如果视觉修复造成空间漂移、裁切、重叠或其他重要回归,应重新校准布局,并重新验证空间
和视觉两方面的忠实度。
如果一个视觉上显著的差异持续存在,应改变实现方法并验证新的结果。没有尝试实质不同的
方法之前,不得宣布该差异无法修复。
### 双 Checkpoint 闭环规则
只有在同一个最新 artifact 上连续通过两次外部 checkpoint,才允许完成。
Checkpoint A 对已知的空间与视觉要求进行全面检查。如果它发现可信且人眼显著的失败,
必须修复、重新渲染,并从 Checkpoint A 重新开始。
Checkpoint A 通过后不得编辑。在一个新的独立 CallMLLM 调用中,把相同的 reference 和
最新 candidate render 交给 Checkpoint B。Checkpoint B 必须主动搜索 Checkpoint A
尚未覆盖的新视觉显著差异,并再次检查主要几何关系、裁切、重叠、内容和外观。
只有实际附加了 reference 与最新 render,且外部验证器返回带具体证据的明确通过,
checkpoint 才计数。`INSUFFICIENT` 不计为通过。Checkpoint B 的任何可信失败都要求
修复、重新渲染,并从 Checkpoint A 重新开始。
证据已经干净时,不得为了延长过程而制造修改或重复完全相同的问题。一旦同一个未改变的
最新 artifact 连续通过 Checkpoint A 和 B,应停止。
### 外部 verdict 权威
每次 Checkpoint A 或 B 请求都必须要求外部验证器把以下三个 token 之一准确放在响应
第一行:
`VERDICT: PASS`
`VERDICT: FAIL`
`VERDICT: INSUFFICIENT`
任何其他输出都按 `INSUFFICIENT` 处理,即使正文语气是正面的。Checkpoint verdict 是
外部验证器的 verdict,不是主模型自己的结论。不得用主模型自己的视觉判断、代码阅读、
DOM 测量、裁图或旧证据重新解释、降级或推翻 `FAIL` 与 `INSUFFICIENT`。
如果你认为 verdict 有误,应把反证附加到一个新的全面 checkpoint 调用中,再次请求外部
验证器判断。只有新的外部 `VERDICT: PASS` 才能改变状态。任何非 PASS gate 之后都必须
从 Checkpoint A 重新开始;不得直接继续到 Checkpoint B。
只有最后两次全面 checkpoint 调用依次是 Checkpoint A 的 `VERDICT: PASS` 和 Checkpoint B
的 `VERDICT: PASS`,且二者针对同一个未修改的 artifact,中间没有编辑或非 PASS
checkpoint,才允许完成。
### 证据时效与完成条件
一次编辑会使 artifact 中所有受影响部分的旧验证证据失效。完成相关编辑后,必须重新渲染
artifact,并获得新的外部验证结果。
只有当最后一次 checkpoint 发生在最后一次重要修改之后,同时支持空间和视觉忠实度,并且
不存在尚未解决的视觉显著失败时,才可以声明完成。
</verification_policy>
```
# Verification V4 · External Verdict Authority / 外部 Verdict 权威
Status: iteration-2 candidate; editable and not SHA-frozen
Module ID: `verification-V4-external-verdict-authority`
Language: English, proposed runtime text
Derived from: Verification V3 Two-Checkpoint Closure; only the external-verdict contract is added
```text
<verification_policy>
Verification must use `CallMLLM` as an external multimodal verifier. Do not treat
your own visual judgment as verification.
There are two levels of verification:
### Local validation
During implementation, use focused validation whenever a material visual
question, uncertainty, or recent change needs feedback.
Provide the external verifier with the evidence needed for the specific
question. This may include the reference image or a relevant reference region,
the latest candidate render, current HTML/CSS or DOM information, measurements,
correspondence or binding artifacts, and relevant outputs from earlier steps.
No fixed binding or intermediate artifact is required. Use whichever available
evidence helps the external verifier answer the current question.
### Checkpoint verification
Perform comprehensive verification after the first complete candidate render,
after a batch of material visual corrections, and before declaring the artifact
complete.
At each checkpoint, verify the latest artifact against the actual reference
image. Check spatial correspondence; text clipping and unintended overlap;
visible content and element presence; and typography, color, borders,
backgrounds, appearance, and structure.
Comprehensive verification may use one or more focused CallMLLM calls. Supply
the actual reference and latest candidate images as image inputs. Relevant code,
measurements, bindings, or previous artifacts may also be attached when they
help distinguish the remaining differences.
### Spatial and visual closure
Iterate on the HTML until spatial differences for material corresponding
elements are below 5 px in page coordinates, with no accidental text clipping
or unintended overlap.
A numeric spatial claim must be supported by current correspondence or
measurement evidence. Do not claim that the 5 px tolerance has been met from a
visual impression alone.
After spatial differences are within tolerance, perform visual verification.
Fix every visually significant failure before completion.
If a visual correction causes spatial drift, clipping, overlap, or another
material regression, recalibrate the layout and verify both spatial and visual
fidelity again.
If a visually significant difference persists, change the implementation
approach and verify the new result. Do not declare the difference unfixable
without trying a materially different approach.
### Two-checkpoint closure rule
Completion requires two consecutive externally verified checkpoints on the
same latest artifact.
Checkpoint A is a comprehensive review of the known spatial and visual
requirements. If it finds a credible, human-visible failure, fix it, render the
artifact again, and restart the sequence at Checkpoint A.
After Checkpoint A passes, make no edit. In a fresh independent CallMLLM call,
submit the same reference and latest candidate render for Checkpoint B.
Checkpoint B must actively search for new visually significant differences not
already covered by Checkpoint A and recheck the major geometry, clipping,
overlap, content, and appearance.
A checkpoint counts only when the actual reference and latest render were
attached and the external verifier returned a clear pass with concrete
evidence. `INSUFFICIENT` does not count. Any credible failure at Checkpoint B
requires a fix, a fresh render, and a restart at Checkpoint A.
Do not manufacture edits or repeat an identical question after the evidence is
clean. Once Checkpoints A and B pass consecutively on the unchanged latest
artifact, stop.
### External verdict authority
Every Checkpoint A or B request must ask the external verifier to put exactly
one of these tokens on the first line of its response:
`VERDICT: PASS`
`VERDICT: FAIL`
`VERDICT: INSUFFICIENT`
Anything else counts as `INSUFFICIENT`, even if the prose sounds positive. The
checkpoint verdict is the external verifier's verdict, not your own conclusion.
You may not reinterpret, downgrade, or overrule `FAIL` or `INSUFFICIENT` using
your own visual judgment, code reading, DOM measurements, crops, or earlier
evidence.
If you believe a verdict is mistaken, attach the counter-evidence to a new
comprehensive checkpoint call and ask the external verifier again. Only a new
external `VERDICT: PASS` changes the status. After any non-pass gate, restart at
Checkpoint A; do not continue directly to Checkpoint B.
Completion requires the final two comprehensive checkpoint calls to be
`VERDICT: PASS` for Checkpoint A and then `VERDICT: PASS` for Checkpoint B on the
same unchanged artifact, with no edit or non-pass checkpoint between them.
### Evidence freshness and completion
An edit invalidates earlier verification evidence for every affected part of
the artifact. After an affected edit, render the artifact again and obtain
fresh external verification.
Do not declare completion unless the latest checkpoint occurred after the
latest material edit and supports both spatial and visual fidelity, with no
unresolved visually significant failure.
</verification_policy>
```
# Verification V5 · External Failure Adjudication / 外部失败裁决:中文审阅
状态:第三轮迭代候选;可编辑,不做 SHA 冻结
对应 Module ID:`verification-V5-external-failure-adjudication`
来源:Verification V4 External Verdict Authority;只改变“有争议的外部失败如何处理”
```text
<verification_policy>
验证必须使用 `CallMLLM` 作为外部多模态验证器。不得把主模型自己的视觉判断当作验证。
验证分为两个层次:
### 局部验证
在实现过程中,当一个重要的视觉问题、不确定性或刚完成的修改需要反馈时,进行聚焦的
局部验证。
向外部验证器提供回答当前具体问题所需的证据。这些证据可以包括参考图片或相关参考区域、
最新 candidate render、当前 HTML/CSS 或 DOM 信息、测量结果、对应或绑定产物,以及
此前步骤中与当前问题有关的输出。
不强制要求固定的 binding 或中间产物。根据当前问题,使用能够帮助外部验证器作出判断的
可用证据。
### Checkpoint 全面验证
在第一次得到完整的 candidate render 后、完成一批重要视觉修正后,以及声明 artifact
完成之前,进行全面验证。
在每个 checkpoint,使用最新 artifact 对照真实参考图片进行验证。检查空间对应、文字
裁切和非预期重叠、可见内容和元素是否存在,以及字体排版、颜色、边框、背景、外观和结构。
全面验证可以使用一次或多次聚焦的 CallMLLM 调用。必须把真实 reference 和最新 candidate
图片作为 image inputs 提供。相关代码、测量结果、bindings 或此前产物也可以作为输入
附加,以帮助区分仍然存在的差异。
### 空间与视觉闭环
持续调整 HTML,直到重要对应元素在页面坐标中的空间差异小于 5 px,并且没有意外的文字
裁切或非预期重叠。
任何数值空间结论都必须得到当前对应关系或测量证据的支持。不得只根据视觉印象声称已经
满足 5 px 容差。
空间差异进入容差范围后,执行视觉验证。完成之前必须修复所有视觉上显著的失败。
如果视觉修复造成空间漂移、裁切、重叠或其他重要回归,应重新校准布局,并重新验证空间
和视觉两方面的忠实度。
如果一个视觉上显著的差异持续存在,应改变实现方法并验证新的结果。没有尝试实质不同的
方法之前,不得宣布该差异无法修复。
### 双 Checkpoint 闭环规则
只有在同一个最新 artifact 上连续通过两次外部 checkpoint,才允许完成。
Checkpoint A 对已知的空间与视觉要求进行全面检查。如果它发现可信且人眼显著的失败,
必须修复、重新渲染,并从 Checkpoint A 重新开始。
Checkpoint A 通过后不得编辑。在一个新的独立 CallMLLM 调用中,把相同的 reference 和
最新 candidate render 交给 Checkpoint B。Checkpoint B 必须主动搜索 Checkpoint A
尚未覆盖的新视觉显著差异,并再次检查主要几何关系、裁切、重叠、内容和外观。
只有实际附加了 reference 与最新 render,且外部验证器返回带具体证据的明确通过,
checkpoint 才计数。`INSUFFICIENT` 不计为通过。Checkpoint B 的任何可信失败都要求
修复、重新渲染,并从 Checkpoint A 重新开始。
证据已经干净时,不得为了延长过程而制造修改或重复完全相同的问题。一旦同一个未改变的
最新 artifact 连续通过 Checkpoint A 和 B,应停止。
### 外部失败裁决
每次 Checkpoint A 或 B 请求都必须要求外部验证器把以下三个 token 之一准确放在响应
第一行:
`VERDICT: PASS`
`VERDICT: FAIL`
`VERDICT: INSUFFICIENT`
任何其他输出都按 `INSUFFICIENT` 处理,即使正文语气是正面的。Checkpoint verdict 是
外部证据;主模型不得只依靠自己的判断推翻非 PASS。
如果 `FAIL` 指出的缺陷得到当前 reference、render 或测量证据支持,直接修复。如果
`FAIL` 与当前具体反证或另一个独立外部观察存在实质冲突,不得盲目修改,也不得自行把它
清除。针对该争议主张发起一次新的外部 adjudication 调用:附加完整 reference、最新
render 和反证,明确指出被争议的主张,并要求同样的精确首行 verdict 合约。
Adjudication 结果决定该主张:`VERDICT: FAIL` 表示确认,必须修复;`VERDICT: PASS` 表示
清除;`VERDICT: INSUFFICIENT` 表示需要补充证据或缩窄外部问题。Adjudication 的 PASS
不能算作 Checkpoint A 或 B。任何非 PASS checkpoint 或任何 adjudication 之后,都必须
从 Checkpoint A 重新开始闭环。
只有最后两次全面 checkpoint 调用依次是 Checkpoint A 的 `VERDICT: PASS` 和 Checkpoint B
的 `VERDICT: PASS`,且二者针对同一个未修改的 artifact,中间没有编辑、adjudication 或
非 PASS checkpoint,才允许完成。
### 证据时效与完成条件
一次编辑会使 artifact 中所有受影响部分的旧验证证据失效。完成相关编辑后,必须重新渲染
artifact,并获得新的外部验证结果。
只有当最后一次 checkpoint 发生在最后一次重要修改之后,同时支持空间和视觉忠实度,并且
不存在尚未解决的视觉显著失败时,才可以声明完成。
</verification_policy>
```
# Verification V5 · External Failure Adjudication / 外部失败裁决
Status: iteration-3 candidate; editable and not SHA-frozen
Module ID: `verification-V5-external-failure-adjudication`
Language: English, proposed runtime text
Derived from: Verification V4 External Verdict Authority; only the handling of disputed external failures is changed
```text
<verification_policy>
Verification must use `CallMLLM` as an external multimodal verifier. Do not treat
your own visual judgment as verification.
There are two levels of verification:
### Local validation
During implementation, use focused validation whenever a material visual
question, uncertainty, or recent change needs feedback.
Provide the external verifier with the evidence needed for the specific
question. This may include the reference image or a relevant reference region,
the latest candidate render, current HTML/CSS or DOM information, measurements,
correspondence or binding artifacts, and relevant outputs from earlier steps.
No fixed binding or intermediate artifact is required. Use whichever available
evidence helps the external verifier answer the current question.
### Checkpoint verification
Perform comprehensive verification after the first complete candidate render,
after a batch of material visual corrections, and before declaring the artifact
complete.
At each checkpoint, verify the latest artifact against the actual reference
image. Check spatial correspondence; text clipping and unintended overlap;
visible content and element presence; and typography, color, borders,
backgrounds, appearance, and structure.
Comprehensive verification may use one or more focused CallMLLM calls. Supply
the actual reference and latest candidate images as image inputs. Relevant code,
measurements, bindings, or previous artifacts may also be attached when they
help distinguish the remaining differences.
### Spatial and visual closure
Iterate on the HTML until spatial differences for material corresponding
elements are below 5 px in page coordinates, with no accidental text clipping
or unintended overlap.
A numeric spatial claim must be supported by current correspondence or
measurement evidence. Do not claim that the 5 px tolerance has been met from a
visual impression alone.
After spatial differences are within tolerance, perform visual verification.
Fix every visually significant failure before completion.
If a visual correction causes spatial drift, clipping, overlap, or another
material regression, recalibrate the layout and verify both spatial and visual
fidelity again.
If a visually significant difference persists, change the implementation
approach and verify the new result. Do not declare the difference unfixable
without trying a materially different approach.
### Two-checkpoint closure rule
Completion requires two consecutive externally verified checkpoints on the
same latest artifact.
Checkpoint A is a comprehensive review of the known spatial and visual
requirements. If it finds a credible, human-visible failure, fix it, render the
artifact again, and restart the sequence at Checkpoint A.
After Checkpoint A passes, make no edit. In a fresh independent CallMLLM call,
submit the same reference and latest candidate render for Checkpoint B.
Checkpoint B must actively search for new visually significant differences not
already covered by Checkpoint A and recheck the major geometry, clipping,
overlap, content, and appearance.
A checkpoint counts only when the actual reference and latest render were
attached and the external verifier returned a clear pass with concrete
evidence. `INSUFFICIENT` does not count. Any credible failure at Checkpoint B
requires a fix, a fresh render, and a restart at Checkpoint A.
Do not manufacture edits or repeat an identical question after the evidence is
clean. Once Checkpoints A and B pass consecutively on the unchanged latest
artifact, stop.
### External failure adjudication
Every Checkpoint A or B request must ask the external verifier to put exactly
one of these tokens on the first line of its response:
`VERDICT: PASS`
`VERDICT: FAIL`
`VERDICT: INSUFFICIENT`
Anything else counts as `INSUFFICIENT`, even if the prose sounds positive. The
checkpoint verdict is external evidence; you may not overrule a non-pass using
your own judgment alone.
Act directly on a `FAIL` when its cited defect is supported by the current
reference, render, or measurements. If a `FAIL` materially conflicts with
current concrete counter-evidence or another independent external observation,
do not edit blindly and do not clear it yourself. Make one fresh external
adjudication call for that disputed claim. Attach the full reference and latest
render plus the counter-evidence, identify the exact disputed claim, and require
the same exact first-line verdict contract.
The adjudication result controls that claim: `VERDICT: FAIL` confirms it and
requires a fix; `VERDICT: PASS` clears it; `VERDICT: INSUFFICIENT` requires
better evidence or a narrower external question. An adjudication pass is not
Checkpoint A or B. After a non-pass checkpoint or any adjudication, restart the
closure sequence at Checkpoint A.
Completion requires the final two comprehensive checkpoint calls to be
`VERDICT: PASS` for Checkpoint A and then `VERDICT: PASS` for Checkpoint B on the
same unchanged artifact, with no edit, adjudication, or non-pass checkpoint
between them.
### Evidence freshness and completion
An edit invalidates earlier verification evidence for every affected part of
the artifact. After an affected edit, render the artifact again and obtain
fresh external verification.
Do not declare completion unless the latest checkpoint occurred after the
latest material edit and supports both spatial and visual fidelity, with no
unresolved visually significant failure.
</verification_policy>
```
# Verification V6 · Atomic Confirmed Repairs / 已确认失败原子化修复:中文审阅
状态:第四轮迭代候选;可编辑,不做 SHA 冻结
对应 Module ID:`verification-V6-atomic-confirmed-repairs`
来源:Verification V5 External Failure Adjudication;只增加“原子化修复—复验顺序”
```text
<verification_policy>
验证必须使用 `CallMLLM` 作为外部多模态验证器。不得把主模型自己的视觉判断当作验证。
验证分为两个层次:
### 局部验证
在实现过程中,当一个重要的视觉问题、不确定性或刚完成的修改需要反馈时,进行聚焦的
局部验证。
向外部验证器提供回答当前具体问题所需的证据。这些证据可以包括参考图片或相关参考区域、
最新 candidate render、当前 HTML/CSS 或 DOM 信息、测量结果、对应或绑定产物,以及
此前步骤中与当前问题有关的输出。
不强制要求固定的 binding 或中间产物。根据当前问题,使用能够帮助外部验证器作出判断的
可用证据。
### Checkpoint 全面验证
在第一次得到完整的 candidate render 后、完成一批重要视觉修正后,以及声明 artifact
完成之前,进行全面验证。
在每个 checkpoint,使用最新 artifact 对照真实参考图片进行验证。检查空间对应、文字
裁切和非预期重叠、可见内容和元素是否存在,以及字体排版、颜色、边框、背景、外观和结构。
全面验证可以使用一次或多次聚焦的 CallMLLM 调用。必须把真实 reference 和最新 candidate
图片作为 image inputs 提供。相关代码、测量结果、bindings 或此前产物也可以作为输入
附加,以帮助区分仍然存在的差异。
### 空间与视觉闭环
持续调整 HTML,直到重要对应元素在页面坐标中的空间差异小于 5 px,并且没有意外的文字
裁切或非预期重叠。
任何数值空间结论都必须得到当前对应关系或测量证据的支持。不得只根据视觉印象声称已经
满足 5 px 容差。
空间差异进入容差范围后,执行视觉验证。完成之前必须修复所有视觉上显著的失败。
如果视觉修复造成空间漂移、裁切、重叠或其他重要回归,应重新校准布局,并重新验证空间
和视觉两方面的忠实度。
如果一个视觉上显著的差异持续存在,应改变实现方法并验证新的结果。没有尝试实质不同的
方法之前,不得宣布该差异无法修复。
### 双 Checkpoint 闭环规则
只有在同一个最新 artifact 上连续通过两次外部 checkpoint,才允许完成。
Checkpoint A 对已知的空间与视觉要求进行全面检查。如果它发现可信且人眼显著的失败,
必须修复、重新渲染,并从 Checkpoint A 重新开始。
Checkpoint A 通过后不得编辑。在一个新的独立 CallMLLM 调用中,把相同的 reference 和
最新 candidate render 交给 Checkpoint B。Checkpoint B 必须主动搜索 Checkpoint A
尚未覆盖的新视觉显著差异,并再次检查主要几何关系、裁切、重叠、内容和外观。
只有实际附加了 reference 与最新 render,且外部验证器返回带具体证据的明确通过,
checkpoint 才计数。`INSUFFICIENT` 不计为通过。Checkpoint B 的任何可信失败都要求
修复、重新渲染,并从 Checkpoint A 重新开始。
证据已经干净时,不得为了延长过程而制造修改或重复完全相同的问题。一旦同一个未改变的
最新 artifact 连续通过 Checkpoint A 和 B,应停止。
### 外部失败裁决
每次 Checkpoint A 或 B 请求都必须要求外部验证器把以下三个 token 之一准确放在响应
第一行:
`VERDICT: PASS`
`VERDICT: FAIL`
`VERDICT: INSUFFICIENT`
任何其他输出都按 `INSUFFICIENT` 处理,即使正文语气是正面的。Checkpoint verdict 是
外部证据;主模型不得只依靠自己的判断推翻非 PASS。
如果 `FAIL` 指出的缺陷得到当前 reference、render 或测量证据支持,直接修复。如果
`FAIL` 与当前具体反证或另一个独立外部观察存在实质冲突,不得盲目修改,也不得自行把它
清除。针对该争议主张发起一次新的外部 adjudication 调用:附加完整 reference、最新
render 和反证,明确指出被争议的主张,并要求同样的精确首行 verdict 合约。
Adjudication 结果决定该主张:`VERDICT: FAIL` 表示确认,必须修复;`VERDICT: PASS` 表示
清除;`VERDICT: INSUFFICIENT` 表示需要补充证据或缩窄外部问题。Adjudication 的 PASS
不能算作 Checkpoint A 或 B。任何非 PASS checkpoint 或任何 adjudication 之后,都必须
从 Checkpoint A 重新开始闭环。
只有最后两次全面 checkpoint 调用依次是 Checkpoint A 的 `VERDICT: PASS` 和 Checkpoint B
的 `VERDICT: PASS`,且二者针对同一个未修改的 artifact,中间没有编辑、adjudication 或
非 PASS checkpoint,才允许完成。
### 已确认失败的原子化修复
当 checkpoint 或 adjudication 同时确认多个对人眼显著的失败时,先按因果范围排序,
一次只修复一个因果相关的问题组。不要把彼此无关的已确认失败合并成一次大修改。
每个问题组修复后,重新渲染最新 candidate,并使用外部局部验证检查已修主张及其可能回归,
再开始下一个已知问题组。局部验证失败时继续修复并复验;通过后才进入下一个已确认问题组。
所有已确认主张清除后,从 Checkpoint A 重新开始全面闭环。
不要把一个连贯的父布局修正拆成装饰性的微小编辑;没有已确认失败时,不得制造修复阶段。
### 证据时效与完成条件
一次编辑会使 artifact 中所有受影响部分的旧验证证据失效。完成相关编辑后,必须重新渲染
artifact,并获得新的外部验证结果。
只有当最后一次 checkpoint 发生在最后一次重要修改之后,同时支持空间和视觉忠实度,并且
不存在尚未解决的视觉显著失败时,才可以声明完成。
</verification_policy>
```
# Verification V6 · Atomic Confirmed Repairs / 已确认失败原子化修复
Status: iteration-4 candidate; editable and not SHA-frozen
Module ID: `verification-V6-atomic-confirmed-repairs`
Language: English, proposed runtime text
Derived from: Verification V5 External Failure Adjudication; only atomic repair-and-recheck sequencing is added
```text
<verification_policy>
Verification must use `CallMLLM` as an external multimodal verifier. Do not treat
your own visual judgment as verification.
There are two levels of verification:
### Local validation
During implementation, use focused validation whenever a material visual
question, uncertainty, or recent change needs feedback.
Provide the external verifier with the evidence needed for the specific
question. This may include the reference image or a relevant reference region,
the latest candidate render, current HTML/CSS or DOM information, measurements,
correspondence or binding artifacts, and relevant outputs from earlier steps.
No fixed binding or intermediate artifact is required. Use whichever available
evidence helps the external verifier answer the current question.
### Checkpoint verification
Perform comprehensive verification after the first complete candidate render,
after a batch of material visual corrections, and before declaring the artifact
complete.
At each checkpoint, verify the latest artifact against the actual reference
image. Check spatial correspondence; text clipping and unintended overlap;
visible content and element presence; and typography, color, borders,
backgrounds, appearance, and structure.
Comprehensive verification may use one or more focused CallMLLM calls. Supply
the actual reference and latest candidate images as image inputs. Relevant code,
measurements, bindings, or previous artifacts may also be attached when they
help distinguish the remaining differences.
### Spatial and visual closure
Iterate on the HTML until spatial differences for material corresponding
elements are below 5 px in page coordinates, with no accidental text clipping
or unintended overlap.
A numeric spatial claim must be supported by current correspondence or
measurement evidence. Do not claim that the 5 px tolerance has been met from a
visual impression alone.
After spatial differences are within tolerance, perform visual verification.
Fix every visually significant failure before completion.
If a visual correction causes spatial drift, clipping, overlap, or another
material regression, recalibrate the layout and verify both spatial and visual
fidelity again.
If a visually significant difference persists, change the implementation
approach and verify the new result. Do not declare the difference unfixable
without trying a materially different approach.
### Two-checkpoint closure rule
Completion requires two consecutive externally verified checkpoints on the
same latest artifact.
Checkpoint A is a comprehensive review of the known spatial and visual
requirements. If it finds a credible, human-visible failure, fix it, render the
artifact again, and restart the sequence at Checkpoint A.
After Checkpoint A passes, make no edit. In a fresh independent CallMLLM call,
submit the same reference and latest candidate render for Checkpoint B.
Checkpoint B must actively search for new visually significant differences not
already covered by Checkpoint A and recheck the major geometry, clipping,
overlap, content, and appearance.
A checkpoint counts only when the actual reference and latest render were
attached and the external verifier returned a clear pass with concrete
evidence. `INSUFFICIENT` does not count. Any credible failure at Checkpoint B
requires a fix, a fresh render, and a restart at Checkpoint A.
Do not manufacture edits or repeat an identical question after the evidence is
clean. Once Checkpoints A and B pass consecutively on the unchanged latest
artifact, stop.
### External failure adjudication
Every Checkpoint A or B request must ask the external verifier to put exactly
one of these tokens on the first line of its response:
`VERDICT: PASS`
`VERDICT: FAIL`
`VERDICT: INSUFFICIENT`
Anything else counts as `INSUFFICIENT`, even if the prose sounds positive. The
checkpoint verdict is external evidence; you may not overrule a non-pass using
your own judgment alone.
Act directly on a `FAIL` when its cited defect is supported by the current
reference, render, or measurements. If a `FAIL` materially conflicts with
current concrete counter-evidence or another independent external observation,
do not edit blindly and do not clear it yourself. Make one fresh external
adjudication call for that disputed claim. Attach the full reference and latest
render plus the counter-evidence, identify the exact disputed claim, and require
the same exact first-line verdict contract.
The adjudication result controls that claim: `VERDICT: FAIL` confirms it and
requires a fix; `VERDICT: PASS` clears it; `VERDICT: INSUFFICIENT` requires
better evidence or a narrower external question. An adjudication pass is not
Checkpoint A or B. After a non-pass checkpoint or any adjudication, restart the
closure sequence at Checkpoint A.
Completion requires the final two comprehensive checkpoint calls to be
`VERDICT: PASS` for Checkpoint A and then `VERDICT: PASS` for Checkpoint B on the
same unchanged artifact, with no edit, adjudication, or non-pass checkpoint
between them.
### Atomic repair checkpoints
When a checkpoint or adjudication confirms multiple human-visible defects, rank
them by causal scope and repair one causally related cluster at a time. Do not
bundle unrelated confirmed failures into one large edit.
After each cluster is repaired, render the latest candidate and run focused
external validation of the repaired claim and its likely regressions before
starting the next known cluster. A failed focused check requires another repair
and recheck; a pass permits the next confirmed cluster. After all confirmed
claims are cleared, restart the comprehensive closure sequence at Checkpoint A.
Do not split one coherent parent-layout correction into cosmetic micro-edits,
and do not manufacture repair stages when no confirmed failure remains.
### Evidence freshness and completion
An edit invalidates earlier verification evidence for every affected part of
the artifact. After an affected edit, render the artifact again and obtain
fresh external verification.
Do not declare completion unless the latest checkpoint occurred after the
latest material edit and supports both spatial and visual fidelity, with no
unresolved visually significant failure.
</verification_policy>
```