Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

create-verification-skill

创建 verification skill

Every serious project needs a scripted way to drive the real app and prove behavior: launch it, exercise a feature the way a user would, and capture evidence. This skill generates that as a project-local skill (.cursor/skills/verify-<app>/) tailored to the repo. You write the generator’s output for the next agent, not for a human: it will be read cold, mid-task, by an agent that has never seen the app.

每个认真的项目都需要脚本化方式驱动真实应用并证明行为:启动它、像用户一样走功能、抓证据。本 skill 把它生成成贴合仓库的项目本地 skill(.cursor/skills/verify-<app>/)。你写生成器输出是给下一个 agent,不是给人:它会在任务中途被从未见过这应用的 agent 冷启动阅读。

1. Interview the repo, not the user

1. 问仓库,别问用户

Answer these from the codebase and only ask the user what you cannot observe:

从代码库回答这些问题,只有观察不到的才问用户:

  • Surface: what does a user actually touch? A web UI, a CLI/TUI, a desktop app, an API, a mobile app, a library? A repo can have several; pick the primary one and note the rest.

  • Run: how does the app start locally? Prefer the repo’s own documented dev command (package scripts, Makefile, README quickstart). Note ports, env vars, seed data, auth.

  • Drive: how can an agent interact with it programmatically? Existing harnesses first — Playwright/Cypress specs, expect scripts, PTY helpers, curl-able endpoints, a debug port. Only then pick a generic recipe: browser/CDP for web and Electron, a tmux/PTY harness for CLI/TUI, plain HTTP for services.

  • Observe: what evidence can be captured? Screenshots, terminal transcripts, response bodies, logs, exit codes, DB state.

  • Isolate: can two instances run side by side (ports, data dirs, profiles)? If not, say so in the generated skill: refusing to double-drive a shared instance beats corrupting the user’s session.

  • Surface: 用户实际碰什么?Web UI、CLI/TUI、桌面应用、API、移动应用、库?仓库可能有多个;选主的,记下其余。

  • Run: 应用本地怎么启动?优先仓库自己文档里的开发命令(package scripts、Makefile、README quickstart)。记下端口、环境变量、seed 数据、认证。

  • Drive: agent 怎么程序化交互?先已有 harness——Playwright/Cypress specs、expect 脚本、PTY helper、可 curl 的端点、debug 端口。再选通用配方:Web/Electron 用 browser/CDP,CLI/TUI 用 tmux/PTY harness,服务用纯 HTTP。

  • Observe: 能抓什么证据?截图、终端 transcript、响应体、日志、退出码、DB 状态。

  • Isolate: 两个实例能否并排跑(端口、数据目录、profile)?不能就在生成的 skill 里说清:拒绝双开共享实例,胜过弄坏用户会话。

If the checkout doesn’t build or start as-is, fix that first (or report it precisely) before generating; a skill written against a broken base teaches wrong steps. When an irrelevant missing asset blocks startup (a static dir the API never serves, a sample config), the generated skill may create it, clearly marked as verification scaffolding, and remove it in cleanup.

checkout 原样建不起来或启不来,先修(或精确报告)再生成;对着坏基底写的 skill 会教错步骤。无关缺失资源挡住启动时(API 从不服务的静态目录、示例配置),生成的 skill 可以创建它,明确标成 verification scaffolding,并在 cleanup 里删掉。

2. Generate the skill

2. 生成 skill

Write .cursor/skills/verify-<app>/SKILL.md with YAML frontmatter (name: verify-<app> and a description that names the app, the surface, and when to reach for it — without frontmatter the skill never registers) and these sections, each grounded in what the interview actually found (no placeholders left):

写 .cursor/skills/verify-<app>/SKILL.md,带 YAML frontmatter(name: verify-<app>,以及点名应用、surface、何时用的 description——没 frontmatter skill 不会注册),以及下列小节,每节都锚定 interview 实际发现(不留占位符):

  • Launch: the exact command that starts the app for verification, and how to tell it’s ready (a log line, a port answering, a prompt). Include teardown. For a short-lived CLI or TUI there is no server to keep alive: launch means build the binary (or install deps) once, then start each drive in its own isolated PTY or tmux session.

  • Doctor: one read-only check that answers “is this instance worth driving?” — process up, right version/build, port owned by us, auth valid. An agent runs this first whenever anything looks off.

  • Drive: the harness recipe with real selectors/commands from this repo, not examples. Prefer stable handles (ARIA labels, data attributes, prompt strings, route paths) over coordinates and tab order.

  • Evidence: what to capture for a proof and where it goes. State the proof standards: exercise the real user path, not internal setters or test-only endpoints; capture the action and the resulting state, not just the final screen; verify side effects (files written, rows inserted, messages sent) alongside what’s visible; mocks only where a production boundary already isolates the external system. When the safe path is a dry-run or test mode, verify what it actually skips by observing (files, network, git refs) rather than trusting its name: some dry-runs still touch the network or open a browser.

  • Cleanup: how to tear down instances the run created. Never kill by process name; kill what you started. Cleanup removes instances and scratch state, never the evidence: proof artifacts survive the teardown, in a location the skill names.

  • Helpers: any script the skill ships is executable and its invocation is shown in the skill body. A helper the reader has to reverse-engineer is not a helper.

  • Launch: 启动应用做验证的精确命令,以及如何判断就绪(日志行、端口应答、prompt)。含 teardown。短命 CLI/TUI 没有要保活的服务器:launch 表示构建二进制(或装依赖)一次,然后每次 drive 在自己隔离的 PTY 或 tmux 会话里启动。

  • Doctor: 一个只读检查,回答「这实例值得开吗?」——进程在、版本/构建对、端口是我们的、认证有效。看起来不对时 agent 先跑这个。

  • Drive: harness 配方,用本仓库真实 selector/命令,不是示例。优先稳定句柄(ARIA label、data 属性、prompt 字符串、路由路径),不要坐标和 Tab 顺序。

  • Evidence: 证明要抓什么、放哪。写明证明标准:走真实用户路径,不是内部 setter 或仅测试端点;抓动作与结果状态,不只最终画面;副作用(写文件、插行、发消息)与可见物一并验证;只有生产边界已隔离外部系统处才用 mock。安全路径是 dry-run 或 test mode 时,靠观察(文件、网络、git ref)验证它实际跳过了什么,别信名字:有些 dry-run 仍碰网络或开浏览器。

  • Cleanup: 如何拆掉本次创建的实例。绝不按进程名杀;只杀你启动的。Cleanup 去掉实例和临时状态,绝不去证据:证明产物在 teardown 后仍在,位置由 skill 点名。

  • Helpers: skill 附带的脚本可执行,调用写法写在 skill 正文。读者得逆向工程的 helper 不算 helper。

3. Seed the feature map

3. 播种 feature map

Create .cursor/skills/verify-<app>/features/README.md plus one file per user-facing feature you can identify (aim for the top 3-5 to start, from routes, commands, menus, or docs). Follow the shape in references/feature-map-example/, with a README index and one file per feature. Each file answers, from the user’s point of view: what the feature is, how to reach it, how to drive it with the harness, and what observable end state proves it works. The four H2s are Sub-features, How to get to it (user POV), Driving it with <harness>, and Gotchas. The map is the repo’s maintained verification source; a proof that drives one convenient entry point is incomplete when the map lists others.

创建 .cursor/skills/verify-<app>/features/README.md,加上你能识别的每个面向用户功能一个文件(起步瞄准前 3–5 个,来自路由、命令、菜单或文档)。形状跟 references/feature-map-example/:README 索引 + 每功能一文件。每文件从用户视角回答:功能是什么、怎么到达、怎么用 harness 驱动、什么可观察终态证明管用。四个 H2:Sub-features、How to get to it (user POV)、Driving it with <harness>、Gotchas。map 是仓库维护的验证源;map 列了其他入口时,只开一个方便入口的证明不完整。

4. Prove the generated skill before handing it over

4. 交出去前先证明生成的 skill

Run its own instructions end to end once: launch, doctor, drive ONE mapped feature (one is enough; the map exists so later runs can cover the rest), capture evidence, clean up. After cleanup, confirm the evidence still exists at the named location — a cleanup that eats the proof fails this step. Fix what fails, and run the generated cleanup after every failed iteration too, so broken attempts don’t strand processes and ports. A generated skill that was never executed is a draft, not a deliverable.

端到端跑一遍它自己的指示:launch、doctor、drive 一个已映射功能(一个够;map 存在是为了后续覆盖其余)、抓证据、cleanup。cleanup 后确认证据仍在点名位置——吃掉证明的 cleanup 本步失败。修失败的,每次失败迭代后也跑生成的 cleanup,免得坏尝试留下进程和端口。从未执行过的生成 skill 是草稿,不是交付物。

5. Offer the maintenance loop

5. 提供维护环

Point the user at /maintain-verification-skill for keeping the map honest as the app changes. Suggest a cadence only if they ask.

指向 /maintain-verification-skill,让 map 随应用变化保持诚实。只有用户问起才建议节奏。