principle-make-operations-idempotent
操作要幂等
Design operations so they converge to the correct state regardless of how many times they run or where they start from. Every state-mutating operation should answer: “What happens if this runs twice? What happens if the previous run crashed halfway?”
设计操作,使无论跑几次、从哪开始,都收敛到正确状态。每个改状态的操作都要能回答:「跑两次会怎样?上次半路崩溃会怎样?」
Why: Commands, lifecycle operations, and processing loops run where crashes, restarts, and retries are normal. If partial state changes the next run’s outcome, every restart becomes a debugging session.
为什么: 命令、生命周期操作、处理循环跑在崩溃/重启/重试是常态的地方。若部分状态会改下次结果,每次重启都变成调试会。
The pattern:
模式:
-
Convergent startup: scan for existing state, clean stale artifacts, adopt live sessions
-
Content-based cleanup: compare by content equivalence, not creation order
-
Self-healing locks: use PID-based stale lock detection
-
Idempotent scheduling: failed work respawns cleanly, fresh input regenerated after each cycle
-
收敛启动:扫描已有状态、清陈旧产物、接管活会话
-
按内容清理:比内容等价,不比创建顺序
-
自愈锁:用基于 PID 的陈旧锁检测
-
幂等调度:失败工作干净重生,每轮后重新生成新鲜输入
The test:
自检:
-
What happens if this runs twice in a row?
-
What happens if the previous run crashed at every possible point?
-
Does re-execution converge to the same end state?
-
连跑两次会怎样?
-
上次在每个可能点崩溃会怎样?
-
再执行是否收敛到同一终态?
If any answer is “it depends on what state was left behind,” the operation needs a reconciliation step.
任一答案是「取决于留下什么状态」,这操作就需要和解步骤。