# mp-plan-reviewer rounds (dispatch-critic) — intent-to-completion

# Plan review — round 1 (dispatch-critic, task kgsyqpjex, 2026-09-03, plan_hash sha256:823c4cf8…)
- verdict: FAIL
- coverage: 13/15 acceptance criteria  (uncovered: §11 `overlap-sequencer`; §10.1/§11 live v9-to-v10 bootstrap stage)
- goal coverage: 4/6 goals served  (unserved: G1 (no required overlap-sequencer evidence); G6 (bootstrap machinery is not scheduled to execute before finish))
- findings:
  - [coverage] §11 / tasks 3, 6, 33, 34 — Task 33 names `test/overlap-sequencer.test.mjs` in an integration note, but no task owns that file or the full conflict/resume/abort/stale-review/locking matrix — fix: add a dedicated integration-test task depending on 3, 7, 30, 31, 33, and 34 that owns and runs the named suite.
  - [coverage] §10.1 / tasks 7, 43–47 — `plan.md` ends after task 29 with no required non-task bootstrap marker, so changing the future sequencer and implementing bootstrap tooling does not cause this v9-run plan to execute the live arm-command-record stage before `mp finish` — fix: render an explicit post-wave bootstrap marker after Wave 8 that drives `status → arm → command → record` through completion before finish.
  - [consistency] tasks 1, 7, 27, 32, 39 — `required_successor` has a schema and several consumers, but no task explicitly records it for `v10-validation`; intent rejection likewise stores correction text without guaranteeing the required-successor event — fix: make task 7 record the current run's `v10-validation` obligation before finish and task 27 emit the linked obligation for both rejection classes.
  - [consistency] tasks 3, 10, 13 — Task 3's generic "goals-load options" wording does not honor the required durable `--interview-waived` routing or guarantee replay/spec-path forwarding and preservation of `assumed_row_missing` — fix: state those exact calls and error behavior in task 3 and add a dependency from task 3 to task 13.
  - [consistency] tasks 3, 16, 20 — Task 3 does not explicitly emit `config show.harness.autoCompactWindow` or export `KNOWN_FLAGS` and `isKnownFlag`, leaving consumers required by tasks 16 and 20 without guaranteed producers — fix: add the exact output field and exports to task 3 and assert them in tasks 6 and 20.
  - [consistency] tasks 7, 9, 14 — The sequencer task says only "quoted bundle data" and "assessment"; it does not pin the critic payload to the verbatim anchor, ledger, and draft or pin assessor calls to implementation/final modes with their distinct evidence tuples — fix: add the exact dispatch payloads, mode names, and binding fields to task 7 and cover them in prompt-contract tests.
  - [consistency] tasks 3, 7, 27 — Task 27 produces intent-correction text for a successor, but neither CLI seeding nor the sequencer is explicitly assigned to load that predecessor event into the successor interview context — fix: assign predecessor-context projection to task 3 and consumption at interview start to task 7, with a black-box case in `test/intent-rejected.test.mjs`.
  - [decomposition] tasks 3, 7, 13, 14, 26–28 — Integration consumers precede interfaces they must consume: task 3 is before task 13, while task 7 is before the final assessor and final-check/archive-push implementations — fix: add dependencies `3 → 13` and `7 → 13,14,26,27,28`.
  - [verify] tasks 1–5, 7, 8, 43 — Verifiers name paths absent from their own or any earlier owning task, including `test/bundle.test.mjs`, `test/bin-masterplan.test.mjs`, `test/*.test.mjs`, both prompt tests, `test/docs-contract.test.mjs`, and the current bundle state; task 2's CLI test is owned only by later task 6 — fix: list focused verifier files with their implementation tasks or declared dependencies, replace wildcard runs with owned suites, and use a temporary fixture state for task 43.
  - [verify] task 1 — A pre-existing bundle test plus a token grep does not prove the new event schemas, rejection behavior, exports, legacy tolerance, or post-archive restrictions — fix: make task 1 own focused bundle tests covering valid and invalid instances of every new event and all completion compatibility cases.
  - [verify] task 3 — Pre-existing CLI tests and an undirected full-suite run cannot prove the many new wiring paths, especially because task 6 adds the relevant CLI assertions only later — fix: add task-3-owned black-box cases for every new verb/flag, waiver forwarding, final-receipt binding, config output, lock ordering, and stored plan-path default.
  - [verify] tasks 4, 5 — Greps plus the full suite do not specifically prove plan-path precedence, planning-mode resume routing, bootstrap-marker exclusion, v2 outcome rendering, or v1 fallback — fix: add focused continue/resume/plan-merge and wave-summary behavioral tests to these tasks.
  - [verify] tasks 9, 14 — Bare greps prove only that selected words occur in the prompts, not the critic schema, quoted-data boundary, two assessor modes, evidence rules, or v1 output shape — fix: add structured prompt-contract tests that parse and assert the complete protocols.
  - [verify] task 41 — The positive idempotent-release smoke test does not exercise invalid versions, unbumped files, foreign tags, or the CHANGELOG-only commit restriction — fix: add isolated negative and file-diff cases for each refusal and mutation contract.
  - [verify] tasks 43, 44 — Syntax checks and broad greps do not prove bootstrap ordering, arm/record bindings, pre/postconditions, pinned-v9 use, fixture substitutions, cleanup, or recovery; task 43's status grep is also tautological and depends on the MAIN-owned live bundle — fix: add worktree-local fixture smoke tests for the driver and rehearsal script, leaving only the credentialed GitHub cycle for the live stage.
- note: Counts use §11's 15 top-level acceptance bullets, with §10's live-stage requirement folded into the bootstrap bullet; all cited task ids and spec refs were confirmed.

## Dispositions applied (apply-review-r1.mjs → plan_hash sha256:191f5caa…, 48 tasks / 11 waves)
All findings folded as fragment edits except: bootstrap stage stays a sequencer-level pre-finish stage (spec §10.1, A33) — stated in plan.index.json meta.solution with G6 goal-check enforcement; rehearsal verify stays structural (fixture exercise owned by the dependent bootstrap suite); assessor-prompt structured tests owned by intent-goals.agents-compat.

# Plan review — round 2 (dispatch-critic, task kk4kd1hs0, 2026-09-03, plan_hash sha256:191f5caa…)
- verdict: FAIL
- coverage: 12/15 acceptance criteria  (uncovered: §11 overlap-sequencer semantic/action matrix; §11 deploy-commit-identity legacy output in runs list/status; §11 resume-brief obligation output in runs list/status)
- goal coverage: 5/6 goals served  (unserved: G6)
- findings:
  - [coverage] G6; tasks 7, 42, 44–48; spec.md §10.1 — These tasks implement or fixture-test bootstrap machinery, but none performs the live release/tag/push/CI/install walk; meta.solution is not the required rendered post-wave marker, so the evidence G6 names is produced only nominally — fix: add a non-dispatchable, rendered post-wave stage consumed by the current orchestrator and require successful `bootstrap status` before finish.
  - [coverage] task 9; spec.md §8 and §11 — The planned overlap suite omits the required in-progress, plan-path, archived, no-overlap, predecessor-link, resumed-recorder, abort, and gated-versus-loose action cases — fix: add those black-box cases to task 9 and assert their durable events or absence.
  - [coverage] tasks 1, 3, 24, 30, 39; spec.md §7.4 and §11 — Legacy classification is tested in bundle, retro, and doctor layers, but no task explicitly proves that `mp runs list` and `mp status` report an archive with no completion field as `legacy` — fix: add both CLI assertions to task 3 or its dependent CLI regression task.
  - [consistency] tasks 3, 33, 40; spec.md §7.4 and §11 — Task 33 requires its obligation projection to be reused by runs list and status, but task 3 only says "enriched runs/status" and owns no explicit obligation-consumer assertions — fix: name the shared projection call in task 3 and test both commands with the exact successor seed command.
  - [coverage] task 3; spec.md §9 — The plan only requires `mp config show` to print `harness.autoCompactWindow`; no task implements or tests the required recommendation being no lower than the run threshold or proves settings remain untouched — fix: extend task 3's config-show contract and black-box tests for recommendation and read-only behavior.
  - [consistency] tasks 17, 21; spec.md §4.4 and §10.1.1 — No task owns `lib/paths.mjs`, although it honors `CLAUDE_CONFIG_DIR` and the inventory will reject environment access outside `readEnv` — fix: assign `lib/paths.mjs` to a task depending on task 17 and route that environment lookup through `readEnv`.
  - [consistency] task 21; spec.md §10.1 — plan.index.json meta.solution says waves 0–9 implement the spec even though task 21 is in wave 10, making the bootstrap-stage boundary internally contradictory — fix: change the narrative to waves 0–10 and explicitly place the stage after task 21.
  - [verify] task 15 — Its greps do not prove the separate modes, mode-specific evidence rules, or v1 output schema claimed by the task; task 16's later verification does not make task 15's own verify_commands adequate — fix: move or duplicate the structured prompt assertions into task 15's verification.
  - [verify] task 45 — Syntax, executable-bit, and bare-grep checks cannot prove the rehearsal walk's fixture routing, pinned-v9 execution, cleanup, audits, or arm/record behavior even though fixture-backed behavioral testing is possible — fix: give task 45 a focused fixture test or merge its implementation verification into the dependent bootstrap-suite task.
- note: none

## Dispositions applied (apply-review-r2.mjs → plan_hash sha256:73a247e3…, 48 tasks / 11 waves)
Findings 2–9 folded as fragment edits (overlap matrix; legacy + obligation output in CLI task and its regression task; config-show recommendation/read-only contract; new config-knobs.env-reads task owning lib/paths.mjs and every other direct env read, owners of lib/continue.mjs, lib/sweep.mjs, lib/runs.mjs, bin/install-pi.mjs convert their own and depend on config-knobs.core; narrative waves corrected; agents-compat merged into assessor-prompts; rehearsal task owns a gh-shim fixture test). Finding 1 held: the live bootstrap stage is sequencer-level per spec §10.1 / A33; narrative states the rule (finish only after bootstrap status reports the pass complete) and the G6 goal-check enforcement.

# Plan review — round 3 (dispatch-critic, task ks5jhfiu5, 2026-09-03, plan_hash sha256:73a247e3…)
- verdict: REVISE
- coverage: 15/15 acceptance criteria  (uncovered: none)
- goal coverage: 6/6 goals served  (unserved: none)
- findings:
  - [consistency] tasks 4, 7 / spec.md §10.1 — The current-run bootstrap schedule exists only in index meta: plan.md ends at task 20, and no current continue/resume consumer is assigned to surface that meta before `mp finish`, leaving the required non-task post-wave handoff implicit — fix: make task 4 consume the indexed post-wave stage or render an explicit non-dispatchable marker after wave 10, while task 7 owns its execution protocol.
  - [verify] tasks 5, 20, 34 — Their verify commands run `test/goals.test.mjs`, `test/knob-contract.test.mjs`, and `test/resume-brief.test.mjs` respectively, although those paths belong to tasks 13, 19, and 33 and the index exposes no declared dependencies permitting them — fix: declare dependencies 5→13, 20→19, and 34→33, or remove the cross-owned verification commands.
  - [verify] task 45 / spec.md §10.1, §11 — The credentialed GitHub rehearsal path is explicitly skipped by the shim test, leaving bare greps as the only proof of PR creation/merge and no behavioral proof of private-repository creation, cleanup, or failure cleanup — fix: use a stateful `gh` shim to execute and assert the full create/PR/merge/delete command sequence while reserving real network execution for the live stage.
  - [decomposition] tasks 3, 6 — Task 3 already claims black-box coverage for every new verb, flag, overlap/config behavior, completion rendering, and stored-plan default in `test/cli-surface.test.mjs`; task 6 repeats that ownership as a test-only sliver — fix: narrow task 3 to implementation-facing smoke cases and leave the exhaustive matrix to task 6, or merge task 6 into task 3.
- note: Coverage counts the 15 top-level §11 acceptance bullets; all four artifacts were complete and all cited task ids and spec references were confirmed.

## Dispositions applied (apply-review-r3.mjs; no fourth critic round — coverage complete, edits mechanical; the plan gate's cross-vendor review sees the final bytes)
- Finding 1: task 4 (planning-consumers) renders index.meta narrative in plan.md (v10); for THIS run the handoff rule lives in index.meta.solution + WORKLOG.md (the 9.10.0 renderer cannot print it); task 7 owns the execution protocol.
- Finding 2: no change — deps 5→13 (intent-goals.goals-v2), 20→19 (config-knobs.contracts), 34→33 (run-awareness.resume-brief) ARE declared in .plan-fragments.json; the merged index drops dep edges, so the reviewer could not see them (a reviewer-input gap, noted for the v10 planning-consumers task as a candidate: carry deps into the index).
- Finding 3: task 45's test uses a stateful gh shim asserting the full create/PR/merge/delete sequence, private flag, and cleanup on success and simulated failure.
- Finding 4: task 3 narrowed to implementation-facing smoke cases; task 6 keeps the exhaustive matrix.

# Plan gate (§3b) — cross-vendor adversarial panel 1 (adversarial-review workflow wf_eca58e83-c62, intensity standard, 2026-09-03)
- Artifacts: spec.md rev 14 + plan.md + plan.index.json at gate hash sha256:8d8e8b3d… (plan_hash 3d019b74…), 256705 chars embedded verbatim (sha256 78e84566…).
- Lanes: openai/gpt-5.6-sol → revise; zhipu/glm-5.2 → reject; three in-repo lenses (approach, failure-case, blast-radius) → revise; synthesis (in-tree verified) → revise. 1,276,626 tokens.
- Verdict: revise — 1 blocker, 8 should-fix, 3 nits; 7 reviewer claims dropped by synthesis (re-litigating spec decisions or misreading the plan). Full record: gate-plan-panel-1.json; digest: gate-plan-panel-1-notes.txt.
- Dispositions (apply-review-r4.mjs → plan_hash 9818b4e0…):
  B1 driver ownership: scripts/bootstrap-v10.mjs co-owned by the rehearsal and all three bootstrap suites; lib/config.mjs co-owned by legacy-migration and env-reads (serialised); bin/masterplan.mjs co-owned by cli-contract and overlap-contract (plus lib/seed-lock.mjs, lib/runs.mjs); test/retro-goals.test.mjs owned and run by retro-reporting.
  S2 task 2 rewritten to the real defect (die() exit 2 collides with the unknown-flag contract; distinct exit code for uncaught errors).
  S3 bootstrap marker: durable `pre_finish_stage_required` event recorded on this bundle (v9 `mp event`), named in the plan narrative and in task 1's generic-event tolerance; renderer change stays in the planning-consumers task.
  S4 push_archive producer: finish-step detects the push (origin/<base> moved to the receipt's base_after) — assigned to the archive-push task with both test rows.
  S5 critic ledger: `mp interview replay` (verbatim ledger) added to the interview library, CLI, and sequencer contracts; prompt-structure test asserts Q/A text.
  S6 rehearsal drives every §10.3 failure row against fixtures. S7 retro-goals test owned. S8 CLAUDE_CODE_SESSION_ID registered (owner-lock observable) and documented. S9 `pushed: no` asserted in the CLI regression task.
  N2 grep alternations replaced by per-term loops (tasks 1, 8). N3 doctor tasks scoped fixture-only; recovery suite re-runs the checks against writer-produced bundles. N1 (label vs task id) not applied: fragment keys are the stable references; ids are assigned at merge.
- Panel 2 (confirmation pass over the revised bytes) runs at light intensity: one cross-vendor lane + one in-repo lens.

# Plan gate — panel 2 (light confirmation pass, adversarial-review workflow wf_d96d8b7e-9fc, 2026-09-03)
- Bytes: plan_hash 9818b4e0… (264512 chars, sha256 f19811e8…). Lane: openai/gpt-5.6-sol → revise; in-repo approach lens → revise; synthesis (in-tree verified) → revise with NO blocker: 2 should-fix, 4 nits; 3 reviewer claims dropped (successor run as a task; pre-step remote snapshot; location errors). 584,005 tokens. Record: gate-plan-panel-2.json / gate-plan-panel-2-notes.txt.
- Dispositions (apply-review-r5.mjs): SF1 pre-finish rule split into (a) through surfaces_live before finish + required_successor recorded, (b) gate step inside branch_finish before merge — narrative, sequencer task, protocol-tests assertion, and a superseding v2 `pre_finish_stage_required` event. SF2 rejection interface gains a required `--successor=<slug>` (CLI task, completion task, doctor matcher note). Nits: pushed_base producer runs on the check-only recovery record (+ finish-replay row); interview flag list copied verbatim into the interview and CLI tasks; task 3 smoke slice and driver step-shape smoke cases added. Soft-enforcement nit: no plan change (spec §10.1/A33/A35).
- Panel 3 = light pass over the final bytes; recorded as the gate review if no blocker.

# §3c alignment audit (dispatch-critic, 2026-09-03, plan_hash c19d76ac…): anchor verbatim; 18/18 covered (A1–A12, G1–G6), 0 narrowed / dropped / contradicted / widened. Digest: alignment-audit.txt. Clauses confirmed under the operator's /goal directive (alignment-clauses.json).
