Bind structured S2_00 handoff to the inline-prompt S2_10 Agent and reseal the parent-child release chain.
8.7 KiB
S2_10 Hybrid Prompt Implementation v1
Objective
Implement the S2_10 legal-resolution-map package defined by stage_2_optimal_update_strategy_v.5-1.md, S2_10_SOW_v.1.md, and S2_10_assets_v.1.md, while retaining the compatible legal/runtime boundaries from S2_10_SOW.md. Produce the versioned authoring YAML Stage_2_S2_10_v.1.yml, deploy its version-free projection and supporting assets under Default_Agent/Stage_2_Clean/, and leave a reproducible offline validation record.
Deliverables
YAML_Prompts/2. Stage_2/Stage_2_S2_10_v.1.yml- The fixed S2_10 physical asset set listed by
S2_10_assets_v.1.md, deployed at its canonical paths - Updated shared release/module manifests required to close the new hashes
- Byte-identical
.py/.txtpairs for every Python asset - Updated
YAML_Prompts/2. Stage_2/MEMORY.md
Scope and Non-Scope
In scope: hybrid inline-static prompt ownership, two-field dynamic-item boundary, schema/binding/workflow/projection/receipt/release updates, fixtures, offline tests, and release resealing. Out of scope: live GPT invocation, native adapter execution, production admission, legal sign-off, and dependencies on Stage 2 v0-v3 YAML specifications.
Known Inputs
stage_2_optimal_update_strategy_v.5-1.mdS2_10_SOW.mdand controlling hybrid deltaS2_10_SOW_v.1.mdS2_10_assets_v.1.mdSKILL.mdtest_code_executor.ipynb- Current
Default_Agent/Stage_2_Clean/S2_00/S2_10 assets
Material Assumptions
- The explicit requested output basename
Stage_2_S2_10_v.1.ymlis the authoring source; the deployment projection remains version-free. S2_10_SOW_v.1.mdcontrols where it narrows or replaces the olderS2_10_SOW.mdexternal-prompt design.- S2_10 contains exactly one direct LLM map task and no runtime Python or Code Executor task. Python files are offline build/validation/test assets only and must have byte-identical
.txtmirrors. - The 36-file fixed S2_10 inventory counts physical files; shared release manifests that require resealing are tracked separately.
Questions That Could Change the Outcome
- Whether the host platform's native LLM adapter supports the declared four-message order and structured-output contract remains a live-admission question, not an offline implementation blocker.
- Model benchmark and Korean-law review receipts remain pending until external evidence exists.
Workstreams and Dependencies
- Freeze the 36-file inventory and retained/replaced/missing ledger.
- In parallel, draft the hybrid authoring YAML and update independent assets.
- Build the version-free projection, then bind hashes into receipts and the child release.
- Reseal shared parent/module manifests without changing S2_00 business logic.
- Run one asset-level review/fix cycle and exactly two YAML review/fix cycles.
- Run offline validation, parity, negative-mutation, and forbidden-dependency checks.
- Record results and residual live-admission gaps in project memory.
Source and Tool Plan
- Use local repository documents and assets only; no web research is required.
- Use
apply_patchfor authored changes and mechanical copy only for.pyto.txtbyte mirrors. - Use repository Python validators/tests in offline mode. Do not invoke live models or adapters.
- Use independent subagents for inventory/design review and YAML validation; batch asset-level review assignments within the four-agent concurrency limit.
Validation Plan
- Parse all YAML/JSON and validate schemas and cross-file hashes.
- Assert one map task, zero deterministic/tool/reduce tasks, exact model/reasoning/verbosity, no tools, and
preflight: false. - Assert inline common/static prompts and exactly two permitted dynamic item placeholders.
- Reject legacy preassembled prompt fields and raw Stage 1 rereads.
- Verify local-reference output schema and adapter-owned permanent-ID rewrite boundary.
- Verify
.py/.txtbyte parity. - Run positive fixtures and required negative mutations.
- Distinguish offline contract completion from live adapter/model/legal admission.
Approval Boundaries
No external writes, live model calls, production admission, or commits are authorized. Existing unrelated worktree changes are preserved.
Progress
- Read repository rules, local YAML skill, notebook guidance, strategy/SOW/assets, and project memory.
- Freeze the exact inventory and delta ledger (
K=36). - Draft and deploy all S2_10 assets.
- Complete one asset validation and fix cycle.
- Complete YAML validation and fix cycle 1.
- Complete YAML validation and fix cycle 2.
- Reseal shared manifests and run full offline validation.
- Update project memory and record residual uncertainty.
Decision Log
| Date/Stage | Decision | Basis | Consequence |
|---|---|---|---|
| 2026-08-30 / intake | Treat S2_10_SOW_v.1.md as the controlling hybrid delta over S2_10_SOW.md. |
It was produced specifically to implement stage_2_optimal_update_strategy_v.5-1.md; the old SOW still assumes externally assembled full prompts. |
Retain compatible legal/output/adapter contracts, replace prompt ownership and item boundaries. |
| 2026-08-30 / intake | Use Stage_2_S2_10_v.1.yml as source and version-free deployment projection. |
Explicit user filename plus existing deployment convention. | Builder/receipts must bind both source and projection identities without overwriting the old versioned source. |
| 2026-08-30 / intake | No runtime Python in S2_10. | Strategy and SOW assign the single legal judgment to one native LLM map task. | Notebook Code Executor rules apply only as safety constraints to offline Python assets; no inline run_code belongs in the S2_10 YAML. |
| 2026-08-30 / implementation | Exclude cluster_case_payload_sha256 from the payload's inner echo and bind it only in the top-level item/output echo. |
A payload containing its own byte hash would require an impossible fixed point. | The inner payload echo has 19 fields; the top-level dispatch/model echo has all 20 fields. |
| 2026-08-30 / asset review | Replace the regression manifest's generic oracle marker with fixture-specific exact oracle IDs and recompute manifest/module/release closure. | The one requested asset-review round found three affected physical assets plus stale module/test counts. | Tests now compare exact per-fixture oracle lists; module counts are builder-derived (76, 12). |
Evidence Ledger
| Claim/Issue | Source or Test | Status | Notes |
|---|---|---|---|
| Hybrid prompt ownership | strategy v5-1 and SOW v1 | accepted | YAML owns stable common/task prompt; item owns case data only. |
| Physical asset count | assets v1 inventory and filesystem audit | verified | K=36; all 36 paths exist at the specified locations. |
| S2_10 offline tests | tests/s2_10 |
PASS | 39/39, including exact fixture oracle seal and .py/.txt parity. |
| S2_00 compatibility | tests/s2_00, both builders in check mode |
PASS | 69/69; parent hash and opaque S2_10 Agent/binding handoff resealed. |
| Live adapter compatibility | external adapter evidence | pending | Must not be inferred from offline tests. |
| Model quality and legal correctness | benchmark/legal review receipts | pending | Receipts remain explicitly non-admitted. |
Risks and Failure Modes
- Hash cycles between parent and child releases; prevent by keeping the parent one-way and not reverse-locking the child from the parent.
- YAML parser normalization changing prompt bytes; use explicit block scalars and hash the parsed canonical strings where specified.
- Dynamic item reintroducing full prompt text or PII; enforce schema and negative fixtures.
- Static tests being mistaken for live execution or legal approval; retain explicit status labels and pending receipts.
- Shared dirty worktree causing accidental scope expansion; edit only identified Stage 2 assets.
Results and Residual Uncertainty
The 36-file hybrid S2_10 package is implemented and offline verified. The authoring and deployment Agent are byte-identical and contain exactly one gpt-5.6-sol/xhigh/medium/responses map task, four ordered prompt blocks, and two dynamic JSON placeholders. Parent release raw hash is 8dcfe060af33e2d86a93f1c277c485b67255500a8cfb7f3024ffc427e802b948; child release raw hash is 5adb88482fc6e54ae7475570facf9a6ae87fa613fce65d867d16d5256604f14d. Both builders pass read-only check mode, S2_10 tests pass 39/39, S2_00 compatibility tests pass 69/69, and all seven Python assets have byte-identical .txt mirrors.
This is not live or legal admission. The platform adapter, model benchmark, official authority/profile release, and Korean-lawyer review receipts remain explicitly pending; the parent remains DRAFT_NOT_EXECUTABLE and production execution remains blocked. No live LLM/MCP adapter or end-to-end matter run was performed.