feat(stage2): integrate hybrid S2_10

Bind structured S2_00 handoff to the inline-prompt S2_10 Agent and reseal the parent-child release chain.
This commit is contained in:
2026-08-31 01:11:40 +09:00
parent 1b7d2d3a09
commit 978e6a6b72
73 changed files with 28204 additions and 6598 deletions
@@ -0,0 +1,58 @@
# S2_00 hybrid handoff v2 implementation plan
## Purpose
Revise the S2_00 specification, asset ledger, executable Agent YAML, and deployed S2_00 assets so that C15 hands S2_10 structured legal context and opaque Agent/binding digests instead of materialized P00/P10 prompt bytes or downstream model settings. Preserve the existing deterministic ingress, conservation, clustering, DAG, slice, atomic-publish, and direct MCP Code Executor architecture.
## Scope
- Create `YAML_Prompts/2. Stage_2/S2_00_SOW_v.2.md`.
- Create `YAML_Prompts/2. Stage_2/S2_00_assets_v.2.md`.
- Create `YAML_Prompts/2. Stage_2/Stage_2_S2_00_v.2.yml`.
- Update affected deployment assets under `YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/` and preserve `.py`/`.txt` byte parity.
- Append a concise completion record to `YAML_Prompts/2. Stage_2/MEMORY.md`.
- Do not depend on Stage 2 v0-v3 YAML specifications.
## Contract decisions
- S2_00 remains one deterministic `code-executor.run_code` task with full inline Python and direct localdocs MCP calls.
- C00, C05, C10, Stage 1 conservation, cluster DAG, cluster slices, output root, and atomic publication semantics remain unchanged.
- C15 no longer materializes or hashes P00/P10 prompt bytes and no longer owns S2_10 model/effort/verbosity.
- Each bundle cohort carries release-sealed `s2_10_agent_sha256`, `s2_10_llm_binding_sha256`, ordered selected-context refs/hash, cluster-slice refs/hash, cohort membership, and an embedded closed `context_materialization_receipt`.
- Deprecated `static_prefix_refs`, `expected_prefix_sha256`, materialized static-prefix fields, and downstream model settings are removed from the new schema version rather than silently coexisting.
- Offline/static verification, downstream S2_10 hybrid admission, live adapter/model execution, and legal sign-off remain distinct statuses.
## Work breakdown and dependencies
1. Inspect the current SOW/assets/YAML and deployed runtime/schema/workflow/cache/binding/fixtures/manifests/builders/tests.
2. Produce SOW v2 and assets v2 from the v5-1 contract; independently review each once and apply one correction pass.
3. Implement the new bundle-plan contract in the authoring YAML and affected deployment assets. Independent schema/workflow/policy/test edits may proceed in parallel; generated runtime mirrors and receipts follow the authoring YAML; release manifests are resealed last.
4. Review every changed logical asset once in maximally parallel waves, then apply one correction pass.
5. Run YAML/JSON parsing, Python compile, `.py`/`.txt` parity, projection checks, targeted S2_00 tests, full S2_00 offline validation, stale-field scans, hash/receipt consistency checks, and `git diff --check`.
6. Update project MEMORY with exact artifact paths, decisions, test evidence, and remaining live/downstream admission boundary.
## Validation acceptance criteria
- New authoring and deployment Agent YAML parse and contain exactly one deterministic Code Executor task with no LLM task.
- Inline Python equals deployed `runtime/s2_00_ingress.py` and `.txt` mirror byte-for-byte.
- New bundle-plan schema and runtime agree on all required and forbidden fields.
- S2_00 tests cover hybrid handoff success and reject legacy static-prefix/model-owner fields.
- Builders/receipts/manifests contain current hashes after all edits.
- No affected S2_00 artifact references Stage 2 v0-v3 YAML.
- Results are reported as offline/static only unless live evidence actually exists.
## Progress
- [x] Read root rules, Stage 2 MEMORY, local YAML skill, Python workflow guidance, and MCP Code Executor notebook.
- [x] Identify the v5-1 hybrid handoff contract and current legacy C15 fields.
- [x] Draft and review SOW v2.
- [x] Draft and review assets v2.
- [x] Implement and review authoring YAML v2.
- [x] Update and review deployed assets.
- [x] Reseal hashes/manifests and run validation.
- [x] Append project MEMORY and hand off results.
## Discoveries and decision log
- The worktree contains unrelated user edits and untracked files. Preserve them and scope all changes to the requested S2_00 v2 package plus this plan and MEMORY.
- Existing S2_10 deployment artifacts are a separate workstream. S2_00 v2 may implement the hybrid handoff contract, but live/downstream admission must remain pending until a matching S2_10 hybrid release is built and sealed.
@@ -0,0 +1,116 @@
# S2_10 Hybrid Prompt Implementation v1
## Objective
Implement the S2_10 legal-resolution-map package defined by `stage_2_optimal_update_strategy_v.5-1.md`, `S2_10_SOW_v.1.md`, and `S2_10_assets_v.1.md`, while retaining the compatible legal/runtime boundaries from `S2_10_SOW.md`. Produce the versioned authoring YAML `Stage_2_S2_10_v.1.yml`, deploy its version-free projection and supporting assets under `Default_Agent/Stage_2_Clean/`, and leave a reproducible offline validation record.
## Deliverables
- `YAML_Prompts/2. Stage_2/Stage_2_S2_10_v.1.yml`
- The fixed S2_10 physical asset set listed by `S2_10_assets_v.1.md`, deployed at its canonical paths
- Updated shared release/module manifests required to close the new hashes
- Byte-identical `.py`/`.txt` pairs for every Python asset
- Updated `YAML_Prompts/2. Stage_2/MEMORY.md`
## Scope and Non-Scope
In scope: hybrid inline-static prompt ownership, two-field dynamic-item boundary, schema/binding/workflow/projection/receipt/release updates, fixtures, offline tests, and release resealing. Out of scope: live GPT invocation, native adapter execution, production admission, legal sign-off, and dependencies on Stage 2 v0-v3 YAML specifications.
## Known Inputs
- `stage_2_optimal_update_strategy_v.5-1.md`
- `S2_10_SOW.md` and controlling hybrid delta `S2_10_SOW_v.1.md`
- `S2_10_assets_v.1.md`
- `SKILL.md`
- `test_code_executor.ipynb`
- Current `Default_Agent/Stage_2_Clean/` S2_00/S2_10 assets
## Material Assumptions
- The explicit requested output basename `Stage_2_S2_10_v.1.yml` is the authoring source; the deployment projection remains version-free.
- `S2_10_SOW_v.1.md` controls where it narrows or replaces the older `S2_10_SOW.md` external-prompt design.
- S2_10 contains exactly one direct LLM map task and no runtime Python or Code Executor task. Python files are offline build/validation/test assets only and must have byte-identical `.txt` mirrors.
- The 36-file fixed S2_10 inventory counts physical files; shared release manifests that require resealing are tracked separately.
## Questions That Could Change the Outcome
- Whether the host platform's native LLM adapter supports the declared four-message order and structured-output contract remains a live-admission question, not an offline implementation blocker.
- Model benchmark and Korean-law review receipts remain pending until external evidence exists.
## Workstreams and Dependencies
1. Freeze the 36-file inventory and retained/replaced/missing ledger.
2. In parallel, draft the hybrid authoring YAML and update independent assets.
3. Build the version-free projection, then bind hashes into receipts and the child release.
4. Reseal shared parent/module manifests without changing S2_00 business logic.
5. Run one asset-level review/fix cycle and exactly two YAML review/fix cycles.
6. Run offline validation, parity, negative-mutation, and forbidden-dependency checks.
7. Record results and residual live-admission gaps in project memory.
## Source and Tool Plan
- Use local repository documents and assets only; no web research is required.
- Use `apply_patch` for authored changes and mechanical copy only for `.py` to `.txt` byte mirrors.
- Use repository Python validators/tests in offline mode. Do not invoke live models or adapters.
- Use independent subagents for inventory/design review and YAML validation; batch asset-level review assignments within the four-agent concurrency limit.
## Validation Plan
- Parse all YAML/JSON and validate schemas and cross-file hashes.
- Assert one map task, zero deterministic/tool/reduce tasks, exact model/reasoning/verbosity, no tools, and `preflight: false`.
- Assert inline common/static prompts and exactly two permitted dynamic item placeholders.
- Reject legacy preassembled prompt fields and raw Stage 1 rereads.
- Verify local-reference output schema and adapter-owned permanent-ID rewrite boundary.
- Verify `.py`/`.txt` byte parity.
- Run positive fixtures and required negative mutations.
- Distinguish offline contract completion from live adapter/model/legal admission.
## Approval Boundaries
No external writes, live model calls, production admission, or commits are authorized. Existing unrelated worktree changes are preserved.
## Progress
- [x] Read repository rules, local YAML skill, notebook guidance, strategy/SOW/assets, and project memory.
- [x] Freeze the exact inventory and delta ledger (`K=36`).
- [x] Draft and deploy all S2_10 assets.
- [x] Complete one asset validation and fix cycle.
- [x] Complete YAML validation and fix cycle 1.
- [x] Complete YAML validation and fix cycle 2.
- [x] Reseal shared manifests and run full offline validation.
- [x] Update project memory and record residual uncertainty.
## Decision Log
| Date/Stage | Decision | Basis | Consequence |
|---|---|---|---|
| 2026-08-30 / intake | Treat `S2_10_SOW_v.1.md` as the controlling hybrid delta over `S2_10_SOW.md`. | It was produced specifically to implement `stage_2_optimal_update_strategy_v.5-1.md`; the old SOW still assumes externally assembled full prompts. | Retain compatible legal/output/adapter contracts, replace prompt ownership and item boundaries. |
| 2026-08-30 / intake | Use `Stage_2_S2_10_v.1.yml` as source and version-free deployment projection. | Explicit user filename plus existing deployment convention. | Builder/receipts must bind both source and projection identities without overwriting the old versioned source. |
| 2026-08-30 / intake | No runtime Python in S2_10. | Strategy and SOW assign the single legal judgment to one native LLM map task. | Notebook Code Executor rules apply only as safety constraints to offline Python assets; no inline `run_code` belongs in the S2_10 YAML. |
| 2026-08-30 / implementation | Exclude `cluster_case_payload_sha256` from the payload's inner echo and bind it only in the top-level item/output echo. | A payload containing its own byte hash would require an impossible fixed point. | The inner payload echo has 19 fields; the top-level dispatch/model echo has all 20 fields. |
| 2026-08-30 / asset review | Replace the regression manifest's generic oracle marker with fixture-specific exact oracle IDs and recompute manifest/module/release closure. | The one requested asset-review round found three affected physical assets plus stale module/test counts. | Tests now compare exact per-fixture oracle lists; module counts are builder-derived (`76`, `12`). |
## Evidence Ledger
| Claim/Issue | Source or Test | Status | Notes |
|---|---|---|---|
| Hybrid prompt ownership | strategy v5-1 and SOW v1 | accepted | YAML owns stable common/task prompt; item owns case data only. |
| Physical asset count | assets v1 inventory and filesystem audit | verified | K=36; all 36 paths exist at the specified locations. |
| S2_10 offline tests | `tests/s2_10` | PASS | 39/39, including exact fixture oracle seal and `.py/.txt` parity. |
| S2_00 compatibility | `tests/s2_00`, both builders in check mode | PASS | 69/69; parent hash and opaque S2_10 Agent/binding handoff resealed. |
| Live adapter compatibility | external adapter evidence | pending | Must not be inferred from offline tests. |
| Model quality and legal correctness | benchmark/legal review receipts | pending | Receipts remain explicitly non-admitted. |
## Risks and Failure Modes
- Hash cycles between parent and child releases; prevent by keeping the parent one-way and not reverse-locking the child from the parent.
- YAML parser normalization changing prompt bytes; use explicit block scalars and hash the parsed canonical strings where specified.
- Dynamic item reintroducing full prompt text or PII; enforce schema and negative fixtures.
- Static tests being mistaken for live execution or legal approval; retain explicit status labels and pending receipts.
- Shared dirty worktree causing accidental scope expansion; edit only identified Stage 2 assets.
## Results and Residual Uncertainty
The 36-file hybrid S2_10 package is implemented and offline verified. The authoring and deployment Agent are byte-identical and contain exactly one `gpt-5.6-sol/xhigh/medium/responses` map task, four ordered prompt blocks, and two dynamic JSON placeholders. Parent release raw hash is `8dcfe060af33e2d86a93f1c277c485b67255500a8cfb7f3024ffc427e802b948`; child release raw hash is `5adb88482fc6e54ae7475570facf9a6ae87fa613fce65d867d16d5256604f14d`. Both builders pass read-only check mode, S2_10 tests pass 39/39, S2_00 compatibility tests pass 69/69, and all seven Python assets have byte-identical `.txt` mirrors.
This is not live or legal admission. The platform adapter, model benchmark, official authority/profile release, and Korean-lawyer review receipts remain explicitly pending; the parent remains `DRAFT_NOT_EXECUTABLE` and production execution remains blocked. No live LLM/MCP adapter or end-to-end matter run was performed.