feat(stage2): integrate hybrid S2_10
Bind structured S2_00 handoff to the inline-prompt S2_10 Agent and reseal the parent-child release chain.
This commit is contained in:
@@ -0,0 +1,116 @@
|
||||
# S2_10 Hybrid Prompt Implementation v1
|
||||
|
||||
## Objective
|
||||
|
||||
Implement the S2_10 legal-resolution-map package defined by `stage_2_optimal_update_strategy_v.5-1.md`, `S2_10_SOW_v.1.md`, and `S2_10_assets_v.1.md`, while retaining the compatible legal/runtime boundaries from `S2_10_SOW.md`. Produce the versioned authoring YAML `Stage_2_S2_10_v.1.yml`, deploy its version-free projection and supporting assets under `Default_Agent/Stage_2_Clean/`, and leave a reproducible offline validation record.
|
||||
|
||||
## Deliverables
|
||||
|
||||
- `YAML_Prompts/2. Stage_2/Stage_2_S2_10_v.1.yml`
|
||||
- The fixed S2_10 physical asset set listed by `S2_10_assets_v.1.md`, deployed at its canonical paths
|
||||
- Updated shared release/module manifests required to close the new hashes
|
||||
- Byte-identical `.py`/`.txt` pairs for every Python asset
|
||||
- Updated `YAML_Prompts/2. Stage_2/MEMORY.md`
|
||||
|
||||
## Scope and Non-Scope
|
||||
|
||||
In scope: hybrid inline-static prompt ownership, two-field dynamic-item boundary, schema/binding/workflow/projection/receipt/release updates, fixtures, offline tests, and release resealing. Out of scope: live GPT invocation, native adapter execution, production admission, legal sign-off, and dependencies on Stage 2 v0-v3 YAML specifications.
|
||||
|
||||
## Known Inputs
|
||||
|
||||
- `stage_2_optimal_update_strategy_v.5-1.md`
|
||||
- `S2_10_SOW.md` and controlling hybrid delta `S2_10_SOW_v.1.md`
|
||||
- `S2_10_assets_v.1.md`
|
||||
- `SKILL.md`
|
||||
- `test_code_executor.ipynb`
|
||||
- Current `Default_Agent/Stage_2_Clean/` S2_00/S2_10 assets
|
||||
|
||||
## Material Assumptions
|
||||
|
||||
- The explicit requested output basename `Stage_2_S2_10_v.1.yml` is the authoring source; the deployment projection remains version-free.
|
||||
- `S2_10_SOW_v.1.md` controls where it narrows or replaces the older `S2_10_SOW.md` external-prompt design.
|
||||
- S2_10 contains exactly one direct LLM map task and no runtime Python or Code Executor task. Python files are offline build/validation/test assets only and must have byte-identical `.txt` mirrors.
|
||||
- The 36-file fixed S2_10 inventory counts physical files; shared release manifests that require resealing are tracked separately.
|
||||
|
||||
## Questions That Could Change the Outcome
|
||||
|
||||
- Whether the host platform's native LLM adapter supports the declared four-message order and structured-output contract remains a live-admission question, not an offline implementation blocker.
|
||||
- Model benchmark and Korean-law review receipts remain pending until external evidence exists.
|
||||
|
||||
## Workstreams and Dependencies
|
||||
|
||||
1. Freeze the 36-file inventory and retained/replaced/missing ledger.
|
||||
2. In parallel, draft the hybrid authoring YAML and update independent assets.
|
||||
3. Build the version-free projection, then bind hashes into receipts and the child release.
|
||||
4. Reseal shared parent/module manifests without changing S2_00 business logic.
|
||||
5. Run one asset-level review/fix cycle and exactly two YAML review/fix cycles.
|
||||
6. Run offline validation, parity, negative-mutation, and forbidden-dependency checks.
|
||||
7. Record results and residual live-admission gaps in project memory.
|
||||
|
||||
## Source and Tool Plan
|
||||
|
||||
- Use local repository documents and assets only; no web research is required.
|
||||
- Use `apply_patch` for authored changes and mechanical copy only for `.py` to `.txt` byte mirrors.
|
||||
- Use repository Python validators/tests in offline mode. Do not invoke live models or adapters.
|
||||
- Use independent subagents for inventory/design review and YAML validation; batch asset-level review assignments within the four-agent concurrency limit.
|
||||
|
||||
## Validation Plan
|
||||
|
||||
- Parse all YAML/JSON and validate schemas and cross-file hashes.
|
||||
- Assert one map task, zero deterministic/tool/reduce tasks, exact model/reasoning/verbosity, no tools, and `preflight: false`.
|
||||
- Assert inline common/static prompts and exactly two permitted dynamic item placeholders.
|
||||
- Reject legacy preassembled prompt fields and raw Stage 1 rereads.
|
||||
- Verify local-reference output schema and adapter-owned permanent-ID rewrite boundary.
|
||||
- Verify `.py`/`.txt` byte parity.
|
||||
- Run positive fixtures and required negative mutations.
|
||||
- Distinguish offline contract completion from live adapter/model/legal admission.
|
||||
|
||||
## Approval Boundaries
|
||||
|
||||
No external writes, live model calls, production admission, or commits are authorized. Existing unrelated worktree changes are preserved.
|
||||
|
||||
## Progress
|
||||
|
||||
- [x] Read repository rules, local YAML skill, notebook guidance, strategy/SOW/assets, and project memory.
|
||||
- [x] Freeze the exact inventory and delta ledger (`K=36`).
|
||||
- [x] Draft and deploy all S2_10 assets.
|
||||
- [x] Complete one asset validation and fix cycle.
|
||||
- [x] Complete YAML validation and fix cycle 1.
|
||||
- [x] Complete YAML validation and fix cycle 2.
|
||||
- [x] Reseal shared manifests and run full offline validation.
|
||||
- [x] Update project memory and record residual uncertainty.
|
||||
|
||||
## Decision Log
|
||||
|
||||
| Date/Stage | Decision | Basis | Consequence |
|
||||
|---|---|---|---|
|
||||
| 2026-08-30 / intake | Treat `S2_10_SOW_v.1.md` as the controlling hybrid delta over `S2_10_SOW.md`. | It was produced specifically to implement `stage_2_optimal_update_strategy_v.5-1.md`; the old SOW still assumes externally assembled full prompts. | Retain compatible legal/output/adapter contracts, replace prompt ownership and item boundaries. |
|
||||
| 2026-08-30 / intake | Use `Stage_2_S2_10_v.1.yml` as source and version-free deployment projection. | Explicit user filename plus existing deployment convention. | Builder/receipts must bind both source and projection identities without overwriting the old versioned source. |
|
||||
| 2026-08-30 / intake | No runtime Python in S2_10. | Strategy and SOW assign the single legal judgment to one native LLM map task. | Notebook Code Executor rules apply only as safety constraints to offline Python assets; no inline `run_code` belongs in the S2_10 YAML. |
|
||||
| 2026-08-30 / implementation | Exclude `cluster_case_payload_sha256` from the payload's inner echo and bind it only in the top-level item/output echo. | A payload containing its own byte hash would require an impossible fixed point. | The inner payload echo has 19 fields; the top-level dispatch/model echo has all 20 fields. |
|
||||
| 2026-08-30 / asset review | Replace the regression manifest's generic oracle marker with fixture-specific exact oracle IDs and recompute manifest/module/release closure. | The one requested asset-review round found three affected physical assets plus stale module/test counts. | Tests now compare exact per-fixture oracle lists; module counts are builder-derived (`76`, `12`). |
|
||||
|
||||
## Evidence Ledger
|
||||
|
||||
| Claim/Issue | Source or Test | Status | Notes |
|
||||
|---|---|---|---|
|
||||
| Hybrid prompt ownership | strategy v5-1 and SOW v1 | accepted | YAML owns stable common/task prompt; item owns case data only. |
|
||||
| Physical asset count | assets v1 inventory and filesystem audit | verified | K=36; all 36 paths exist at the specified locations. |
|
||||
| S2_10 offline tests | `tests/s2_10` | PASS | 39/39, including exact fixture oracle seal and `.py/.txt` parity. |
|
||||
| S2_00 compatibility | `tests/s2_00`, both builders in check mode | PASS | 69/69; parent hash and opaque S2_10 Agent/binding handoff resealed. |
|
||||
| Live adapter compatibility | external adapter evidence | pending | Must not be inferred from offline tests. |
|
||||
| Model quality and legal correctness | benchmark/legal review receipts | pending | Receipts remain explicitly non-admitted. |
|
||||
|
||||
## Risks and Failure Modes
|
||||
|
||||
- Hash cycles between parent and child releases; prevent by keeping the parent one-way and not reverse-locking the child from the parent.
|
||||
- YAML parser normalization changing prompt bytes; use explicit block scalars and hash the parsed canonical strings where specified.
|
||||
- Dynamic item reintroducing full prompt text or PII; enforce schema and negative fixtures.
|
||||
- Static tests being mistaken for live execution or legal approval; retain explicit status labels and pending receipts.
|
||||
- Shared dirty worktree causing accidental scope expansion; edit only identified Stage 2 assets.
|
||||
|
||||
## Results and Residual Uncertainty
|
||||
|
||||
The 36-file hybrid S2_10 package is implemented and offline verified. The authoring and deployment Agent are byte-identical and contain exactly one `gpt-5.6-sol/xhigh/medium/responses` map task, four ordered prompt blocks, and two dynamic JSON placeholders. Parent release raw hash is `8dcfe060af33e2d86a93f1c277c485b67255500a8cfb7f3024ffc427e802b948`; child release raw hash is `5adb88482fc6e54ae7475570facf9a6ae87fa613fce65d867d16d5256604f14d`. Both builders pass read-only check mode, S2_10 tests pass 39/39, S2_00 compatibility tests pass 69/69, and all seven Python assets have byte-identical `.txt` mirrors.
|
||||
|
||||
This is not live or legal admission. The platform adapter, model benchmark, official authority/profile release, and Korean-lawyer review receipts remain explicitly pending; the parent remains `DRAFT_NOT_EXECUTABLE` and production execution remains blocked. No live LLM/MCP adapter or end-to-end matter run was performed.
|
||||
Reference in New Issue
Block a user