Files
Liti-agent-Development/Case_02_Comparison_Research/plans/s2-10-hybrid-implementation-v1.md
T
jhogyu 978e6a6b72 feat(stage2): integrate hybrid S2_10
Bind structured S2_00 handoff to the inline-prompt S2_10 Agent and reseal the parent-child release chain.
2026-08-31 01:11:40 +09:00

8.7 KiB

S2_10 Hybrid Prompt Implementation v1

Objective

Implement the S2_10 legal-resolution-map package defined by stage_2_optimal_update_strategy_v.5-1.md, S2_10_SOW_v.1.md, and S2_10_assets_v.1.md, while retaining the compatible legal/runtime boundaries from S2_10_SOW.md. Produce the versioned authoring YAML Stage_2_S2_10_v.1.yml, deploy its version-free projection and supporting assets under Default_Agent/Stage_2_Clean/, and leave a reproducible offline validation record.

Deliverables

  • YAML_Prompts/2. Stage_2/Stage_2_S2_10_v.1.yml
  • The fixed S2_10 physical asset set listed by S2_10_assets_v.1.md, deployed at its canonical paths
  • Updated shared release/module manifests required to close the new hashes
  • Byte-identical .py/.txt pairs for every Python asset
  • Updated YAML_Prompts/2. Stage_2/MEMORY.md

Scope and Non-Scope

In scope: hybrid inline-static prompt ownership, two-field dynamic-item boundary, schema/binding/workflow/projection/receipt/release updates, fixtures, offline tests, and release resealing. Out of scope: live GPT invocation, native adapter execution, production admission, legal sign-off, and dependencies on Stage 2 v0-v3 YAML specifications.

Known Inputs

  • stage_2_optimal_update_strategy_v.5-1.md
  • S2_10_SOW.md and controlling hybrid delta S2_10_SOW_v.1.md
  • S2_10_assets_v.1.md
  • SKILL.md
  • test_code_executor.ipynb
  • Current Default_Agent/Stage_2_Clean/ S2_00/S2_10 assets

Material Assumptions

  • The explicit requested output basename Stage_2_S2_10_v.1.yml is the authoring source; the deployment projection remains version-free.
  • S2_10_SOW_v.1.md controls where it narrows or replaces the older S2_10_SOW.md external-prompt design.
  • S2_10 contains exactly one direct LLM map task and no runtime Python or Code Executor task. Python files are offline build/validation/test assets only and must have byte-identical .txt mirrors.
  • The 36-file fixed S2_10 inventory counts physical files; shared release manifests that require resealing are tracked separately.

Questions That Could Change the Outcome

  • Whether the host platform's native LLM adapter supports the declared four-message order and structured-output contract remains a live-admission question, not an offline implementation blocker.
  • Model benchmark and Korean-law review receipts remain pending until external evidence exists.

Workstreams and Dependencies

  1. Freeze the 36-file inventory and retained/replaced/missing ledger.
  2. In parallel, draft the hybrid authoring YAML and update independent assets.
  3. Build the version-free projection, then bind hashes into receipts and the child release.
  4. Reseal shared parent/module manifests without changing S2_00 business logic.
  5. Run one asset-level review/fix cycle and exactly two YAML review/fix cycles.
  6. Run offline validation, parity, negative-mutation, and forbidden-dependency checks.
  7. Record results and residual live-admission gaps in project memory.

Source and Tool Plan

  • Use local repository documents and assets only; no web research is required.
  • Use apply_patch for authored changes and mechanical copy only for .py to .txt byte mirrors.
  • Use repository Python validators/tests in offline mode. Do not invoke live models or adapters.
  • Use independent subagents for inventory/design review and YAML validation; batch asset-level review assignments within the four-agent concurrency limit.

Validation Plan

  • Parse all YAML/JSON and validate schemas and cross-file hashes.
  • Assert one map task, zero deterministic/tool/reduce tasks, exact model/reasoning/verbosity, no tools, and preflight: false.
  • Assert inline common/static prompts and exactly two permitted dynamic item placeholders.
  • Reject legacy preassembled prompt fields and raw Stage 1 rereads.
  • Verify local-reference output schema and adapter-owned permanent-ID rewrite boundary.
  • Verify .py/.txt byte parity.
  • Run positive fixtures and required negative mutations.
  • Distinguish offline contract completion from live adapter/model/legal admission.

Approval Boundaries

No external writes, live model calls, production admission, or commits are authorized. Existing unrelated worktree changes are preserved.

Progress

  • Read repository rules, local YAML skill, notebook guidance, strategy/SOW/assets, and project memory.
  • Freeze the exact inventory and delta ledger (K=36).
  • Draft and deploy all S2_10 assets.
  • Complete one asset validation and fix cycle.
  • Complete YAML validation and fix cycle 1.
  • Complete YAML validation and fix cycle 2.
  • Reseal shared manifests and run full offline validation.
  • Update project memory and record residual uncertainty.

Decision Log

Date/Stage Decision Basis Consequence
2026-08-30 / intake Treat S2_10_SOW_v.1.md as the controlling hybrid delta over S2_10_SOW.md. It was produced specifically to implement stage_2_optimal_update_strategy_v.5-1.md; the old SOW still assumes externally assembled full prompts. Retain compatible legal/output/adapter contracts, replace prompt ownership and item boundaries.
2026-08-30 / intake Use Stage_2_S2_10_v.1.yml as source and version-free deployment projection. Explicit user filename plus existing deployment convention. Builder/receipts must bind both source and projection identities without overwriting the old versioned source.
2026-08-30 / intake No runtime Python in S2_10. Strategy and SOW assign the single legal judgment to one native LLM map task. Notebook Code Executor rules apply only as safety constraints to offline Python assets; no inline run_code belongs in the S2_10 YAML.
2026-08-30 / implementation Exclude cluster_case_payload_sha256 from the payload's inner echo and bind it only in the top-level item/output echo. A payload containing its own byte hash would require an impossible fixed point. The inner payload echo has 19 fields; the top-level dispatch/model echo has all 20 fields.
2026-08-30 / asset review Replace the regression manifest's generic oracle marker with fixture-specific exact oracle IDs and recompute manifest/module/release closure. The one requested asset-review round found three affected physical assets plus stale module/test counts. Tests now compare exact per-fixture oracle lists; module counts are builder-derived (76, 12).

Evidence Ledger

Claim/Issue Source or Test Status Notes
Hybrid prompt ownership strategy v5-1 and SOW v1 accepted YAML owns stable common/task prompt; item owns case data only.
Physical asset count assets v1 inventory and filesystem audit verified K=36; all 36 paths exist at the specified locations.
S2_10 offline tests tests/s2_10 PASS 39/39, including exact fixture oracle seal and .py/.txt parity.
S2_00 compatibility tests/s2_00, both builders in check mode PASS 69/69; parent hash and opaque S2_10 Agent/binding handoff resealed.
Live adapter compatibility external adapter evidence pending Must not be inferred from offline tests.
Model quality and legal correctness benchmark/legal review receipts pending Receipts remain explicitly non-admitted.

Risks and Failure Modes

  • Hash cycles between parent and child releases; prevent by keeping the parent one-way and not reverse-locking the child from the parent.
  • YAML parser normalization changing prompt bytes; use explicit block scalars and hash the parsed canonical strings where specified.
  • Dynamic item reintroducing full prompt text or PII; enforce schema and negative fixtures.
  • Static tests being mistaken for live execution or legal approval; retain explicit status labels and pending receipts.
  • Shared dirty worktree causing accidental scope expansion; edit only identified Stage 2 assets.

Results and Residual Uncertainty

The 36-file hybrid S2_10 package is implemented and offline verified. The authoring and deployment Agent are byte-identical and contain exactly one gpt-5.6-sol/xhigh/medium/responses map task, four ordered prompt blocks, and two dynamic JSON placeholders. Parent release raw hash is 8dcfe060af33e2d86a93f1c277c485b67255500a8cfb7f3024ffc427e802b948; child release raw hash is 5adb88482fc6e54ae7475570facf9a6ae87fa613fce65d867d16d5256604f14d. Both builders pass read-only check mode, S2_10 tests pass 39/39, S2_00 compatibility tests pass 69/69, and all seven Python assets have byte-identical .txt mirrors.

This is not live or legal admission. The platform adapter, model benchmark, official authority/profile release, and Korean-lawyer review receipts remain explicitly pending; the parent remains DRAFT_NOT_EXECUTABLE and production execution remains blocked. No live LLM/MCP adapter or end-to-end matter run was performed.