feat(stage1-part2): special-law 배선·seed admission 도입 및 안정화 (Aug 21~22 디버깅 사이클)

- prompt_compiler에 SLP(특별법 프로파일) 배선, prompt_composition_policy 갱신
- seed_admission_policy.v1 신설: R0 전체 스키마 검증을 4단 수용 등급으로 전환
- special_law_profiles/ 3종 + profile_ids.json 카탈로그 배포
- stage_1_part_2_v.8.yml: worker 모델 gemini-3.1-pro-preview 복원, 투영 정렬,
  slice event 식별자 정본화, S0 enum 어휘 명시, 시드 신선도 실물 근거화
- 개정 전 판본 8종 outdated/ 보존, 문제 분석 문서 5종 수록

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-23 21:51:29 +09:00
co-authored by Claude Fable 5
parent 31035d5757
commit 85fca071c0
30 changed files with 32532 additions and 185 deletions
@@ -0,0 +1 @@
{"catalog_id":"special_law_profiles","entries":[{"attached_domain_id":"E-20","profile_id":"SLP-IP","status":"active"},{"attached_domain_id":"E-21","profile_id":"SLP-MEDIA-MEDIATION","status":"active"},{"attached_domain_id":"E-06","profile_id":"SLP-PRODUCT-LIABILITY","status":"active"}],"expected_count":3,"schema_version":"stage1_assembly_reference_catalog.v1"}
@@ -0,0 +1,52 @@
# SLP-IP — 지식재산 특별법 profile
`profile_id`: `SLP-IP` · 부착 도메인: E-20 지식재산(`ip_claims`)
조립 위치: 공통 계약 → 의존 도메인(E-01·E-05) → E-20 overlay → **이 profile** → runtime guard
## 권한 경계
- 이 조각은 E-20 overlay를 좁힐 뿐 입력·출력·법리 범위를 넓히지 않는다. 새 출력 key와 새 review key를 만들지 않는다.
- `final_conclusion_forbidden`은 유지된다. 권리 유효·귀속, 보호범위 포함 여부, 침해 성립, 손해액을 확정하지 않는다.
- 법령 조문 번호·존속기간·법정손해배상 한도·요율을 이 조각에 상수로 두지 않는다. 전부 SG-04 law-version 후보로 넘긴다.
- **권리 종류를 먼저 특정하지 못하면 요건 층을 단일화하지 않는다.** 특허·실용신안, 상표, 디자인, 저작권, 부정경쟁·영업비밀은 각기 다른 특별법 체계이고 성립·보호범위·침해 판단·구제·손해 산정 특칙이 서로 다르므로, 종류 미특정 시 후보를 종류별로 병렬 보존하고 review로 표시한다.
## 1. 식별 단서 — source-backed only
"특허·저작권·상표"라는 단어만 있는 자료는 활성화 근거가 아니다(E-20 `E20-N01`). 다음이 있을 때 이 층을 연다.
1. **권리 발생 방식**: 등록으로 발생하는 권리(특허·실용신안·상표·디자인)는 등록원부·공보 원천을, 무방식으로 발생하는 권리(저작권)는 창작·공표·저작자 표시 원천을, 영업비밀·성과 도용은 비밀관리성·경제적 유용성·비공지성 또는 성과의 형성 원천을 각각 요구한다.
2. **권리자 지위**: 원시취득·승계·직무발명 또는 업무상저작물·공동보유·양도·전용실시권·통상실시권·이용허락 중 어느 구조인지 원천으로 특정한다.
3. **피고 실시행위**: 생산·사용·양도·대여·수입·전시·복제·전송·표장 사용·부정취득·부정사용 중 어떤 행위 유형인지와 그 기간·경로.
4. **구제수단 축**: 금지·폐기·신용회복·손해배상·부당이득·실시료 상당액 중 무엇을 구하는지. 요건이 다르므로 합치지 않는다.
## 2. 고유 요건 슬롯·증거 component
E-20 슬롯에 연결하며 새 슬롯을 만들지 않는다.
- `e20.el01` — 권리 특정·존속. 등록권리는 등록번호·권리범위 문언·존속 상태를, 저작물은 창작성 표현 부분의 특정을 후보로 남긴다. 존속기간 계산은 이 조각이 하지 않고 기산 사실만 보존한다.
- `e20.el02` — 귀속·양도·실시허락. 등록원부와 계약이 불일치하면 어느 쪽도 지우지 말고 `E20_SOURCE_CONFLICT_REVIEW`와 함께 병존시킨다.
- `e20.el03`·`e20.el04` — 비교 구조는 권리 종류별로 요소가 다르다. 특허·실용신안은 청구항 구성요소 대비, 상표는 표장·지정상품 유사와 출처 혼동 우려, 디자인은 전체적 심미감, 저작권은 의거성과 실질적 유사성, 영업비밀은 비밀관리성과 부정취득·사용 태양. **다른 종류의 판단 요소를 교차 적용하지 않는다.**
- `e20.op01`·`e20.op02` — 항변을 별개 후보로 둔다: 권리 무효 사유, 권리 소진, 선사용, 자유실시·공지기술, 허락·묵시적 이용허락, 권리남용, 저작권의 제한·인용 사유, 의거성 부인. 성립 여부는 판단하지 않는다.
- **트랙 경계**: 등록권리의 무효·권리범위 확인은 심판·심결취소 트랙에 속하고 이 민사 트랙과 다르다. 무효 주장이 원천에 있으면 판단하지 말고 `TRACK_BOUNDARY_REVIEW`로 경계를 남긴다. 형사 고소·행정조치 접점도 같은 방식으로 표시한다.
- 증거 component는 `e20.right_identity`·`e20.ownership_license`·`e20.infringing_act` 등 등록 어휘를 그대로 쓴다.
## 3. review code 발행 규칙
등록된 코드만 쓰고, 발행 경로는 공통 계약이 정한 후보별 `review_items`와 root `unknown_or_unrouted_reviews`뿐이다.
- 권리 종류가 특정되지 않았거나 종류별 특별법 원천이 없다 → `E20_SPECIAL_LAW_PROFILE_REQUIRED`
- 권리 특정·존속 원천이 부족하다 → `E20_RIGHT_IDENTITY_VALIDITY_REVIEW`
- 귀속·실시허락 원천이 부족하거나 충돌한다 → `E20_OWNERSHIP_LICENSE_REVIEW`
- 비교 대상·범위 원천이 부족하다 → `E20_SCOPE_COMPARISON_REVIEW`
- 실시행위 원천이 부족하다 → `E20_INFRINGEMENT_ACT_REVIEW`
- 구제수단별 요건이 섞였거나 미분화다 → `E20_REMEDY_BOUNDARY_REVIEW`
- 무효·심판·형사·행정 트랙 접점이 있다 → `TRACK_BOUNDARY_REVIEW`
- 손해 추정·실시료 규정군이 필요한데 formula가 pin되지 않았다 → `DEFERRED_CALCULATION_TRACK_REVIEW`
## 4. 계산·법령 version 경계
- CE-01·CE-R1은 deferred다. 손해 추정·실시료 상당액·법정손해배상 선택지는 **operand와 근거 원천만** 보존하고 산정을 실행하지 않는다.
- SG-04로 pin할 것: 권리 종류별 적용 법률과 행위기간별 버전, 존속·기간 규정, 손해 추정·법정손해배상 규정의 버전 간 차이, 판례 변경 review.
- **E-20 extension은 profile 전용 key를 선언하지 않는다.** 이 profile의 후보는 선언된 `ip_right_candidates`·`ownership_license_candidates`·`infringing_act_candidates`·`scope_comparison_candidates`·`calculation_operand_candidates`·`boundary_routes`에 담고, 새 key를 만들지 않는다.
최종 법률결론과 최종 청구액을 확정하지 않는다. 원천 참조와 미해결 review code를 보존한다.
@@ -0,0 +1,50 @@
# SLP-MEDIA-MEDIATION — 언론피해 구제 특별법 profile
`profile_id`: `SLP-MEDIA-MEDIATION` · 부착 도메인: E-21 언론·인격권(`media_personality_rights`)
조립 위치: 공통 계약 → 의존 도메인 → E-21 overlay → **이 profile** → runtime guard
## 권한 경계
- 이 조각은 E-21 overlay를 좁힐 뿐 입력·출력·법리 범위를 넓히지 않는다. 새 출력 key와 새 review key를 만들지 않는다.
- `final_conclusion_forbidden`은 유지된다. 위법성, 진실성·상당성, 정정·반론 청구권 성립, 기간 준수 여부를 확정하지 않는다.
- 청구기간·제척기간의 일수와 법령 조문 번호를 이 조각에 상수로 두지 않는다. 기산 사실만 후보로 남기고 기간 규범은 SG-04로 넘긴다.
- 인격권 침해 일반(E-21 본체)과 이 특별법 층은 다르다. **매체가 이 특별법의 적용 대상인지 먼저 확인하지 못하면 구제수단 요건을 단일화하지 않는다.**
## 1. 식별 단서 — source-backed only
1. **적용 대상 매체 판정**: 신문·잡지 등 정기간행물, 방송, 뉴스통신, 인터넷신문, 인터넷뉴스서비스 등 이 특별법이 정한 언론·매체에 해당하는지를 보이는 원천. 개인 게시물·일반 게시판·사적 전파는 이 층이 아니라 일반 인격권·정보통신 영역의 경계 사안으로 표시한다.
2. **보도·게시물 특정**: 원문과 수정본, 제목·본문·이미지·영상, 게재 매체·게시 주체·게시 시각·URL·캡처를 버전별로 보존한다. 정정 대상은 "보도 내용 중 사실적 주장"이므로 의견 표명 부분과 구분해 특정한다.
3. **피해자 특정**: 성명·직함·사진·정황 등으로 피해자가 특정·식별 가능한지의 원천. 집단 표시의 경우 개별 구성원 식별 가능성 사실을 따로 남긴다.
4. **구제수단 축**: 정정보도, 반론보도, 추후보도, 손해배상은 요건·기산점·상대방이 서로 다르다. **합치지 않고 축별로 분리 보존한다.**
## 2. 고유 요건 슬롯·증거 component
E-21 슬롯에 연결하며 새 슬롯을 만들지 않는다.
- `e21.el01` — 보도 원본·버전·게시 주체·시점. 정정·삭제·수정 이력 자체가 요건 사실이므로 변경 전후를 모두 보존한다.
- `e21.el02` — 식별가능성과 전파범위. 열람·구독·전재·포털 노출 등 전파 경로 원천을 별도 후보로 둔다.
- `e21.el03` — 사실적 주장과 의견 표명의 구분, 진실성·상당성·공익성 관련 사실. 구제수단별로 요구되는 바가 다르다: 반론보도는 보도 내용의 진실 여부와 무관하게 청구 가능한 구조이고, 추후보도는 형사절차 관련 보도 후의 결과 확정 사실을 기산 사실로 한다. **각 축의 요건 사실을 서로 옮겨 쓰지 않는다.**
- `e21.el04`·`e21.el05` — 명예·신용·초상·사생활·개인정보 침해와 동의, 그리고 구제수단별 기간 기산 사실(보도를 안 날, 보도가 있은 날, 형사절차 결과 확정일 등)을 축별로 분리해 후보로 남긴다. 도과 판단은 하지 않는다.
- `e21.op01`·`e21.op02` — 항변을 별개 후보로 둔다: 진실성·상당성, 공익성, 의견·논평, 동의, 이미 공개된 사실, 피해자 비식별, 정정보도 요건 흠결, 기간 경과 주장. 성립 여부는 판단하지 않는다.
- **트랙 경계**: 언론중재위원회의 조정·중재 절차와 법원 소송은 서로 다른 트랙이고 전치·병행 여부가 사건마다 다르다. 조정 신청·중재 합의·직권조정 관련 원천이 있으면 절차 판단을 하지 말고 경계 review로 남긴다. 형사 명예훼손·개인정보 규제 접점도 같다.
## 3. review code 발행 규칙
등록된 코드만 쓰고, 발행 경로는 공통 계약이 정한 후보별 `review_items`와 root `unknown_or_unrouted_reviews`뿐이다.
- 적용 대상 매체 판정 원천이 없거나 이 층의 필수 원천이 빠졌다 → `E21_SPECIAL_LAW_PROFILE_REQUIRED`
- 보도 원본·버전·게시 시각이 불완전하거나 충돌한다 → `E21_PUBLICATION_VERSION_REVIEW`
- 피해자 특정·식별가능성 원천이 부족하다 → `E21_IDENTIFIABILITY_REVIEW`
- 사실 적시와 의견의 구분, 진실성·상당성 원천이 부족하다 → `E21_FACT_OPINION_TRUTH_REVIEW`
- 공익성 형량 자료가 한쪽만 있다 → `E21_PUBLIC_INTEREST_BALANCING_REVIEW`
- 동의·사생활·개인정보 관련 원천이 부족하다 → `E21_CONSENT_PRIVACY_REVIEW`
- 구제수단별 기간 기산 사실이 없거나 축이 섞였다 → `E21_REMEDY_PERIOD_REVIEW`
- 조정·중재·형사·행정 트랙 접점이 있다 → `TRACK_BOUNDARY_REVIEW`
## 4. 계산·법령 version 경계
- CE-05에는 operand와 `ready|partial|blocked` 상태만 보낸다. 위자료 액수·산정 기준을 이 조각이 정하지 않는다.
- SG-04로 pin할 것: 대상 매체 유형별 적용 법률, 구제수단별 청구기간 규범과 기준시점, 시행일·경과규정, 판례 변경 review.
- **E-21 extension은 profile 전용 key를 선언하지 않는다.** 이 profile의 후보는 선언된 `publication_candidates`·`subject_content_candidates`·`personality_privacy_candidates`·`remedy_timing_candidates`·`calculation_operand_candidates`·`boundary_routes`에 담고, 새 key를 만들지 않는다.
최종 법률결론과 최종 청구액을 확정하지 않는다. 원천 참조와 미해결 review code를 보존한다.
@@ -0,0 +1,51 @@
# SLP-PRODUCT-LIABILITY — 제조물책임 특별법 profile
`profile_id`: `SLP-PRODUCT-LIABILITY` · 부착 도메인: E-06 전문가·제조물 책임(`professional_liability`)
조립 위치: 공통 계약 → 의존 도메인(E-01·E-05·X3) → E-06 overlay → **이 profile** → runtime guard
## 권한 경계
- 이 조각은 E-06 overlay를 좁힐 뿐 입력·출력·법리 범위를 넓히지 않는다. 새 출력 key와 새 review key를 만들지 않는다.
- `final_conclusion_forbidden`은 유지된다. 특별법 적용 여부, 결함 인정, 면책 성립, 기간 도과를 확정하지 않는다.
- 법령 조문 번호·법정률·기간·한도·추정 규정의 수치를 이 조각에 상수로 두지 않는다. 전부 SG-04 law-version 후보로 넘긴다.
- 제조물 단서가 있는 사건에만 이 층을 연다. 의료·전문직 과실 사건에 기계적으로 부착하지 않으며, 단순 품질불만·수선 요구만 있는 자료는 E-06 `E06-N02` 경계에 따라 monitor로 둔다.
## 1. 식별 단서 — source-backed only
제품명·브랜드 단독 언급은 활성화 근거가 아니다. 다음 원천 조합이 있을 때만 요건 층을 연다.
1. **제조물성**: 제조·가공된 동산임을 보이는 원천(다른 동산·부동산의 일부를 이루는 경우 포함). 미가공 1차산물·부동산 자체·용역·정보는 경계 사안으로 표시한다.
2. **책임주체**: 제조업자, 자기를 제조업자로 표시하거나 오인하게 할 표시를 한 자, 수입한 자, 제조업자를 알 수 없을 때의 공급자. 지위마다 요건이 다르므로 합치지 않는다.
3. **손해**: 정상적인 사용 상태에서 발생한 생명·신체·재산 손해. **제조물 자체에만 생긴 손해**는 이 profile이 아니라 계약·하자담보 leaf 소관이므로 경계 표시와 함께 보존한다.
4. **결함 유형**: 제조상·설계상·표시상(설명·지시·경고) 중 어느 유형인지 원천에 따라 분기한다. 유형 미특정은 후보 삭제 사유가 아니라 review 사유다.
## 2. 고유 요건 슬롯·증거 component
E-06 슬롯에 연결하며 새 슬롯을 만들지 않는다.
- `e06.el08.product_identity_chain` — 제조물 동일성과 공급·유통·사용·보관 계보. 계보 단절은 대체원인 항변과 연결되는 사실이므로 단절 자체를 후보로 남긴다.
- `e06.el09.product_defect` — 유형별 사실 요소를 구분해 수집한다. 제조상은 설계·제조방법 기준으로부터의 이탈, 설계상은 합리적 대체설계의 채용 가능성과 미채용, 표시상은 합리적 설명·지시·경고의 결여. 유형 간 요소를 뒤섞지 않는다.
- `e06.el04.damage_change`·`e06.el05.causal_mechanism` — 증명 완화 구조를 뒷받침하는 세 사실을 **각각 독립 후보로** 슬롯화한다: ① 정상적 사용 상태에서 손해가 발생했다는 사정, ② 그 손해가 제조업자의 배타적 지배영역에서 비롯되었다는 사정, ③ 그러한 손해가 결함 없이는 통상 발생하지 않는다는 사정. 추정 규정을 적용해 결함·인과를 인정하는 판단은 하지 않는다.
- `e06.op01`~ 반대사실·항변 — 공급 당시 결함 부존재, 공급 당시의 과학·기술 수준으로 결함 발견 불가(개발위험), 법령이 정한 기준 준수로 인한 결함, 원재료·부품 제조업자의 설계·제작 지시 종속을 각각 별개 후보로 둔다. 결함 방지 조치를 하지 않았다는 반대 사정도 함께 보존하되 면책 성립 여부는 판단하지 않는다.
- `e06.el10.law_version_dates` — 기간 축이 둘(손해와 책임주체를 안 날 기준, 제조물 공급일 기준)이므로 기산 사실을 축별로 분리해 후보로 남기고 도과 판단은 하지 않는다.
- 증거 component는 `e06.identity_record`·`e06.primary_elements`·`e06.opposing_record`·`e06.timeline_record` 어휘를 그대로 쓴다.
## 3. review code 발행 규칙
등록된 코드만 쓰고, 발행 경로는 공통 계약이 정한 후보별 `review_items`와 root `unknown_or_unrouted_reviews`뿐이다.
- 제조물 단서가 있으나 이 층의 필수 원천이 없다 → `SPECIAL_LAW_PROFILE_REQUIRED_REVIEW`
- 결함 유형·표시 내용·사용 상태 원천이 충돌한다 → `E06_DISCLOSURE_OR_PRODUCT_DEFECT_CONFLICT_REVIEW`
- 필수 슬롯이 비었다 → `E06_REQUIRED_ELEMENT_GAP_REVIEW`
- 기산 사실·적용 법률 버전이 미확정이다 → `E06_LAW_VERSION_OR_PERIOD_INPUT_REVIEW`
- 제조물 자체 손해만 있거나 계약·하자담보 구성과 경합한다 → `E06_SPECIAL_LAW_OR_TRACK_BOUNDARY_REVIEW`(필요 시 `TRACK_BOUNDARY_REVIEW` 병행)
- 리콜·안전규제·행정처분 접점이 있다 → `PUBLIC_LAW_NEXUS_REVIEW`
- 결함·사용 상태가 상담록 진술로만 지지된다 → `E06_MEETING_ONLY_SUPPORT_REVIEW`
## 4. 계산·법령 version 경계
- CE-03·CE-05에는 operand와 `ready|partial|blocked` 상태만 보낸다. 산식·법정률·기간·한도를 이 조각이 정하지 않는다.
- SG-04로 pin할 것: 제조물책임 특별법과 일반 불법행위 구성의 적용 법률 후보, 공급일·사고일·인지일 등 사건 기준일별 버전, 경과규정 적용 여부, 추정·면책 규정의 버전 간 차이.
- E-06 extension은 `special_law_profile_candidates`를 선언한다. 이 profile의 후보는 그 key에 담고 새 key를 만들지 않는다.
최종 법률결론과 최종 청구액을 확정하지 않는다. 원천 참조와 미해결 review code를 보존한다.
@@ -0,0 +1 @@
{"special_law_profile_registry_index":{"catalog_id":"special_law_profiles","contract_guards":{"final_conclusion_forbidden":true,"meeting_only_evidence_promotion_forbidden":true,"source_membership_required":true,"strict_json_output":true,"unknown_values_require_review":true},"entries":[{"attached_domain_ids":["E-20"],"label_ko":"지식재산 특별법 profile","profile_id":"SLP-IP","prompt_overlay_path":"SLP-IP/prompt_overlay.md","prompt_overlay_sha256":"7ed12d8183992733e79feb25622e6980642308d0705b8409f6a465394d3b7f13","status":"active"},{"attached_domain_ids":["E-21"],"label_ko":"언론피해 구제 특별법 profile","profile_id":"SLP-MEDIA-MEDIATION","prompt_overlay_path":"SLP-MEDIA-MEDIATION/prompt_overlay.md","prompt_overlay_sha256":"04ea71d164111934039640486ba113c22793c95168318c38974897d90ffe5a0e","status":"active"},{"attached_domain_ids":["E-06"],"label_ko":"제조물책임 특별법 profile","profile_id":"SLP-PRODUCT-LIABILITY","prompt_overlay_path":"SLP-PRODUCT-LIABILITY/prompt_overlay.md","prompt_overlay_sha256":"78281b81af2e61647024cbd72f9245d14bd3a6750413d16eb9aa093e6778d491","status":"active"}],"expected_count":3,"generated_at":"2026-08-21T00:00:00+09:00","id_policy":{"final_conclusion_forbidden":true,"law_constants_in_prompt_forbidden":true,"profile_id_pattern":"^SLP-[A-Z][A-Z0-9-]{1,63}$","version_pinning_signal_id":"SG-04"},"registry_version":"Stage1.Assembly.2026-08-21.v1","schema_version":"stage1_special_law_profile_registry_index.v1"}}
@@ -133,22 +133,8 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl
"mutually exclusive prompt markers coexist",
present,
)
guard = policy.get("deterministic_prompt_size_guard", {})
maximum_bytes = int(guard.get("max_utf8_bytes", 96000))
maximum_scalars = int(guard.get("max_unicode_scalars", 24000))
scalar_count = len(compiled)
estimate = math.ceil(len(compiled_bytes) / 4)
exceeded = []
if len(compiled_bytes) > maximum_bytes:
exceeded.append("max_utf8_bytes")
if scalar_count > maximum_scalars:
exceeded.append("max_unicode_scalars")
if exceeded:
raise RuntimeContractError(
str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")),
"compiled prompt exceeds deterministic byte/scalar guard",
{"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded},
)
manifest = {
"schema_version": "compiled_prompt_manifest.v1",
"compiled_prompt_sha256": sha256_bytes(compiled_bytes),
@@ -156,12 +142,9 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl
"deterministic_prompt_size_guard": {
"utf8_bytes": len(compiled_bytes),
"unicode_scalars": scalar_count,
"max_utf8_bytes": maximum_bytes,
"max_unicode_scalars": maximum_scalars,
"not_a_tokenizer": True,
"model_context_fit_not_proven": True,
"legacy_advisory_estimate": estimate,
"pass": not exceeded,
},
"normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator},
"fragments": [
@@ -246,8 +229,19 @@ def collect_domain_fragments(
profile_map = profile_paths or {}
for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key):
if profile not in profile_map:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}")
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id))
raise RuntimeContractError(
"PROMPT_REQUIRED_FRAGMENT_MISSING",
f"special-law profile prompt path not supplied by caller: {profile}",
{"domain_id": domain_id, "declared_profile": profile, "supplied_profile_ids": sorted(profile_map)},
)
profile_path = Path(profile_map[profile]).resolve()
if not profile_path.is_file():
raise RuntimeContractError(
"PROMPT_REQUIRED_FRAGMENT_MISSING",
f"special-law profile prompt file missing: {profile}",
{"domain_id": domain_id, "declared_profile": profile, "resolved_path": str(profile_path)},
)
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(profile_path), domain_id))
if runtime_guard_path:
specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve())))
specs.extend(extra_specs or [])
@@ -133,22 +133,8 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl
"mutually exclusive prompt markers coexist",
present,
)
guard = policy.get("deterministic_prompt_size_guard", {})
maximum_bytes = int(guard.get("max_utf8_bytes", 96000))
maximum_scalars = int(guard.get("max_unicode_scalars", 24000))
scalar_count = len(compiled)
estimate = math.ceil(len(compiled_bytes) / 4)
exceeded = []
if len(compiled_bytes) > maximum_bytes:
exceeded.append("max_utf8_bytes")
if scalar_count > maximum_scalars:
exceeded.append("max_unicode_scalars")
if exceeded:
raise RuntimeContractError(
str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")),
"compiled prompt exceeds deterministic byte/scalar guard",
{"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded},
)
manifest = {
"schema_version": "compiled_prompt_manifest.v1",
"compiled_prompt_sha256": sha256_bytes(compiled_bytes),
@@ -156,12 +142,9 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl
"deterministic_prompt_size_guard": {
"utf8_bytes": len(compiled_bytes),
"unicode_scalars": scalar_count,
"max_utf8_bytes": maximum_bytes,
"max_unicode_scalars": maximum_scalars,
"not_a_tokenizer": True,
"model_context_fit_not_proven": True,
"legacy_advisory_estimate": estimate,
"pass": not exceeded,
},
"normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator},
"fragments": [
@@ -246,8 +229,19 @@ def collect_domain_fragments(
profile_map = profile_paths or {}
for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key):
if profile not in profile_map:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}")
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id))
raise RuntimeContractError(
"PROMPT_REQUIRED_FRAGMENT_MISSING",
f"special-law profile prompt path not supplied by caller: {profile}",
{"domain_id": domain_id, "declared_profile": profile, "supplied_profile_ids": sorted(profile_map)},
)
profile_path = Path(profile_map[profile]).resolve()
if not profile_path.is_file():
raise RuntimeContractError(
"PROMPT_REQUIRED_FRAGMENT_MISSING",
f"special-law profile prompt file missing: {profile}",
{"domain_id": domain_id, "declared_profile": profile, "resolved_path": str(profile_path)},
)
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(profile_path), domain_id))
if runtime_guard_path:
specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve())))
specs.extend(extra_specs or [])
@@ -22,8 +22,6 @@
"separator": "\n\n"
},
"deterministic_prompt_size_guard": {
"max_utf8_bytes": 96000,
"max_unicode_scalars": 24000,
"overflow_fail_code": "PROMPT_SIZE_GUARD_EXCEEDED",
"not_a_tokenizer": true,
"model_context_fit_not_proven": true,
@@ -0,0 +1,412 @@
{
"schema_version": "stage1_seed_admission_policy.v1",
"status": "ACTIVE",
"purpose": "R0 가 이미 수행하는 전체 스키마 검증(worker_output_validator -> schema_subset_validator)의 결과를 등급으로 나누어 처리하는 규칙을 코드 밖에 선언한다. 구조 위반은 막고, 어휘·형태 위반은 정본으로 정규화하거나 정당한 자유 공간으로 수확한 뒤 검토로 남긴다. 값을 지어내지 않는다 — 표에 없는 낱말과 대조 불가능한 참조는 복구하지 않고 격리한다.",
"authority": {
"seed_contract": "task_c_bo_stage_b_domain_bo_seed.v3",
"validator": "Default_Agent/stage1_runtime/worker_output_validator.txt",
"schema": "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json",
"case_kind_names_forbidden": true,
"note": "검증은 신설이 아니다. R0 가 이미 전체 스키마로 돌리고 결과를 guard_warnings 로 강등하고 있었다(13도메인 1,151건 실측). 이 정책은 그 결과에 처분을 준다. 회차 실측(13도메인): before 1,151건 -> 복구 후 잔여를 후보 격리로 흡수. 키 표류(component_id->slot_id, fact/value->excerpt)는 rename_sibling 으로 이름만 바로잡는다. 워커가 echo 한 sha 는 신선도의 근거가 아니다 — 실물 두 값(계획서 sha ↔ R0 가 읽은 문서 sha)이 판정하고, echo 불일치는 검토로 남긴다. echo 실패는 둘로 가른다 — 다른 슬라이스의 sha 를 echo 하면 워커가 엉뚱한 슬라이스로 작업한 것이라 실물 대조로 잡히지 않는 중대 위반(block)이고, 그 밖은 전사 실수(review)다. 슬라이스를 못 읽어 실물 대조 자체가 불가한 도메인은 강등하지 않는다(quarantine)."
},
"enforcement": "enforce",
"enforcement_modes": {
"shadow": "복구·격리를 계산하고 기록하되 seed 를 바꾸지 않고 아무것도 막지 않는다. 첫 전환 회차용.",
"enforce": "복구를 적용하고 격리·중단 판정을 실제로 내린다."
},
"repair_order": [
"repair_normalize",
"repair_derive",
"repair_harvest",
"repair_evict"
],
"max_repair_passes": 12,
"disposition_by_code": {
"SCHEMA_ENUM": "repair_normalize",
"SCHEMA_REQUIRED": "repair_derive",
"SCHEMA_ADDITIONAL_PROPERTY": "repair_harvest",
"SCHEMA_PATTERN": "repair_harvest",
"SCHEMA_TYPE": "repair_harvest",
"SCHEMA_MIN_LENGTH": "repair_harvest",
"SCHEMA_MAX_LENGTH": "repair_harvest",
"SCHEMA_UNIQUE_ITEMS": "repair_harvest",
"WORKER_COMPLETION_RECEIPT_MISMATCH": "repair_derive",
"UNREGISTERED_EFFECT_REVIEW_MISSING": "review",
"WORKER_DECLARATION_MISMATCH": "review",
"WORKER_NOT_R0_ELIGIBLE": "review",
"WORKER_SEED_ID_DUPLICATE": "quarantine",
"WORKER_SOURCE_MEMBERSHIP_FAILED": "quarantine",
"WORKER_ROOT_INVALID": "block",
"WORKER_SCHEMA_VERSION_INVALID": "block",
"WORKER_DOMAIN_MISMATCH": "block",
"WORKER_TASK_INSTANCE_MISMATCH": "block",
"WORKER_CONTRACT_GUARD_FAILED": "block",
"WORKER_FINAL_FIELD_PRESENT": "block",
"CALCULATION_REVIEW_MISSING": "review",
"CALCULATION_DOMAIN_NOT_DECLARED": "review",
"EVIDENCE_SLOT_NOT_DECLARED": "review",
"UNKNOWN_VALUE_REVIEW_MISSING": "review",
"WORKER_BO_TYPE_UNREGISTERED": "review",
"WORKER_DEPENDENCY_UNREGISTERED": "review",
"WORKER_DOMAIN_UNREGISTERED": "review",
"WORKER_SLICE_GUARD_MISMATCH": "review",
"WORKER_SLICE_HASH_MISMATCH": "review",
"WORKER_REGISTRY_HASH_MISMATCH": "review",
"SLICE_ROOT_INVALID": "review",
"ACTIVATION_EXPECTED_SET_INVALID": "review",
"DOMAIN_ID_CATALOG_INVALID": "review",
"DOMAIN_ID_CATALOG_DUPLICATE": "review",
"DOMAIN_ID_CATALOG_ID_INVALID": "review",
"FINAL_CONCLUSION_FIELD_FORBIDDEN": "block",
"VALIDATOR_CRASHED": "quarantine",
"VALIDATOR_CRASHED_ROOT": "block",
"SEED_ECHO_SHA_MISMATCH": "review",
"SEED_ECHO_FOREIGN_SLICE": "block",
"SEED_ECHO_SHA_UNVERIFIABLE": "quarantine",
"SEED_ECHO_SHA_UNEXPLAINED": "quarantine"
},
"unknown_code_disposition": "quarantine",
"surplus_sink": {
"candidate_path": "extensions.worker_surplus",
"by_container": {
"bo_seed_candidates[].element_fact_candidates[]": "payload",
"bo_seed_candidates[].opposing_fact_candidates[]": "payload",
"bo_seed_candidates[].defense_candidates[]": "payload"
},
"note": "컨테이너 지역 payload 에는 잎 이름 그대로 담는다(하류 어댑터가 payload[\"status\"] 처럼 짧은 이름으로 읽는다). 후보 수준 보관은 {path,value} 덧붙이기 전용 목록이라 어떤 값도 덮이지 않는다.",
"candidate_path_is_list": true
},
"enum_synonyms": {
"status": {
"ready": "READY",
"ok": "READY",
"complete": "READY",
"ready_with_review": "READY_WITH_REVIEW",
"review": "READY_WITH_REVIEW",
"no_support": "NO_SUPPORT",
"none": "NO_SUPPORT",
"empty": "NO_SUPPORT"
},
"bo_seed_candidates[].evidence_slot_status[].status": {
"supported": "filled",
"proven": "filled",
"sufficient": "filled",
"satisfied": "filled",
"met": "filled",
"complete": "filled",
"partially_supported": "partial",
"partial_support": "partial",
"insufficient": "partial",
"weak": "partial",
"missing_required": "missing",
"unsupported": "missing",
"absent": "missing",
"none": "missing",
"not_found": "missing",
"conflict": "conflicted",
"contradicted": "conflicted",
"disputed": "conflicted"
},
"bo_seed_candidates[].calculation_requests[].completeness": {
"complete": "ready",
"ok": "ready",
"sufficient": "ready",
"incomplete": "partial",
"partial_input": "partial",
"blocked_by_missing": "blocked",
"unavailable": "blocked",
"postponed": "deferred",
"later": "deferred"
},
"bo_seed_candidates[].review_items[].severity": {
"information": "info",
"informational": "info",
"note": "info",
"soft_warning": "review",
"warning": "review",
"minor": "review",
"hard": "hard_warning",
"major": "hard_warning",
"severe": "hard_warning",
"error": "block",
"critical": "block",
"blocker": "block"
},
"unknown_or_unrouted_reviews[].severity": {
"information": "info",
"informational": "info",
"note": "info",
"soft_warning": "review",
"warning": "review",
"minor": "review",
"hard": "hard_warning",
"major": "hard_warning",
"severe": "hard_warning",
"error": "block",
"critical": "block",
"blocker": "block"
},
"bo_seed_candidates[].element_fact_candidates[].source_kind": {
"evidence_index": "evidence",
"document": "evidence",
"doc": "evidence",
"event": "event_candidate",
"event_id": "event_candidate",
"meeting": "meeting_clause",
"clause": "meeting_clause",
"consultation": "meeting_clause"
},
"bo_seed_candidates[].opposing_fact_candidates[].source_kind": {
"evidence_index": "evidence",
"document": "evidence",
"doc": "evidence",
"event": "event_candidate",
"event_id": "event_candidate",
"meeting": "meeting_clause",
"clause": "meeting_clause",
"consultation": "meeting_clause"
},
"bo_seed_candidates[].defense_candidates[].source_kind": {
"evidence_index": "evidence",
"document": "evidence",
"doc": "evidence",
"event": "event_candidate",
"event_id": "event_candidate",
"meeting": "meeting_clause",
"clause": "meeting_clause",
"consultation": "meeting_clause"
}
},
"enum_case_fold": true,
"enum_unmapped_disposition": "repair_evict",
"required_derivation": {
"bo_seed_candidates[].element_fact_candidates[].source_id": {
"rule": "first_universe_member_of_sibling_array",
"sibling_candidates": [
"source_refs",
"refs",
"evidence_refs",
"source_ref"
]
},
"bo_seed_candidates[].element_fact_candidates[].source_kind": {
"rule": "universe_kind_of_sibling",
"sibling": "source_id",
"fallback_sibling_arrays": [
"source_refs",
"refs",
"evidence_refs",
"source_ref"
],
"note": "형제 source_id 가 먼저 채워지지 않았을 수 있어 참조 배열 폴백을 둔다."
},
"bo_seed_candidates[].opposing_fact_candidates[].source_id": {
"rule": "first_universe_member_of_sibling_array",
"sibling_candidates": [
"source_refs",
"refs",
"evidence_refs",
"source_ref"
]
},
"bo_seed_candidates[].opposing_fact_candidates[].source_kind": {
"rule": "universe_kind_of_sibling",
"sibling": "source_id",
"fallback_sibling_arrays": [
"source_refs",
"refs",
"evidence_refs",
"source_ref"
]
},
"bo_seed_candidates[].defense_candidates[].source_id": {
"rule": "first_universe_member_of_sibling_array",
"sibling_candidates": [
"source_refs",
"refs",
"evidence_refs",
"source_ref"
]
},
"bo_seed_candidates[].defense_candidates[].source_kind": {
"rule": "universe_kind_of_sibling",
"sibling": "source_id",
"fallback_sibling_arrays": [
"source_refs",
"refs",
"evidence_refs",
"source_ref"
]
},
"bo_seed_candidates[].calculation_requests[].calculation_domain": {
"rule": "sibling_matching_pattern",
"sibling_candidates": [
"calculation_id",
"calculation_code",
"domain"
],
"pattern": "^CE-[0-9]{2}$",
"fallback_rename_sibling": [
"calculation_id",
"calculation_code",
"domain",
"ce_id"
]
},
"bo_seed_candidates[].calculation_requests[].completeness": {
"rule": "constant",
"value": "partial",
"note": "워커가 완결성을 말하지 않았다. 낙관도 비관도 아닌 중간값을 두고 검토로 올린다."
},
"bo_seed_candidates[].calculation_requests[].source_refs": {
"rule": "constant",
"value": [],
"note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다."
},
"bo_seed_candidates[].evidence_slot_status[].source_refs": {
"rule": "constant",
"value": [],
"note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다."
},
"bo_seed_candidates[].review_items[].severity": {
"rule": "constant",
"value": "review"
},
"bo_seed_candidates[].review_items[].review_code": {
"rule": "constant",
"value": "WORKER_REVIEW_UNSPECIFIED"
},
"bo_seed_candidates[].review_items[].reason": {
"rule": "constant",
"value": "워커가 사유를 적지 않았다. R0 수용 단계에서 채운 자리이므로 Stage2 가 원문을 다시 본다."
},
"bo_seed_candidates[].review_items[].source_refs": {
"rule": "constant",
"value": [],
"note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다."
},
"bo_seed_candidates[].legal_effect_candidates[].source_refs": {
"rule": "constant",
"value": [],
"note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다."
},
"bo_seed_candidates[].party_roles[].party_refs": {
"rule": "constant",
"value": []
},
"completion_receipt.task_instance_id": {
"rule": "from_context",
"key": "task_instance_id"
},
"completion_receipt.domain_id": {
"rule": "from_context",
"key": "domain_id"
},
"completion_receipt.seed_count": {
"rule": "from_context",
"key": "seed_count"
},
"completion_receipt.emitted_seed_ids": {
"rule": "from_context",
"key": "emitted_seed_ids"
},
"completion_receipt.r0_membership_key": {
"rule": "from_context",
"key": "r0_membership_key"
},
"bo_seed_candidates[].evidence_slot_status[].slot_id": {
"rule": "rename_sibling",
"sibling_candidates": [
"component_id",
"slot",
"element_slot_id",
"registry_component_id",
"slot_key"
],
"note": "워커가 같은 자리를 component_id 로 부르는 표류가 실측됐다. 이름만 바로잡고 값은 보존한다."
},
"bo_seed_candidates[].element_fact_candidates[].excerpt": {
"rule": "rename_sibling",
"sibling_candidates": [
"fact",
"value",
"content",
"text",
"statement"
]
},
"bo_seed_candidates[].opposing_fact_candidates[].excerpt": {
"rule": "rename_sibling",
"sibling_candidates": [
"fact",
"value",
"content",
"text",
"statement"
]
},
"bo_seed_candidates[].defense_candidates[].excerpt": {
"rule": "rename_sibling",
"sibling_candidates": [
"fact",
"value",
"content",
"text",
"statement",
"defense"
]
},
"unknown_or_unrouted_reviews[].severity": {
"rule": "constant",
"value": "review"
},
"unknown_or_unrouted_reviews[].source_refs": {
"rule": "constant",
"value": [],
"note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다."
},
"unknown_or_unrouted_reviews[].review_code": {
"rule": "rename_sibling",
"sibling_candidates": [
"code",
"issue_code",
"review_id"
]
},
"unknown_or_unrouted_reviews[].reason": {
"rule": "rename_sibling",
"sibling_candidates": [
"message",
"note",
"description",
"detail"
]
}
},
"derivation_failed_disposition": "repair_evict",
"quarantine_budget": {
"max_quarantined_candidate_ratio": 0.34,
"min_admitted_candidates_run": 1,
"note": "격리는 후보 단위다. 한 도메인의 후보가 전부 격리되면 그 도메인은 NO_SUPPORT 로 내려가고(이는 v3 계약상 적법한 상태다) 회차는 계속된다. 회차 전체 격리 비율이 한계를 넘으면 모델 잡음이 아니라 계약 파손이므로 그때 경성 중단한다."
},
"review": {
"issue_type": "seed_admission_repair",
"quarantine_issue_type": "seed_admission_quarantine",
"severity": "SOFT_WARNING",
"quarantine_severity": "HARD_WARNING",
"downstream_owner": "Stage2",
"record_fields": [
"candidate_ref",
"path",
"code",
"action",
"received",
"applied",
"rule_id"
],
"note": "모든 복구는 원래 값(received)과 적용 값(applied)을 함께 남긴다. 조용한 덮어쓰기를 금지하기 위한 것이며, Stage2 는 이 기록만으로 무엇이 바뀌었는지 재구성할 수 있어야 한다. 실질 서술(excerpt·value·reason 등)을 담은 채 밀려난 항목은 기계적 키 이동과 등급을 나눈다 — 같은 SOFT_WARNING 에 섞으면 소수가 다수에 묻힌다.",
"evicted_with_content_issue_type": "seed_admission_content_evicted",
"evicted_with_content_severity": "HARD_WARNING"
},
"residual_root_disposition": "review"
}
@@ -0,0 +1,322 @@
#!/usr/bin/env python3
"""Deterministically compose common, domain, dependency, and special-law prompts."""
from __future__ import annotations
import argparse
import json
import math
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Any
from registry_loader import load_registry
from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json
@dataclass(frozen=True)
class FragmentSpec:
fragment_id: str
category: str
path: str
source_domain_id: str | None = None
def _priority(policy: dict[str, Any]) -> dict[str, int]:
out: dict[str, int] = {}
for index, item in enumerate(policy.get("fragment_priority", [])):
if isinstance(item, str):
out[item] = index * 10
elif isinstance(item, dict) and item.get("category"):
out[str(item["category"])] = int(item.get("rank", index * 10))
if not out:
raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty")
return out
def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str:
codes = policy.get("fail_codes")
return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback
def _normalize_fragment(path: str) -> tuple[str, bytes, str]:
source = Path(path)
try:
text = source.read_text(encoding="utf-8")
except FileNotFoundError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc
except UnicodeDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc
if source.suffix.lower() == ".json":
try:
value = json.loads(text)
except json.JSONDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc
normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n"
else:
normalized = normalize_text(text)
data = normalized.encode("utf-8")
return normalized, data, sha256_bytes(data)
def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]:
priorities = _priority(policy)
required_categories = [str(item) for item in policy.get("required_categories", [])]
categories = {spec.category for spec in specs}
missing_categories = [item for item in required_categories if item not in categories]
if missing_categories:
raise RuntimeContractError(
_fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"),
"required prompt fragment category missing",
missing_categories,
)
unknown = sorted(categories - set(priorities))
if unknown:
raise RuntimeContractError(
_fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"),
"fragment category has no priority",
unknown,
)
prepared: list[dict[str, Any]] = []
by_id: dict[str, str] = {}
for spec in specs:
text, data, digest = _normalize_fragment(spec.path)
previous = by_id.get(spec.fragment_id)
if previous is not None and previous != digest:
raise RuntimeContractError(
_fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"),
f"same fragment_id has different normalized hashes: {spec.fragment_id}",
{"first": previous, "second": digest},
)
by_id[spec.fragment_id] = digest
prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest})
prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path))
seen_sha: set[str] = set()
kept: list[dict[str, Any]] = []
omitted: list[dict[str, Any]] = []
for item in prepared:
spec = item["spec"]
source_path = Path(spec.path).resolve()
release_root = Path(__file__).resolve().parents[1]
try:
logical_path = source_path.relative_to(release_root).as_posix()
except ValueError as exc:
raise RuntimeContractError(
"PROMPT_FRAGMENT_PATH_INVALID",
"prompt fragment must resolve under the release root",
source_path.name,
) from exc
public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id}
if item["sha256"] in seen_sha:
omitted.append({**public, "reason": "duplicate_normalized_sha256"})
continue
seen_sha.add(item["sha256"])
kept.append(item)
separator = str(policy.get("byte_contract", {}).get("separator", "\n\n"))
compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n"
compiled_bytes = compiled.encode("utf-8")
required_markers = [str(item) for item in policy.get("required_markers", [])]
missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()]
if missing_markers:
raise RuntimeContractError(
_fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"),
"compiled prompt lacks required guard marker",
missing_markers,
)
for conflict in policy.get("conflict_markers", []):
if isinstance(conflict, list) and len(conflict) > 1:
present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()]
if len(present) > 1:
raise RuntimeContractError(
_fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"),
"mutually exclusive prompt markers coexist",
present,
)
guard = policy.get("deterministic_prompt_size_guard", {})
maximum_bytes = int(guard.get("max_utf8_bytes", 96000))
maximum_scalars = int(guard.get("max_unicode_scalars", 24000))
scalar_count = len(compiled)
estimate = math.ceil(len(compiled_bytes) / 4)
exceeded = []
if len(compiled_bytes) > maximum_bytes:
exceeded.append("max_utf8_bytes")
if scalar_count > maximum_scalars:
exceeded.append("max_unicode_scalars")
if exceeded:
raise RuntimeContractError(
str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")),
"compiled prompt exceeds deterministic byte/scalar guard",
{"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded},
)
manifest = {
"schema_version": "compiled_prompt_manifest.v1",
"compiled_prompt_sha256": sha256_bytes(compiled_bytes),
"compiled_prompt_bytes": len(compiled_bytes),
"deterministic_prompt_size_guard": {
"utf8_bytes": len(compiled_bytes),
"unicode_scalars": scalar_count,
"max_utf8_bytes": maximum_bytes,
"max_unicode_scalars": maximum_scalars,
"not_a_tokenizer": True,
"model_context_fit_not_proven": True,
"legacy_advisory_estimate": estimate,
"pass": not exceeded,
},
"normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator},
"fragments": [
{
"fragment_id": item["spec"].fragment_id,
"category": item["spec"].category,
"path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()),
"sha256": item["sha256"],
"hash_kind": "file_sha256",
"source_domain_id": item["spec"].source_domain_id,
}
for item in kept
],
"deduplicated_fragments": omitted,
"required_markers_verified": required_markers,
}
return compiled, manifest
def _path_from_config(config_record: dict[str, Any], value: str) -> str:
config_path = config_record.get("config_path")
base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", "."))
path = Path(value)
return str(path if path.is_absolute() else (base / path).resolve())
def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]:
config = config_record["config"]
values: list[Any] = []
if isinstance(config.get("prompt_fragments"), list):
values.extend(config["prompt_fragments"])
for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"):
if isinstance(config.get(key), str):
values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"})
index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {}
if not values and isinstance(index_entry.get("prompt_overlay_path"), str):
values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"})
specs: list[FragmentSpec] = []
for index, value in enumerate(values, 1):
if isinstance(value, str):
path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}"
elif isinstance(value, dict) and isinstance(value.get("path"), str):
path_value = value["path"]
fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}")
else:
continue
specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id))
return specs
def collect_domain_fragments(
domain_id: str,
registry: dict[str, Any],
*,
common_contract_path: str,
profile_paths: dict[str, str] | None = None,
runtime_guard_path: str | None = None,
extra_specs: list[FragmentSpec] | None = None,
) -> list[FragmentSpec]:
entries = registry["entries"]
if domain_id not in entries:
raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}")
specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))]
visited: set[str] = set()
def add_dependencies(current: str) -> None:
for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key):
if dependency not in entries:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}")
if dependency in visited:
continue
visited.add(dependency)
add_dependencies(dependency)
specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency))
add_dependencies(domain_id)
domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id)
if not domain_specs:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}")
specs.extend(domain_specs)
profiles = entries[domain_id]["config"].get("special_law_profiles", [])
profile_map = profile_paths or {}
for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key):
if profile not in profile_map:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}")
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id))
if runtime_guard_path:
specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve())))
specs.extend(extra_specs or [])
return specs
def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]:
out: dict[str, str] = {}
for value in values:
if separator not in value:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value)
key, path = value.split(separator, 1)
out[key] = path
return out
def _parse_extra(values: list[str]) -> list[FragmentSpec]:
out: list[FragmentSpec] = []
for value in values:
parts = value.split(":", 2)
if len(parts) != 3:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value)
out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve())))
return out
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--registry-index", required=True)
parser.add_argument("--domain-id", required=True)
parser.add_argument("--policy", required=True)
parser.add_argument("--common-contract", required=True)
parser.add_argument("--runtime-guard")
parser.add_argument("--profile", action="append", default=[])
parser.add_argument("--fragment", action="append", default=[])
parser.add_argument("--output-prompt", required=True)
parser.add_argument("--output-manifest", required=True)
args = parser.parse_args(argv)
try:
registry = load_registry(args.registry_index)
specs = collect_domain_fragments(
args.domain_id,
registry,
common_contract_path=args.common_contract,
profile_paths=_parse_mapping(args.profile),
runtime_guard_path=args.runtime_guard,
extra_specs=_parse_extra(args.fragment),
)
policy = load_json(args.policy)
compiled, manifest = compile_fragments(specs, policy)
prompt_path = Path(args.output_prompt)
prompt_path.parent.mkdir(parents=True, exist_ok=True)
prompt_path.write_text(compiled, encoding="utf-8", newline="")
from runtime_common import sha256_file
manifest.update({
"domain_id": args.domain_id,
"compiled_prompt_path": str(prompt_path),
"registry_index_sha256": registry["index_sha256"],
"composition_policy_path": args.policy,
"composition_policy_sha256": sha256_file(args.policy),
})
write_json(args.output_manifest, manifest)
emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]})
return 0
except RuntimeContractError as exc:
emit_cli_result({"status": "FAILED", "error": exc.as_dict()})
return 1
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,322 @@
#!/usr/bin/env python3
"""Deterministically compose common, domain, dependency, and special-law prompts."""
from __future__ import annotations
import argparse
import json
import math
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Any
from registry_loader import load_registry
from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json
@dataclass(frozen=True)
class FragmentSpec:
fragment_id: str
category: str
path: str
source_domain_id: str | None = None
def _priority(policy: dict[str, Any]) -> dict[str, int]:
out: dict[str, int] = {}
for index, item in enumerate(policy.get("fragment_priority", [])):
if isinstance(item, str):
out[item] = index * 10
elif isinstance(item, dict) and item.get("category"):
out[str(item["category"])] = int(item.get("rank", index * 10))
if not out:
raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty")
return out
def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str:
codes = policy.get("fail_codes")
return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback
def _normalize_fragment(path: str) -> tuple[str, bytes, str]:
source = Path(path)
try:
text = source.read_text(encoding="utf-8")
except FileNotFoundError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc
except UnicodeDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc
if source.suffix.lower() == ".json":
try:
value = json.loads(text)
except json.JSONDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc
normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n"
else:
normalized = normalize_text(text)
data = normalized.encode("utf-8")
return normalized, data, sha256_bytes(data)
def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]:
priorities = _priority(policy)
required_categories = [str(item) for item in policy.get("required_categories", [])]
categories = {spec.category for spec in specs}
missing_categories = [item for item in required_categories if item not in categories]
if missing_categories:
raise RuntimeContractError(
_fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"),
"required prompt fragment category missing",
missing_categories,
)
unknown = sorted(categories - set(priorities))
if unknown:
raise RuntimeContractError(
_fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"),
"fragment category has no priority",
unknown,
)
prepared: list[dict[str, Any]] = []
by_id: dict[str, str] = {}
for spec in specs:
text, data, digest = _normalize_fragment(spec.path)
previous = by_id.get(spec.fragment_id)
if previous is not None and previous != digest:
raise RuntimeContractError(
_fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"),
f"same fragment_id has different normalized hashes: {spec.fragment_id}",
{"first": previous, "second": digest},
)
by_id[spec.fragment_id] = digest
prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest})
prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path))
seen_sha: set[str] = set()
kept: list[dict[str, Any]] = []
omitted: list[dict[str, Any]] = []
for item in prepared:
spec = item["spec"]
source_path = Path(spec.path).resolve()
release_root = Path(__file__).resolve().parents[1]
try:
logical_path = source_path.relative_to(release_root).as_posix()
except ValueError as exc:
raise RuntimeContractError(
"PROMPT_FRAGMENT_PATH_INVALID",
"prompt fragment must resolve under the release root",
source_path.name,
) from exc
public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id}
if item["sha256"] in seen_sha:
omitted.append({**public, "reason": "duplicate_normalized_sha256"})
continue
seen_sha.add(item["sha256"])
kept.append(item)
separator = str(policy.get("byte_contract", {}).get("separator", "\n\n"))
compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n"
compiled_bytes = compiled.encode("utf-8")
required_markers = [str(item) for item in policy.get("required_markers", [])]
missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()]
if missing_markers:
raise RuntimeContractError(
_fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"),
"compiled prompt lacks required guard marker",
missing_markers,
)
for conflict in policy.get("conflict_markers", []):
if isinstance(conflict, list) and len(conflict) > 1:
present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()]
if len(present) > 1:
raise RuntimeContractError(
_fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"),
"mutually exclusive prompt markers coexist",
present,
)
guard = policy.get("deterministic_prompt_size_guard", {})
maximum_bytes = int(guard.get("max_utf8_bytes", 96000))
maximum_scalars = int(guard.get("max_unicode_scalars", 24000))
scalar_count = len(compiled)
estimate = math.ceil(len(compiled_bytes) / 4)
exceeded = []
if len(compiled_bytes) > maximum_bytes:
exceeded.append("max_utf8_bytes")
if scalar_count > maximum_scalars:
exceeded.append("max_unicode_scalars")
if exceeded:
raise RuntimeContractError(
str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")),
"compiled prompt exceeds deterministic byte/scalar guard",
{"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded},
)
manifest = {
"schema_version": "compiled_prompt_manifest.v1",
"compiled_prompt_sha256": sha256_bytes(compiled_bytes),
"compiled_prompt_bytes": len(compiled_bytes),
"deterministic_prompt_size_guard": {
"utf8_bytes": len(compiled_bytes),
"unicode_scalars": scalar_count,
"max_utf8_bytes": maximum_bytes,
"max_unicode_scalars": maximum_scalars,
"not_a_tokenizer": True,
"model_context_fit_not_proven": True,
"legacy_advisory_estimate": estimate,
"pass": not exceeded,
},
"normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator},
"fragments": [
{
"fragment_id": item["spec"].fragment_id,
"category": item["spec"].category,
"path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()),
"sha256": item["sha256"],
"hash_kind": "file_sha256",
"source_domain_id": item["spec"].source_domain_id,
}
for item in kept
],
"deduplicated_fragments": omitted,
"required_markers_verified": required_markers,
}
return compiled, manifest
def _path_from_config(config_record: dict[str, Any], value: str) -> str:
config_path = config_record.get("config_path")
base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", "."))
path = Path(value)
return str(path if path.is_absolute() else (base / path).resolve())
def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]:
config = config_record["config"]
values: list[Any] = []
if isinstance(config.get("prompt_fragments"), list):
values.extend(config["prompt_fragments"])
for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"):
if isinstance(config.get(key), str):
values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"})
index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {}
if not values and isinstance(index_entry.get("prompt_overlay_path"), str):
values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"})
specs: list[FragmentSpec] = []
for index, value in enumerate(values, 1):
if isinstance(value, str):
path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}"
elif isinstance(value, dict) and isinstance(value.get("path"), str):
path_value = value["path"]
fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}")
else:
continue
specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id))
return specs
def collect_domain_fragments(
domain_id: str,
registry: dict[str, Any],
*,
common_contract_path: str,
profile_paths: dict[str, str] | None = None,
runtime_guard_path: str | None = None,
extra_specs: list[FragmentSpec] | None = None,
) -> list[FragmentSpec]:
entries = registry["entries"]
if domain_id not in entries:
raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}")
specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))]
visited: set[str] = set()
def add_dependencies(current: str) -> None:
for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key):
if dependency not in entries:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}")
if dependency in visited:
continue
visited.add(dependency)
add_dependencies(dependency)
specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency))
add_dependencies(domain_id)
domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id)
if not domain_specs:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}")
specs.extend(domain_specs)
profiles = entries[domain_id]["config"].get("special_law_profiles", [])
profile_map = profile_paths or {}
for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key):
if profile not in profile_map:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}")
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id))
if runtime_guard_path:
specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve())))
specs.extend(extra_specs or [])
return specs
def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]:
out: dict[str, str] = {}
for value in values:
if separator not in value:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value)
key, path = value.split(separator, 1)
out[key] = path
return out
def _parse_extra(values: list[str]) -> list[FragmentSpec]:
out: list[FragmentSpec] = []
for value in values:
parts = value.split(":", 2)
if len(parts) != 3:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value)
out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve())))
return out
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--registry-index", required=True)
parser.add_argument("--domain-id", required=True)
parser.add_argument("--policy", required=True)
parser.add_argument("--common-contract", required=True)
parser.add_argument("--runtime-guard")
parser.add_argument("--profile", action="append", default=[])
parser.add_argument("--fragment", action="append", default=[])
parser.add_argument("--output-prompt", required=True)
parser.add_argument("--output-manifest", required=True)
args = parser.parse_args(argv)
try:
registry = load_registry(args.registry_index)
specs = collect_domain_fragments(
args.domain_id,
registry,
common_contract_path=args.common_contract,
profile_paths=_parse_mapping(args.profile),
runtime_guard_path=args.runtime_guard,
extra_specs=_parse_extra(args.fragment),
)
policy = load_json(args.policy)
compiled, manifest = compile_fragments(specs, policy)
prompt_path = Path(args.output_prompt)
prompt_path.parent.mkdir(parents=True, exist_ok=True)
prompt_path.write_text(compiled, encoding="utf-8", newline="")
from runtime_common import sha256_file
manifest.update({
"domain_id": args.domain_id,
"compiled_prompt_path": str(prompt_path),
"registry_index_sha256": registry["index_sha256"],
"composition_policy_path": args.policy,
"composition_policy_sha256": sha256_file(args.policy),
})
write_json(args.output_manifest, manifest)
emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]})
return 0
except RuntimeContractError as exc:
emit_cli_result({"status": "FAILED", "error": exc.as_dict()})
return 1
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,46 @@
{
"schema_version": "prompt_composition_policy.v2",
"fragment_priority": [
{"category": "common_contract", "rank": 10},
{"category": "common_dependency", "rank": 20},
{"category": "domain", "rank": 30},
{"category": "special_law_profile", "rank": 40},
{"category": "runtime_guard", "rank": 50}
],
"required_categories": ["common_contract", "domain"],
"deduplication": {
"algorithm": "sha256",
"scope": "normalized_fragment_bytes",
"same_sha_action": "keep_first_by_priority_then_id",
"same_fragment_id_different_sha": "PROMPT_FRAGMENT_ID_HASH_CONFLICT"
},
"byte_contract": {
"encoding": "UTF-8",
"line_endings": "LF",
"fragment_trailing_newline": "exactly_one",
"compiled_trailing_newline": "exactly_one",
"separator": "\n\n"
},
"deterministic_prompt_size_guard": {
"max_utf8_bytes": 96000,
"max_unicode_scalars": 24000,
"overflow_fail_code": "PROMPT_SIZE_GUARD_EXCEEDED",
"not_a_tokenizer": true,
"model_context_fit_not_proven": true,
"legacy_advisory_estimate": {
"formula": "ceil(utf8_bytes/4)",
"execution_authority": false
}
},
"required_markers": ["final_conclusion_forbidden"],
"conflict_markers": [],
"fail_codes": {
"required_fragment_missing": "PROMPT_REQUIRED_FRAGMENT_MISSING",
"unknown_priority_category": "PROMPT_PRIORITY_UNKNOWN",
"fragment_id_hash_conflict": "PROMPT_FRAGMENT_ID_HASH_CONFLICT",
"prompt_size_guard_exceeded": "PROMPT_SIZE_GUARD_EXCEEDED",
"required_guard_missing": "PROMPT_FINAL_CONCLUSION_GUARD_MISSING",
"exclusive_marker_conflict": "PROMPT_EXCLUSIVE_MARKER_CONFLICT",
"fragment_path_invalid": "PROMPT_FRAGMENT_PATH_INVALID"
}
}
@@ -0,0 +1,54 @@
# special_law_profiles 자산 부재 문제 — 판단서
작성일 2026-08-21 · 대상 릴리스 `v.7/extension_research/Default_Agent` · 계기 `stage_1_part_2_v.8.yml` 첫 task 실패(`PROMPT_REQUIRED_FRAGMENT_MISSING`)
---
## 0. 실측 요약 (이 문서의 근거)
| 확인 항목 | 실측 결과 | 근거 |
|---|---|---|
| `special_law_profiles/` 디렉터리 | 릴리스에 **없음** | `Default_Agent` 하위 0건 |
| SLP 실사용 ID | **3종** — `SLP-PRODUCT-LIABILITY`(E-06), `SLP-IP`(E-20), `SLP-MEDIA-MEDIATION`(E-21) | 각 `domains/*/domain_config.json` |
| 그 외 SLP 이름 | `SLP-AUTO-DAMAGE`·`SLP-STATE-COMPENSATION` 은 **스키마 주석 안에만** 존재 | `domain_slice.schema.v2.json:101` |
| 조립 시 하드 실패 지점 | profile ID가 `profile_paths` 에 없으면 무조건 raise | `stage1_runtime/prompt_compiler.py:230-232` |
| 런타임이 기대하는 경로 | `Default_Agent/special_law_profiles/<SLP-ID>/prompt_overlay.md` | `Claude_YAML/Stage_1_Registry_Runtime_v1.yml:473, 492-494` |
| fanout 이 profile ID를 얻는 곳 | `registry["entries"][did]["config"]` = 실제 `domain_config.json` | `Stage_1_Registry_Runtime_v1.yml:311` |
| `runtime_manifest.json` | 432 entries, `special_law` 문자열 **0건** | manifest 전문 검색 |
| 조각 경로 제약 | 릴리스 루트(`Default_Agent/`) 밖이면 `PROMPT_FRAGMENT_PATH_INVALID` | `prompt_compiler.py` `compile_fragments()` |
즉 **설정은 요구하고, 런타임은 경로까지 정해 두었으며, 실물 파일만 없다.** 원인 분석의 "자산이 릴리스에 없다"는 진단은 이 릴리스에서도 그대로 성립한다.
---
## 1. 질문 1 — `special_law_profiles` 자산을 어떻게 할 것인가
**정공법(자산 신설)으로 가되, 범위를 선언된 3종으로 한정한다.** `Default_Agent/special_law_profiles/{SLP-PRODUCT-LIABILITY, SLP-IP, SLP-MEDIA-MEDIATION}/prompt_overlay.md` 를 만들고, 카탈로그 `special_law_profiles/_registry_index.json` 을 함께 두어 `registry_validator` 의 `REFERENCE_CATALOG_NOT_PROVIDED` 경고를 해소한 뒤, `runtime_manifest.json` 에 3~4건을 등재하고 `runtime_artifact_count` 를 증분한다. `prompt_compiler.py:230` 의 raise 를 완화하는 2안은 채택하지 않는다 — 그 raise 는 "선언했으면 반드시 조립한다"는 fail-closed 계약이고, E-06 의 `seed_prompt_overlay.md:5, :34` 가 "제조물 사건은 `SLP-PRODUCT-LIABILITY` 없이 종결하지 않는다"고 프롬프트 본문에서 명시하고 있어, 프로필을 건너뛰고 review_items 에만 표시하면 프롬프트가 스스로 금지한 상태(특별법 검토 없는 제조물 분석)를 산출물로 내보내게 된다. 크기 가드 제거와 성격이 다른 이유가 여기 있다 — 크기 가드는 토크나이저가 아닌 추정치에 근거한 보수적 차단이었지만, 이 raise 는 법리 완결성 계약이다. 콘텐츠 부담도 실제로는 작다: `ensemble_v2.md` 상 특별법은 "도메인 payload 플래그"이고 시행일·경과규정·버전 pinning 은 SG-04(`governing_law_version_signals.json`)가 코어 signal 로 흡수하므로, 프로필 overlay 는 **법령 조문·요율·기간 상수를 담지 않고** ① 특별법 후보 식별 단서, ② 그 특별법 고유 요건 슬롯과 증거 component, ③ 발행할 review code, ④ 버전 판단은 SG-04 로 넘긴다는 경계 선언만 담는 얇은 조각이면 충분하다(도메인 overlay 와 동일 톤, `final_conclusion_forbidden` 유지). 스키마 주석이 말하는 "9종"은 릴리스 어디에도 실체·목록이 없으므로 추종하지 말고, 나머지 6종은 해당 도메인이 `domain_config.json` 에서 실제로 선언할 때 같은 형식으로 증설한다.
### 구현 시 지켜야 할 계약 (실측 기반)
1. 경로는 `Default_Agent/special_law_profiles/<ID>/prompt_overlay.md` 고정 — YAML 이 `PROFILE_DIR + "/" + profile_id + "/prompt_overlay.md"` 로 조립한다.
2. 릴리스 루트 밖 경로 금지(`compile_fragments` 의 `relative_to(release_root)` 검사).
3. 3개 파일의 정규화 sha256 이 서로 달라야 한다 — 동일하면 `deduplication` 이 뒤엣것을 `duplicate_normalized_sha256` 로 탈락시킨다.
4. UTF-8·LF·후행 개행 1개(`byte_contract`). 조각 카테고리는 `special_law_profile`(rank 40)로 이미 정책에 등록되어 있어 정책 수정은 불필요하다.
5. `runtime_manifest.json` 은 `{"path", "sha256"}` 배열이며 `runtime_artifact_count` 와 길이가 일치해야 한다.
### 문자열 절단(`SLP-PRODUCT-LIABI`, 17자)에 대하여
v.7 릴리스 전체를 검색해도 절단형은 0건이고, `domain_config.json` 과 E-06 overlay 모두 21자 정식 ID를 쓴다. 스키마 패턴도 절단형을 그대로 통과시키므로(길이 하한만 있음) 스키마가 자른 것도 아니다. 이 릴리스에서는 원인을 설명할 근거가 없으므로, 실행 워크스페이스의 `main.py` 생성 단계(Stage YAML)에서 확인하는 것이 맞다. 다만 **절단을 고쳐도 파일이 없으면 동일하게 실패**하므로 자산 신설이 선행 조건이다.
---
## 2. 질문 2 — ★ Insight 1·2번은 실행을 위해 고쳐야 하는가
**둘 다 stage 1 실행을 막는 결함이 아니다 — 1번은 문서 드리프트, 2번은 이미 수정된 과거 이력의 잔재다.** 1번(인덱스의 `special_law_profiles` 가 전부 `[]` 인데 E-06·E-20·E-21 config 에는 값이 있음)은, `registry_loader.load_registry()` 가 엔트리를 `config`(= `domain_config.json` 실제 로드본)와 `index_entry`(인덱스 원문)로 분리해 담고 `prompt_compiler.py:227`, `domain_slice_compiler.py:482`, fanout 빌더(`Stage_1_Registry_Runtime_v1.yml:311`)가 **모두 `config` 만** 읽기 때문에 실행 경로에서 소비되지 않는다. `registry_validator` 도 `prompt_overlay` 에 대해서만 인덱스↔config 불일치(`PROMPT_OVERLAY_REFERENCE_DIVERGENCE`)를 검사할 뿐 이 필드는 대조하지 않는다. 따라서 지금 실패는 인덱스의 `[]` 때문이 아니라 값이 정상적으로 흘러가 파일을 찾는 단계에서 나는 것이며, 인덱스 정합화는 실행 차단 해소와 무관한 위생 작업이다(다만 인덱스만 보고 판단하는 사람·도구를 오도하므로 자산 신설과 함께 채워 두는 편이 낫다 — 인덱스 수정 시 `config_sha256` 재계산 대상은 아니고 인덱스 자체 해시만 갱신되면 된다). 2번(스키마 패턴이 소문자라 SLP slice 가 거부된다는 주석)은 **주석만 남고 패턴은 이미 고쳐져 있다**: 현재 `domain_slice.schema.v2.json:104` 의 패턴은 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|[a-z][a-z0-9_]{1,127})$` 이고, 실측 결과 `SLP-PRODUCT-LIABILITY`·`SLP-IP`·`SLP-MEDIA-MEDIATION` 3종 모두 매치된다. 즉 고칠 대상은 스키마 패턴이 아니라 사실과 어긋난 `$comment` 문구이며(더불어 그 주석이 존재한다고 전제하는 `special_law_profiles/_registry_index.json` 이 실재하지 않는다는 점까지 함께 정리해야 한다), 이는 실행 차단이 아닌 문서 정확성 문제다.
### 정리 — 실행 재개에 필요한 것 / 아닌 것
| 항목 | 실행 차단 | 조치 |
|---|---|---|
| `special_law_profiles/<ID>/prompt_overlay.md` 3종 부재 | **예 (유일한 차단)** | 신설 |
| `runtime_manifest.json` 미등재 | 아니오(실행은 되나 릴리스 무결성 미봉인) | 자산 신설과 동시에 등재 |
| `_registry_index.json` 의 `special_law_profiles: []` | 아니오 | 위생 차원에서 값 채움 |
| `domain_slice.schema.v2.json:101` 주석 | 아니오 | 주석 수정(패턴은 이미 정상) |
| `SLP-PRODUCT-LIABI` 절단 | 파일이 생긴 뒤 재확인 | 실행 워크스페이스 `main.py` 생성부 점검 |
@@ -0,0 +1,74 @@
# special_law_profiles 그리고 insight 답변 분석
작성일 2026-08-21 · 검증 대상 `extension_research/assets_special_law_problem.md`(이하 "v.1") · 방법: v.1의 모든 사실 주장과 권고를 릴리스 실물(`Default_Agent/`)·런타임 코드(`stage1_runtime/`)·참조 런타임 YAML(`Claude_YAML/Stage_1_Registry_Runtime_v1.yml`)에 대해 명령 단위로 재실측
**검증 범위의 경계**: 실제 실패한 실행본 `stage_1_part_2_v.8.yml` 은 이 저장소에 없다(`/Users/jsahn/Works/작업결과/v2` 소재). 따라서 "런타임이 기대하는 경로" 류의 주장은 이 폴더의 참조본 YAML 근거이며, 실행본이 동일 규약을 쓴다는 것은 에러 문구(`special-law profile prompt path missing`)가 `prompt_compiler.py:232` 의 raise 문자열과 일치한다는 점으로 뒷받침된다. 이 경계는 v.1이 명시하지 않았던 것으로, 본 리포트에서 명시한다.
---
## 1. 질문 1 — special_law_profiles 자산 처리 방향
### 1.1 검증 내용
**유지 판정(재실측으로 확인된 주장) 5건**
| v.1 주장 | 재실측 결과 | 판정 |
|---|---|---|
| `special_law_profiles/` 실물 부재가 유일한 실행 차단 | `Default_Agent` 하위 해당 디렉터리·파일 0건, `prompt_compiler.py:231-232` 는 `profile_map` 에 키 없으면 무조건 raise | **유지** |
| 신설 범위는 선언된 3종 한정 | config 선언은 E-06·E-20·E-21 의 3종뿐. `SLP-AUTO-DAMAGE`·`SLP-STATE-COMPENSATION` 은 `domain_slice.schema.v2.json:101` **주석 문자열 안에만** 존재 | **유지** |
| raise 완화(2안) 기각 | v.1 은 E-06 근거만 들었으나, 재실측 결과 **3개 도메인 전부** 프롬프트 본문이 합성을 강제한다: E-06 overlay:5·:34("SLP-PRODUCT-LIABILITY 없이 종결하지 않는다"), E-20 overlay:13("**SLP-IP를 반드시 합성**하고 … SG-04로 pin한다"), E-21 overlay:13("**SLP-MEDIA-MEDIATION을 반드시 합성**…"). 완화 시 세 도메인 모두 자기 프롬프트가 금지한 산출물을 내게 된다 | **유지·강화** |
| 경로 계약 `special_law_profiles/<ID>/prompt_overlay.md` + 릴리스 루트 하위 강제 | 참조본 YAML:473(`PROFILE_DIR`)·:492-494(`--profile <ID>=<DIR>/<ID>/prompt_overlay.md`), `compile_fragments` 의 `relative_to(release_root)` 검사 확인 | **유지** |
| overlay 는 조문·요율·기간 상수 없는 얇은 조각(버전 판단은 SG-04 위임) | `ensemble_v2.md:22`("special_law는 도메인 payload 플래그")·`:175`(SG-04 가 "적용 법률·특별법 후보, 시행일·경과규정 … version pinning" 흡수) 재확인. E-20·E-21 overlay:13 도 "SG-04로 pin"을 명시해 위임 구조가 이미 프롬프트에 내장돼 있음 | **유지** |
**정정 판정(v.1 이 부정확했던 주장) 3건**
1. **"3개 파일의 정규화 sha256 이 서로 달라야 한다"(v.1 계약 3) → 하드 계약 아님.** dedup(`duplicate_normalized_sha256`)의 스코프는 `compile_fragments()` **1회 호출 = 도메인 1개 컴파일 내부**다. 각 도메인은 SLP 를 1종만 선언하므로 한 컴파일에 `special_law_profile` 조각은 최대 1개이고, 서로 다른 도메인의 profile 파일끼리는 애초에 같은 dedup 집합에 들어가지 않는다. 실제 제약은 "같은 컴파일 안의 다른 조각(공통 계약·도메인 overlay 등)과 정규화 바이트가 같으면 탈락" 뿐이며 현실적으로 발생하지 않는다. 3개 파일 내용을 서로 다르게 쓰는 것은 여전히 옳지만(내용상 당연), 위반 시 조립이 깨지는 계약은 아니다 — **권장으로 강등**.
2. **"카탈로그 `special_law_profiles/_registry_index.json` 을 함께 두어 `REFERENCE_CATALOG_NOT_PROVIDED` 경고 해소" → 파일 생성만으로는 해소되지 않는다.** `registry_validator` 의 카탈로그 주입 경로는 두 가지뿐이다: ① 도메인 인덱스의 `reference_catalogs`(또는 `reference_sets`) 맵 — 현재 없음, ② CLI `--reference-catalog KIND=PATH`. 파일을 만들어 두기만 하면 validator 는 그 존재를 모른다. 그리고 재실측 결과 **참조본 A0 는 카탈로그의 정본 위치를 이미 확정해 두었다**: `--reference-catalog profile=Default_Agent/platform/reference_catalogs/profile_ids.json`(YAML A0 호출부). 그 디렉터리에 실재하는 것은 `calculation_ids.json`·`domain_ids.json` 뿐이고 `profile_ids.json` 은 없다. 즉 카탈로그의 1차 신설 위치는 v.1 이 말한 `special_law_profiles/_registry_index.json` 이 아니라 **`platform/reference_catalogs/profile_ids.json`** 이며(형식은 `calculation_ids.json` 과 동형: `{"catalog_id": …, "entries": [{"profile_id": "SLP-…", …}]}` — validator 의 `_catalog_ids` 가 `profile_id` 키를 수집한다), `special_law_profiles/_registry_index.json` 은 별도 목적(자산 인덱스 + sha 봉인, 아래 종합 답변)으로 두는 이원 구조가 도메인 registry 와 동형이다.
3. **"`runtime_manifest.json` 에 3~4건 등재하고 count 증분"(v.1 계약 5 포함) → 현행 장부 관례와 어긋나는 권고.** 재실측 결과 manifest 의 `domains/` 51건은 `module_role_projection.json` 25 + `structure_types.json` 25 + `_common/common_worker_contract.md` 1 이 전부로, **`domain_config.json`(26개)과 `seed_prompt_overlay.md`(26개)는 하나도 등재돼 있지 않다.** registry 자산의 무결성 봉인은 manifest 가 아니라 registry 인덱스가 담당한다(`config_sha256` 는 `load_registry` 가 로드 시 강제 대조, `prompt_overlay_sha256` 는 `registry_validator` 가 대조). 같은 성격의 자산인 SLP overlay 만 manifest 에 싣는 것은 장부 비대칭이며, 봉인의 동형 위치는 **SLP 자체 인덱스에 `prompt_overlay_sha256` 를 기재**하는 것이다. 또한 `runtime_artifact_count` 를 소비·대조하는 코드는 릴리스·YAML 어디에도 없다(grep 0건) — "길이와 일치해야 한다"는 실행 계약이 아니라 장부 자체 일관성 관례다. manifest 등재는 **선택 사항으로 강등**.
**신규 발견(v.1 누락, 별건 플래그) 2건**
- 참조본 A0 가 요구하는 자산 5개 중 **4개가 릴리스에 없다**: `platform/schemas/registry_index.schema.json`, `platform/schemas/domain_config.schema.json`, `platform/reference_catalogs/signal_ids.json`, `platform/reference_catalogs/profile_ids.json`. `load_json` 은 파일 부재 시 `FILE_NOT_FOUND` 로 raise 하므로 **참조본 A0 를 그대로 실행하면 SLP 문제에 도달하기 전에 죽는다.** 실행본 v.8 은 A1(프롬프트 조립)까지 도달했으므로 A0 요구가 참조본과 다르다는 뜻이다 — SLP 와 별개의 참조본↔릴리스 정합 결손으로 기록해 둔다.
- 도메인 인덱스의 `review_items` 는 `ASSEMBLY_INDEPENDENT_REVIEW_PENDING` 1건뿐, SLP 자산 미구축은 기록돼 있지 않다. `closure_gate.pass: true`·`missing_paths: []` 와 함께, **설계 장부가 이 결손을 인지하지 못한 채 닫혀 있었다**는 방증이다(전체 상태 표지는 `program_release_status: STAGE1_NOT_RELEASE_READY` 로 정직).
### 1.2 개선 항목을 반영한 종합 답변
정공법(자산 신설) 결론과 3종 한정, raise 완화 기각, 얇은 조각 설계는 그대로 유지한다 — 기각 논거는 오히려 강해졌다(합성 강제 문구가 E-06 만이 아니라 E-20:13·E-21:13 까지 3개 도메인 전부의 프롬프트 본문에 있다). 다만 신설 범위를 v.1 의 "overlay 3개 + 카탈로그 1개 + manifest 등재"에서 다음 **4점 세트**로 정정한다. ① `Default_Agent/special_law_profiles/{SLP-PRODUCT-LIABILITY, SLP-IP, SLP-MEDIA-MEDIATION}/prompt_overlay.md` 3종 — 유일한 실행 차단 해소분이며, 조문·요율·기간 상수 없이 식별 단서·요건 슬롯·review code·SG-04 경계 선언만 담는 얇은 조각(UTF-8·LF·후행 개행 1, `special_law_profile` rank 40 은 정책에 기등록). ② `special_law_profiles/_registry_index.json` — profile_id·prompt_overlay_path·**prompt_overlay_sha256** 을 기재하는 자산 인덱스. sha 봉인의 동형 위치는 manifest 가 아니라 여기다(도메인 overlay 의 봉인이 도메인 인덱스에 있듯이). 스키마 주석이 전제한 파일이기도 하다. ③ `platform/reference_catalogs/profile_ids.json` — 참조본 A0 가 `--reference-catalog profile=` 로 이미 경로를 확정해 둔 검증용 ID 카탈로그(`calculation_ids.json` 과 동형, `profile_id` 키). 카탈로그는 파일만 만들면 발견되지 않고 validator 호출부 또는 인덱스 `reference_catalogs` 에 배선돼야 한다는 점이 v.1 에서 빠져 있었다. ④ `runtime_manifest.json` 등재는 **하지 않거나, 하려면 장부 정책을 먼저 정한다** — 현행 manifest 는 domain_config·seed_prompt_overlay 를 싣지 않으므로 SLP overlay 만 실으면 비대칭이 된다. "3개 파일 sha 상이" 는 계약이 아니라 권장(dedup 은 도메인별 컴파일 내부 스코프)으로 정정한다. 절단 문자열(`SLP-PRODUCT-LIABI`) 판단은 v.1 그대로 유지한다: 릴리스 내 절단형 0건, 스키마도 통과시키므로 원인은 실행 워크스페이스 `main.py` 생성부에서 찾아야 하고, 어느 쪽이든 자산 신설이 선행 조건이다.
---
## 2. 질문 2 — ★ Insight 1·2번이 실행을 위해 수정해야 하는 문제인가
### 2.1 검증 내용
**Insight 1(인덱스 `[]` ↔ config 값 불일치) — v.1 의 "실행 비차단·위생 작업" 결론 유지, 근거 보강, 표현 1건 정정.**
- 소비 경로 재확인: `load_registry` 는 엔트리를 `config`(domain_config.json 실물 로드본)와 `index_entry`(인덱스 원문)로 분리 저장하고(`registry_loader.py:111-117`), `prompt_compiler.py:228`·`domain_slice_compiler.py:482`·fanout 빌더(참조본 YAML:311) 는 전부 `config` 만 읽는다. `registry_validator` 의 `REFERENCE_FIELDS` 순회도 `config.get(field)` 만 본다. 인덱스의 `special_law_profiles: []` 를 읽는 코드는 **0곳** — 실행 비차단 확정.
- 해시 영향 재실측: 인덱스의 이 필드를 채워도 ① 각 `domain_config.json` 은 무변경이므로 `load_registry` 가 강제 대조하는 `config_sha256` 은 그대로 유효하고, ② 인덱스 파일 자체의 해시는 로드 시점에 `sha256_file` 로 **자기계산**될 뿐 기대값 장부가 어디에도 없다(`runtime_manifest.json` 에 `_registry_index.json` 미등재 실측). **v.1 의 "인덱스 자체 해시만 갱신되면 된다"는 표현은 정정한다 — 갱신할 해시 장부 자체가 존재하지 않으므로, 채움은 아무 해시도 깨뜨리지 않고 아무 갱신도 요구하지 않는다.** 유일한 유보: 실행본 A0 가 `--index-schema` 검증을 수행한다면 채운 값이 그 스키마와 정합해야 하는데, 참조본이 가리키는 `registry_index.schema.json` 실물이 없어 현재 확인 불가.
- 같은 유형의 더 큰 사례(신규): 인덱스가 E-06·E-20·E-21 에 선언한 `profile_schema_path`·`extension_schema_path`·`fixture_manifest_path`·`effect_projection_schema_path` 의 실물이 **12건 전부 MISSING** 이다(실측). 그러나 어떤 실행 경로도 이 4경로를 로드하지 않으므로 역시 비차단 — "인덱스 선언 ≠ 실물"은 SLP 필드 하나가 아니라 인덱스 전반의 상태이며, 위생 작업의 실제 범위는 Insight 1 이 지적한 것보다 넓다.
**Insight 2(스키마 소문자 패턴 결함) — v.1 의 "이미 수정됨·주석만 잔재" 결론 유지, 실효성 보강.**
- 패턴 재검증: `domain_slice.schema.v2.json:104` 의 현재 패턴 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|[a-z][a-z0-9_]{1,127})$` 에 3종 SLP ID 전부 매치(re 실측). 주석(:101)이 서술하는 "소문자 패턴이라 거부" 상태는 현재 파일에 존재하지 않는다.
- 실효성 보강(신규): 이 패턴은 장식이 아니라 **실행 경로에서 실제 소비된다** — 참조본 A2 가 `--slice-schema` 로 이 스키마를 `domain_slice_compiler` 에 전달하고(YAML:654·:686), `schema_subset_validator` 는 `pattern` 키워드를 지원하며 위반 시 `SCHEMA_PATTERN` 을 발행한다(:136-139). 즉 "패턴이 이미 고쳐져 있다"는 사실은 slice 검증 통과 여부를 실제로 좌우하는 확인이고, 수정 대상은 사실과 어긋난 `$comment` 문구(그리고 그 주석이 전제하는 `special_law_profiles/_registry_index.json` 의 실물화 — 질문 1 의 ② 신설로 자연 해소)뿐이다.
### 2.2 개선 항목을 반영한 종합 답변
Insight 1·2번 모두 **stage 1 YAML 실행을 위해 수정해야 하는 문제가 아니라는 v.1 의 결론은 재검증에서 그대로 성립한다.** 1번의 인덱스 `special_law_profiles: []` 는 실행 경로 어디에서도 읽히지 않고(모든 소비자가 `config` 측을 읽는다), 채워 넣더라도 `config_sha256` 은 무변경·인덱스 자체는 기대값 해시 장부가 아예 없어 어떤 봉인도 깨지 않는다 — v.1 이 말한 "인덱스 해시 갱신"은 불필요한 것이 아니라 **개념적으로 존재하지 않는 절차**였다는 점만 정정한다. 아울러 위생 작업의 실제 범위는 이 필드 하나가 아니다: 같은 인덱스가 선언한 profile·extension·fixture·effect_projection 4경로의 실물이 세 도메인 12건 전부 부재한 상태이므로, 인덱스 정합화를 하려면 이 선언들까지 함께 실물화하거나 제거·보류 표기해야 장부가 다시 참말이 된다. 2번은 스키마 패턴이 이미 SLP 대문자 어휘를 받도록 수정돼 있고 그 패턴이 A2 의 `--slice-schema` 경유로 실제 slice 검증에 소비됨을 확인했으므로, 남은 작업은 실행과 무관한 문서 정확성 — 과거 결함 서술로 굳어 있는 `$comment` 를 현재형 사실("어휘는 SLP 카탈로그 소관, 패턴은 수용")로 고쳐 쓰는 것 — 뿐이다. 요컨대 에이전트가 stage 1 을 다시 실행 가능하게 만드는 유일한 수정은 질문 1 의 profile 자산 신설이고, Insight 1·2 는 그 작업에 얹어 처리하면 되는 장부·문서 위생이다.
---
## 3. 총괄 — v.1 대비 변경 요약
| 항목 | v.1 | v.2 판정 |
|---|---|---|
| 자산 신설(정공법)·3종 한정·raise 유지 | 채택 | **유지** (합성 강제 문구 3개 도메인 전부로 논거 강화) |
| overlay 얇은 조각 설계(SG-04 위임) | 채택 | **유지** |
| 카탈로그 위치·배선 | `special_law_profiles/_registry_index.json` "두면 경고 해소" | **정정**: 1차 정본은 `platform/reference_catalogs/profile_ids.json`(A0 가 경로 기확정) + validator 배선 필요. SLP 자체 인덱스는 sha 봉인용으로 별도 |
| sha 봉인 위치 | `runtime_manifest.json` 등재 | **정정**: manifest 는 config·overlay 를 원래 싣지 않음(51건 실측). 봉인은 SLP 인덱스의 `prompt_overlay_sha256` 이 동형. manifest 등재·count 는 선택(외부 소비 0) |
| "3개 파일 sha 상이" 계약 | 하드 계약 | **정정**: dedup 은 도메인별 컴파일 내부 스코프 — 권장으로 강등 |
| Insight 1 실행 비차단 | 채택 | **유지** + "인덱스 해시 갱신" 표현 정정(갱신할 장부 없음) + 위생 범위 확장(선언 4경로 12건 MISSING) |
| Insight 2 실행 비차단 | 채택 | **유지** + 패턴의 실행 소비 확인(A2 `--slice-schema` + `SCHEMA_PATTERN`)으로 실효성 보강 |
| (신규) 참조본 A0 요구 자산 4종 부재 | — | **별건 플래그**: registry_index.schema·domain_config.schema·signal_ids·profile_ids 부재 — 참조본 그대로면 A0 가 `FILE_NOT_FOUND` 로 SLP 이전에 실패 |
@@ -0,0 +1,74 @@
# special_law_profiles 그리고 insight 답변 분석
작성일 2026-08-21 · 검증 대상 `extension_research/assets_special_law_problem.md`(이하 "v.1") · 방법: v.1의 모든 사실 주장과 권고를 릴리스 실물(`Default_Agent/`)·런타임 코드(`stage1_runtime/`)·참조 런타임 YAML(`Claude_YAML/Stage_1_Registry_Runtime_v1.yml`)에 대해 명령 단위로 재실측
**검증 범위의 경계**: 실제 실패한 실행본 `stage_1_part_2_v.8.yml` 은 이 저장소에 없다(`/Users/jsahn/Works/작업결과/v2` 소재). 따라서 "런타임이 기대하는 경로" 류의 주장은 이 폴더의 참조본 YAML 근거이며, 실행본이 동일 규약을 쓴다는 것은 에러 문구(`special-law profile prompt path missing`)가 `prompt_compiler.py:232` 의 raise 문자열과 일치한다는 점으로 뒷받침된다. 이 경계는 v.1이 명시하지 않았던 것으로, 본 리포트에서 명시한다.
---
## 1. 질문 1 — special_law_profiles 자산 처리 방향
### 1.1 검증 내용
**유지 판정(재실측으로 확인된 주장) 5건**
| v.1 주장 | 재실측 결과 | 판정 |
|---|---|---|
| `special_law_profiles/` 실물 부재가 유일한 실행 차단 | `Default_Agent` 하위 해당 디렉터리·파일 0건, `prompt_compiler.py:231-232` 는 `profile_map` 에 키 없으면 무조건 raise | **유지** |
| 신설 범위는 선언된 3종 한정 | config 선언은 E-06·E-20·E-21 의 3종뿐. `SLP-AUTO-DAMAGE`·`SLP-STATE-COMPENSATION` 은 `domain_slice.schema.v2.json:101` **주석 문자열 안에만** 존재 | **유지** |
| raise 완화(2안) 기각 | v.1 은 E-06 근거만 들었으나, 재실측 결과 **3개 도메인 전부** 프롬프트 본문이 합성을 강제한다: E-06 overlay:5·:34("SLP-PRODUCT-LIABILITY 없이 종결하지 않는다"), E-20 overlay:13("**SLP-IP를 반드시 합성**하고 … SG-04로 pin한다"), E-21 overlay:13("**SLP-MEDIA-MEDIATION을 반드시 합성**…"). 완화 시 세 도메인 모두 자기 프롬프트가 금지한 산출물을 내게 된다 | **유지·강화** |
| 경로 계약 `special_law_profiles/<ID>/prompt_overlay.md` + 릴리스 루트 하위 강제 | 참조본 YAML:473(`PROFILE_DIR`)·:492-494(`--profile <ID>=<DIR>/<ID>/prompt_overlay.md`), `compile_fragments` 의 `relative_to(release_root)` 검사 확인 | **유지** |
| overlay 는 조문·요율·기간 상수 없는 얇은 조각(버전 판단은 SG-04 위임) | `ensemble_v2.md:22`("special_law는 도메인 payload 플래그")·`:175`(SG-04 가 "적용 법률·특별법 후보, 시행일·경과규정 … version pinning" 흡수) 재확인. E-20·E-21 overlay:13 도 "SG-04로 pin"을 명시해 위임 구조가 이미 프롬프트에 내장돼 있음 | **유지** |
**정정 판정(v.1 이 부정확했던 주장) 3건**
1. **"3개 파일의 정규화 sha256 이 서로 달라야 한다"(v.1 계약 3) → 하드 계약 아님.** dedup(`duplicate_normalized_sha256`)의 스코프는 `compile_fragments()` **1회 호출 = 도메인 1개 컴파일 내부**다. 각 도메인은 SLP 를 1종만 선언하므로 한 컴파일에 `special_law_profile` 조각은 최대 1개이고, 서로 다른 도메인의 profile 파일끼리는 애초에 같은 dedup 집합에 들어가지 않는다. 실제 제약은 "같은 컴파일 안의 다른 조각(공통 계약·도메인 overlay 등)과 정규화 바이트가 같으면 탈락" 뿐이며 현실적으로 발생하지 않는다. 3개 파일 내용을 서로 다르게 쓰는 것은 여전히 옳지만(내용상 당연), 위반 시 조립이 깨지는 계약은 아니다 — **권장으로 강등**.
2. **"카탈로그 `special_law_profiles/_registry_index.json` 을 함께 두어 `REFERENCE_CATALOG_NOT_PROVIDED` 경고 해소" → 파일 생성만으로는 해소되지 않는다.** `registry_validator` 의 카탈로그 주입 경로는 두 가지뿐이다: ① 도메인 인덱스의 `reference_catalogs`(또는 `reference_sets`) 맵 — 현재 없음, ② CLI `--reference-catalog KIND=PATH`. 파일을 만들어 두기만 하면 validator 는 그 존재를 모른다. 그리고 재실측 결과 **참조본 A0 는 카탈로그의 정본 위치를 이미 확정해 두었다**: `--reference-catalog profile=Default_Agent/platform/reference_catalogs/profile_ids.json`(YAML A0 호출부). 그 디렉터리에 실재하는 것은 `calculation_ids.json`·`domain_ids.json` 뿐이고 `profile_ids.json` 은 없다. 즉 카탈로그의 1차 신설 위치는 v.1 이 말한 `special_law_profiles/_registry_index.json` 이 아니라 **`platform/reference_catalogs/profile_ids.json`** 이며(형식은 `calculation_ids.json` 과 동형: `{"catalog_id": …, "entries": [{"profile_id": "SLP-…", …}]}` — validator 의 `_catalog_ids` 가 `profile_id` 키를 수집한다), `special_law_profiles/_registry_index.json` 은 별도 목적(자산 인덱스 + sha 봉인, 아래 종합 답변)으로 두는 이원 구조가 도메인 registry 와 동형이다.
3. **"`runtime_manifest.json` 에 3~4건 등재하고 count 증분"(v.1 계약 5 포함) → 현행 장부 관례와 어긋나는 권고.** 재실측 결과 manifest 의 `domains/` 51건은 `module_role_projection.json` 25 + `structure_types.json` 25 + `_common/common_worker_contract.md` 1 이 전부로, **`domain_config.json`(26개)과 `seed_prompt_overlay.md`(26개)는 하나도 등재돼 있지 않다.** registry 자산의 무결성 봉인은 manifest 가 아니라 registry 인덱스가 담당한다(`config_sha256` 는 `load_registry` 가 로드 시 강제 대조, `prompt_overlay_sha256` 는 `registry_validator` 가 대조). 같은 성격의 자산인 SLP overlay 만 manifest 에 싣는 것은 장부 비대칭이며, 봉인의 동형 위치는 **SLP 자체 인덱스에 `prompt_overlay_sha256` 를 기재**하는 것이다. 또한 `runtime_artifact_count` 를 소비·대조하는 코드는 릴리스·YAML 어디에도 없다(grep 0건) — "길이와 일치해야 한다"는 실행 계약이 아니라 장부 자체 일관성 관례다. manifest 등재는 **선택 사항으로 강등**.
**신규 발견(v.1 누락, 별건 플래그) 2건**
- 참조본 A0 가 요구하는 자산 5개 중 **4개가 릴리스에 없다**: `platform/schemas/registry_index.schema.json`, `platform/schemas/domain_config.schema.json`, `platform/reference_catalogs/signal_ids.json`, `platform/reference_catalogs/profile_ids.json`. `load_json` 은 파일 부재 시 `FILE_NOT_FOUND` 로 raise 하므로 **참조본 A0 를 그대로 실행하면 SLP 문제에 도달하기 전에 죽는다.** 실행본 v.8 은 A1(프롬프트 조립)까지 도달했으므로 A0 요구가 참조본과 다르다는 뜻이다 — SLP 와 별개의 참조본↔릴리스 정합 결손으로 기록해 둔다.
- 도메인 인덱스의 `review_items` 는 `ASSEMBLY_INDEPENDENT_REVIEW_PENDING` 1건뿐, SLP 자산 미구축은 기록돼 있지 않다. `closure_gate.pass: true`·`missing_paths: []` 와 함께, **설계 장부가 이 결손을 인지하지 못한 채 닫혀 있었다**는 방증이다(전체 상태 표지는 `program_release_status: STAGE1_NOT_RELEASE_READY` 로 정직).
### 1.2 개선 항목을 반영한 종합 답변
정공법(자산 신설) 결론과 3종 한정, raise 완화 기각, 얇은 조각 설계는 그대로 유지한다 — 기각 논거는 오히려 강해졌다(합성 강제 문구가 E-06 만이 아니라 E-20:13·E-21:13 까지 3개 도메인 전부의 프롬프트 본문에 있다). 다만 신설 범위를 v.1 의 "overlay 3개 + 카탈로그 1개 + manifest 등재"에서 다음 **4점 세트**로 정정한다. ① `Default_Agent/special_law_profiles/{SLP-PRODUCT-LIABILITY, SLP-IP, SLP-MEDIA-MEDIATION}/prompt_overlay.md` 3종 — 유일한 실행 차단 해소분이며, 조문·요율·기간 상수 없이 식별 단서·요건 슬롯·review code·SG-04 경계 선언만 담는 얇은 조각(UTF-8·LF·후행 개행 1, `special_law_profile` rank 40 은 정책에 기등록). ② `special_law_profiles/_registry_index.json` — profile_id·prompt_overlay_path·**prompt_overlay_sha256** 을 기재하는 자산 인덱스. sha 봉인의 동형 위치는 manifest 가 아니라 여기다(도메인 overlay 의 봉인이 도메인 인덱스에 있듯이). 스키마 주석이 전제한 파일이기도 하다. ③ `platform/reference_catalogs/profile_ids.json` — 참조본 A0 가 `--reference-catalog profile=` 로 이미 경로를 확정해 둔 검증용 ID 카탈로그(`calculation_ids.json` 과 동형, `profile_id` 키). 카탈로그는 파일만 만들면 발견되지 않고 validator 호출부 또는 인덱스 `reference_catalogs` 에 배선돼야 한다는 점이 v.1 에서 빠져 있었다. ④ `runtime_manifest.json` 등재는 **하지 않거나, 하려면 장부 정책을 먼저 정한다** — 현행 manifest 는 domain_config·seed_prompt_overlay 를 싣지 않으므로 SLP overlay 만 실으면 비대칭이 된다. "3개 파일 sha 상이" 는 계약이 아니라 권장(dedup 은 도메인별 컴파일 내부 스코프)으로 정정한다. 절단 문자열(`SLP-PRODUCT-LIABI`) 판단은 v.1 그대로 유지한다: 릴리스 내 절단형 0건, 스키마도 통과시키므로 원인은 실행 워크스페이스 `main.py` 생성부에서 찾아야 하고, 어느 쪽이든 자산 신설이 선행 조건이다.
---
## 2. 질문 2 — ★ Insight 1·2번이 실행을 위해 수정해야 하는 문제인가
### 2.1 검증 내용
**Insight 1(인덱스 `[]` ↔ config 값 불일치) — v.1 의 "실행 비차단·위생 작업" 결론 유지, 근거 보강, 표현 1건 정정.**
- 소비 경로 재확인: `load_registry` 는 엔트리를 `config`(domain_config.json 실물 로드본)와 `index_entry`(인덱스 원문)로 분리 저장하고(`registry_loader.py:111-117`), `prompt_compiler.py:228`·`domain_slice_compiler.py:482`·fanout 빌더(참조본 YAML:311) 는 전부 `config` 만 읽는다. `registry_validator` 의 `REFERENCE_FIELDS` 순회도 `config.get(field)` 만 본다. 인덱스의 `special_law_profiles: []` 를 읽는 코드는 **0곳** — 실행 비차단 확정.
- 해시 영향 재실측: 인덱스의 이 필드를 채워도 ① 각 `domain_config.json` 은 무변경이므로 `load_registry` 가 강제 대조하는 `config_sha256` 은 그대로 유효하고, ② 인덱스 파일 자체의 해시는 로드 시점에 `sha256_file` 로 **자기계산**될 뿐 기대값 장부가 어디에도 없다(`runtime_manifest.json` 에 `_registry_index.json` 미등재 실측). **v.1 의 "인덱스 자체 해시만 갱신되면 된다"는 표현은 정정한다 — 갱신할 해시 장부 자체가 존재하지 않으므로, 채움은 아무 해시도 깨뜨리지 않고 아무 갱신도 요구하지 않는다.** 유일한 유보: 실행본 A0 가 `--index-schema` 검증을 수행한다면 채운 값이 그 스키마와 정합해야 하는데, 참조본이 가리키는 `registry_index.schema.json` 실물이 없어 현재 확인 불가.
- 같은 유형의 더 큰 사례(신규): 인덱스가 E-06·E-20·E-21 에 선언한 `profile_schema_path`·`extension_schema_path`·`fixture_manifest_path`·`effect_projection_schema_path` 의 실물이 **12건 전부 MISSING** 이다(실측). 그러나 어떤 실행 경로도 이 4경로를 로드하지 않으므로 역시 비차단 — "인덱스 선언 ≠ 실물"은 SLP 필드 하나가 아니라 인덱스 전반의 상태이며, 위생 작업의 실제 범위는 Insight 1 이 지적한 것보다 넓다.
**Insight 2(스키마 소문자 패턴 결함) — v.1 의 "이미 수정됨·주석만 잔재" 결론 유지, 실효성 보강.**
- 패턴 재검증: `domain_slice.schema.v2.json:104` 의 현재 패턴 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|[a-z][a-z0-9_]{1,127})$` 에 3종 SLP ID 전부 매치(re 실측). 주석(:101)이 서술하는 "소문자 패턴이라 거부" 상태는 현재 파일에 존재하지 않는다.
- 실효성 보강(신규): 이 패턴은 장식이 아니라 **실행 경로에서 실제 소비된다** — 참조본 A2 가 `--slice-schema` 로 이 스키마를 `domain_slice_compiler` 에 전달하고(YAML:654·:686), `schema_subset_validator` 는 `pattern` 키워드를 지원하며 위반 시 `SCHEMA_PATTERN` 을 발행한다(:136-139). 즉 "패턴이 이미 고쳐져 있다"는 사실은 slice 검증 통과 여부를 실제로 좌우하는 확인이고, 수정 대상은 사실과 어긋난 `$comment` 문구(그리고 그 주석이 전제하는 `special_law_profiles/_registry_index.json` 의 실물화 — 질문 1 의 ② 신설로 자연 해소)뿐이다.
### 2.2 개선 항목을 반영한 종합 답변
Insight 1·2번 모두 **stage 1 YAML 실행을 위해 수정해야 하는 문제가 아니라는 v.1 의 결론은 재검증에서 그대로 성립한다.** 1번의 인덱스 `special_law_profiles: []` 는 실행 경로 어디에서도 읽히지 않고(모든 소비자가 `config` 측을 읽는다), 채워 넣더라도 `config_sha256` 은 무변경·인덱스 자체는 기대값 해시 장부가 아예 없어 어떤 봉인도 깨지 않는다 — v.1 이 말한 "인덱스 해시 갱신"은 불필요한 것이 아니라 **개념적으로 존재하지 않는 절차**였다는 점만 정정한다. 아울러 위생 작업의 실제 범위는 이 필드 하나가 아니다: 같은 인덱스가 선언한 profile·extension·fixture·effect_projection 4경로의 실물이 세 도메인 12건 전부 부재한 상태이므로, 인덱스 정합화를 하려면 이 선언들까지 함께 실물화하거나 제거·보류 표기해야 장부가 다시 참말이 된다. 2번은 스키마 패턴이 이미 SLP 대문자 어휘를 받도록 수정돼 있고 그 패턴이 A2 의 `--slice-schema` 경유로 실제 slice 검증에 소비됨을 확인했으므로, 남은 작업은 실행과 무관한 문서 정확성 — 과거 결함 서술로 굳어 있는 `$comment` 를 현재형 사실("어휘는 SLP 카탈로그 소관, 패턴은 수용")로 고쳐 쓰는 것 — 뿐이다. 요컨대 에이전트가 stage 1 을 다시 실행 가능하게 만드는 유일한 수정은 질문 1 의 profile 자산 신설이고, Insight 1·2 는 그 작업에 얹어 처리하면 되는 장부·문서 위생이다.
---
## 3. 총괄 — v.1 대비 변경 요약
| 항목 | v.1 | v.2 판정 |
|---|---|---|
| 자산 신설(정공법)·3종 한정·raise 유지 | 채택 | **유지** (합성 강제 문구 3개 도메인 전부로 논거 강화) |
| overlay 얇은 조각 설계(SG-04 위임) | 채택 | **유지** |
| 카탈로그 위치·배선 | `special_law_profiles/_registry_index.json` "두면 경고 해소" | **정정**: 1차 정본은 `platform/reference_catalogs/profile_ids.json`(A0 가 경로 기확정) + validator 배선 필요. SLP 자체 인덱스는 sha 봉인용으로 별도 |
| sha 봉인 위치 | `runtime_manifest.json` 등재 | **정정**: manifest 는 config·overlay 를 원래 싣지 않음(51건 실측). 봉인은 SLP 인덱스의 `prompt_overlay_sha256` 이 동형. manifest 등재·count 는 선택(외부 소비 0) |
| "3개 파일 sha 상이" 계약 | 하드 계약 | **정정**: dedup 은 도메인별 컴파일 내부 스코프 — 권장으로 강등 |
| Insight 1 실행 비차단 | 채택 | **유지** + "인덱스 해시 갱신" 표현 정정(갱신할 장부 없음) + 위생 범위 확장(선언 4경로 12건 MISSING) |
| Insight 2 실행 비차단 | 채택 | **유지** + 패턴의 실행 소비 확인(A2 `--slice-schema` + `SCHEMA_PATTERN`)으로 실효성 보강 |
| (신규) 참조본 A0 요구 자산 4종 부재 | — | **별건 플래그**: registry_index.schema·domain_config.schema·signal_ids·profile_ids 부재 — 참조본 그대로면 A0 가 `FILE_NOT_FOUND` 로 SLP 이전에 실패 |
@@ -0,0 +1,59 @@
# Default_Agent 자산의 프롬프트 크기 limit 삭제 기록
작업일 2026-08-20 · 대상 `v.7/extension_research/Default_Agent/` · 구본 보관 `v.7/extension_research/assets_outdated/`
## 1. 요청과 범위
`Default_Agent/` 자산(json·md·py·txt)에서 `utf8_bytes`·`unicode_scalars`에 걸린 limit만 삭제하고, 그 외 내용은 일절 손대지 않는다. 개정본은 제자리 overwrite, 구본은 `assets_outdated/`에 `_old`를 붙여 보존.
전수 탐색 결과 limit 보유 파일은 **3개뿐**이었다. `.md` 29종에는 이 limit을 서술한 문서가 없었다(검색어: `utf8_bytes` `unicode_scalars` `96000` `24000` `size_guard` `상한` `최대` 등).
| 파일 | 구본 sha(앞16) | 개정본 sha(앞16) | 행수 |
|---|---|---|---|
| `stage1_runtime/prompt_composition_policy.json` | `06dbc4de60ab3052` | `8d3f003ecc5e53eb` | 46 → 44 |
| `stage1_runtime/prompt_compiler.py` | `6922aa5274cea35a` | `d2f0256444f581d4` | 322 → 305 |
| `stage1_runtime/prompt_compiler.txt` | `6922aa5274cea35a` | `d2f0256444f581d4` | 322 → 305 |
`.py`와 `.txt`는 개정 전후 모두 바이트 동일(동일 파일의 확장자 이본).
## 2. 삭제 내역
**정책 JSON — 2행.** `deterministic_prompt_size_guard`에서 상한값 2개만 제거.
```
- "max_utf8_bytes": 96000,
- "max_unicode_scalars": 24000,
```
**컴파일러 py/txt — 17행.** 강제(enforcement) 경로 전부 + 매니페스트의 상한 반영 필드.
- 상한 조회 3행 — `guard = policy.get("deterministic_prompt_size_guard", {})`, `maximum_bytes`, `maximum_scalars`
- 판정·차단 11행 — `exceeded = []`, 두 비교문, `raise RuntimeContractError(... PROMPT_SIZE_GUARD_EXCEEDED ...)` 블록
- 매니페스트 3행 — `max_utf8_bytes`, `max_unicode_scalars`, `"pass": not exceeded`
정책 JSON의 상한을 지워도 컴파일러가 `guard.get("max_utf8_bytes", 96000)` 형태로 **하드코딩 기본값을 갖고 있었기 때문에**, JSON만 고쳤다면 limit이 96,000/24,000으로 그대로 살아 있었을 것이다. 컴파일러 개정이 필수였던 이유다.
## 3. 남긴 것과 근거
| 남긴 것 | 근거 |
|---|---|
| `scalar_count`, `estimate`, 매니페스트 `utf8_bytes`·`unicode_scalars`·`legacy_advisory_estimate` | **계측**이지 limit이 아니다. 관측성 유지 |
| `not_a_tokenizer`, `model_context_fit_not_proven` | 성격 고지 문구 |
| JSON `overflow_fail_code`, `fail_codes.prompt_size_guard_exceeded` | 실패코드 **레지스트리 항목**이지 수치 상한이 아니다. 현재 도달 불가(사문)지만 삭제는 "limit 외 무단절" 제약 위반 |
`"pass": not exceeded`는 limit이 아니라 판정 결과지만, 삭제된 `exceeded`에만 의존해 `NameError` 없이 존치가 불가능했고 `True` 하드코딩은 내용 **추가**가 되므로 삭제가 유일한 준수 선택지였다.
## 4. 검증
편집은 인덱스 사전 검증 후 행 삭제 방식으로만 수행(추가·수정 0). 자체 확인에 이어 독립 sub-agent 1회차 검증 **PASS, 결함 0** — 반복 불필요.
- **삭제 전용 확인** — `git diff --numstat`이 세 파일 모두 `0` insertions(`0 17`, `0 17`, `0 2`), 변경(`c`) 헝크 0
- **부수 피해 없음** — 구본에서 해당 행만 지운 결과가 개정본과 sha256 완전 일치(세 파일 모두)
- **유효성** — JSON 파싱 OK, `py_compile` OK, 삭제 변수(`maximum_bytes`·`maximum_scalars`·`exceeded`·`guard`) 잔존 참조 0
- **기능 실증** — sub-agent가 scratchpad 사본에 한글 40,000자(≈120KB) 조각을 실제로 통과시켜 대조: 구본은 `PROMPT_SIZE_GUARD_EXCEEDED` 예외 발생, 개정본은 정상 완료(`utf8_bytes=120041`, `unicode_scalars=40041` 계측만 기록). limit이 실제로 더 이상 구속하지 않음을 실행으로 확인
- **구본 무결성** — 백업 3본이 `git HEAD` 및 `runtime_manifest.json` 기록과 sha256 일치(3중 대조)
## 5. 후속 확인이 필요한 사항 (이번 회차에서 손대지 않음)
1. **`Default_Agent/runtime_manifest.json` 시효 만료.** 1644–1653행이 세 파일의 **개정 전** sha256을 기록하고 있어 현재 실물과 불일치한다. 무결성 재계산 검사를 돌리면 3건이 실패한다. 갱신은 limit 외 내용 변경이라 제약상 수행하지 않았다 — 매니페스트 재생성 여부는 결정이 필요하다. 현재 `Default_Agent/` 내에 이 파일을 **읽는** 코드는 없어 즉시 장애는 아니다.
2. **형제 배포 트리는 limit 보유 상태 유지.** `ver_8_yaml_candidates/Default_Agent_Stage_1/`, `P1-T2_to_T9_선결작업수행결과/00_배포트리/Default_Agent_Stage_1/`, `선행구축/Default_Agent_Stage_1/` 등에 같은 컴파일러 사본이 있고 limit이 살아 있다. 지정 경로가 `Default_Agent/`였으므로 범위 밖으로 두었다 — 배포 트리까지 일괄 적용할지 확인이 필요하다.
@@ -0,0 +1,203 @@
# special_law_profile 문제 해결 전략서
작성일 2026-08-21 · 근거 문서 `extension_research/assets_special_law_problem_v.2.md` §1.2(4점 세트) · 대상 릴리스 `v.7/extension_research/Default_Agent`
## 0. 목적·범위·전제
`stage_1_part_2_v.8.yml` 실행을 막는 유일한 원인인 `special_law_profiles` 자산 부재를 v.2 §1.2의 확정 전략 — ① overlay 3종 신설, ② SLP 자산 인덱스 신설(sha 봉인), ③ `profile_ids.json` 카탈로그 신설·배선, ④ manifest 등재는 관례 정합 범위만 — 대로 해소하기 위한 실행 계획이다. `prompt_compiler.py:231-232` 의 raise 는 완화하지 않는다(3개 도메인 overlay 가 프롬프트 본문에서 profile 합성을 강제하는 법리 완결성 계약). 참조본 A0 요구 자산 4종 부재(registry_index.schema 등)는 **별건**으로 이 전략서 범위 밖이며, §2의 G 작업에서 실행 워크스페이스 확인 항목으로만 다룬다.
**manifest 정책은 결정 사항이 아니라 관례에서 도출된다(실측)**: 현행 `runtime_manifest.json` 은 `platform/reference_catalogs/calculation_ids.json`·`domain_ids.json` 을 **등재**하고, `domain_config.json`·`seed_prompt_overlay.md`·`domains/_registry_index.json` 은 **미등재**다. 따라서 ③ `profile_ids.json` 은 등재(+1건, count 432→433), ① overlay 와 ② SLP 인덱스는 미등재가 동형이다.
## 1. 작업 흐름도 — DAG (UML activity diagram 표기)
`●` 시작 / `◉` 종료 / `━━` fork·join 동기화 바(바 아래 병렬, 바에서 합류) / `[ ]` activity / `◇` 결정(게이트) / `(선택)` 비필수 병렬 레인
```
●
│
▼
[T0] 파라미터 고정
│ (ID 3종·경로 규약·byte contract·rank 40 재확인)
│
━━━━━━━━━━━━━━━━━━━━━━ fork ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
│ │ │ │ │ │
▼ ▼ ▼ ▼ ▼ ▼
[A1] [A2] [A3] [C] [G] [H](선택)
overlay overlay overlay profile_ids 절단 원인 인덱스·주석
SLP- SLP-IP SLP-MEDIA- .json 작성 조사(실행 위생(Insight
PRODUCT- MEDIATION +manifest 워크스페이스, 1·2 잔여)
LIABILITY 등재(+1) 읽기 전용)
│ │ │ │ │ │
━━━━ join(A1·A2·A3) ━━━━ │ │ │
▼ │ │ │
[B] special_law_profiles/ │ │ │
_registry_index.json │ │ │
(3종 sha256 봉인) │ │ │
│ │ │ │
━━━━━━━━━ join(B·C) ━━━━━━━━━━━━━━━━━ │ │
▼ │ │
━━━━━━━━━━━━━━━ fork ━━━━━━━━━━━━━━━ │ │
│ │ │ │ │
▼ ▼ ▼ │ │
[E1a] [E1b] [E1c] │ │
E-06 컴파일 E-20 컴파일 E-21 컴파일 │ │
스모크 스모크 스모크 │ │
│ │ │ │ │
━━━━━━━━━ join(E1a·E1b·E1c) ━━━ │ │
▼ │ │
[E2] registry_validator │ │
카탈로그 배선 스모크 │ │
▼ │ │
[E3] sha 봉인 대조 │ │
│ │ │
━━━━━━━━━━━━━ join(E3·G·H) ━━━━━━━━━━━━━━━━━━━━━━━━━ │
▼ (H는 미착수여도 통과) ━━━━━━
[F] 기록·문서 반영
│
▼
◇ 게이트: E1~E3 전건 PASS ∧ G 결론 확보
│ yes
▼
[R] stage_1_part_2_v.8.yml 재실행 (실행 워크스페이스)
│
▼
◉
```
**병렬/직렬 구분과 임계 경로**
| 구분 | 작업 | 근거 |
|---|---|---|
| 병렬 1군 (T0 직후 동시 착수) | A1·A2·A3, C, G, H | 상호 입력 의존 없음. C는 ID 목록만 필요(overlay 내용 불요), G는 외부 워크스페이스 읽기 전용 |
| 직렬 | A(3건 전부)→B | B의 sha256 은 overlay **최종 바이트**에서 재계산해야 하므로 A 완료가 선행 |
| 직렬 | (B∧C)→E1→E2→E3 | E1 은 자산 실물 필요, E2 는 카탈로그 필요, E3 은 B 의 봉인값 필요 |
| 합류 | (E3∧G)→F→게이트→R | 절단 원인 미확인 상태로 재실행하면 같은 A1 단계에서 원인 불명 재실패 위험 |
| 임계 경로 | T0→A(최장 overlay)→B→E1→E2→E3→F→R | 법률 콘텐츠 작성(A)이 최장 작업. C·G·H 는 임계 경로 밖 |
## 2. 작업별 세부 내역
### T0 — 파라미터 고정 (직렬 관문, 즉시 완료 가능)
- **목적**: 신설 자산이 지켜야 할 계약값을 문서·코드 실측값으로 고정해 A~E 전 작업의 공통 전제로 삼는다.
- **세부 내역**: ① 대상 ID 3종 확정 — `SLP-PRODUCT-LIABILITY`(E-06), `SLP-IP`(E-20), `SLP-MEDIA-MEDIATION`(E-21). ② 경로 규약 — `Default_Agent/special_law_profiles/<ID>/prompt_overlay.md`(참조본 YAML:473·:492 의 `--profile <ID>=<PROFILE_DIR>/<ID>/prompt_overlay.md` 조립식과 일치, 릴리스 루트 하위 필수). ③ byte contract — UTF-8·LF·후행 개행 정확히 1(`prompt_composition_policy.json` `byte_contract`). ④ 조각 카테고리 `special_law_profile` rank 40 기등록 확인 — **정책 파일 수정 없음**. ⑤ 스키마 패턴 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|…)$` 에 3종 매치 기확인 — **스키마 수정 없음**.
- **완료 기준**: 위 5항이 본 전략서에 기록됨(본 절로 충족).
### A1·A2·A3 — profile overlay 3종 작성 (상호 병렬, 임계 경로)
- **목적**: 유일한 실행 차단 해소분. 도메인 overlay 와 동일 톤의 "얇은 조각" 프롬프트 3건.
- **입력**: 각 도메인의 `seed_prompt_overlay.md`·`domain_config.json`(요건 슬롯·review code 어휘), `ensemble_v2.md` §4.2(SG-04 위임 규약).
- **공통 작성 지침(4절 구성)**: ① **식별 단서** — 이 특별법 체계가 적용 후보가 되는 원천(source-backed) 단서. ② **고유 요건 슬롯·증거 component** — 일반법 골격(각 도메인 overlay 소관)과 겹치지 않는 특별법 고유 쟁점의 슬롯화. ③ **발행할 review code** — 요건 미충족·경계 사안에서 침묵 대신 남길 표지. ④ **SG-04 경계 선언** — 적용 법률·버전·시행일·경과규정 판단은 operand 와 후보만 담아 SG-04 로 위임. **금지**: 법령 조문 번호·요율·기간·한도 상수 기재(도메인 overlay 의 "법정률·기간을 상수로 두지 않는다" 규율과 동일), 최종 결론(`final_conclusion_forbidden` 톤 유지 — 다만 마커 자체는 common contract 가 충족하므로 overlay 에 의무 문구는 아님).
- **개별 내역**:
- **A1 `SLP-PRODUCT-LIABILITY`**: 결함 3유형(제조·설계·표시) 후보 슬롯, 정상 사용·유통 당시 상태·대체원인 증거 component, 결함·인과 추정 구조는 operand 상태로만, 면책 항변 슬롯, E-06 config 의 `E06-N02`(단순 품질불만은 monitor) 경계와 정합. E-06 overlay:5·:34 의 "profile 없이 종결 금지" 규칙의 수신 측임을 명시.
- **A2 `SLP-IP`**: 권리 종류별(특허·저작권·상표·디자인·영업비밀) 특별법 후보 식별, 권리 유효·귀속·보호범위 대 피고 실시행위 대비 슬롯(E-20 config 의 scope/infringement review code 어휘 재사용), 손해액 산정 특칙은 계산 도메인(CE-01·CE-R1)·SG-04 위임. E-20 overlay:13 의 "반드시 합성 + SG-04 pin" 과 정합.
- **A3 `SLP-MEDIA-MEDIATION`**: 언론 보도·매체 유형별 구제수단(정정·반론·추후보도·손해배상) 후보 슬롯, 조정·중재 전치 여부는 절차 경계 review 로만 표지(X3 횡단 소관 침범 금지), 인격권·표현의 자유 형량은 결론 금지 준수. E-21 overlay:13 과 정합.
- **완료 기준(각 건)**: 4절 구성 충족, 금지 사항 0건, UTF-8·LF·후행 개행 1, 상수 검색(`grep -E '제[0-9]+조|[0-9]+%|[0-9]+년'`) 0건.
- **리스크**: 도메인 overlay 와 내용 중복이 크면 정규화 sha 동일로 같은 컴파일 내 dedup 탈락 가능(이론상) — 특별법 고유 층만 담으면 발생하지 않음. 크기 가드는 advisory 전환 상태라 차단 없음.
### B — `special_law_profiles/_registry_index.json` 작성 (A 완료 후 직렬)
- **목적**: SLP 자산 인덱스 + sha 봉인. 도메인 registry 의 `_registry_index.json` 과 동형 구조(봉인의 정위치는 manifest 가 아니라 자산군 자체 인덱스 — v.2 정정 ③). `domain_slice.schema.v2.json:101` 주석이 전제한 파일의 실물화이기도 하다.
- **세부 내역**: 엔트리 3건 — `profile_id`·`prompt_overlay_path`(인덱스 기준 상대경로)·`prompt_overlay_sha256`(**실물 파일 바이트의 sha256_file 값, 옮겨 적기 금지·재계산 원칙** — MEMORY 2026-08-20 교훈 1). `schema_version`·`catalog_id` 헤더 포함. 현재 이 인덱스를 읽는 코드는 없음(사람·후속 validator 용 장부) — validator 확장은 별건.
- **완료 기준**: JSON 파싱 OK, 3건 sha 를 독립 재계산으로 대조 일치(E3 에서 재확인).
### C — `platform/reference_catalogs/profile_ids.json` 작성 (A 와 병렬)
- **목적**: `registry_validator` 참조 무결성 검증용 ID 카탈로그. 참조본 A0 가 `--reference-catalog profile=<이 경로>` 로 이미 확정해 둔 정본 위치(v.2 정정 ②).
- **세부 내역**: `calculation_ids.json` 동형 — `{"catalog_id": "special_law_profiles", "entries": [{"profile_id": "SLP-…", "status": …}, ×3]}`. validator 의 `_catalog_ids` 가 `profile_id` 키를 수집하므로 이 키명이 계약. **배선 주의**: 파일 생성만으로는 발견되지 않는다 — 참조본 A0 는 기배선, 실행본 v.8 의 배선 여부는 G 의 확인 항목. (도메인 인덱스 `reference_catalogs` 맵 신설은 인덱스 수정을 수반하므로 채택하지 않음.) manifest 등재 +1건(카탈로그는 등재군 관례), `runtime_artifact_count` 432→433 — 값 치환은 앵커 기반 행 편집(MEMORY 2026-08-20 교훈 2).
- **완료 기준**: E2 에서 `special_law_profiles` 필드의 `REFERENCE_CATALOG_NOT_PROVIDED` 경고 소멸·`REFERENCE_NOT_FOUND` 0건.
### G — 절단 문자열 원인 조사 (독립 병렬, 외부 워크스페이스, 읽기 전용)
- **목적**: 에러의 `SLP-PRODUCT-LIABI`(17자) 절단 원인 규명. 릴리스 내 절단형 0건 기확인 — 원인은 실행 워크스페이스 측.
- **세부 내역**: `/Users/jsahn/Works/작업결과/v2` 의 Stage YAML 에서 ① `SLP-PRODUCT-LIABI` 문자열 검색, ② `main.py` 생성부의 `{{item_json}}` 직렬화·슬라이싱·길이 제한 검사, ③ v.8 A1 의 `--profile` 조립식이 T0 경로 규약과 일치하는지, ④ v.8 A0 에 profile 카탈로그 배선이 있는지(참조본과 달리 없을 수 있음 — 없으면 E2 는 릴리스 측 검증으로만 유효). **YAML 은 읽기만 하고 수정하지 않는다.**
- **완료 기준**: 절단 발생 지점 특정 또는 "워크스페이스 파일 교체로 소멸" 확인. 게이트(R) 선행 조건.
### H — 장부·문서 위생 (선택 병렬, Insight 1·2 잔여)
- **세부 내역**: ① 도메인 인덱스 E-06·E-20·E-21 엔트리의 `special_law_profiles: []` 를 config 값으로 채움 — 어떤 해시도 깨지 않음(config 무변경·인덱스는 기대값 장부 부재, v.2 §2.1 실측). ② `domain_slice.schema.v2.json:101` `$comment` 를 현재형 사실로 재작성(B 완료로 "실물 부재" 전제도 해소됨). ③ 인덱스 선언 4경로(profile_schema·extension_schema·fixture_manifest·effect_projection_schema) 12건 MISSING 은 실물화가 아니라 **기록만**(별건 결정 대상).
- **완료 기준**: 착수 시에만 적용. 미착수여도 게이트 통과에 영향 없음.
### E1a·E1b·E1c — 컴파일 스모크 (B∧C 후, 상호 병렬)
- **세부 내역(각 도메인)**: `prompt_compiler.main([--registry-index, --domain-id E-06|E-20|E-21, --policy, --common-contract, --profile <ID>=Default_Agent/special_law_profiles/<ID>/prompt_overlay.md, --output-prompt/<scratch>, --output-manifest/<scratch>])` — 실행본 A1 의 호출 형태 재현. 산출 manifest 에서 ① rc 0, ② `special_law_profile` 카테고리 조각 존재·`source_domain_id` 일치, ③ `omitted` 에 profile 조각 없음, ④ `final_conclusion_forbidden` 마커 충족 확인. 회귀 대조로 SLP 비선언 도메인 1건(예: E-05) 컴파일 rc 0 확인.
- **완료 기준**: 3건 전부 4항 PASS. 산출물은 scratch 에만 기록.
### E2 — validator 배선 스모크 (E1 후 직렬)
- **세부 내역**: `registry_validator.main([--index, --reference-catalog profile=platform/reference_catalogs/profile_ids.json, --output/<scratch>])`. 판정: `special_law_profiles` 필드의 closure 가 `catalog: "special_law_profiles"`… 아님 — `REFERENCE_FIELDS` 매핑상 kind 는 `special_law_profiles` 이므로 CLI 는 `--reference-catalog special_law_profiles=<path>` 로 전달해야 한다(**참조본 A0 의 `profile=` kind 표기는 매핑 키와 불일치 — 실측 후 kind 문자열을 `special_law_profiles` 로 맞춰 검증하고, 이 불일치는 F 에서 별건 기록**). 기대: 해당 필드 경고 소멸, `REFERENCE_NOT_FOUND` 0, status 는 다른 결손(calculation·signal 카탈로그 미제공 경고 등)과 무관하게 errors 0 유지.
- **완료 기준**: E-06·E-20·E-21 의 `special_law_profiles` closure 가 `missing: []` 로 기록.
### E3 — sha 봉인 대조 (E2 후 직렬)
- **세부 내역**: B 인덱스의 3개 `prompt_overlay_sha256` 을 실물에서 전수 재계산 대조 + C 등재분 포함 manifest 전수 재계산(기대: 일치 156·불일치 0·부재 277 — 기존 155 에 profile_ids.json +1). MEMORY 2026-08-20 교훈 3(전수 재계산으로 상태 증명) 적용.
- **완료 기준**: 불일치 0.
### F — 기록·문서 반영 (E3∧G∧H 합류 후)
- **세부 내역**: ① `MEMORY.md` 압축 요약 추가. ② v.2 리포트에 후속 상태 부기(신설 자산 목록·검증 결과). ③ 별건 플래그 이관 기록 — 참조본 A0 자산 4종 부재, E2 에서 확인된 kind 표기(`profile=` vs `special_law_profiles=`) 불일치, H-③ 4경로 MISSING. 커밋은 사용자 지시 시에만.
### R — 재실행 게이트 (◇ 통과 후, 실행 워크스페이스)
- **선행 조건**: E1~E3 전건 PASS ∧ G 원인 결론. 워크스페이스 반영은 "신설 자산 복사" 방식(기존 4파일 교체 방식과 동일 절차)으로 하고, `stage_1_part_2_v.8.yml` 재실행으로 A1 통과를 확인한다.
## 3. 완료 판정 총괄 (DoD)
| # | 판정 항목 | 판정 방법 |
|---|---|---|
| 1 | overlay 3종 실재·계약 준수 | A 완료 기준 4항 |
| 2 | 조립 통과 | E1 3건 rc 0 + manifest 4항 |
| 3 | 참조 무결성 closure | E2 `missing: []` |
| 4 | 봉인 정합 | E3 불일치 0 (manifest 156/0/277) |
| 5 | 절단 원인 결론 | G 보고 |
| 6 | 장부 기록 | F 완료 (별건 3건 이관 포함) |
| 7 | 실행 복구 | R 에서 part 2 A1 통과 |
---
# 4. 실행 결과 (2026-08-21 실행)
전략서 §1 DAG 중 **T0 → A1·A2·A3 → B → C → E1 → E2 → E3 → F 를 실행 완료**했다. G 는 이 머신에서 수행 불가, H 는 보류, R 은 실행 워크스페이스 소관으로 미착수다.
## 4.1 신설·변경 자산
| 자산 | 상태 | bytes | sha256 |
|---|---|---|---|
| `special_law_profiles/SLP-PRODUCT-LIABILITY/prompt_overlay.md` | 신설 | 5652 | `78281b81af2e61647024cbd72f9245d14bd3a6750413d16eb9aa093e6778d491` |
| `special_law_profiles/SLP-IP/prompt_overlay.md` | 신설 | 5720 | `7ed12d8183992733e79feb25622e6980642308d0705b8409f6a465394d3b7f13` |
| `special_law_profiles/SLP-MEDIA-MEDIATION/prompt_overlay.md` | 신설 | 5605 | `04ea71d164111934039640486ba113c22793c95168318c38974897d90ffe5a0e` |
| `special_law_profiles/_registry_index.json` | 신설(봉인 장부) | 1478 | `8491e715c47eaa7f0fe749a6356c8c654535f382aadcc08e59c8f3eb93420177` |
| `platform/reference_catalogs/profile_ids.json` | 신설(검증 카탈로그) | 363 | `84f3538229df73e2d7ca58eb32da3d9b38ab1037edb67d4ac43e31e51c20aea3` |
| `runtime_manifest.json` | 변경 | — | `profile_ids.json` 1건 등재, `runtime_artifact_count` 432→433 |
**A 작성 시 반영한 실측 제약 1건(전략서 §2-A 에 없던 것)**: `routing/extension_payload_key_declarations.v1.json` 확인 결과 `special_law_profile_candidates` key 를 선언하는 도메인은 **E-06 뿐**이고, E-20·E-21 은 선언하지 않으며 세 도메인 모두 `additional_properties_allowed: false` 다. 따라서 세 overlay 에 같은 출력 key 를 지시하면 E-20·E-21 worker 출력이 스키마 위반이 된다. 각 overlay 의 §4 는 도메인별 선언 key 만 지목하도록 작성했다(E-06 → `special_law_profile_candidates`, E-20 → `ip_right_candidates` 등 7종, E-21 → `publication_candidates` 등 7종). review 발행 경로도 공통 계약이 정한 후보별 `review_items`·root `unknown_or_unrouted_reviews` 로 한정했다(`review_flags` 등 별도 key 지시 없음).
`_registry_index.json`·overlay 3종은 manifest 에 등재하지 않았다 — 현행 manifest 가 `domain_config.json`·`seed_prompt_overlay.md`·`domains/_registry_index.json` 을 싣지 않는 관례와 동형이며, 봉인은 SLP 인덱스의 `prompt_overlay_sha256` 이 담당한다. 카탈로그(`profile_ids.json`)만 등재한 것은 `calculation_ids.json`·`domain_ids.json` 2/2 가 등재돼 있는 관례를 따른 것이다.
## 4.2 검증 결과 (DoD 대조)
| # | 판정 항목 | 결과 |
|---|---|---|
| 1 | overlay 계약 준수 | **PASS** — 법령 조문·요율·기간 상수 grep 0건, UTF-8·LF·후행 개행 1 (3건 전부) |
| 2 | 조립 통과(E1) | **PASS** — E-06·E-20·E-21 `rc=0`, `special_law_profile` 조각이 rank 40 위치에 1건씩 부착(`source_domain_id` 일치), `deduplicated_fragments` 빈 배열, `final_conclusion_forbidden` 마커 검증. 컴파일 크기 24,960 / 21,841 / 21,807 bytes. 회귀 대조로 SLP 비선언 도메인 E-05 도 `rc=0`(profile 조각 0건) |
| 3 | 참조 무결성(E2) | **PASS** — `--reference-catalog special_law_profiles=…` 주입 시 3개 도메인 closure `missing: []`, `REFERENCE_CATALOG_NOT_PROVIDED` 경고 54→51 감소, errors 0 |
| 4 | 봉인 정합(E3) | **PASS** — SLP 인덱스 3건 실물 재계산 대조 불일치 0, manifest 전수 재계산 **일치 156 · 불일치 0 · 파일없음 277**(직전 회차 155/0/277 에 카탈로그 1건 증가분과 일치) |
| 5 | 절단 원인(G) | **수행 불가** — `/Users/jsahn` 경로가 이 머신에 존재하지 않음(`No such file or directory`). 실행 워크스페이스 보유자만 확인 가능 |
| 6 | 장부 기록(F) | **완료** — 본 절, v.2 리포트 부기, `MEMORY.md` |
| 7 | 실행 복구(R) | **미착수** — 실행 워크스페이스 소관 |
## 4.3 확정된 별건 3건
1. **validator 카탈로그 kind 표기 불일치(실측 확정)**: `registry_validator.REFERENCE_FIELDS` 는 `special_law_profiles` 필드를 kind `special_law_profiles` 로 매핑한다. 참조본 A0 의 `--reference-catalog profile=…` 표기로 실제 주입해 보면 경고가 54 로 유지되고 closure 가 `catalog: null / unverified` 로 남는다 — **즉 그 표기로는 카탈로그가 배선되지 않는다.** `special_law_profiles=` 로 고쳐야 검증이 성립한다. (같은 이유로 `calculation`·`signal` 표기도 각각 `computations`·`signals` 여야 하는지 별도 확인 필요 — `REFERENCE_FIELDS` 값은 `computations`·`signals` 다.)
2. **참조본 A0 요구 자산 잔여 부재 3종**: `platform/schemas/registry_index.schema.json`, `platform/schemas/domain_config.schema.json`, `platform/reference_catalogs/signal_ids.json`. (4종 중 `profile_ids.json` 은 이번 회차로 해소.) `load_json` 이 부재 시 `FILE_NOT_FOUND` 로 raise 하므로 참조본 A0 를 그대로 실행하면 여전히 A0 에서 정지한다.
3. **인덱스 선언 4경로 12건 MISSING**: E-06·E-20·E-21 의 `profile_schema_path`·`extension_schema_path`·`fixture_manifest_path`·`effect_projection_schema_path` 실물 부재. 실행 경로에서 로드되지 않아 비차단.
## 4.4 H(장부·문서 위생) 보류 사유
- **H-① 인덱스 `special_law_profiles` 채움 보류**: 실행 워크스페이스의 `registry_index.schema.json` 을 아직 보지 못했다. 그 스키마가 slice 스키마와 같은 소문자 패턴 결함을 갖고 있으면, 현재 `[]` 로 통과하던 A0 인덱스 검증이 값을 채우는 순간 `REGISTRY_INDEX_SCHEMA_FAILED` 로 **새로 깨진다**. 실행 이득이 0인 반면 회귀 위험이 있으므로 그 스키마 확인 후로 미룬다.
- **H-② 스키마 `$comment` 수정 보류**: `domain_slice.schema.v2.json` 은 manifest 등재 자산이라 수정 시 sha 갱신이 따라붙고, 디버깅 진행 중 실행 워크스페이스 사본과 바이트 드리프트를 만든다. 실행 무관한 문서 수정이므로 H-① 과 함께 일괄 처리한다.
## 4.5 다음 행동 (R, 사용자 소관)
1. 신설 5개 파일(`special_law_profiles/` 4 + `platform/reference_catalogs/profile_ids.json`)과 갱신된 `runtime_manifest.json` 을 실행 워크스페이스 릴리스 트리에 반영한다.
2. `stage_1_part_2_v.8.yml` 을 재실행해 A1(프롬프트 조립) 통과를 확인한다.
3. 실패가 재현되면 에러의 profile ID 표기를 확인한다 — 절단형(`SLP-PRODUCT-LIABI`)이 그대로 나오면 §4.3-1 과 무관한 워크스페이스 `main.py` 생성부 문제이므로 G 를 그때 수행한다.
@@ -0,0 +1,305 @@
#!/usr/bin/env python3
"""Deterministically compose common, domain, dependency, and special-law prompts."""
from __future__ import annotations
import argparse
import json
import math
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Any
from registry_loader import load_registry
from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json
@dataclass(frozen=True)
class FragmentSpec:
fragment_id: str
category: str
path: str
source_domain_id: str | None = None
def _priority(policy: dict[str, Any]) -> dict[str, int]:
out: dict[str, int] = {}
for index, item in enumerate(policy.get("fragment_priority", [])):
if isinstance(item, str):
out[item] = index * 10
elif isinstance(item, dict) and item.get("category"):
out[str(item["category"])] = int(item.get("rank", index * 10))
if not out:
raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty")
return out
def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str:
codes = policy.get("fail_codes")
return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback
def _normalize_fragment(path: str) -> tuple[str, bytes, str]:
source = Path(path)
try:
text = source.read_text(encoding="utf-8")
except FileNotFoundError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc
except UnicodeDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc
if source.suffix.lower() == ".json":
try:
value = json.loads(text)
except json.JSONDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc
normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n"
else:
normalized = normalize_text(text)
data = normalized.encode("utf-8")
return normalized, data, sha256_bytes(data)
def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]:
priorities = _priority(policy)
required_categories = [str(item) for item in policy.get("required_categories", [])]
categories = {spec.category for spec in specs}
missing_categories = [item for item in required_categories if item not in categories]
if missing_categories:
raise RuntimeContractError(
_fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"),
"required prompt fragment category missing",
missing_categories,
)
unknown = sorted(categories - set(priorities))
if unknown:
raise RuntimeContractError(
_fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"),
"fragment category has no priority",
unknown,
)
prepared: list[dict[str, Any]] = []
by_id: dict[str, str] = {}
for spec in specs:
text, data, digest = _normalize_fragment(spec.path)
previous = by_id.get(spec.fragment_id)
if previous is not None and previous != digest:
raise RuntimeContractError(
_fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"),
f"same fragment_id has different normalized hashes: {spec.fragment_id}",
{"first": previous, "second": digest},
)
by_id[spec.fragment_id] = digest
prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest})
prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path))
seen_sha: set[str] = set()
kept: list[dict[str, Any]] = []
omitted: list[dict[str, Any]] = []
for item in prepared:
spec = item["spec"]
source_path = Path(spec.path).resolve()
release_root = Path(__file__).resolve().parents[1]
try:
logical_path = source_path.relative_to(release_root).as_posix()
except ValueError as exc:
raise RuntimeContractError(
"PROMPT_FRAGMENT_PATH_INVALID",
"prompt fragment must resolve under the release root",
source_path.name,
) from exc
public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id}
if item["sha256"] in seen_sha:
omitted.append({**public, "reason": "duplicate_normalized_sha256"})
continue
seen_sha.add(item["sha256"])
kept.append(item)
separator = str(policy.get("byte_contract", {}).get("separator", "\n\n"))
compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n"
compiled_bytes = compiled.encode("utf-8")
required_markers = [str(item) for item in policy.get("required_markers", [])]
missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()]
if missing_markers:
raise RuntimeContractError(
_fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"),
"compiled prompt lacks required guard marker",
missing_markers,
)
for conflict in policy.get("conflict_markers", []):
if isinstance(conflict, list) and len(conflict) > 1:
present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()]
if len(present) > 1:
raise RuntimeContractError(
_fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"),
"mutually exclusive prompt markers coexist",
present,
)
scalar_count = len(compiled)
estimate = math.ceil(len(compiled_bytes) / 4)
manifest = {
"schema_version": "compiled_prompt_manifest.v1",
"compiled_prompt_sha256": sha256_bytes(compiled_bytes),
"compiled_prompt_bytes": len(compiled_bytes),
"deterministic_prompt_size_guard": {
"utf8_bytes": len(compiled_bytes),
"unicode_scalars": scalar_count,
"not_a_tokenizer": True,
"model_context_fit_not_proven": True,
"legacy_advisory_estimate": estimate,
},
"normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator},
"fragments": [
{
"fragment_id": item["spec"].fragment_id,
"category": item["spec"].category,
"path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()),
"sha256": item["sha256"],
"hash_kind": "file_sha256",
"source_domain_id": item["spec"].source_domain_id,
}
for item in kept
],
"deduplicated_fragments": omitted,
"required_markers_verified": required_markers,
}
return compiled, manifest
def _path_from_config(config_record: dict[str, Any], value: str) -> str:
config_path = config_record.get("config_path")
base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", "."))
path = Path(value)
return str(path if path.is_absolute() else (base / path).resolve())
def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]:
config = config_record["config"]
values: list[Any] = []
if isinstance(config.get("prompt_fragments"), list):
values.extend(config["prompt_fragments"])
for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"):
if isinstance(config.get(key), str):
values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"})
index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {}
if not values and isinstance(index_entry.get("prompt_overlay_path"), str):
values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"})
specs: list[FragmentSpec] = []
for index, value in enumerate(values, 1):
if isinstance(value, str):
path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}"
elif isinstance(value, dict) and isinstance(value.get("path"), str):
path_value = value["path"]
fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}")
else:
continue
specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id))
return specs
def collect_domain_fragments(
domain_id: str,
registry: dict[str, Any],
*,
common_contract_path: str,
profile_paths: dict[str, str] | None = None,
runtime_guard_path: str | None = None,
extra_specs: list[FragmentSpec] | None = None,
) -> list[FragmentSpec]:
entries = registry["entries"]
if domain_id not in entries:
raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}")
specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))]
visited: set[str] = set()
def add_dependencies(current: str) -> None:
for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key):
if dependency not in entries:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}")
if dependency in visited:
continue
visited.add(dependency)
add_dependencies(dependency)
specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency))
add_dependencies(domain_id)
domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id)
if not domain_specs:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}")
specs.extend(domain_specs)
profiles = entries[domain_id]["config"].get("special_law_profiles", [])
profile_map = profile_paths or {}
for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key):
if profile not in profile_map:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}")
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id))
if runtime_guard_path:
specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve())))
specs.extend(extra_specs or [])
return specs
def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]:
out: dict[str, str] = {}
for value in values:
if separator not in value:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value)
key, path = value.split(separator, 1)
out[key] = path
return out
def _parse_extra(values: list[str]) -> list[FragmentSpec]:
out: list[FragmentSpec] = []
for value in values:
parts = value.split(":", 2)
if len(parts) != 3:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value)
out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve())))
return out
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--registry-index", required=True)
parser.add_argument("--domain-id", required=True)
parser.add_argument("--policy", required=True)
parser.add_argument("--common-contract", required=True)
parser.add_argument("--runtime-guard")
parser.add_argument("--profile", action="append", default=[])
parser.add_argument("--fragment", action="append", default=[])
parser.add_argument("--output-prompt", required=True)
parser.add_argument("--output-manifest", required=True)
args = parser.parse_args(argv)
try:
registry = load_registry(args.registry_index)
specs = collect_domain_fragments(
args.domain_id,
registry,
common_contract_path=args.common_contract,
profile_paths=_parse_mapping(args.profile),
runtime_guard_path=args.runtime_guard,
extra_specs=_parse_extra(args.fragment),
)
policy = load_json(args.policy)
compiled, manifest = compile_fragments(specs, policy)
prompt_path = Path(args.output_prompt)
prompt_path.parent.mkdir(parents=True, exist_ok=True)
prompt_path.write_text(compiled, encoding="utf-8", newline="")
from runtime_common import sha256_file
manifest.update({
"domain_id": args.domain_id,
"compiled_prompt_path": str(prompt_path),
"registry_index_sha256": registry["index_sha256"],
"composition_policy_path": args.policy,
"composition_policy_sha256": sha256_file(args.policy),
})
write_json(args.output_manifest, manifest)
emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]})
return 0
except RuntimeContractError as exc:
emit_cli_result({"status": "FAILED", "error": exc.as_dict()})
return 1
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,305 @@
#!/usr/bin/env python3
"""Deterministically compose common, domain, dependency, and special-law prompts."""
from __future__ import annotations
import argparse
import json
import math
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Any
from registry_loader import load_registry
from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json
@dataclass(frozen=True)
class FragmentSpec:
fragment_id: str
category: str
path: str
source_domain_id: str | None = None
def _priority(policy: dict[str, Any]) -> dict[str, int]:
out: dict[str, int] = {}
for index, item in enumerate(policy.get("fragment_priority", [])):
if isinstance(item, str):
out[item] = index * 10
elif isinstance(item, dict) and item.get("category"):
out[str(item["category"])] = int(item.get("rank", index * 10))
if not out:
raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty")
return out
def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str:
codes = policy.get("fail_codes")
return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback
def _normalize_fragment(path: str) -> tuple[str, bytes, str]:
source = Path(path)
try:
text = source.read_text(encoding="utf-8")
except FileNotFoundError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc
except UnicodeDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc
if source.suffix.lower() == ".json":
try:
value = json.loads(text)
except json.JSONDecodeError as exc:
raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc
normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n"
else:
normalized = normalize_text(text)
data = normalized.encode("utf-8")
return normalized, data, sha256_bytes(data)
def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]:
priorities = _priority(policy)
required_categories = [str(item) for item in policy.get("required_categories", [])]
categories = {spec.category for spec in specs}
missing_categories = [item for item in required_categories if item not in categories]
if missing_categories:
raise RuntimeContractError(
_fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"),
"required prompt fragment category missing",
missing_categories,
)
unknown = sorted(categories - set(priorities))
if unknown:
raise RuntimeContractError(
_fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"),
"fragment category has no priority",
unknown,
)
prepared: list[dict[str, Any]] = []
by_id: dict[str, str] = {}
for spec in specs:
text, data, digest = _normalize_fragment(spec.path)
previous = by_id.get(spec.fragment_id)
if previous is not None and previous != digest:
raise RuntimeContractError(
_fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"),
f"same fragment_id has different normalized hashes: {spec.fragment_id}",
{"first": previous, "second": digest},
)
by_id[spec.fragment_id] = digest
prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest})
prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path))
seen_sha: set[str] = set()
kept: list[dict[str, Any]] = []
omitted: list[dict[str, Any]] = []
for item in prepared:
spec = item["spec"]
source_path = Path(spec.path).resolve()
release_root = Path(__file__).resolve().parents[1]
try:
logical_path = source_path.relative_to(release_root).as_posix()
except ValueError as exc:
raise RuntimeContractError(
"PROMPT_FRAGMENT_PATH_INVALID",
"prompt fragment must resolve under the release root",
source_path.name,
) from exc
public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id}
if item["sha256"] in seen_sha:
omitted.append({**public, "reason": "duplicate_normalized_sha256"})
continue
seen_sha.add(item["sha256"])
kept.append(item)
separator = str(policy.get("byte_contract", {}).get("separator", "\n\n"))
compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n"
compiled_bytes = compiled.encode("utf-8")
required_markers = [str(item) for item in policy.get("required_markers", [])]
missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()]
if missing_markers:
raise RuntimeContractError(
_fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"),
"compiled prompt lacks required guard marker",
missing_markers,
)
for conflict in policy.get("conflict_markers", []):
if isinstance(conflict, list) and len(conflict) > 1:
present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()]
if len(present) > 1:
raise RuntimeContractError(
_fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"),
"mutually exclusive prompt markers coexist",
present,
)
scalar_count = len(compiled)
estimate = math.ceil(len(compiled_bytes) / 4)
manifest = {
"schema_version": "compiled_prompt_manifest.v1",
"compiled_prompt_sha256": sha256_bytes(compiled_bytes),
"compiled_prompt_bytes": len(compiled_bytes),
"deterministic_prompt_size_guard": {
"utf8_bytes": len(compiled_bytes),
"unicode_scalars": scalar_count,
"not_a_tokenizer": True,
"model_context_fit_not_proven": True,
"legacy_advisory_estimate": estimate,
},
"normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator},
"fragments": [
{
"fragment_id": item["spec"].fragment_id,
"category": item["spec"].category,
"path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()),
"sha256": item["sha256"],
"hash_kind": "file_sha256",
"source_domain_id": item["spec"].source_domain_id,
}
for item in kept
],
"deduplicated_fragments": omitted,
"required_markers_verified": required_markers,
}
return compiled, manifest
def _path_from_config(config_record: dict[str, Any], value: str) -> str:
config_path = config_record.get("config_path")
base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", "."))
path = Path(value)
return str(path if path.is_absolute() else (base / path).resolve())
def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]:
config = config_record["config"]
values: list[Any] = []
if isinstance(config.get("prompt_fragments"), list):
values.extend(config["prompt_fragments"])
for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"):
if isinstance(config.get(key), str):
values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"})
index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {}
if not values and isinstance(index_entry.get("prompt_overlay_path"), str):
values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"})
specs: list[FragmentSpec] = []
for index, value in enumerate(values, 1):
if isinstance(value, str):
path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}"
elif isinstance(value, dict) and isinstance(value.get("path"), str):
path_value = value["path"]
fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}")
else:
continue
specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id))
return specs
def collect_domain_fragments(
domain_id: str,
registry: dict[str, Any],
*,
common_contract_path: str,
profile_paths: dict[str, str] | None = None,
runtime_guard_path: str | None = None,
extra_specs: list[FragmentSpec] | None = None,
) -> list[FragmentSpec]:
entries = registry["entries"]
if domain_id not in entries:
raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}")
specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))]
visited: set[str] = set()
def add_dependencies(current: str) -> None:
for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key):
if dependency not in entries:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}")
if dependency in visited:
continue
visited.add(dependency)
add_dependencies(dependency)
specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency))
add_dependencies(domain_id)
domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id)
if not domain_specs:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}")
specs.extend(domain_specs)
profiles = entries[domain_id]["config"].get("special_law_profiles", [])
profile_map = profile_paths or {}
for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key):
if profile not in profile_map:
raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}")
specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id))
if runtime_guard_path:
specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve())))
specs.extend(extra_specs or [])
return specs
def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]:
out: dict[str, str] = {}
for value in values:
if separator not in value:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value)
key, path = value.split(separator, 1)
out[key] = path
return out
def _parse_extra(values: list[str]) -> list[FragmentSpec]:
out: list[FragmentSpec] = []
for value in values:
parts = value.split(":", 2)
if len(parts) != 3:
raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value)
out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve())))
return out
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--registry-index", required=True)
parser.add_argument("--domain-id", required=True)
parser.add_argument("--policy", required=True)
parser.add_argument("--common-contract", required=True)
parser.add_argument("--runtime-guard")
parser.add_argument("--profile", action="append", default=[])
parser.add_argument("--fragment", action="append", default=[])
parser.add_argument("--output-prompt", required=True)
parser.add_argument("--output-manifest", required=True)
args = parser.parse_args(argv)
try:
registry = load_registry(args.registry_index)
specs = collect_domain_fragments(
args.domain_id,
registry,
common_contract_path=args.common_contract,
profile_paths=_parse_mapping(args.profile),
runtime_guard_path=args.runtime_guard,
extra_specs=_parse_extra(args.fragment),
)
policy = load_json(args.policy)
compiled, manifest = compile_fragments(specs, policy)
prompt_path = Path(args.output_prompt)
prompt_path.parent.mkdir(parents=True, exist_ok=True)
prompt_path.write_text(compiled, encoding="utf-8", newline="")
from runtime_common import sha256_file
manifest.update({
"domain_id": args.domain_id,
"compiled_prompt_path": str(prompt_path),
"registry_index_sha256": registry["index_sha256"],
"composition_policy_path": args.policy,
"composition_policy_sha256": sha256_file(args.policy),
})
write_json(args.output_manifest, manifest)
emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]})
return 0
except RuntimeContractError as exc:
emit_cli_result({"status": "FAILED", "error": exc.as_dict()})
return 1
if __name__ == "__main__":
sys.exit(main())
@@ -238,6 +238,7 @@ Agent:
POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json"
SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json"
FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json"
SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json"
# F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다.
# S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로
@@ -253,6 +254,7 @@ Agent:
"Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0
"Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러
"Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책
"Default_Agent/stage1_runtime/seed_admission_policy.v1.json", # R0 — 시드 수용 정책
)
RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1"
# registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다.
@@ -464,6 +466,39 @@ Agent:
stage_text(overlay_logical, read_raw(overlay_logical))
except Exception as exc:
warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc))
# 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는
# domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서
# 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도
# PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다).
# 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다.
profile_paths = {}
try:
slp_index_raw = read_raw(SPECIAL_LAW_INDEX)
except Exception as exc:
warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc))
else:
stage_text(SPECIAL_LAW_INDEX, slp_index_raw)
slp_doc = json.loads(slp_index_raw)
slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc)
slp_base = posixpath.dirname(SPECIAL_LAW_INDEX)
for entry in slp_index.get("entries") or []:
profile_id = entry.get("profile_id")
overlay_path = entry.get("prompt_overlay_path")
if not isinstance(profile_id, str) or not profile_id:
continue
if not isinstance(overlay_path, str) or not overlay_path:
warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id))
continue
overlay_logical = unicodedata.normalize(
"NFC", overlay_path if overlay_path.startswith("Default_Agent/")
else posixpath.join(slp_base, overlay_path))
try:
stage_text(overlay_logical, read_raw(overlay_logical))
except Exception as exc:
warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc))
continue
profile_paths[profile_id] = overlay_logical
stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT))
stage_text(POLICY, read_raw(POLICY))
@@ -536,7 +571,7 @@ Agent:
for domain_id in expected_runnable:
specs = prompt_compiler.collect_domain_fragments(
domain_id, registry, common_contract_path=COMMON_CONTRACT,
extra_specs=vocabulary_specs)
profile_paths=profile_paths, extra_specs=vocabulary_specs)
text, manifest_row = prompt_compiler.compile_fragments(specs, policy)
rel = "%s/%s.md" % (PROMPT_DIR, domain_id)
stage_text(rel, text)
@@ -579,9 +614,48 @@ Agent:
if isinstance(domain_id, str) and domain_id and codes:
screening_calc[domain_id] = sorted(set(codes))
# P-2a — stage_a_context 를 slice 컴파일보다 먼저 만든다. slice 의 source_universe 가
# 인용할 event 식별자의 정본이 event_candidate_map 이기 때문이다. 순서가 뒤였을 때
# A0 는 원시 evidence_event_candidates 문서를 넘겼고, domain_slice_compiler._records 가
# 그 문서-단위 items(30건)를 후보로 오인해 candidate_id 를 못 찾아
# event:unidentified:NNNN 로 대체했다. 그 값은 source_universe_manifest 의
# event_candidate_ids(EVT-...-NN, 84건)와 교집합이 0 이라, 워커가 규율을 지켜
# slice 안의 것만 인용해도 R0 가 "source_refs outside Stage A universe" 로 차단했다.
meeting_raw = read_raw(MEETING)
created_at_utc = utc_now()
input_digests = {
MEETING: sha_text(meeting_raw),
EVIDENCE: sha_text(read_raw(EVIDENCE)),
EVENTS: sha_text(read_raw(EVENTS)),
SCREENING: seal["screening_sha256"],
ACTIVATION_MANIFEST: seal["activation_manifest_sha256"],
REGISTRY_INDEX: seal["registry_index_sha256"],
}
stage_a = stage_a_context_builder.build_stage_a_context(
meeting_text=meeting_raw,
evidence_obj=evidence_document,
event_obj=events_document,
input_digests_sha256=input_digests,
created_at_utc=created_at_utc,
digest_guard=seal,
expected_runnable_domain_ids=expected_runnable)
source_manifest = stage_a_context_builder.build_source_universe_manifest(
stage_a, input_digests_sha256=input_digests,
registry_index_sha256=seal["registry_index_sha256"])
# 맵을 통째로 넘기지 않는다 — domain_slice_compiler._records 는 dict 를 받으면
# by_evidence_index(값이 id 리스트)에 먼저 걸려 빈 목록을 돌려준다. 후보 레코드
# 목록으로 평탄화해 넘겨야 _record_id 가 candidate_id 를 찾아 EVT-...-NN 을 쓴다.
event_candidate_records = [
row for row in ((stage_a.get("event_candidate_map") or {}).get("by_candidate_id") or {}).values()
if isinstance(row, dict)]
if not event_candidate_records:
raise RuntimeError(json.dumps({
"reason_code": "PART2_EVENT_CANDIDATE_MAP_EMPTY",
"message": "stage_a event_candidate_map.by_candidate_id 가 비었다. slice 의 event 식별자를 정본으로 실을 수 없다.",
}, ensure_ascii=False))
try:
result = domain_slice_compiler.compile_domain_slices(
manifest, registry, evidence_document, events_document,
manifest, registry, evidence_document, event_candidate_records,
manifest_sha256=seal["activation_manifest_sha256"],
evidence_sha256=sha_text(read_raw(EVIDENCE)),
events_sha256=sha_text(read_raw(EVENTS)),
@@ -641,27 +715,7 @@ Agent:
# P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다.
# v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다.
meeting_raw = read_raw(MEETING)
created_at_utc = utc_now()
input_digests = {
MEETING: sha_text(meeting_raw),
EVIDENCE: sha_text(read_raw(EVIDENCE)),
EVENTS: sha_text(read_raw(EVENTS)),
SCREENING: seal["screening_sha256"],
ACTIVATION_MANIFEST: seal["activation_manifest_sha256"],
REGISTRY_INDEX: seal["registry_index_sha256"],
}
stage_a = stage_a_context_builder.build_stage_a_context(
meeting_text=meeting_raw,
evidence_obj=evidence_document,
event_obj=events_document,
input_digests_sha256=input_digests,
created_at_utc=created_at_utc,
digest_guard=seal,
expected_runnable_domain_ids=expected_runnable)
source_manifest = stage_a_context_builder.build_source_universe_manifest(
stage_a, input_digests_sha256=input_digests,
registry_index_sha256=seal["registry_index_sha256"])
# 계산은 P-2a 에서 이미 끝났다(slice 가 같은 식별자를 써야 하므로 앞당겼다). 여기서는 기록만 한다.
write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a}))
write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest))
# P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다.
@@ -707,6 +761,10 @@ Agent:
"expected_runnable_domain_ids": expected_runnable,
"slice_count": len(slice_hashes),
"fanout_instance_count": len(plan_root.get("task_instances") or []),
# 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은
# 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다
# (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다.
"dynamic_fanout": plan_root.get("task_instances") or [],
"digest_guard": seal,
"errors": ERRORS, "warnings": WARNINGS}
@@ -723,13 +781,17 @@ Agent:
- "{{item.compiled_prompt_path}}"
- "{{item.slice_path}}"
llm_provider: google
llm_model: 'gemini-3.1-flash-lite'
llm_model: 'gemini-3.1-pro-preview'
llm_reasoning: high
llm_verbosity: medium
llm_verbosity: low
use_tools:
- localdocs
cache_control:
mode: auto
ttl: 15m
prompt: |-
prompts:
- role: user
content: |-
<TASK_META>
<TASK_NAME>Task_C_B_domain_worker</TASK_NAME>
<ROLE>
@@ -798,6 +860,7 @@ Agent:
{
"seed_id": "<도메인슬러그-001 꼴>",
"bo_type": "<slice.allowed_legal_effect_bo_types 안의 값>",
"juristic_act_type": "<법률행위 유형 문자열 또는 null>",
"source_refs": [],
"registry_component_ids": [],
"element_fact_candidates": [],
@@ -807,8 +870,12 @@ Agent:
"calculation_requests": [],
"dependency_refs": [],
"legal_effect_candidates": [],
"party_roles": [],
"time_facts": [],
"object_refs": [],
"amount_facts": [],
"review_items": [],
"extensions": {}
"extensions": {"domain_payload": {"action_summary": null, "action_type": null}}
}
],
"unknown_or_unrouted_reviews": [],
@@ -828,6 +895,51 @@ Agent:
- `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다.
- 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다.
억지로 채우는 것이 비워 두는 것보다 나쁘다.
- 아래 자리들은 BO 호환면 투영(`bo_surface_projection_policy.v1`)이 읽는 1순위 출처다.
비워 두면 BO.json 의 해당 칸이 폴백 값으로 채워지고 schema_field_fallback 검토가 발행된다.
slice 의 source_universe 안에 근거가 있으면 채운다. 근거가 없으면 비워 두고 사유를 남긴다 —
추측으로 채우지 않는다. 사건종류 이름을 값으로 쓰지 않는다.
· `juristic_act_type` : 법률행위 유형 문자열 1개(없으면 null). -> JuristicAct.label
· `extensions.domain_payload.action_summary` : 이 후보가 무엇인지 한 문장. -> Action
· `extensions.domain_payload.action_type` : "법률행위(legal acts)" 또는 "사실행위(factual acts)". -> ActionType
· `legal_effect_candidates[]` : {"type_id": "<소문자_스네이크>", "source_refs": [], "registered": true|false}. -> Legal_Keywords
· `time_facts[]` : {"fact_type": "<소문자_스네이크>", "value": "<시점 문자열 또는 null>", "source_refs": []}. -> BehaviorTime · TimeText
· `object_refs[]` : 목적물 식별자 문자열. -> core_field_base.Object
· `amount_facts[]` : {"amount_type": "<소문자_스네이크>", "decimal_value": "<숫자 문자열 또는 null>", "currency": "KRW", "source_refs": []}. -> amount
· `party_roles[]` : {"role": "<소문자_스네이크>", "party_refs": []}. 투영 대상은 아니나 스키마 필드다.
- `slice_sha256` 와 `compiled_prompt_sha256` 은 골격이 준 값을 **글자 그대로** 옮기는 자리다.
슬라이스 본문을 뒤져 비슷한 이름을 찾지 않는다. 슬라이스에서 읽어 오는 sha 는
`registry_index_sha256` 과 `domain_config_sha256` 둘뿐이다.
특히 `activation.activation_manifest_sha256` 은 전 도메인이 같은 값이라 혼동하기 쉽다 —
그것을 집으면 R0 가 전사 실수로 기록하고 이 산출에 검토 표시를 남긴다.
다른 도메인의 sha 를 집으면 "남의 자료로 작업했다"로 판정되어 회차가 중단된다.
그 둘 어디에도 해당하지 않는 값(어디서 왔는지 설명되지 않는 sha)을 적으면
더 무겁게 다뤄져 이 후보가 격리된다. 값을 지어내는 것이 가장 나쁘다.
- 아래 다섯 어휘는 스키마가 고정한 것이다. 다른 낱말을 쓰면 S0 신호 게이트가 경성으로 막는다.
R0 의 검증기는 스키마의 부분집합만 보므로 여기서 틀려도 그 단계에서는 걸리지 않는다.
· seed 최상위 `status` : READY | READY_WITH_REVIEW | NO_SUPPORT | BLOCKED | FAILED
· `evidence_slot_status[].status` : filled | partial | missing | conflicted
(요건 슬롯을 뒷받침하는 근거가 충분하면 filled, 일부만이면 partial,
없으면 missing, 상충 근거가 함께 있으면 conflicted)
· `calculation_requests[].completeness` : ready | partial | blocked | deferred
· `review_items[].severity` 와 `unknown_or_unrouted_reviews[].severity` : info | review | hard_warning | block
· 세 후보 배열의 `source_kind` : meeting_clause | event_candidate | evidence | bo | fact
| signal | registry | law_version | calculation | other
- 아래 네 객체는 `additionalProperties: false` 다. 적힌 키 말고는 **한 개도** 넣지 않는다.
필수 키를 빠뜨리거나 임의 키를 더하면 스키마 위반이다.
· `element_fact_candidates[]` · `opposing_fact_candidates[]` · `defense_candidates[]` :
{"source_id": "<slice source_universe 의 source_id>", "source_kind": "<위 어휘>",
"excerpt": "<선택: 근거 문구>", "payload": {}}
— 필수는 source_id · source_kind 둘이다. `slot_id` 나 `fact` 같은 키는 이 배열에 없다.
슬롯 판정은 `evidence_slot_status[]` 가 맡는다.
· `evidence_slot_status[]` :
{"slot_id": "<요건 슬롯 id>", "status": "<위 어휘>", "source_refs": [], "review_code": null}
· `calculation_requests[]` :
{"calculation_domain": "<CE-** 계산 도메인 코드>", "completeness": "<위 어휘>",
"source_refs": [], "review_code": null}
· `review_items[]` :
{"review_code": "<대문자_스네이크>", "severity": "<위 어휘>",
"reason": "<왜 검토가 필요한지 한 문장>", "source_refs": []}
</OUTPUT_SCHEMA>
<VALIDATION_GATES>
@@ -861,6 +973,7 @@ Agent:
# publisher + domain_join + PostB_1 통합 결정적 reducer.
# Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3)
from __future__ import annotations
import copy
import hashlib
import itertools
import json
@@ -881,6 +994,11 @@ Agent:
FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json"
SLICE_DIR = "runtime/domain_slices"
SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json"
# R0-6 — 전체 스키마 검증은 신설이 아니다. worker_output_validator 가 이미
# schema_subset_validator 로 전 계약을 검사하고 있었고(13도메인 1,151건 실측),
# R0 가 그 결과를 guard_warnings 로 강등하고 있었다. 이 정책은 그 결과에 처분을 준다.
SEED_ADMISSION_POLICY_PATH = "Default_Agent/stage1_runtime/seed_admission_policy.v1.json"
SEED_ADMISSION_SCHEMA_VERSION = "stage1_seed_admission_policy.v1"
SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3"
SLICE_ROOT_KEY = "stage_b_domain_slice"
# R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5.
@@ -1195,6 +1313,12 @@ Agent:
norm = _dict(policy.get("f0_normalization"))
bo_type = cand.get("bo_type")
if not isinstance(bo_type, str):
# 워커가 문자열이 아닌 값을 넣으면 집합 비교가 터진다(dict 면 unhashable).
# 수용 단계가 걸러야 하지만 투영은 마지막 방어선이라 여기서도 막는다 —
# 여기서 죽으면 회차 전체가 죽고, 원인이 워커 값이라는 것도 드러나지 않는다.
note("schema_field_fallback", "BOType", "non_string:%s" % type(bo_type).__name__)
bo_type = norm.get("bo_type_default")
if allowed_bo_types and bo_type not in allowed_bo_types:
note("legal_effect_uncertain", "BOType", "bo_type")
@@ -1297,7 +1421,7 @@ Agent:
}
# ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ----------
def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None:
def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]], echo_mismatches: list[dict[str, Any]] | None = None) -> None:
"""v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다.
구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11),
@@ -1319,11 +1443,20 @@ Agent:
raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates")
elif status not in ("READY", "READY_WITH_REVIEW"):
raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}")
# R0-7 — 신선도를 워커의 자기 신고로 판정하지 않는다.
# echo 된 sha 는 모델이 옮겨 적은 값이라 전사 실수가 곧 회차 사망이 된다.
# 실측: E-00 이 slice_sha256 자리에 슬라이스 안의 activation_manifest_sha256 을 집었고,
# 그 값은 14개 슬라이스 전부에 같은 값으로 들어 있어 어느 도메인에서든 재발할 수 있다.
# 진짜 신선도는 모델을 거치지 않는 실물 두 값(계획서 sha ↔ R0 가 직접 읽은 문서 sha)이
# 말한다. 그 대조는 호출부가 하고(PART2_SLICE_STALE), 여기서는 echo 불일치를
# 수용 정책 관할로 넘길 수 있도록 모아만 둔다.
_echo = echo_mismatches if echo_mismatches is not None else []
for key in ("slice_sha256", "compiled_prompt_sha256"):
want = plan_row.get(key)
if isinstance(want, str) and want:
if seed_obj.get(key) != want:
raise ValueError(f"{domain_id}: stale seed output: {key} mismatch")
_echo.append({"domain_id": domain_id, "key": key,
"expected": want, "received": seed_obj.get(key)})
else:
warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"})
@@ -1344,6 +1477,450 @@ Agent:
if cand.get("ReasonRefs") not in ([], None):
raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent")
# ------------------------------------------------------------------
# R0-6 — 시드 수용(admission). 검증 결과에 등급을 준다.
# 구조 위반은 막고(block), 어휘·형태 위반은 정규화·수확으로 살리고(repair),
# 살릴 수 없는 것은 후보 하나만 격리한다(quarantine). 회차 전체를 죽이지 않는다.
# 값을 지어내지 않는다 — 표에 없는 낱말과 universe 에 없는 참조는 복구 대상이 아니다.
# ------------------------------------------------------------------
PATH_INDEX_RE = re.compile(r"\[(\d+)\]")
def _norm_path(path: str) -> str:
"""검증기 경로를 정책 키 모양으로 바꾼다. 인덱스는 [] 로 접는다."""
body = path.replace("$.stage_b_domain_bo_seed_output.", "").replace("$.stage_b_domain_bo_seed_output", "")
return PATH_INDEX_RE.sub("[]", body).lstrip(".")
def _path_tokens(path: str) -> list[Any]:
body = path.replace("$.stage_b_domain_bo_seed_output.", "").replace("$.stage_b_domain_bo_seed_output", "")
tokens: list[Any] = []
for part in body.lstrip(".").split("."):
if not part:
continue
name = part.split("[", 1)[0]
if name:
tokens.append(name)
for hit in PATH_INDEX_RE.findall(part):
tokens.append(int(hit))
return tokens
def _resolve_parent(root: Any, tokens: list[Any]) -> tuple[Any, Any]:
"""마지막 토큰의 부모 컨테이너와 그 토큰을 돌려준다. 못 찾으면 (None, None)."""
node = root
for tok in tokens[:-1]:
if isinstance(tok, int):
if not isinstance(node, list) or tok >= len(node):
return (None, None)
node = node[tok]
else:
if not isinstance(node, dict) or tok not in node:
return (None, None)
node = node[tok]
return (node, tokens[-1]) if tokens else (None, None)
def _candidate_index(tokens: list[Any]) -> int | None:
for pos, tok in enumerate(tokens):
if tok == "bo_seed_candidates" and pos + 1 < len(tokens) and isinstance(tokens[pos + 1], int):
return tokens[pos + 1]
return None
def _sink_for(seed_obj: dict[str, Any], tokens: list[Any], policy: dict[str, Any], norm: str):
"""초과 정보를 담을 자리 두 곳을 돌려준다: (컨테이너 지역 payload, 후보 수준 보관 목록).
지역 payload 의 키는 **잎 이름 그대로** 쓴다 — 하류 어댑터가
payload["status"] 처럼 짧은 이름으로 읽기 때문이다. 정규경로를 키로 쓰면
바이트는 남아도 아무도 읽지 못한다.
후보 수준 보관은 dict 가 아니라 **덧붙이기 전용 목록**이다. 같은 정규경로의
두 번째 값이 첫 번째를 조용히 덮는 사고를 구조적으로 막는다.
"""
sink_cfg = _dict(policy.get("surplus_sink"))
local: dict[str, Any] | None = None
by_container = _dict(sink_cfg.get("by_container"))
container_key = norm.rsplit(".", 1)[0] if "." in norm else norm
local_name = by_container.get(container_key)
if local_name:
parent, _ = _resolve_parent(seed_obj, tokens)
if isinstance(parent, dict):
bucket = parent.get(local_name)
if not isinstance(bucket, dict):
bucket = {}
parent[local_name] = bucket
local = bucket
shelf: list[Any] | None = None
cand_idx = _candidate_index(tokens)
cands = seed_obj.get("bo_seed_candidates")
if cand_idx is not None and isinstance(cands, list) and cand_idx < len(cands) and isinstance(cands[cand_idx], dict):
node: Any = cands[cand_idx]
steps = str(sink_cfg.get("candidate_path") or "extensions.worker_surplus").split(".")
for step in steps[:-1]:
nxt = node.get(step)
if not isinstance(nxt, dict):
nxt = {}
node[step] = nxt
node = nxt
leaf = steps[-1]
if not isinstance(node.get(leaf), list):
node[leaf] = []
shelf = node[leaf]
return (local, shelf)
def _universe_kind(value: str, universe: dict[str, set]) -> str | None:
if value in universe["source_evidence_indexes"]:
return "evidence"
if value in universe["source_event_candidate_ids"]:
return "event_candidate"
if value in universe["source_meeting_clause_ids"]:
return "meeting_clause"
return None
def _derive_required(seed_obj, tokens, norm, rule, universe, ctx):
"""선언된 규칙으로만 채운다. 규칙이 없거나 실패하면 (False, None)."""
parent, key = _resolve_parent(seed_obj, tokens)
if not isinstance(parent, dict):
return (False, None)
how = str(rule.get("rule") or "")
if how == "constant":
value = rule.get("value")
return (True, copy.deepcopy(value)) if "value" in rule else (False, None)
if how == "from_context":
ckey = str(rule.get("key") or "")
return (True, copy.deepcopy(ctx[ckey])) if ckey in ctx else (False, None)
if how == "first_universe_member_of_sibling_array":
known = (universe["source_evidence_indexes"] | universe["source_event_candidate_ids"]
| universe["source_meeting_clause_ids"])
for sib in rule.get("sibling_candidates") or []:
raw = parent.get(sib)
values = raw if isinstance(raw, list) else ([raw] if isinstance(raw, str) else [])
for item in values:
if isinstance(item, str) and item in known:
return (True, item)
return (False, None)
if how == "universe_kind_of_sibling":
sib = parent.get(str(rule.get("sibling") or ""))
kind = _universe_kind(sib, universe) if isinstance(sib, str) else None
if not kind:
# 형제가 아직 채워지지 않았을 수 있다. 같은 원소의 참조 배열에서 직접 읽는다.
for name in rule.get("fallback_sibling_arrays") or []:
raw = parent.get(name)
values = raw if isinstance(raw, list) else ([raw] if isinstance(raw, str) else [])
for item in values:
if isinstance(item, str):
kind = _universe_kind(item, universe)
if kind:
break
if kind:
break
return (True, kind) if kind else (False, None)
if how == "rename_sibling":
# 워커가 같은 뜻을 다른 이름으로 적었을 때 이름만 바로잡는다. 값은 그대로 옮긴다.
# 옮긴 뒤 원래 키를 지운다 — 남겨 두면 다음 패스에서 초과 속성으로 다시 걸린다.
for sib in rule.get("sibling_candidates") or []:
if sib in parent and parent.get(sib) not in (None, ""):
return (True, parent.pop(sib))
return (False, None)
if how == "sibling_matching_pattern":
pat = re.compile(str(rule.get("pattern") or "^$"))
for sib in rule.get("sibling_candidates") or []:
value = parent.get(sib)
if isinstance(value, str) and pat.fullmatch(value):
return (True, value)
for sib in rule.get("fallback_rename_sibling") or []:
if sib in parent and parent.get(sib) not in (None, ""):
return (True, parent.pop(sib))
return (False, None)
return (False, None)
def _promote_alias(parent, key, norm, policy, record):
"""수확 직전에 한 번 더 본다 — 이 키가 선언된 자리의 다른 이름일 뿐인가.
그렇다면 자유 공간으로 밀어 넣지 않고 제 자리로 올린다. 워커가 excerpt 를
value·content 로 부르는 표류가 실측됐고, 그 값은 하류가 실제로 읽는 칸이다.
"""
if not isinstance(parent, dict) or not isinstance(key, str):
return False
container = norm.rsplit(".", 1)[0] if "." in norm else ""
for target_path, rule in _dict(policy.get("required_derivation")).items():
if str(rule.get("rule") or "") != "rename_sibling":
continue
if target_path.rsplit(".", 1)[0] != container:
continue
target = target_path.rsplit(".", 1)[-1]
if target in parent and parent.get(target) not in (None, ""):
continue
if key in (rule.get("sibling_candidates") or []):
parent[target] = parent.pop(key)
record["received"] = _clip(parent[target])
record["applied"] = "promoted_to:%s" % target
record["rule_id"] = "alias_promotion:%s" % target_path
return True
return False
def _harvest(seed_obj, tokens, policy, norm, record):
"""규약 밖 값을 버리지 않고 자유 공간으로 옮긴다. 옮긴 사실을 기록한다."""
parent, key = _resolve_parent(seed_obj, tokens)
if parent is None:
return False
if _promote_alias(parent, key, norm, policy, record):
return True
local, shelf = _sink_for(seed_obj, tokens, policy, norm)
if isinstance(parent, list) and isinstance(key, int):
if key >= len(parent):
return False
moved = parent.pop(key)
leaf = None
elif isinstance(parent, dict):
if key not in parent:
return False
moved = parent.pop(key)
leaf = key
else:
return False
placed = []
# 1) 컨테이너 지역 payload — 하류가 읽는 짧은 이름으로. 이미 있으면 덮지 않는다.
if isinstance(local, dict) and isinstance(leaf, str) and leaf not in local:
local[leaf] = moved
placed.append("payload.%s" % leaf)
# 2) 후보 수준 보관 — 덧붙이기 전용이라 어떤 값도 덮이지 않는다. 원래 경로를 함께 남긴다.
if isinstance(shelf, list):
# D11 — norm 은 인덱스를 접으므로 같은 컨테이너의 두 값이 구별되지 않는다.
# 원본 경로를 함께 남겨야 부모 원소와의 결합(예: role ↔ party_refs)을 복원할 수 있다.
shelf.append({"path": norm, "source_path": record.get("path"), "value": moved})
placed.append("worker_surplus[]")
if not placed:
# 보관할 자리가 없으면 지우지 않는다. 되돌려 놓고 실패로 돌려주면
# 이 위반은 잔여로 남아 정책이 정한 처분(기본 review)으로 간다.
# 뿌리 수준 배열(unknown_or_unrouted_reviews 등)이 여기 해당한다 —
# 후보에 매이지 않아 보관처가 없는데, 그렇다고 사건 자료를 버릴 수는 없다.
if isinstance(parent, list) and isinstance(key, int):
parent.insert(key, moved)
elif isinstance(parent, dict) and isinstance(key, str):
parent[key] = moved
return False
record["received"] = _clip(moved)
record["applied"] = "+".join(placed)
# 기계적 키 이동과, 실질 서술을 담은 채 **객체째** 밀려난 것은 검토 무게가 다르다.
# 후자를 같은 등급에 섞으면 1,200건 속 몇 건을 사람이 찾아내야 한다.
# 판정은 dict 로 좁힌다 — 스칼라 한 개의 자리 이동(당사자 이름·slot_id 문자열)은
# 원소가 사라진 것이 아니라 키가 옮겨진 것이라 무게가 다르다.
if isinstance(moved, dict):
for _k in ("excerpt", "value", "fact", "content", "reason", "statement"):
_v = moved.get(_k)
if isinstance(_v, str) and _v.strip():
record["content_bearing"] = True
break
return True
def _evict(seed_obj, tokens, policy, norm, record, quarantined=None):
"""복구 불가한 객체 하나를 배열에서 들어내 보관한다. 후보 자체는 삭제하지 않는다.
D4·D8 — 후보를 pop 하면 ① 예산에 잡히지 않아 소리 없이 사라지고
② 뒤 후보의 인덱스가 밀려 이미 기록한 격리 표시가 다른 후보를 가리킨다.
후보 수준이면 삭제 대신 격리로 돌린다.
"""
trimmed = list(tokens)
while trimmed and not isinstance(trimmed[-1], int):
trimmed.pop()
if not trimmed:
return False
if len(trimmed) == 2 and trimmed[0] == "bo_seed_candidates":
if quarantined is None:
return False
quarantined.add(trimmed[1])
record["applied"] = "quarantined_candidate"
return True
return _harvest(seed_obj, trimmed, policy, _norm_path("$.stage_b_domain_bo_seed_output." + norm), record)
def _clip(value: Any, limit: int = 200) -> Any:
try:
text = json.dumps(value, ensure_ascii=False)
except Exception:
text = str(value)
return text if len(text) <= limit else text[:limit] + "…"
def admit_seed(seed_obj, domain_id, schema, slice_doc, universe, policy, records, plan_row=None, validator=None):
"""검증 -> 처분 -> 복구를 수렴할 때까지 돌리고, 남은 것은 후보 격리로 넘긴다.
돌려주는 것: (수용된 seed, 격리된 후보 인덱스 집합, 경성 중단 사유 목록)
"""
# D1 — 검증기는 인자로 받는다. main() 지역 이름을 전역처럼 읽으면 NameError 로 즉사한다.
if validator is None:
return (seed_obj, set(), [])
def _raw_validate(obj):
return validator.validate_worker_output(
{"stage_b_domain_bo_seed_output": obj}, schema=schema,
expected_domain_id=domain_id, slice_document=slice_doc)
def _validate(obj):
"""검증기 자체가 터질 수 있다(예: bo_type 이 dict 면 unhashable).
D9 — 예외를 코드로 바꾸기만 하면 부족하다. 그 오류의 path 는 뿌리라
후보를 지목하지 못하고, 처분이 뿌리 잔여(review)로 강등돼 문제 후보가
그대로 투영으로 흘러 project_to_bo_surface 에서 다시 죽는다.
그래서 **어느 후보가 터뜨렸는지 후보 단위로 좁혀** path 에 인덱스를 실어 준다.
"""
try:
return _raw_validate(obj)
except Exception as exc:
detail = "%s: %s" % (type(exc).__name__, str(exc)[:160])
errors = []
cands = _list(obj.get("bo_seed_candidates"))
for idx in range(len(cands)):
probe = dict(obj)
probe["bo_seed_candidates"] = [cands[idx]]
try:
_raw_validate(probe)
except Exception:
errors.append({"code": "VALIDATOR_CRASHED",
"path": "$.stage_b_domain_bo_seed_output.bo_seed_candidates[%d]" % idx,
"message": detail})
if not errors:
# 후보를 좁히지 못했다. 뿌리 문제이므로 명시적으로 끊는다 —
# review 로 강등해 투영에서 죽게 두는 것이 최악이다.
errors.append({"code": "VALIDATOR_CRASHED_ROOT",
"path": "$.stage_b_domain_bo_seed_output", "message": detail})
return {"errors": errors, "warnings": []}
disposition = _dict(policy.get("disposition_by_code"))
unknown_disp = str(policy.get("unknown_code_disposition") or "quarantine")
synonyms = _dict(policy.get("enum_synonyms"))
case_fold = bool(policy.get("enum_case_fold"))
enum_unmapped = str(policy.get("enum_unmapped_disposition") or "repair_evict")
derivations = _dict(policy.get("required_derivation"))
derive_failed = str(policy.get("derivation_failed_disposition") or "repair_evict")
order = [str(v) for v in (policy.get("repair_order") or [])]
max_passes = int(policy.get("max_repair_passes") or 3)
blocks: list[dict[str, Any]] = []
quarantined: set[int] = set()
for _pass in range(max_passes):
cands = _list(seed_obj.get("bo_seed_candidates"))
# task_instance_id 는 지어내지 않는다 — A0 의 fan-out 계획 행이 든 값을 쓴다.
ctx = {"domain_id": domain_id, "seed_count": len(cands),
"emitted_seed_ids": [str(_dict(c).get("seed_id") or "") for c in cands]}
_plan_tid = _dict(plan_row).get("task_instance_id")
if isinstance(_plan_tid, str) and _plan_tid:
ctx["task_instance_id"] = _plan_tid
# r0_membership_key 는 R0 가 자기 관측으로 결정적으로 만든다(워커가 알 수 없는 값이다).
ctx["r0_membership_key"] = "%s:%s" % (
domain_id, hashlib.sha256(json.dumps(ctx["emitted_seed_ids"], ensure_ascii=False,
sort_keys=True).encode("utf-8")).hexdigest()[:32])
report = _validate(seed_obj)
errors = [e for e in (report.get("errors") or []) if isinstance(e, dict)]
if not errors:
break
buckets: dict[str, list[dict[str, Any]]] = {}
for err in errors:
disp = str(disposition.get(str(err.get("code"))) or unknown_disp)
buckets.setdefault(disp, []).append(err)
for err in buckets.get("block") or []:
blocks.append({"code": err.get("code"), "path": err.get("path"), "message": err.get("message")})
if blocks:
return (seed_obj, quarantined, blocks)
for err in buckets.get("review") or []:
records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"),
"action": "review", "received": None, "applied": None,
"rule_id": "disposition:review"})
progressed = False
for action in order:
# 원소를 들어내는 처분(harvest·evict)만 내림차순으로 돈다 — 앞 인덱스를 먼저
# 지우면 뒤 경로가 밀리기 때문이다. 값을 채우는 처분(normalize·derive)은
# 오름차순이어야 한다: source_kind 는 형제 source_id 를 읽으므로 순서가 뒤집히면
# 형제가 아직 없어 파생이 실패하고, 그 실패가 요건사실 원소를 통째로 들어낸다.
_removes = action in ("repair_harvest", "repair_evict")
for err in sorted(buckets.get(action) or [],
key=lambda e: PATH_INDEX_RE.sub(
lambda m: "[%05d]" % int(m.group(1)), str(e.get("path") or "")),
reverse=_removes):
path = str(err.get("path") or "")
tokens = _path_tokens(path)
if not tokens:
continue
norm = _norm_path(path)
rec = {"domain_id": domain_id, "code": err.get("code"), "path": path,
"action": action, "received": None, "applied": None, "rule_id": None}
done = False
if action == "repair_normalize":
parent, key = _resolve_parent(seed_obj, tokens)
table = _dict(synonyms.get(norm))
if isinstance(parent, dict) and key in parent and table:
raw = parent.get(key)
probe = raw.casefold() if (case_fold and isinstance(raw, str)) else raw
mapped = table.get(probe) if isinstance(probe, str) else None
if mapped is not None:
rec["received"], rec["applied"] = _clip(raw), mapped
rec["rule_id"] = "enum_synonyms:%s" % norm
parent[key] = mapped
done = True
if not done and enum_unmapped == "repair_evict":
rec["rule_id"] = "enum_unmapped:%s" % norm
done = _evict(seed_obj, tokens, policy, norm, rec, quarantined)
elif action == "repair_derive":
rule = _dict(derivations.get(norm))
if rule:
ok, value = _derive_required(seed_obj, tokens, norm, rule, universe, ctx)
if ok:
parent, key = _resolve_parent(seed_obj, tokens)
if isinstance(parent, dict):
rec["received"], rec["applied"] = None, _clip(value)
rec["rule_id"] = "required_derivation:%s" % norm
parent[key] = value
done = True
if not done and derive_failed == "repair_evict":
rec["rule_id"] = "derivation_failed:%s" % norm
done = _evict(seed_obj, tokens, policy, norm, rec, quarantined)
elif action == "repair_harvest":
rec["rule_id"] = "surplus_sink:%s" % norm
done = _harvest(seed_obj, tokens, policy, norm, rec)
elif action == "repair_evict":
rec["rule_id"] = "evict:%s" % norm
done = _evict(seed_obj, tokens, policy, norm, rec, quarantined)
elif action == "quarantine":
idx = _candidate_index(tokens)
if idx is not None:
quarantined.add(idx)
rec["rule_id"] = "quarantine:%s" % norm
done = True
if done:
cand_idx = _candidate_index(tokens)
if cand_idx is not None:
cand = _list(seed_obj.get("bo_seed_candidates"))
if cand_idx < len(cand):
rec["candidate_ref"] = _dict(cand[cand_idx]).get("candidate_ref") or _dict(cand[cand_idx]).get("seed_id")
records.append(rec)
progressed = True
if progressed:
break
if not progressed:
break
# 수렴하지 않고 남은 위반은 후보 단위로 격리한다. 회차는 계속된다.
report = _validate(seed_obj)
for err in (report.get("errors") or []):
if not isinstance(err, dict):
continue
# review 로 수용하기로 선언된 코드는 남아 있는 것이 정상이다. 다시 격리하지 않는다.
if str(disposition.get(str(err.get("code"))) or "") == "review":
records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"),
"action": "review", "received": None, "applied": None,
"rule_id": "disposition:review(residual)"})
continue
tokens = _path_tokens(str(err.get("path") or ""))
idx = _candidate_index(tokens)
if idx is None:
root_disp = str(policy.get("residual_root_disposition") or "review")
if root_disp == "block":
blocks.append({"code": err.get("code"), "path": err.get("path"),
"message": err.get("message"), "note": "root-level residue"})
else:
records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"),
"action": "review", "received": None, "applied": None,
"rule_id": "residual_root"})
else:
quarantined.add(idx)
records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"),
"action": "quarantine", "received": None, "applied": None,
"rule_id": "residual_after_repair"})
return (seed_obj, quarantined, blocks)
# ---------- PostB_1 이식: sort key / duplicate keys / schema risk ----------
def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]:
provenance = _dict(seed.get("provenance"))
@@ -1451,6 +2028,17 @@ Agent:
# 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 —
# 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5).
admission_policy = _dict(json.loads(_verify_asset(SEED_ADMISSION_POLICY_PATH)))
if admission_policy.get("schema_version") != SEED_ADMISSION_SCHEMA_VERSION:
raise RuntimeError("PART2_SEED_ADMISSION_POLICY_INVALID")
admission_enforcing = str(admission_policy.get("enforcement") or "enforce") == "enforce"
admission_records: list[dict[str, Any]] = []
declared_candidate_total = 0
admission_blocks: list[dict[str, Any]] = []
quarantined_by_domain: dict[str, set] = {}
admitted_total = 0
quarantined_total = 0
seed_objects: dict[str, dict[str, Any]] = {}
projected_candidates: dict[str, list[dict[str, Any]]] = {}
review_handoff_items: list[dict[str, Any]] = []
@@ -1463,33 +2051,124 @@ Agent:
seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output"))
if not seed_obj:
raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing")
validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings)
_plan_row = PLAN_ROWS.get(domain_id) or {}
_echo_mismatches: list[dict[str, Any]] = []
validate_seed_object(seed_obj, domain_id, _plan_row, guard_warnings, _echo_mismatches)
# 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와
# BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다.
# 원문도 함께 든다 — 신선도 판정의 실물 근거가 이 바이트열의 해시다.
# D3 — 원문 읽기가 실패해도 문서 자체는 기존 경로로 다시 시도한다. read_raw 는 봉투가
# 다르면 raise 하므로, 그것 때문에 slice_doc 까지 잃으면 allowed_legal_effect_bo_types 가
# 조용히 빈 집합이 된다.
_slice_raw = None
slice_doc = None
try:
_slice_raw = read_raw("%s/%s.json" % (SLICE_DIR, domain_id))
slice_doc = json.loads(_slice_raw)
except Exception:
_slice_raw = None
try:
slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id))
except Exception:
slice_doc = None
# R0-7a — 신선도의 실물 근거. 계획서가 든 sha 와 R0 가 지금 읽은 문서의 sha 를 맞춘다.
# 둘 다 모델을 거치지 않으므로 여기서 어긋나면 정말로 낡았거나 배포가 어긋난 것이다.
_want_slice = _plan_row.get("slice_sha256")
if isinstance(_want_slice, str) and _want_slice and _slice_raw is not None:
# sha_text 는 A0 블록 전용이다. R0 는 이 블록의 관용대로 인라인으로 센다.
_actual_slice = hashlib.sha256(_slice_raw.encode("utf-8")).hexdigest()
if _actual_slice != _want_slice:
raise RuntimeError(json.dumps({
"reason_code": "PART2_SLICE_STALE",
"message": "fan-out 계획이 든 slice sha 와 실제 slice 문서가 다르다. 시드가 아니라 슬라이스가 어긋났다.",
"domain_id": domain_id, "expected": _want_slice, "actual": _actual_slice,
}, ensure_ascii=False))
# R0-7b — echo 불일치의 등급은 정책이 정한다. 다만 두 실패는 성격이 다르므로 코드를 가른다.
# · 다른 슬라이스의 sha 를 echo 했다 = 워커가 **엉뚱한 슬라이스로 작업**했다는 뜻이다.
# 실물 대조로는 원리상 못 잡는다(그 도메인 슬라이스는 멀쩡하다). 중대 위반이다.
# · 그 밖(예: activation_manifest_sha256 을 집음)은 전사 실수다. 검토로 족하다.
# D2 — 슬라이스를 못 읽어 실물 대조를 못 한 도메인은 신선도 방어가 0 이 된다.
# 그 경우에는 echo 를 검토로 강등하지 않는다.
# 판정은 slice 에 한정하지 않는다 — 다른 도메인의 compiled_prompt sha 를 집은 것도
# 그 도메인 자료를 읽었다는 뜻이라 같은 무게다. 두 필드를 합쳐 "남의 자료" 집합을 만든다.
_foreign_material_shas = set()
_own_material_shas = {str(_plan_row.get(_mk)) for _mk in ("slice_sha256", "compiled_prompt_sha256")
if _plan_row.get(_mk)}
for _d2, _r in PLAN_ROWS.items():
if _d2 == domain_id or not isinstance(_r, dict):
continue
for _mk in ("slice_sha256", "compiled_prompt_sha256"):
if _r.get(_mk):
_foreign_material_shas.add(str(_r.get(_mk)))
# D7 — 두 도메인의 자료가 우연히 같아지면(예: 빈 도메인 둘) 정상 산출이 남의 것으로 오판된다.
# 자기 값을 빼서 원천 차단한다. 현 데이터에는 중복이 없지만 데이터의 성질에 기대지 않는다.
_foreign_material_shas -= _own_material_shas
# D8 — 우리가 양성으로 아는 유일한 오집합은 activation_manifest_sha256 이다(값으로 식별된다).
# 그것만 전사 실수로 보고, 설명되지 않는 값은 미상으로 두어 더 무겁게 다룬다.
_known_slip_sha = str(_dict(_dict(_dict(slice_doc).get(SLICE_ROOT_KEY)).get("activation")).get(
"activation_manifest_sha256") or "")
_disp_map = _dict(admission_policy.get("disposition_by_code"))
for _em in _echo_mismatches:
_rcv = str(_em.get("received"))
if _rcv in _foreign_material_shas:
_code = "SEED_ECHO_FOREIGN_SLICE"
elif _slice_raw is None:
# 원문이 없으면 실물 대조가 불가능하다. 이 상태에서는 known slip 이라도 강등하지 않는다 —
# activation_manifest_sha256 은 전 도메인이 공유하는 값이라 어느 슬라이스를 읽었는지에
# 대해 아무것도 말해 주지 않는다. 그것을 근거로 가볍게 볼 수 없다.
_code = "SEED_ECHO_SHA_UNVERIFIABLE"
elif _known_slip_sha and _rcv == _known_slip_sha:
_code = "SEED_ECHO_SHA_MISMATCH"
else:
_code = "SEED_ECHO_SHA_UNEXPLAINED"
_disp = str(_disp_map.get(_code) or ("block" if _code == "SEED_ECHO_FOREIGN_SLICE"
else "review" if _code == "SEED_ECHO_SHA_MISMATCH"
else "quarantine"))
if _disp == "block":
raise RuntimeError(json.dumps(dict(
{"reason_code": "PART2_%s" % _code}, **_em), ensure_ascii=False))
admission_records.append({
"domain_id": domain_id, "code": _code,
"path": "$.stage_b_domain_bo_seed_output.%s" % _em["key"],
"action": _disp, "received": _em.get("received"), "applied": None,
"rule_id": "echo_sha_not_trusted:%s" % _em["key"]})
slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {}
allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types")))
allowed_bo_types_by_domain[domain_id] = allowed_bo_types
# R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다.
# R0-6 — 검증 결과를 경고로 흘리지 않고 등급대로 처분한다.
# shadow 모드에서는 계산·기록만 하고 seed 를 바꾸지 않는다(첫 전환 회차용).
# D13 — 회계 누적은 검증기 유무와 무관하다. 분기 안에 두면 검증기 반입이
# 실패한 회차에서 declared=0 · admitted=N 이 되어 헛경보(lost 음수)가 난다.
declared_candidate_total += len(_list(seed_obj.get("bo_seed_candidates")))
if worker_validator is not None:
report = worker_validator.validate_worker_output(
{"stage_b_domain_bo_seed_output": seed_obj},
schema=seed_schema,
expected_domain_id=domain_id,
slice_document=slice_doc)
for item in report.get("errors") or []:
guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id,
"detail": item})
for item in report.get("warnings") or []:
for item in (worker_validator.validate_worker_output(
{"stage_b_domain_bo_seed_output": seed_obj}, schema=seed_schema,
expected_domain_id=domain_id, slice_document=slice_doc).get("warnings") or []):
guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id,
"detail": item})
work_obj = seed_obj if admission_enforcing else copy.deepcopy(seed_obj)
admitted_obj, quarantined_idx, blocks = admit_seed(
work_obj, domain_id, seed_schema, slice_doc, universe,
admission_policy, admission_records, PLAN_ROWS.get(domain_id) or {},
worker_validator)
if blocks:
admission_blocks.extend({"domain_id": domain_id, **b} for b in blocks)
if admission_enforcing:
seed_obj = admitted_obj
quarantined_by_domain[domain_id] = quarantined_idx
else:
quarantined_by_domain[domain_id] = set()
for b in blocks:
guard_warnings.append({"code": "SEED_ADMISSION_SHADOW_BLOCK",
"domain_id": domain_id, "detail": b})
cands = _list(seed_obj.get("bo_seed_candidates"))
projected: list[dict[str, Any]] = []
projection_reviews: list[dict[str, Any]] = []
skip_idx = quarantined_by_domain.get(domain_id) or set()
for idx, cand in enumerate(cands):
if idx in skip_idx:
quarantined_total += 1
continue
if not isinstance(cand, dict):
raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object")
cand = ensure_candidate_ref(cand, domain_id, idx)
@@ -1552,6 +2231,70 @@ Agent:
entry["action_source"] = note_item.get("source")
review_handoff_items.append(entry)
# R0-6b — 수용 결산. 격리는 후보 단위이고, 회차 중단은 예산을 넘을 때만이다.
admitted_total = sum(len(v) for v in projected_candidates.values())
# D4 — 수용 단계에서 후보가 사라지는 경로는 격리 하나뿐이어야 한다.
# 원본 후보 수와 (수용 + 격리)가 맞지 않으면 소리 없이 없어진 것이 있다는 뜻이다.
if admission_enforcing and declared_candidate_total != admitted_total + quarantined_total:
guard_warnings.append({"code": "SEED_ADMISSION_ACCOUNTING_DRIFT",
"declared": declared_candidate_total, "admitted": admitted_total,
"quarantined": quarantined_total,
"lost": declared_candidate_total - admitted_total - quarantined_total})
if admission_blocks and admission_enforcing:
raise RuntimeError(json.dumps({
"reason_code": "PART2_SEED_ADMISSION_BLOCKED",
"message": "시드 뿌리 계약이 깨졌다. 복구 대상이 아니다.",
"blocks": admission_blocks[:20], "block_count": len(admission_blocks),
}, ensure_ascii=False))
budget = _dict(admission_policy.get("quarantine_budget"))
seen_total = admitted_total + quarantined_total
ratio = (quarantined_total / seen_total) if seen_total else 0.0
max_ratio = budget.get("max_quarantined_candidate_ratio")
min_admitted = budget.get("min_admitted_candidates_run")
if admission_enforcing and isinstance(max_ratio, (int, float)) and seen_total and ratio > float(max_ratio):
raise RuntimeError(json.dumps({
"reason_code": "PART2_SEED_ADMISSION_BUDGET_EXCEEDED",
"message": "격리 비율이 한계를 넘었다. 모델 잡음이 아니라 계약 파손으로 본다.",
"quarantined": quarantined_total, "seen": seen_total,
"ratio": round(ratio, 4), "max_ratio": max_ratio,
}, ensure_ascii=False))
if admission_enforcing and isinstance(min_admitted, int) and admitted_total < min_admitted:
raise RuntimeError(json.dumps({
"reason_code": "PART2_SEED_ADMISSION_EMPTY",
"message": "수용된 후보가 없다.", "admitted": admitted_total,
}, ensure_ascii=False))
_adm_review = _dict(admission_policy.get("review"))
_adm_counter = 0
for rec in admission_records:
_adm_counter += 1
_is_q = rec.get("action") == "quarantine"
_is_c = bool(rec.get("content_bearing")) and not _is_q
review_handoff_items.append({
"review_id": "R0:admission:%03d" % _adm_counter,
"source_domain": rec.get("domain_id"),
"severity": str((_adm_review.get("quarantine_severity") if _is_q
else _adm_review.get("evicted_with_content_severity") if _is_c
else _adm_review.get("severity")) or "SOFT_WARNING"),
"issue_type": str((_adm_review.get("quarantine_issue_type") if _is_q
else _adm_review.get("evicted_with_content_issue_type") if _is_c
else _adm_review.get("issue_type")) or "seed_admission_repair"),
"source_review_code": rec.get("code"),
"source_event_candidate_ids": [],
"source_evidence_indexes": [],
"source_meeting_clause_ids": [],
"downstream_owner": str(_adm_review.get("downstream_owner") or "Stage2"),
"template_note": "수용 단계 처분: %s · 경로 %s · 규칙 %s · 원값 %s -> 적용 %s (후보 %s)" % (
rec.get("action"), rec.get("path"), rec.get("rule_id"),
rec.get("received"), rec.get("applied"), rec.get("candidate_ref")),
})
guard_warnings.append({"code": "SEED_ADMISSION_SUMMARY",
"enforcement": admission_policy.get("enforcement"),
"repairs": sum(1 for r in admission_records if str(r.get("action") or "").startswith("repair")),
"reviews": sum(1 for r in admission_records if r.get("action") == "review"),
"quarantined_candidates": quarantined_total,
"admitted_candidates": admitted_total,
"quarantine_ratio": round(ratio, 4)})
# 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선).
# 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다.
input_candidate_total = 0
@@ -2507,6 +3250,13 @@ Agent:
action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal"))
# F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은
# claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다.
if not isinstance(bo_type, str):
# 워커가 문자열이 아닌 값을 넣으면 집합 비교 자체가 터진다(unhashable).
# 수용 단계가 걸러야 하지만, 투영은 마지막 방어선이라 여기서도 막는다.
normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType",
"received": type(bo_type).__name__,
"fallback": f0_norm.get("bo_type_default")})
bo_type = f0_norm.get("bo_type_default")
if bo_type not in bo_types:
normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")})
bo_type = f0_norm.get("bo_type_default")
@@ -2282,6 +2282,14 @@ stage_1_2_3_assets_distribution_location.md`의 `§8. 특기사항 (판독 중
==================================
┌────────────────────────────────────┐
│ Stage 1 -- Part 4 │
│ YAML Docs Update Strategy │
└────────────────────────────────────┘
┌────────────────────────────────────────────┐
│ Stage 1 -- Part 1|2||3 │
│ yaml 실행 체크하면서 동시 진행 │
└────────────────────────────────────────────┘