diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/platform/reference_catalogs/profile_ids.json b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/platform/reference_catalogs/profile_ids.json new file mode 100644 index 00000000..a712eef3 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/platform/reference_catalogs/profile_ids.json @@ -0,0 +1 @@ +{"catalog_id":"special_law_profiles","entries":[{"attached_domain_id":"E-20","profile_id":"SLP-IP","status":"active"},{"attached_domain_id":"E-21","profile_id":"SLP-MEDIA-MEDIATION","status":"active"},{"attached_domain_id":"E-06","profile_id":"SLP-PRODUCT-LIABILITY","status":"active"}],"expected_count":3,"schema_version":"stage1_assembly_reference_catalog.v1"} diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-IP/prompt_overlay.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-IP/prompt_overlay.md new file mode 100644 index 00000000..5f726256 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-IP/prompt_overlay.md @@ -0,0 +1,52 @@ +# SLP-IP — 지식재산 특별법 profile + +`profile_id`: `SLP-IP` · 부착 도메인: E-20 지식재산(`ip_claims`) +조립 위치: 공통 계약 → 의존 도메인(E-01·E-05) → E-20 overlay → **이 profile** → runtime guard + +## 권한 경계 + +- 이 조각은 E-20 overlay를 좁힐 뿐 입력·출력·법리 범위를 넓히지 않는다. 새 출력 key와 새 review key를 만들지 않는다. +- `final_conclusion_forbidden`은 유지된다. 권리 유효·귀속, 보호범위 포함 여부, 침해 성립, 손해액을 확정하지 않는다. +- 법령 조문 번호·존속기간·법정손해배상 한도·요율을 이 조각에 상수로 두지 않는다. 전부 SG-04 law-version 후보로 넘긴다. +- **권리 종류를 먼저 특정하지 못하면 요건 층을 단일화하지 않는다.** 특허·실용신안, 상표, 디자인, 저작권, 부정경쟁·영업비밀은 각기 다른 특별법 체계이고 성립·보호범위·침해 판단·구제·손해 산정 특칙이 서로 다르므로, 종류 미특정 시 후보를 종류별로 병렬 보존하고 review로 표시한다. + +## 1. 식별 단서 — source-backed only + +"특허·저작권·상표"라는 단어만 있는 자료는 활성화 근거가 아니다(E-20 `E20-N01`). 다음이 있을 때 이 층을 연다. + +1. **권리 발생 방식**: 등록으로 발생하는 권리(특허·실용신안·상표·디자인)는 등록원부·공보 원천을, 무방식으로 발생하는 권리(저작권)는 창작·공표·저작자 표시 원천을, 영업비밀·성과 도용은 비밀관리성·경제적 유용성·비공지성 또는 성과의 형성 원천을 각각 요구한다. +2. **권리자 지위**: 원시취득·승계·직무발명 또는 업무상저작물·공동보유·양도·전용실시권·통상실시권·이용허락 중 어느 구조인지 원천으로 특정한다. +3. **피고 실시행위**: 생산·사용·양도·대여·수입·전시·복제·전송·표장 사용·부정취득·부정사용 중 어떤 행위 유형인지와 그 기간·경로. +4. **구제수단 축**: 금지·폐기·신용회복·손해배상·부당이득·실시료 상당액 중 무엇을 구하는지. 요건이 다르므로 합치지 않는다. + +## 2. 고유 요건 슬롯·증거 component + +E-20 슬롯에 연결하며 새 슬롯을 만들지 않는다. + +- `e20.el01` — 권리 특정·존속. 등록권리는 등록번호·권리범위 문언·존속 상태를, 저작물은 창작성 표현 부분의 특정을 후보로 남긴다. 존속기간 계산은 이 조각이 하지 않고 기산 사실만 보존한다. +- `e20.el02` — 귀속·양도·실시허락. 등록원부와 계약이 불일치하면 어느 쪽도 지우지 말고 `E20_SOURCE_CONFLICT_REVIEW`와 함께 병존시킨다. +- `e20.el03`·`e20.el04` — 비교 구조는 권리 종류별로 요소가 다르다. 특허·실용신안은 청구항 구성요소 대비, 상표는 표장·지정상품 유사와 출처 혼동 우려, 디자인은 전체적 심미감, 저작권은 의거성과 실질적 유사성, 영업비밀은 비밀관리성과 부정취득·사용 태양. **다른 종류의 판단 요소를 교차 적용하지 않는다.** +- `e20.op01`·`e20.op02` — 항변을 별개 후보로 둔다: 권리 무효 사유, 권리 소진, 선사용, 자유실시·공지기술, 허락·묵시적 이용허락, 권리남용, 저작권의 제한·인용 사유, 의거성 부인. 성립 여부는 판단하지 않는다. +- **트랙 경계**: 등록권리의 무효·권리범위 확인은 심판·심결취소 트랙에 속하고 이 민사 트랙과 다르다. 무효 주장이 원천에 있으면 판단하지 말고 `TRACK_BOUNDARY_REVIEW`로 경계를 남긴다. 형사 고소·행정조치 접점도 같은 방식으로 표시한다. +- 증거 component는 `e20.right_identity`·`e20.ownership_license`·`e20.infringing_act` 등 등록 어휘를 그대로 쓴다. + +## 3. review code 발행 규칙 + +등록된 코드만 쓰고, 발행 경로는 공통 계약이 정한 후보별 `review_items`와 root `unknown_or_unrouted_reviews`뿐이다. + +- 권리 종류가 특정되지 않았거나 종류별 특별법 원천이 없다 → `E20_SPECIAL_LAW_PROFILE_REQUIRED` +- 권리 특정·존속 원천이 부족하다 → `E20_RIGHT_IDENTITY_VALIDITY_REVIEW` +- 귀속·실시허락 원천이 부족하거나 충돌한다 → `E20_OWNERSHIP_LICENSE_REVIEW` +- 비교 대상·범위 원천이 부족하다 → `E20_SCOPE_COMPARISON_REVIEW` +- 실시행위 원천이 부족하다 → `E20_INFRINGEMENT_ACT_REVIEW` +- 구제수단별 요건이 섞였거나 미분화다 → `E20_REMEDY_BOUNDARY_REVIEW` +- 무효·심판·형사·행정 트랙 접점이 있다 → `TRACK_BOUNDARY_REVIEW` +- 손해 추정·실시료 규정군이 필요한데 formula가 pin되지 않았다 → `DEFERRED_CALCULATION_TRACK_REVIEW` + +## 4. 계산·법령 version 경계 + +- CE-01·CE-R1은 deferred다. 손해 추정·실시료 상당액·법정손해배상 선택지는 **operand와 근거 원천만** 보존하고 산정을 실행하지 않는다. +- SG-04로 pin할 것: 권리 종류별 적용 법률과 행위기간별 버전, 존속·기간 규정, 손해 추정·법정손해배상 규정의 버전 간 차이, 판례 변경 review. +- **E-20 extension은 profile 전용 key를 선언하지 않는다.** 이 profile의 후보는 선언된 `ip_right_candidates`·`ownership_license_candidates`·`infringing_act_candidates`·`scope_comparison_candidates`·`calculation_operand_candidates`·`boundary_routes`에 담고, 새 key를 만들지 않는다. + +최종 법률결론과 최종 청구액을 확정하지 않는다. 원천 참조와 미해결 review code를 보존한다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-MEDIA-MEDIATION/prompt_overlay.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-MEDIA-MEDIATION/prompt_overlay.md new file mode 100644 index 00000000..baf37dd4 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-MEDIA-MEDIATION/prompt_overlay.md @@ -0,0 +1,50 @@ +# SLP-MEDIA-MEDIATION — 언론피해 구제 특별법 profile + +`profile_id`: `SLP-MEDIA-MEDIATION` · 부착 도메인: E-21 언론·인격권(`media_personality_rights`) +조립 위치: 공통 계약 → 의존 도메인 → E-21 overlay → **이 profile** → runtime guard + +## 권한 경계 + +- 이 조각은 E-21 overlay를 좁힐 뿐 입력·출력·법리 범위를 넓히지 않는다. 새 출력 key와 새 review key를 만들지 않는다. +- `final_conclusion_forbidden`은 유지된다. 위법성, 진실성·상당성, 정정·반론 청구권 성립, 기간 준수 여부를 확정하지 않는다. +- 청구기간·제척기간의 일수와 법령 조문 번호를 이 조각에 상수로 두지 않는다. 기산 사실만 후보로 남기고 기간 규범은 SG-04로 넘긴다. +- 인격권 침해 일반(E-21 본체)과 이 특별법 층은 다르다. **매체가 이 특별법의 적용 대상인지 먼저 확인하지 못하면 구제수단 요건을 단일화하지 않는다.** + +## 1. 식별 단서 — source-backed only + +1. **적용 대상 매체 판정**: 신문·잡지 등 정기간행물, 방송, 뉴스통신, 인터넷신문, 인터넷뉴스서비스 등 이 특별법이 정한 언론·매체에 해당하는지를 보이는 원천. 개인 게시물·일반 게시판·사적 전파는 이 층이 아니라 일반 인격권·정보통신 영역의 경계 사안으로 표시한다. +2. **보도·게시물 특정**: 원문과 수정본, 제목·본문·이미지·영상, 게재 매체·게시 주체·게시 시각·URL·캡처를 버전별로 보존한다. 정정 대상은 "보도 내용 중 사실적 주장"이므로 의견 표명 부분과 구분해 특정한다. +3. **피해자 특정**: 성명·직함·사진·정황 등으로 피해자가 특정·식별 가능한지의 원천. 집단 표시의 경우 개별 구성원 식별 가능성 사실을 따로 남긴다. +4. **구제수단 축**: 정정보도, 반론보도, 추후보도, 손해배상은 요건·기산점·상대방이 서로 다르다. **합치지 않고 축별로 분리 보존한다.** + +## 2. 고유 요건 슬롯·증거 component + +E-21 슬롯에 연결하며 새 슬롯을 만들지 않는다. + +- `e21.el01` — 보도 원본·버전·게시 주체·시점. 정정·삭제·수정 이력 자체가 요건 사실이므로 변경 전후를 모두 보존한다. +- `e21.el02` — 식별가능성과 전파범위. 열람·구독·전재·포털 노출 등 전파 경로 원천을 별도 후보로 둔다. +- `e21.el03` — 사실적 주장과 의견 표명의 구분, 진실성·상당성·공익성 관련 사실. 구제수단별로 요구되는 바가 다르다: 반론보도는 보도 내용의 진실 여부와 무관하게 청구 가능한 구조이고, 추후보도는 형사절차 관련 보도 후의 결과 확정 사실을 기산 사실로 한다. **각 축의 요건 사실을 서로 옮겨 쓰지 않는다.** +- `e21.el04`·`e21.el05` — 명예·신용·초상·사생활·개인정보 침해와 동의, 그리고 구제수단별 기간 기산 사실(보도를 안 날, 보도가 있은 날, 형사절차 결과 확정일 등)을 축별로 분리해 후보로 남긴다. 도과 판단은 하지 않는다. +- `e21.op01`·`e21.op02` — 항변을 별개 후보로 둔다: 진실성·상당성, 공익성, 의견·논평, 동의, 이미 공개된 사실, 피해자 비식별, 정정보도 요건 흠결, 기간 경과 주장. 성립 여부는 판단하지 않는다. +- **트랙 경계**: 언론중재위원회의 조정·중재 절차와 법원 소송은 서로 다른 트랙이고 전치·병행 여부가 사건마다 다르다. 조정 신청·중재 합의·직권조정 관련 원천이 있으면 절차 판단을 하지 말고 경계 review로 남긴다. 형사 명예훼손·개인정보 규제 접점도 같다. + +## 3. review code 발행 규칙 + +등록된 코드만 쓰고, 발행 경로는 공통 계약이 정한 후보별 `review_items`와 root `unknown_or_unrouted_reviews`뿐이다. + +- 적용 대상 매체 판정 원천이 없거나 이 층의 필수 원천이 빠졌다 → `E21_SPECIAL_LAW_PROFILE_REQUIRED` +- 보도 원본·버전·게시 시각이 불완전하거나 충돌한다 → `E21_PUBLICATION_VERSION_REVIEW` +- 피해자 특정·식별가능성 원천이 부족하다 → `E21_IDENTIFIABILITY_REVIEW` +- 사실 적시와 의견의 구분, 진실성·상당성 원천이 부족하다 → `E21_FACT_OPINION_TRUTH_REVIEW` +- 공익성 형량 자료가 한쪽만 있다 → `E21_PUBLIC_INTEREST_BALANCING_REVIEW` +- 동의·사생활·개인정보 관련 원천이 부족하다 → `E21_CONSENT_PRIVACY_REVIEW` +- 구제수단별 기간 기산 사실이 없거나 축이 섞였다 → `E21_REMEDY_PERIOD_REVIEW` +- 조정·중재·형사·행정 트랙 접점이 있다 → `TRACK_BOUNDARY_REVIEW` + +## 4. 계산·법령 version 경계 + +- CE-05에는 operand와 `ready|partial|blocked` 상태만 보낸다. 위자료 액수·산정 기준을 이 조각이 정하지 않는다. +- SG-04로 pin할 것: 대상 매체 유형별 적용 법률, 구제수단별 청구기간 규범과 기준시점, 시행일·경과규정, 판례 변경 review. +- **E-21 extension은 profile 전용 key를 선언하지 않는다.** 이 profile의 후보는 선언된 `publication_candidates`·`subject_content_candidates`·`personality_privacy_candidates`·`remedy_timing_candidates`·`calculation_operand_candidates`·`boundary_routes`에 담고, 새 key를 만들지 않는다. + +최종 법률결론과 최종 청구액을 확정하지 않는다. 원천 참조와 미해결 review code를 보존한다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-PRODUCT-LIABILITY/prompt_overlay.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-PRODUCT-LIABILITY/prompt_overlay.md new file mode 100644 index 00000000..b28f5232 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/SLP-PRODUCT-LIABILITY/prompt_overlay.md @@ -0,0 +1,51 @@ +# SLP-PRODUCT-LIABILITY — 제조물책임 특별법 profile + +`profile_id`: `SLP-PRODUCT-LIABILITY` · 부착 도메인: E-06 전문가·제조물 책임(`professional_liability`) +조립 위치: 공통 계약 → 의존 도메인(E-01·E-05·X3) → E-06 overlay → **이 profile** → runtime guard + +## 권한 경계 + +- 이 조각은 E-06 overlay를 좁힐 뿐 입력·출력·법리 범위를 넓히지 않는다. 새 출력 key와 새 review key를 만들지 않는다. +- `final_conclusion_forbidden`은 유지된다. 특별법 적용 여부, 결함 인정, 면책 성립, 기간 도과를 확정하지 않는다. +- 법령 조문 번호·법정률·기간·한도·추정 규정의 수치를 이 조각에 상수로 두지 않는다. 전부 SG-04 law-version 후보로 넘긴다. +- 제조물 단서가 있는 사건에만 이 층을 연다. 의료·전문직 과실 사건에 기계적으로 부착하지 않으며, 단순 품질불만·수선 요구만 있는 자료는 E-06 `E06-N02` 경계에 따라 monitor로 둔다. + +## 1. 식별 단서 — source-backed only + +제품명·브랜드 단독 언급은 활성화 근거가 아니다. 다음 원천 조합이 있을 때만 요건 층을 연다. + +1. **제조물성**: 제조·가공된 동산임을 보이는 원천(다른 동산·부동산의 일부를 이루는 경우 포함). 미가공 1차산물·부동산 자체·용역·정보는 경계 사안으로 표시한다. +2. **책임주체**: 제조업자, 자기를 제조업자로 표시하거나 오인하게 할 표시를 한 자, 수입한 자, 제조업자를 알 수 없을 때의 공급자. 지위마다 요건이 다르므로 합치지 않는다. +3. **손해**: 정상적인 사용 상태에서 발생한 생명·신체·재산 손해. **제조물 자체에만 생긴 손해**는 이 profile이 아니라 계약·하자담보 leaf 소관이므로 경계 표시와 함께 보존한다. +4. **결함 유형**: 제조상·설계상·표시상(설명·지시·경고) 중 어느 유형인지 원천에 따라 분기한다. 유형 미특정은 후보 삭제 사유가 아니라 review 사유다. + +## 2. 고유 요건 슬롯·증거 component + +E-06 슬롯에 연결하며 새 슬롯을 만들지 않는다. + +- `e06.el08.product_identity_chain` — 제조물 동일성과 공급·유통·사용·보관 계보. 계보 단절은 대체원인 항변과 연결되는 사실이므로 단절 자체를 후보로 남긴다. +- `e06.el09.product_defect` — 유형별 사실 요소를 구분해 수집한다. 제조상은 설계·제조방법 기준으로부터의 이탈, 설계상은 합리적 대체설계의 채용 가능성과 미채용, 표시상은 합리적 설명·지시·경고의 결여. 유형 간 요소를 뒤섞지 않는다. +- `e06.el04.damage_change`·`e06.el05.causal_mechanism` — 증명 완화 구조를 뒷받침하는 세 사실을 **각각 독립 후보로** 슬롯화한다: ① 정상적 사용 상태에서 손해가 발생했다는 사정, ② 그 손해가 제조업자의 배타적 지배영역에서 비롯되었다는 사정, ③ 그러한 손해가 결함 없이는 통상 발생하지 않는다는 사정. 추정 규정을 적용해 결함·인과를 인정하는 판단은 하지 않는다. +- `e06.op01`~ 반대사실·항변 — 공급 당시 결함 부존재, 공급 당시의 과학·기술 수준으로 결함 발견 불가(개발위험), 법령이 정한 기준 준수로 인한 결함, 원재료·부품 제조업자의 설계·제작 지시 종속을 각각 별개 후보로 둔다. 결함 방지 조치를 하지 않았다는 반대 사정도 함께 보존하되 면책 성립 여부는 판단하지 않는다. +- `e06.el10.law_version_dates` — 기간 축이 둘(손해와 책임주체를 안 날 기준, 제조물 공급일 기준)이므로 기산 사실을 축별로 분리해 후보로 남기고 도과 판단은 하지 않는다. +- 증거 component는 `e06.identity_record`·`e06.primary_elements`·`e06.opposing_record`·`e06.timeline_record` 어휘를 그대로 쓴다. + +## 3. review code 발행 규칙 + +등록된 코드만 쓰고, 발행 경로는 공통 계약이 정한 후보별 `review_items`와 root `unknown_or_unrouted_reviews`뿐이다. + +- 제조물 단서가 있으나 이 층의 필수 원천이 없다 → `SPECIAL_LAW_PROFILE_REQUIRED_REVIEW` +- 결함 유형·표시 내용·사용 상태 원천이 충돌한다 → `E06_DISCLOSURE_OR_PRODUCT_DEFECT_CONFLICT_REVIEW` +- 필수 슬롯이 비었다 → `E06_REQUIRED_ELEMENT_GAP_REVIEW` +- 기산 사실·적용 법률 버전이 미확정이다 → `E06_LAW_VERSION_OR_PERIOD_INPUT_REVIEW` +- 제조물 자체 손해만 있거나 계약·하자담보 구성과 경합한다 → `E06_SPECIAL_LAW_OR_TRACK_BOUNDARY_REVIEW`(필요 시 `TRACK_BOUNDARY_REVIEW` 병행) +- 리콜·안전규제·행정처분 접점이 있다 → `PUBLIC_LAW_NEXUS_REVIEW` +- 결함·사용 상태가 상담록 진술로만 지지된다 → `E06_MEETING_ONLY_SUPPORT_REVIEW` + +## 4. 계산·법령 version 경계 + +- CE-03·CE-05에는 operand와 `ready|partial|blocked` 상태만 보낸다. 산식·법정률·기간·한도를 이 조각이 정하지 않는다. +- SG-04로 pin할 것: 제조물책임 특별법과 일반 불법행위 구성의 적용 법률 후보, 공급일·사고일·인지일 등 사건 기준일별 버전, 경과규정 적용 여부, 추정·면책 규정의 버전 간 차이. +- E-06 extension은 `special_law_profile_candidates`를 선언한다. 이 profile의 후보는 그 key에 담고 새 key를 만들지 않는다. + +최종 법률결론과 최종 청구액을 확정하지 않는다. 원천 참조와 미해결 review code를 보존한다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/_registry_index.json b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/_registry_index.json new file mode 100644 index 00000000..a37dac00 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/special_law_profiles/_registry_index.json @@ -0,0 +1 @@ +{"special_law_profile_registry_index":{"catalog_id":"special_law_profiles","contract_guards":{"final_conclusion_forbidden":true,"meeting_only_evidence_promotion_forbidden":true,"source_membership_required":true,"strict_json_output":true,"unknown_values_require_review":true},"entries":[{"attached_domain_ids":["E-20"],"label_ko":"지식재산 특별법 profile","profile_id":"SLP-IP","prompt_overlay_path":"SLP-IP/prompt_overlay.md","prompt_overlay_sha256":"7ed12d8183992733e79feb25622e6980642308d0705b8409f6a465394d3b7f13","status":"active"},{"attached_domain_ids":["E-21"],"label_ko":"언론피해 구제 특별법 profile","profile_id":"SLP-MEDIA-MEDIATION","prompt_overlay_path":"SLP-MEDIA-MEDIATION/prompt_overlay.md","prompt_overlay_sha256":"04ea71d164111934039640486ba113c22793c95168318c38974897d90ffe5a0e","status":"active"},{"attached_domain_ids":["E-06"],"label_ko":"제조물책임 특별법 profile","profile_id":"SLP-PRODUCT-LIABILITY","prompt_overlay_path":"SLP-PRODUCT-LIABILITY/prompt_overlay.md","prompt_overlay_sha256":"78281b81af2e61647024cbd72f9245d14bd3a6750413d16eb9aa093e6778d491","status":"active"}],"expected_count":3,"generated_at":"2026-08-21T00:00:00+09:00","id_policy":{"final_conclusion_forbidden":true,"law_constants_in_prompt_forbidden":true,"profile_id_pattern":"^SLP-[A-Z][A-Z0-9-]{1,63}$","version_pinning_signal_id":"SG-04"},"registry_version":"Stage1.Assembly.2026-08-21.v1","schema_version":"stage1_special_law_profile_registry_index.v1"}} diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.py b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.py index cdcc341c..c7def471 100644 --- a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.py +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.py @@ -133,22 +133,8 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl "mutually exclusive prompt markers coexist", present, ) - guard = policy.get("deterministic_prompt_size_guard", {}) - maximum_bytes = int(guard.get("max_utf8_bytes", 96000)) - maximum_scalars = int(guard.get("max_unicode_scalars", 24000)) scalar_count = len(compiled) estimate = math.ceil(len(compiled_bytes) / 4) - exceeded = [] - if len(compiled_bytes) > maximum_bytes: - exceeded.append("max_utf8_bytes") - if scalar_count > maximum_scalars: - exceeded.append("max_unicode_scalars") - if exceeded: - raise RuntimeContractError( - str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")), - "compiled prompt exceeds deterministic byte/scalar guard", - {"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded}, - ) manifest = { "schema_version": "compiled_prompt_manifest.v1", "compiled_prompt_sha256": sha256_bytes(compiled_bytes), @@ -156,12 +142,9 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl "deterministic_prompt_size_guard": { "utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, - "max_utf8_bytes": maximum_bytes, - "max_unicode_scalars": maximum_scalars, "not_a_tokenizer": True, "model_context_fit_not_proven": True, "legacy_advisory_estimate": estimate, - "pass": not exceeded, }, "normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator}, "fragments": [ @@ -246,8 +229,19 @@ def collect_domain_fragments( profile_map = profile_paths or {} for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key): if profile not in profile_map: - raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}") - specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id)) + raise RuntimeContractError( + "PROMPT_REQUIRED_FRAGMENT_MISSING", + f"special-law profile prompt path not supplied by caller: {profile}", + {"domain_id": domain_id, "declared_profile": profile, "supplied_profile_ids": sorted(profile_map)}, + ) + profile_path = Path(profile_map[profile]).resolve() + if not profile_path.is_file(): + raise RuntimeContractError( + "PROMPT_REQUIRED_FRAGMENT_MISSING", + f"special-law profile prompt file missing: {profile}", + {"domain_id": domain_id, "declared_profile": profile, "resolved_path": str(profile_path)}, + ) + specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(profile_path), domain_id)) if runtime_guard_path: specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve()))) specs.extend(extra_specs or []) diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.txt b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.txt index cdcc341c..c7def471 100644 --- a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.txt +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_compiler.txt @@ -133,22 +133,8 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl "mutually exclusive prompt markers coexist", present, ) - guard = policy.get("deterministic_prompt_size_guard", {}) - maximum_bytes = int(guard.get("max_utf8_bytes", 96000)) - maximum_scalars = int(guard.get("max_unicode_scalars", 24000)) scalar_count = len(compiled) estimate = math.ceil(len(compiled_bytes) / 4) - exceeded = [] - if len(compiled_bytes) > maximum_bytes: - exceeded.append("max_utf8_bytes") - if scalar_count > maximum_scalars: - exceeded.append("max_unicode_scalars") - if exceeded: - raise RuntimeContractError( - str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")), - "compiled prompt exceeds deterministic byte/scalar guard", - {"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded}, - ) manifest = { "schema_version": "compiled_prompt_manifest.v1", "compiled_prompt_sha256": sha256_bytes(compiled_bytes), @@ -156,12 +142,9 @@ def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tupl "deterministic_prompt_size_guard": { "utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, - "max_utf8_bytes": maximum_bytes, - "max_unicode_scalars": maximum_scalars, "not_a_tokenizer": True, "model_context_fit_not_proven": True, "legacy_advisory_estimate": estimate, - "pass": not exceeded, }, "normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator}, "fragments": [ @@ -246,8 +229,19 @@ def collect_domain_fragments( profile_map = profile_paths or {} for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key): if profile not in profile_map: - raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}") - specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id)) + raise RuntimeContractError( + "PROMPT_REQUIRED_FRAGMENT_MISSING", + f"special-law profile prompt path not supplied by caller: {profile}", + {"domain_id": domain_id, "declared_profile": profile, "supplied_profile_ids": sorted(profile_map)}, + ) + profile_path = Path(profile_map[profile]).resolve() + if not profile_path.is_file(): + raise RuntimeContractError( + "PROMPT_REQUIRED_FRAGMENT_MISSING", + f"special-law profile prompt file missing: {profile}", + {"domain_id": domain_id, "declared_profile": profile, "resolved_path": str(profile_path)}, + ) + specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(profile_path), domain_id)) if runtime_guard_path: specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve()))) specs.extend(extra_specs or []) diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_composition_policy.json b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_composition_policy.json index 11e7a955..91542603 100644 --- a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_composition_policy.json +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/prompt_composition_policy.json @@ -22,8 +22,6 @@ "separator": "\n\n" }, "deterministic_prompt_size_guard": { - "max_utf8_bytes": 96000, - "max_unicode_scalars": 24000, "overflow_fail_code": "PROMPT_SIZE_GUARD_EXCEEDED", "not_a_tokenizer": true, "model_context_fit_not_proven": true, diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/seed_admission_policy.v1.json b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/seed_admission_policy.v1.json new file mode 100644 index 00000000..faea9c32 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/Default_Agent/stage1_runtime/seed_admission_policy.v1.json @@ -0,0 +1,412 @@ +{ + "schema_version": "stage1_seed_admission_policy.v1", + "status": "ACTIVE", + "purpose": "R0 가 이미 수행하는 전체 스키마 검증(worker_output_validator -> schema_subset_validator)의 결과를 등급으로 나누어 처리하는 규칙을 코드 밖에 선언한다. 구조 위반은 막고, 어휘·형태 위반은 정본으로 정규화하거나 정당한 자유 공간으로 수확한 뒤 검토로 남긴다. 값을 지어내지 않는다 — 표에 없는 낱말과 대조 불가능한 참조는 복구하지 않고 격리한다.", + "authority": { + "seed_contract": "task_c_bo_stage_b_domain_bo_seed.v3", + "validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + "schema": "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", + "case_kind_names_forbidden": true, + "note": "검증은 신설이 아니다. R0 가 이미 전체 스키마로 돌리고 결과를 guard_warnings 로 강등하고 있었다(13도메인 1,151건 실측). 이 정책은 그 결과에 처분을 준다. 회차 실측(13도메인): before 1,151건 -> 복구 후 잔여를 후보 격리로 흡수. 키 표류(component_id->slot_id, fact/value->excerpt)는 rename_sibling 으로 이름만 바로잡는다. 워커가 echo 한 sha 는 신선도의 근거가 아니다 — 실물 두 값(계획서 sha ↔ R0 가 읽은 문서 sha)이 판정하고, echo 불일치는 검토로 남긴다. echo 실패는 둘로 가른다 — 다른 슬라이스의 sha 를 echo 하면 워커가 엉뚱한 슬라이스로 작업한 것이라 실물 대조로 잡히지 않는 중대 위반(block)이고, 그 밖은 전사 실수(review)다. 슬라이스를 못 읽어 실물 대조 자체가 불가한 도메인은 강등하지 않는다(quarantine)." + }, + "enforcement": "enforce", + "enforcement_modes": { + "shadow": "복구·격리를 계산하고 기록하되 seed 를 바꾸지 않고 아무것도 막지 않는다. 첫 전환 회차용.", + "enforce": "복구를 적용하고 격리·중단 판정을 실제로 내린다." + }, + "repair_order": [ + "repair_normalize", + "repair_derive", + "repair_harvest", + "repair_evict" + ], + "max_repair_passes": 12, + "disposition_by_code": { + "SCHEMA_ENUM": "repair_normalize", + "SCHEMA_REQUIRED": "repair_derive", + "SCHEMA_ADDITIONAL_PROPERTY": "repair_harvest", + "SCHEMA_PATTERN": "repair_harvest", + "SCHEMA_TYPE": "repair_harvest", + "SCHEMA_MIN_LENGTH": "repair_harvest", + "SCHEMA_MAX_LENGTH": "repair_harvest", + "SCHEMA_UNIQUE_ITEMS": "repair_harvest", + "WORKER_COMPLETION_RECEIPT_MISMATCH": "repair_derive", + "UNREGISTERED_EFFECT_REVIEW_MISSING": "review", + "WORKER_DECLARATION_MISMATCH": "review", + "WORKER_NOT_R0_ELIGIBLE": "review", + "WORKER_SEED_ID_DUPLICATE": "quarantine", + "WORKER_SOURCE_MEMBERSHIP_FAILED": "quarantine", + "WORKER_ROOT_INVALID": "block", + "WORKER_SCHEMA_VERSION_INVALID": "block", + "WORKER_DOMAIN_MISMATCH": "block", + "WORKER_TASK_INSTANCE_MISMATCH": "block", + "WORKER_CONTRACT_GUARD_FAILED": "block", + "WORKER_FINAL_FIELD_PRESENT": "block", + "CALCULATION_REVIEW_MISSING": "review", + "CALCULATION_DOMAIN_NOT_DECLARED": "review", + "EVIDENCE_SLOT_NOT_DECLARED": "review", + "UNKNOWN_VALUE_REVIEW_MISSING": "review", + "WORKER_BO_TYPE_UNREGISTERED": "review", + "WORKER_DEPENDENCY_UNREGISTERED": "review", + "WORKER_DOMAIN_UNREGISTERED": "review", + "WORKER_SLICE_GUARD_MISMATCH": "review", + "WORKER_SLICE_HASH_MISMATCH": "review", + "WORKER_REGISTRY_HASH_MISMATCH": "review", + "SLICE_ROOT_INVALID": "review", + "ACTIVATION_EXPECTED_SET_INVALID": "review", + "DOMAIN_ID_CATALOG_INVALID": "review", + "DOMAIN_ID_CATALOG_DUPLICATE": "review", + "DOMAIN_ID_CATALOG_ID_INVALID": "review", + "FINAL_CONCLUSION_FIELD_FORBIDDEN": "block", + "VALIDATOR_CRASHED": "quarantine", + "VALIDATOR_CRASHED_ROOT": "block", + "SEED_ECHO_SHA_MISMATCH": "review", + "SEED_ECHO_FOREIGN_SLICE": "block", + "SEED_ECHO_SHA_UNVERIFIABLE": "quarantine", + "SEED_ECHO_SHA_UNEXPLAINED": "quarantine" + }, + "unknown_code_disposition": "quarantine", + "surplus_sink": { + "candidate_path": "extensions.worker_surplus", + "by_container": { + "bo_seed_candidates[].element_fact_candidates[]": "payload", + "bo_seed_candidates[].opposing_fact_candidates[]": "payload", + "bo_seed_candidates[].defense_candidates[]": "payload" + }, + "note": "컨테이너 지역 payload 에는 잎 이름 그대로 담는다(하류 어댑터가 payload[\"status\"] 처럼 짧은 이름으로 읽는다). 후보 수준 보관은 {path,value} 덧붙이기 전용 목록이라 어떤 값도 덮이지 않는다.", + "candidate_path_is_list": true + }, + "enum_synonyms": { + "status": { + "ready": "READY", + "ok": "READY", + "complete": "READY", + "ready_with_review": "READY_WITH_REVIEW", + "review": "READY_WITH_REVIEW", + "no_support": "NO_SUPPORT", + "none": "NO_SUPPORT", + "empty": "NO_SUPPORT" + }, + "bo_seed_candidates[].evidence_slot_status[].status": { + "supported": "filled", + "proven": "filled", + "sufficient": "filled", + "satisfied": "filled", + "met": "filled", + "complete": "filled", + "partially_supported": "partial", + "partial_support": "partial", + "insufficient": "partial", + "weak": "partial", + "missing_required": "missing", + "unsupported": "missing", + "absent": "missing", + "none": "missing", + "not_found": "missing", + "conflict": "conflicted", + "contradicted": "conflicted", + "disputed": "conflicted" + }, + "bo_seed_candidates[].calculation_requests[].completeness": { + "complete": "ready", + "ok": "ready", + "sufficient": "ready", + "incomplete": "partial", + "partial_input": "partial", + "blocked_by_missing": "blocked", + "unavailable": "blocked", + "postponed": "deferred", + "later": "deferred" + }, + "bo_seed_candidates[].review_items[].severity": { + "information": "info", + "informational": "info", + "note": "info", + "soft_warning": "review", + "warning": "review", + "minor": "review", + "hard": "hard_warning", + "major": "hard_warning", + "severe": "hard_warning", + "error": "block", + "critical": "block", + "blocker": "block" + }, + "unknown_or_unrouted_reviews[].severity": { + "information": "info", + "informational": "info", + "note": "info", + "soft_warning": "review", + "warning": "review", + "minor": "review", + "hard": "hard_warning", + "major": "hard_warning", + "severe": "hard_warning", + "error": "block", + "critical": "block", + "blocker": "block" + }, + "bo_seed_candidates[].element_fact_candidates[].source_kind": { + "evidence_index": "evidence", + "document": "evidence", + "doc": "evidence", + "event": "event_candidate", + "event_id": "event_candidate", + "meeting": "meeting_clause", + "clause": "meeting_clause", + "consultation": "meeting_clause" + }, + "bo_seed_candidates[].opposing_fact_candidates[].source_kind": { + "evidence_index": "evidence", + "document": "evidence", + "doc": "evidence", + "event": "event_candidate", + "event_id": "event_candidate", + "meeting": "meeting_clause", + "clause": "meeting_clause", + "consultation": "meeting_clause" + }, + "bo_seed_candidates[].defense_candidates[].source_kind": { + "evidence_index": "evidence", + "document": "evidence", + "doc": "evidence", + "event": "event_candidate", + "event_id": "event_candidate", + "meeting": "meeting_clause", + "clause": "meeting_clause", + "consultation": "meeting_clause" + } + }, + "enum_case_fold": true, + "enum_unmapped_disposition": "repair_evict", + "required_derivation": { + "bo_seed_candidates[].element_fact_candidates[].source_id": { + "rule": "first_universe_member_of_sibling_array", + "sibling_candidates": [ + "source_refs", + "refs", + "evidence_refs", + "source_ref" + ] + }, + "bo_seed_candidates[].element_fact_candidates[].source_kind": { + "rule": "universe_kind_of_sibling", + "sibling": "source_id", + "fallback_sibling_arrays": [ + "source_refs", + "refs", + "evidence_refs", + "source_ref" + ], + "note": "형제 source_id 가 먼저 채워지지 않았을 수 있어 참조 배열 폴백을 둔다." + }, + "bo_seed_candidates[].opposing_fact_candidates[].source_id": { + "rule": "first_universe_member_of_sibling_array", + "sibling_candidates": [ + "source_refs", + "refs", + "evidence_refs", + "source_ref" + ] + }, + "bo_seed_candidates[].opposing_fact_candidates[].source_kind": { + "rule": "universe_kind_of_sibling", + "sibling": "source_id", + "fallback_sibling_arrays": [ + "source_refs", + "refs", + "evidence_refs", + "source_ref" + ] + }, + "bo_seed_candidates[].defense_candidates[].source_id": { + "rule": "first_universe_member_of_sibling_array", + "sibling_candidates": [ + "source_refs", + "refs", + "evidence_refs", + "source_ref" + ] + }, + "bo_seed_candidates[].defense_candidates[].source_kind": { + "rule": "universe_kind_of_sibling", + "sibling": "source_id", + "fallback_sibling_arrays": [ + "source_refs", + "refs", + "evidence_refs", + "source_ref" + ] + }, + "bo_seed_candidates[].calculation_requests[].calculation_domain": { + "rule": "sibling_matching_pattern", + "sibling_candidates": [ + "calculation_id", + "calculation_code", + "domain" + ], + "pattern": "^CE-[0-9]{2}$", + "fallback_rename_sibling": [ + "calculation_id", + "calculation_code", + "domain", + "ce_id" + ] + }, + "bo_seed_candidates[].calculation_requests[].completeness": { + "rule": "constant", + "value": "partial", + "note": "워커가 완결성을 말하지 않았다. 낙관도 비관도 아닌 중간값을 두고 검토로 올린다." + }, + "bo_seed_candidates[].calculation_requests[].source_refs": { + "rule": "constant", + "value": [], + "note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다." + }, + "bo_seed_candidates[].evidence_slot_status[].source_refs": { + "rule": "constant", + "value": [], + "note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다." + }, + "bo_seed_candidates[].review_items[].severity": { + "rule": "constant", + "value": "review" + }, + "bo_seed_candidates[].review_items[].review_code": { + "rule": "constant", + "value": "WORKER_REVIEW_UNSPECIFIED" + }, + "bo_seed_candidates[].review_items[].reason": { + "rule": "constant", + "value": "워커가 사유를 적지 않았다. R0 수용 단계에서 채운 자리이므로 Stage2 가 원문을 다시 본다." + }, + "bo_seed_candidates[].review_items[].source_refs": { + "rule": "constant", + "value": [], + "note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다." + }, + "bo_seed_candidates[].legal_effect_candidates[].source_refs": { + "rule": "constant", + "value": [], + "note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다." + }, + "bo_seed_candidates[].party_roles[].party_refs": { + "rule": "constant", + "value": [] + }, + "completion_receipt.task_instance_id": { + "rule": "from_context", + "key": "task_instance_id" + }, + "completion_receipt.domain_id": { + "rule": "from_context", + "key": "domain_id" + }, + "completion_receipt.seed_count": { + "rule": "from_context", + "key": "seed_count" + }, + "completion_receipt.emitted_seed_ids": { + "rule": "from_context", + "key": "emitted_seed_ids" + }, + "completion_receipt.r0_membership_key": { + "rule": "from_context", + "key": "r0_membership_key" + }, + "bo_seed_candidates[].evidence_slot_status[].slot_id": { + "rule": "rename_sibling", + "sibling_candidates": [ + "component_id", + "slot", + "element_slot_id", + "registry_component_id", + "slot_key" + ], + "note": "워커가 같은 자리를 component_id 로 부르는 표류가 실측됐다. 이름만 바로잡고 값은 보존한다." + }, + "bo_seed_candidates[].element_fact_candidates[].excerpt": { + "rule": "rename_sibling", + "sibling_candidates": [ + "fact", + "value", + "content", + "text", + "statement" + ] + }, + "bo_seed_candidates[].opposing_fact_candidates[].excerpt": { + "rule": "rename_sibling", + "sibling_candidates": [ + "fact", + "value", + "content", + "text", + "statement" + ] + }, + "bo_seed_candidates[].defense_candidates[].excerpt": { + "rule": "rename_sibling", + "sibling_candidates": [ + "fact", + "value", + "content", + "text", + "statement", + "defense" + ] + }, + "unknown_or_unrouted_reviews[].severity": { + "rule": "constant", + "value": "review" + }, + "unknown_or_unrouted_reviews[].source_refs": { + "rule": "constant", + "value": [], + "note": "워커가 출처를 적지 않았다. 빈 배열은 \"출처 없음\"이 아니라 \"워커 미기재\"를 뜻하며 매 적용마다 검토가 발행된다." + }, + "unknown_or_unrouted_reviews[].review_code": { + "rule": "rename_sibling", + "sibling_candidates": [ + "code", + "issue_code", + "review_id" + ] + }, + "unknown_or_unrouted_reviews[].reason": { + "rule": "rename_sibling", + "sibling_candidates": [ + "message", + "note", + "description", + "detail" + ] + } + }, + "derivation_failed_disposition": "repair_evict", + "quarantine_budget": { + "max_quarantined_candidate_ratio": 0.34, + "min_admitted_candidates_run": 1, + "note": "격리는 후보 단위다. 한 도메인의 후보가 전부 격리되면 그 도메인은 NO_SUPPORT 로 내려가고(이는 v3 계약상 적법한 상태다) 회차는 계속된다. 회차 전체 격리 비율이 한계를 넘으면 모델 잡음이 아니라 계약 파손이므로 그때 경성 중단한다." + }, + "review": { + "issue_type": "seed_admission_repair", + "quarantine_issue_type": "seed_admission_quarantine", + "severity": "SOFT_WARNING", + "quarantine_severity": "HARD_WARNING", + "downstream_owner": "Stage2", + "record_fields": [ + "candidate_ref", + "path", + "code", + "action", + "received", + "applied", + "rule_id" + ], + "note": "모든 복구는 원래 값(received)과 적용 값(applied)을 함께 남긴다. 조용한 덮어쓰기를 금지하기 위한 것이며, Stage2 는 이 기록만으로 무엇이 바뀌었는지 재구성할 수 있어야 한다. 실질 서술(excerpt·value·reason 등)을 담은 채 밀려난 항목은 기계적 키 이동과 등급을 나눈다 — 같은 SOFT_WARNING 에 섞으면 소수가 다수에 묻힌다.", + "evicted_with_content_issue_type": "seed_admission_content_evicted", + "evicted_with_content_severity": "HARD_WARNING" + }, + "residual_root_disposition": "review" +} diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_compiler_old.py b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_compiler_old.py new file mode 100644 index 00000000..cdcc341c --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_compiler_old.py @@ -0,0 +1,322 @@ +#!/usr/bin/env python3 +"""Deterministically compose common, domain, dependency, and special-law prompts.""" + +from __future__ import annotations + +import argparse +import json +import math +import sys +from dataclasses import dataclass +from pathlib import Path +from typing import Any + +from registry_loader import load_registry +from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json + + +@dataclass(frozen=True) +class FragmentSpec: + fragment_id: str + category: str + path: str + source_domain_id: str | None = None + + +def _priority(policy: dict[str, Any]) -> dict[str, int]: + out: dict[str, int] = {} + for index, item in enumerate(policy.get("fragment_priority", [])): + if isinstance(item, str): + out[item] = index * 10 + elif isinstance(item, dict) and item.get("category"): + out[str(item["category"])] = int(item.get("rank", index * 10)) + if not out: + raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty") + return out + + +def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str: + codes = policy.get("fail_codes") + return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback + + +def _normalize_fragment(path: str) -> tuple[str, bytes, str]: + source = Path(path) + try: + text = source.read_text(encoding="utf-8") + except FileNotFoundError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc + except UnicodeDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc + if source.suffix.lower() == ".json": + try: + value = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc + normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n" + else: + normalized = normalize_text(text) + data = normalized.encode("utf-8") + return normalized, data, sha256_bytes(data) + + +def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]: + priorities = _priority(policy) + required_categories = [str(item) for item in policy.get("required_categories", [])] + categories = {spec.category for spec in specs} + missing_categories = [item for item in required_categories if item not in categories] + if missing_categories: + raise RuntimeContractError( + _fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"), + "required prompt fragment category missing", + missing_categories, + ) + unknown = sorted(categories - set(priorities)) + if unknown: + raise RuntimeContractError( + _fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"), + "fragment category has no priority", + unknown, + ) + prepared: list[dict[str, Any]] = [] + by_id: dict[str, str] = {} + for spec in specs: + text, data, digest = _normalize_fragment(spec.path) + previous = by_id.get(spec.fragment_id) + if previous is not None and previous != digest: + raise RuntimeContractError( + _fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"), + f"same fragment_id has different normalized hashes: {spec.fragment_id}", + {"first": previous, "second": digest}, + ) + by_id[spec.fragment_id] = digest + prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest}) + prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path)) + seen_sha: set[str] = set() + kept: list[dict[str, Any]] = [] + omitted: list[dict[str, Any]] = [] + for item in prepared: + spec = item["spec"] + source_path = Path(spec.path).resolve() + release_root = Path(__file__).resolve().parents[1] + try: + logical_path = source_path.relative_to(release_root).as_posix() + except ValueError as exc: + raise RuntimeContractError( + "PROMPT_FRAGMENT_PATH_INVALID", + "prompt fragment must resolve under the release root", + source_path.name, + ) from exc + public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id} + if item["sha256"] in seen_sha: + omitted.append({**public, "reason": "duplicate_normalized_sha256"}) + continue + seen_sha.add(item["sha256"]) + kept.append(item) + separator = str(policy.get("byte_contract", {}).get("separator", "\n\n")) + compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n" + compiled_bytes = compiled.encode("utf-8") + required_markers = [str(item) for item in policy.get("required_markers", [])] + missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()] + if missing_markers: + raise RuntimeContractError( + _fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"), + "compiled prompt lacks required guard marker", + missing_markers, + ) + for conflict in policy.get("conflict_markers", []): + if isinstance(conflict, list) and len(conflict) > 1: + present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()] + if len(present) > 1: + raise RuntimeContractError( + _fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"), + "mutually exclusive prompt markers coexist", + present, + ) + guard = policy.get("deterministic_prompt_size_guard", {}) + maximum_bytes = int(guard.get("max_utf8_bytes", 96000)) + maximum_scalars = int(guard.get("max_unicode_scalars", 24000)) + scalar_count = len(compiled) + estimate = math.ceil(len(compiled_bytes) / 4) + exceeded = [] + if len(compiled_bytes) > maximum_bytes: + exceeded.append("max_utf8_bytes") + if scalar_count > maximum_scalars: + exceeded.append("max_unicode_scalars") + if exceeded: + raise RuntimeContractError( + str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")), + "compiled prompt exceeds deterministic byte/scalar guard", + {"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded}, + ) + manifest = { + "schema_version": "compiled_prompt_manifest.v1", + "compiled_prompt_sha256": sha256_bytes(compiled_bytes), + "compiled_prompt_bytes": len(compiled_bytes), + "deterministic_prompt_size_guard": { + "utf8_bytes": len(compiled_bytes), + "unicode_scalars": scalar_count, + "max_utf8_bytes": maximum_bytes, + "max_unicode_scalars": maximum_scalars, + "not_a_tokenizer": True, + "model_context_fit_not_proven": True, + "legacy_advisory_estimate": estimate, + "pass": not exceeded, + }, + "normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator}, + "fragments": [ + { + "fragment_id": item["spec"].fragment_id, + "category": item["spec"].category, + "path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()), + "sha256": item["sha256"], + "hash_kind": "file_sha256", + "source_domain_id": item["spec"].source_domain_id, + } + for item in kept + ], + "deduplicated_fragments": omitted, + "required_markers_verified": required_markers, + } + return compiled, manifest + + +def _path_from_config(config_record: dict[str, Any], value: str) -> str: + config_path = config_record.get("config_path") + base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", ".")) + path = Path(value) + return str(path if path.is_absolute() else (base / path).resolve()) + + +def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]: + config = config_record["config"] + values: list[Any] = [] + if isinstance(config.get("prompt_fragments"), list): + values.extend(config["prompt_fragments"]) + for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"): + if isinstance(config.get(key), str): + values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"}) + index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {} + if not values and isinstance(index_entry.get("prompt_overlay_path"), str): + values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"}) + specs: list[FragmentSpec] = [] + for index, value in enumerate(values, 1): + if isinstance(value, str): + path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}" + elif isinstance(value, dict) and isinstance(value.get("path"), str): + path_value = value["path"] + fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}") + else: + continue + specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id)) + return specs + + +def collect_domain_fragments( + domain_id: str, + registry: dict[str, Any], + *, + common_contract_path: str, + profile_paths: dict[str, str] | None = None, + runtime_guard_path: str | None = None, + extra_specs: list[FragmentSpec] | None = None, +) -> list[FragmentSpec]: + entries = registry["entries"] + if domain_id not in entries: + raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}") + specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))] + visited: set[str] = set() + + def add_dependencies(current: str) -> None: + for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key): + if dependency not in entries: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}") + if dependency in visited: + continue + visited.add(dependency) + add_dependencies(dependency) + specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency)) + + add_dependencies(domain_id) + domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id) + if not domain_specs: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}") + specs.extend(domain_specs) + profiles = entries[domain_id]["config"].get("special_law_profiles", []) + profile_map = profile_paths or {} + for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key): + if profile not in profile_map: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}") + specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id)) + if runtime_guard_path: + specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve()))) + specs.extend(extra_specs or []) + return specs + + +def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]: + out: dict[str, str] = {} + for value in values: + if separator not in value: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value) + key, path = value.split(separator, 1) + out[key] = path + return out + + +def _parse_extra(values: list[str]) -> list[FragmentSpec]: + out: list[FragmentSpec] = [] + for value in values: + parts = value.split(":", 2) + if len(parts) != 3: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value) + out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve()))) + return out + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--registry-index", required=True) + parser.add_argument("--domain-id", required=True) + parser.add_argument("--policy", required=True) + parser.add_argument("--common-contract", required=True) + parser.add_argument("--runtime-guard") + parser.add_argument("--profile", action="append", default=[]) + parser.add_argument("--fragment", action="append", default=[]) + parser.add_argument("--output-prompt", required=True) + parser.add_argument("--output-manifest", required=True) + args = parser.parse_args(argv) + try: + registry = load_registry(args.registry_index) + specs = collect_domain_fragments( + args.domain_id, + registry, + common_contract_path=args.common_contract, + profile_paths=_parse_mapping(args.profile), + runtime_guard_path=args.runtime_guard, + extra_specs=_parse_extra(args.fragment), + ) + policy = load_json(args.policy) + compiled, manifest = compile_fragments(specs, policy) + prompt_path = Path(args.output_prompt) + prompt_path.parent.mkdir(parents=True, exist_ok=True) + prompt_path.write_text(compiled, encoding="utf-8", newline="") + from runtime_common import sha256_file + + manifest.update({ + "domain_id": args.domain_id, + "compiled_prompt_path": str(prompt_path), + "registry_index_sha256": registry["index_sha256"], + "composition_policy_path": args.policy, + "composition_policy_sha256": sha256_file(args.policy), + }) + write_json(args.output_manifest, manifest) + emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]}) + return 0 + except RuntimeContractError as exc: + emit_cli_result({"status": "FAILED", "error": exc.as_dict()}) + return 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_compiler_old.txt b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_compiler_old.txt new file mode 100644 index 00000000..cdcc341c --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_compiler_old.txt @@ -0,0 +1,322 @@ +#!/usr/bin/env python3 +"""Deterministically compose common, domain, dependency, and special-law prompts.""" + +from __future__ import annotations + +import argparse +import json +import math +import sys +from dataclasses import dataclass +from pathlib import Path +from typing import Any + +from registry_loader import load_registry +from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json + + +@dataclass(frozen=True) +class FragmentSpec: + fragment_id: str + category: str + path: str + source_domain_id: str | None = None + + +def _priority(policy: dict[str, Any]) -> dict[str, int]: + out: dict[str, int] = {} + for index, item in enumerate(policy.get("fragment_priority", [])): + if isinstance(item, str): + out[item] = index * 10 + elif isinstance(item, dict) and item.get("category"): + out[str(item["category"])] = int(item.get("rank", index * 10)) + if not out: + raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty") + return out + + +def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str: + codes = policy.get("fail_codes") + return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback + + +def _normalize_fragment(path: str) -> tuple[str, bytes, str]: + source = Path(path) + try: + text = source.read_text(encoding="utf-8") + except FileNotFoundError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc + except UnicodeDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc + if source.suffix.lower() == ".json": + try: + value = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc + normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n" + else: + normalized = normalize_text(text) + data = normalized.encode("utf-8") + return normalized, data, sha256_bytes(data) + + +def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]: + priorities = _priority(policy) + required_categories = [str(item) for item in policy.get("required_categories", [])] + categories = {spec.category for spec in specs} + missing_categories = [item for item in required_categories if item not in categories] + if missing_categories: + raise RuntimeContractError( + _fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"), + "required prompt fragment category missing", + missing_categories, + ) + unknown = sorted(categories - set(priorities)) + if unknown: + raise RuntimeContractError( + _fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"), + "fragment category has no priority", + unknown, + ) + prepared: list[dict[str, Any]] = [] + by_id: dict[str, str] = {} + for spec in specs: + text, data, digest = _normalize_fragment(spec.path) + previous = by_id.get(spec.fragment_id) + if previous is not None and previous != digest: + raise RuntimeContractError( + _fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"), + f"same fragment_id has different normalized hashes: {spec.fragment_id}", + {"first": previous, "second": digest}, + ) + by_id[spec.fragment_id] = digest + prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest}) + prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path)) + seen_sha: set[str] = set() + kept: list[dict[str, Any]] = [] + omitted: list[dict[str, Any]] = [] + for item in prepared: + spec = item["spec"] + source_path = Path(spec.path).resolve() + release_root = Path(__file__).resolve().parents[1] + try: + logical_path = source_path.relative_to(release_root).as_posix() + except ValueError as exc: + raise RuntimeContractError( + "PROMPT_FRAGMENT_PATH_INVALID", + "prompt fragment must resolve under the release root", + source_path.name, + ) from exc + public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id} + if item["sha256"] in seen_sha: + omitted.append({**public, "reason": "duplicate_normalized_sha256"}) + continue + seen_sha.add(item["sha256"]) + kept.append(item) + separator = str(policy.get("byte_contract", {}).get("separator", "\n\n")) + compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n" + compiled_bytes = compiled.encode("utf-8") + required_markers = [str(item) for item in policy.get("required_markers", [])] + missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()] + if missing_markers: + raise RuntimeContractError( + _fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"), + "compiled prompt lacks required guard marker", + missing_markers, + ) + for conflict in policy.get("conflict_markers", []): + if isinstance(conflict, list) and len(conflict) > 1: + present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()] + if len(present) > 1: + raise RuntimeContractError( + _fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"), + "mutually exclusive prompt markers coexist", + present, + ) + guard = policy.get("deterministic_prompt_size_guard", {}) + maximum_bytes = int(guard.get("max_utf8_bytes", 96000)) + maximum_scalars = int(guard.get("max_unicode_scalars", 24000)) + scalar_count = len(compiled) + estimate = math.ceil(len(compiled_bytes) / 4) + exceeded = [] + if len(compiled_bytes) > maximum_bytes: + exceeded.append("max_utf8_bytes") + if scalar_count > maximum_scalars: + exceeded.append("max_unicode_scalars") + if exceeded: + raise RuntimeContractError( + str(guard.get("overflow_fail_code", "PROMPT_SIZE_GUARD_EXCEEDED")), + "compiled prompt exceeds deterministic byte/scalar guard", + {"utf8_bytes": len(compiled_bytes), "unicode_scalars": scalar_count, "exceeded_limits": exceeded}, + ) + manifest = { + "schema_version": "compiled_prompt_manifest.v1", + "compiled_prompt_sha256": sha256_bytes(compiled_bytes), + "compiled_prompt_bytes": len(compiled_bytes), + "deterministic_prompt_size_guard": { + "utf8_bytes": len(compiled_bytes), + "unicode_scalars": scalar_count, + "max_utf8_bytes": maximum_bytes, + "max_unicode_scalars": maximum_scalars, + "not_a_tokenizer": True, + "model_context_fit_not_proven": True, + "legacy_advisory_estimate": estimate, + "pass": not exceeded, + }, + "normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator}, + "fragments": [ + { + "fragment_id": item["spec"].fragment_id, + "category": item["spec"].category, + "path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()), + "sha256": item["sha256"], + "hash_kind": "file_sha256", + "source_domain_id": item["spec"].source_domain_id, + } + for item in kept + ], + "deduplicated_fragments": omitted, + "required_markers_verified": required_markers, + } + return compiled, manifest + + +def _path_from_config(config_record: dict[str, Any], value: str) -> str: + config_path = config_record.get("config_path") + base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", ".")) + path = Path(value) + return str(path if path.is_absolute() else (base / path).resolve()) + + +def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]: + config = config_record["config"] + values: list[Any] = [] + if isinstance(config.get("prompt_fragments"), list): + values.extend(config["prompt_fragments"]) + for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"): + if isinstance(config.get(key), str): + values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"}) + index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {} + if not values and isinstance(index_entry.get("prompt_overlay_path"), str): + values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"}) + specs: list[FragmentSpec] = [] + for index, value in enumerate(values, 1): + if isinstance(value, str): + path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}" + elif isinstance(value, dict) and isinstance(value.get("path"), str): + path_value = value["path"] + fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}") + else: + continue + specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id)) + return specs + + +def collect_domain_fragments( + domain_id: str, + registry: dict[str, Any], + *, + common_contract_path: str, + profile_paths: dict[str, str] | None = None, + runtime_guard_path: str | None = None, + extra_specs: list[FragmentSpec] | None = None, +) -> list[FragmentSpec]: + entries = registry["entries"] + if domain_id not in entries: + raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}") + specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))] + visited: set[str] = set() + + def add_dependencies(current: str) -> None: + for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key): + if dependency not in entries: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}") + if dependency in visited: + continue + visited.add(dependency) + add_dependencies(dependency) + specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency)) + + add_dependencies(domain_id) + domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id) + if not domain_specs: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}") + specs.extend(domain_specs) + profiles = entries[domain_id]["config"].get("special_law_profiles", []) + profile_map = profile_paths or {} + for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key): + if profile not in profile_map: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}") + specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id)) + if runtime_guard_path: + specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve()))) + specs.extend(extra_specs or []) + return specs + + +def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]: + out: dict[str, str] = {} + for value in values: + if separator not in value: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value) + key, path = value.split(separator, 1) + out[key] = path + return out + + +def _parse_extra(values: list[str]) -> list[FragmentSpec]: + out: list[FragmentSpec] = [] + for value in values: + parts = value.split(":", 2) + if len(parts) != 3: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value) + out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve()))) + return out + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--registry-index", required=True) + parser.add_argument("--domain-id", required=True) + parser.add_argument("--policy", required=True) + parser.add_argument("--common-contract", required=True) + parser.add_argument("--runtime-guard") + parser.add_argument("--profile", action="append", default=[]) + parser.add_argument("--fragment", action="append", default=[]) + parser.add_argument("--output-prompt", required=True) + parser.add_argument("--output-manifest", required=True) + args = parser.parse_args(argv) + try: + registry = load_registry(args.registry_index) + specs = collect_domain_fragments( + args.domain_id, + registry, + common_contract_path=args.common_contract, + profile_paths=_parse_mapping(args.profile), + runtime_guard_path=args.runtime_guard, + extra_specs=_parse_extra(args.fragment), + ) + policy = load_json(args.policy) + compiled, manifest = compile_fragments(specs, policy) + prompt_path = Path(args.output_prompt) + prompt_path.parent.mkdir(parents=True, exist_ok=True) + prompt_path.write_text(compiled, encoding="utf-8", newline="") + from runtime_common import sha256_file + + manifest.update({ + "domain_id": args.domain_id, + "compiled_prompt_path": str(prompt_path), + "registry_index_sha256": registry["index_sha256"], + "composition_policy_path": args.policy, + "composition_policy_sha256": sha256_file(args.policy), + }) + write_json(args.output_manifest, manifest) + emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]}) + return 0 + except RuntimeContractError as exc: + emit_cli_result({"status": "FAILED", "error": exc.as_dict()}) + return 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_composition_policy_old.json b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_composition_policy_old.json new file mode 100644 index 00000000..11e7a955 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_outdated/prompt_composition_policy_old.json @@ -0,0 +1,46 @@ +{ + "schema_version": "prompt_composition_policy.v2", + "fragment_priority": [ + {"category": "common_contract", "rank": 10}, + {"category": "common_dependency", "rank": 20}, + {"category": "domain", "rank": 30}, + {"category": "special_law_profile", "rank": 40}, + {"category": "runtime_guard", "rank": 50} + ], + "required_categories": ["common_contract", "domain"], + "deduplication": { + "algorithm": "sha256", + "scope": "normalized_fragment_bytes", + "same_sha_action": "keep_first_by_priority_then_id", + "same_fragment_id_different_sha": "PROMPT_FRAGMENT_ID_HASH_CONFLICT" + }, + "byte_contract": { + "encoding": "UTF-8", + "line_endings": "LF", + "fragment_trailing_newline": "exactly_one", + "compiled_trailing_newline": "exactly_one", + "separator": "\n\n" + }, + "deterministic_prompt_size_guard": { + "max_utf8_bytes": 96000, + "max_unicode_scalars": 24000, + "overflow_fail_code": "PROMPT_SIZE_GUARD_EXCEEDED", + "not_a_tokenizer": true, + "model_context_fit_not_proven": true, + "legacy_advisory_estimate": { + "formula": "ceil(utf8_bytes/4)", + "execution_authority": false + } + }, + "required_markers": ["final_conclusion_forbidden"], + "conflict_markers": [], + "fail_codes": { + "required_fragment_missing": "PROMPT_REQUIRED_FRAGMENT_MISSING", + "unknown_priority_category": "PROMPT_PRIORITY_UNKNOWN", + "fragment_id_hash_conflict": "PROMPT_FRAGMENT_ID_HASH_CONFLICT", + "prompt_size_guard_exceeded": "PROMPT_SIZE_GUARD_EXCEEDED", + "required_guard_missing": "PROMPT_FINAL_CONCLUSION_GUARD_MISSING", + "exclusive_marker_conflict": "PROMPT_EXCLUSIVE_MARKER_CONFLICT", + "fragment_path_invalid": "PROMPT_FRAGMENT_PATH_INVALID" + } +} diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem.md new file mode 100644 index 00000000..7734d1b7 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem.md @@ -0,0 +1,54 @@ +# special_law_profiles 자산 부재 문제 — 판단서 + +작성일 2026-08-21 · 대상 릴리스 `v.7/extension_research/Default_Agent` · 계기 `stage_1_part_2_v.8.yml` 첫 task 실패(`PROMPT_REQUIRED_FRAGMENT_MISSING`) + +--- + +## 0. 실측 요약 (이 문서의 근거) + +| 확인 항목 | 실측 결과 | 근거 | +|---|---|---| +| `special_law_profiles/` 디렉터리 | 릴리스에 **없음** | `Default_Agent` 하위 0건 | +| SLP 실사용 ID | **3종** — `SLP-PRODUCT-LIABILITY`(E-06), `SLP-IP`(E-20), `SLP-MEDIA-MEDIATION`(E-21) | 각 `domains/*/domain_config.json` | +| 그 외 SLP 이름 | `SLP-AUTO-DAMAGE`·`SLP-STATE-COMPENSATION` 은 **스키마 주석 안에만** 존재 | `domain_slice.schema.v2.json:101` | +| 조립 시 하드 실패 지점 | profile ID가 `profile_paths` 에 없으면 무조건 raise | `stage1_runtime/prompt_compiler.py:230-232` | +| 런타임이 기대하는 경로 | `Default_Agent/special_law_profiles//prompt_overlay.md` | `Claude_YAML/Stage_1_Registry_Runtime_v1.yml:473, 492-494` | +| fanout 이 profile ID를 얻는 곳 | `registry["entries"][did]["config"]` = 실제 `domain_config.json` | `Stage_1_Registry_Runtime_v1.yml:311` | +| `runtime_manifest.json` | 432 entries, `special_law` 문자열 **0건** | manifest 전문 검색 | +| 조각 경로 제약 | 릴리스 루트(`Default_Agent/`) 밖이면 `PROMPT_FRAGMENT_PATH_INVALID` | `prompt_compiler.py` `compile_fragments()` | + +즉 **설정은 요구하고, 런타임은 경로까지 정해 두었으며, 실물 파일만 없다.** 원인 분석의 "자산이 릴리스에 없다"는 진단은 이 릴리스에서도 그대로 성립한다. + +--- + +## 1. 질문 1 — `special_law_profiles` 자산을 어떻게 할 것인가 + +**정공법(자산 신설)으로 가되, 범위를 선언된 3종으로 한정한다.** `Default_Agent/special_law_profiles/{SLP-PRODUCT-LIABILITY, SLP-IP, SLP-MEDIA-MEDIATION}/prompt_overlay.md` 를 만들고, 카탈로그 `special_law_profiles/_registry_index.json` 을 함께 두어 `registry_validator` 의 `REFERENCE_CATALOG_NOT_PROVIDED` 경고를 해소한 뒤, `runtime_manifest.json` 에 3~4건을 등재하고 `runtime_artifact_count` 를 증분한다. `prompt_compiler.py:230` 의 raise 를 완화하는 2안은 채택하지 않는다 — 그 raise 는 "선언했으면 반드시 조립한다"는 fail-closed 계약이고, E-06 의 `seed_prompt_overlay.md:5, :34` 가 "제조물 사건은 `SLP-PRODUCT-LIABILITY` 없이 종결하지 않는다"고 프롬프트 본문에서 명시하고 있어, 프로필을 건너뛰고 review_items 에만 표시하면 프롬프트가 스스로 금지한 상태(특별법 검토 없는 제조물 분석)를 산출물로 내보내게 된다. 크기 가드 제거와 성격이 다른 이유가 여기 있다 — 크기 가드는 토크나이저가 아닌 추정치에 근거한 보수적 차단이었지만, 이 raise 는 법리 완결성 계약이다. 콘텐츠 부담도 실제로는 작다: `ensemble_v2.md` 상 특별법은 "도메인 payload 플래그"이고 시행일·경과규정·버전 pinning 은 SG-04(`governing_law_version_signals.json`)가 코어 signal 로 흡수하므로, 프로필 overlay 는 **법령 조문·요율·기간 상수를 담지 않고** ① 특별법 후보 식별 단서, ② 그 특별법 고유 요건 슬롯과 증거 component, ③ 발행할 review code, ④ 버전 판단은 SG-04 로 넘긴다는 경계 선언만 담는 얇은 조각이면 충분하다(도메인 overlay 와 동일 톤, `final_conclusion_forbidden` 유지). 스키마 주석이 말하는 "9종"은 릴리스 어디에도 실체·목록이 없으므로 추종하지 말고, 나머지 6종은 해당 도메인이 `domain_config.json` 에서 실제로 선언할 때 같은 형식으로 증설한다. + +### 구현 시 지켜야 할 계약 (실측 기반) + +1. 경로는 `Default_Agent/special_law_profiles//prompt_overlay.md` 고정 — YAML 이 `PROFILE_DIR + "/" + profile_id + "/prompt_overlay.md"` 로 조립한다. +2. 릴리스 루트 밖 경로 금지(`compile_fragments` 의 `relative_to(release_root)` 검사). +3. 3개 파일의 정규화 sha256 이 서로 달라야 한다 — 동일하면 `deduplication` 이 뒤엣것을 `duplicate_normalized_sha256` 로 탈락시킨다. +4. UTF-8·LF·후행 개행 1개(`byte_contract`). 조각 카테고리는 `special_law_profile`(rank 40)로 이미 정책에 등록되어 있어 정책 수정은 불필요하다. +5. `runtime_manifest.json` 은 `{"path", "sha256"}` 배열이며 `runtime_artifact_count` 와 길이가 일치해야 한다. + +### 문자열 절단(`SLP-PRODUCT-LIABI`, 17자)에 대하여 + +v.7 릴리스 전체를 검색해도 절단형은 0건이고, `domain_config.json` 과 E-06 overlay 모두 21자 정식 ID를 쓴다. 스키마 패턴도 절단형을 그대로 통과시키므로(길이 하한만 있음) 스키마가 자른 것도 아니다. 이 릴리스에서는 원인을 설명할 근거가 없으므로, 실행 워크스페이스의 `main.py` 생성 단계(Stage YAML)에서 확인하는 것이 맞다. 다만 **절단을 고쳐도 파일이 없으면 동일하게 실패**하므로 자산 신설이 선행 조건이다. + +--- + +## 2. 질문 2 — ★ Insight 1·2번은 실행을 위해 고쳐야 하는가 + +**둘 다 stage 1 실행을 막는 결함이 아니다 — 1번은 문서 드리프트, 2번은 이미 수정된 과거 이력의 잔재다.** 1번(인덱스의 `special_law_profiles` 가 전부 `[]` 인데 E-06·E-20·E-21 config 에는 값이 있음)은, `registry_loader.load_registry()` 가 엔트리를 `config`(= `domain_config.json` 실제 로드본)와 `index_entry`(인덱스 원문)로 분리해 담고 `prompt_compiler.py:227`, `domain_slice_compiler.py:482`, fanout 빌더(`Stage_1_Registry_Runtime_v1.yml:311`)가 **모두 `config` 만** 읽기 때문에 실행 경로에서 소비되지 않는다. `registry_validator` 도 `prompt_overlay` 에 대해서만 인덱스↔config 불일치(`PROMPT_OVERLAY_REFERENCE_DIVERGENCE`)를 검사할 뿐 이 필드는 대조하지 않는다. 따라서 지금 실패는 인덱스의 `[]` 때문이 아니라 값이 정상적으로 흘러가 파일을 찾는 단계에서 나는 것이며, 인덱스 정합화는 실행 차단 해소와 무관한 위생 작업이다(다만 인덱스만 보고 판단하는 사람·도구를 오도하므로 자산 신설과 함께 채워 두는 편이 낫다 — 인덱스 수정 시 `config_sha256` 재계산 대상은 아니고 인덱스 자체 해시만 갱신되면 된다). 2번(스키마 패턴이 소문자라 SLP slice 가 거부된다는 주석)은 **주석만 남고 패턴은 이미 고쳐져 있다**: 현재 `domain_slice.schema.v2.json:104` 의 패턴은 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|[a-z][a-z0-9_]{1,127})$` 이고, 실측 결과 `SLP-PRODUCT-LIABILITY`·`SLP-IP`·`SLP-MEDIA-MEDIATION` 3종 모두 매치된다. 즉 고칠 대상은 스키마 패턴이 아니라 사실과 어긋난 `$comment` 문구이며(더불어 그 주석이 존재한다고 전제하는 `special_law_profiles/_registry_index.json` 이 실재하지 않는다는 점까지 함께 정리해야 한다), 이는 실행 차단이 아닌 문서 정확성 문제다. + +### 정리 — 실행 재개에 필요한 것 / 아닌 것 + +| 항목 | 실행 차단 | 조치 | +|---|---|---| +| `special_law_profiles//prompt_overlay.md` 3종 부재 | **예 (유일한 차단)** | 신설 | +| `runtime_manifest.json` 미등재 | 아니오(실행은 되나 릴리스 무결성 미봉인) | 자산 신설과 동시에 등재 | +| `_registry_index.json` 의 `special_law_profiles: []` | 아니오 | 위생 차원에서 값 채움 | +| `domain_slice.schema.v2.json:101` 주석 | 아니오 | 주석 수정(패턴은 이미 정상) | +| `SLP-PRODUCT-LIABI` 절단 | 파일이 생긴 뒤 재확인 | 실행 워크스페이스 `main.py` 생성부 점검 | diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem_v.2.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem_v.2.md new file mode 100644 index 00000000..4ac0220d --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem_v.2.md @@ -0,0 +1,74 @@ +# special_law_profiles 그리고 insight 답변 분석 + +작성일 2026-08-21 · 검증 대상 `extension_research/assets_special_law_problem.md`(이하 "v.1") · 방법: v.1의 모든 사실 주장과 권고를 릴리스 실물(`Default_Agent/`)·런타임 코드(`stage1_runtime/`)·참조 런타임 YAML(`Claude_YAML/Stage_1_Registry_Runtime_v1.yml`)에 대해 명령 단위로 재실측 + +**검증 범위의 경계**: 실제 실패한 실행본 `stage_1_part_2_v.8.yml` 은 이 저장소에 없다(`/Users/jsahn/Works/작업결과/v2` 소재). 따라서 "런타임이 기대하는 경로" 류의 주장은 이 폴더의 참조본 YAML 근거이며, 실행본이 동일 규약을 쓴다는 것은 에러 문구(`special-law profile prompt path missing`)가 `prompt_compiler.py:232` 의 raise 문자열과 일치한다는 점으로 뒷받침된다. 이 경계는 v.1이 명시하지 않았던 것으로, 본 리포트에서 명시한다. + +--- + +## 1. 질문 1 — special_law_profiles 자산 처리 방향 + +### 1.1 검증 내용 + +**유지 판정(재실측으로 확인된 주장) 5건** + +| v.1 주장 | 재실측 결과 | 판정 | +|---|---|---| +| `special_law_profiles/` 실물 부재가 유일한 실행 차단 | `Default_Agent` 하위 해당 디렉터리·파일 0건, `prompt_compiler.py:231-232` 는 `profile_map` 에 키 없으면 무조건 raise | **유지** | +| 신설 범위는 선언된 3종 한정 | config 선언은 E-06·E-20·E-21 의 3종뿐. `SLP-AUTO-DAMAGE`·`SLP-STATE-COMPENSATION` 은 `domain_slice.schema.v2.json:101` **주석 문자열 안에만** 존재 | **유지** | +| raise 완화(2안) 기각 | v.1 은 E-06 근거만 들었으나, 재실측 결과 **3개 도메인 전부** 프롬프트 본문이 합성을 강제한다: E-06 overlay:5·:34("SLP-PRODUCT-LIABILITY 없이 종결하지 않는다"), E-20 overlay:13("**SLP-IP를 반드시 합성**하고 … SG-04로 pin한다"), E-21 overlay:13("**SLP-MEDIA-MEDIATION을 반드시 합성**…"). 완화 시 세 도메인 모두 자기 프롬프트가 금지한 산출물을 내게 된다 | **유지·강화** | +| 경로 계약 `special_law_profiles//prompt_overlay.md` + 릴리스 루트 하위 강제 | 참조본 YAML:473(`PROFILE_DIR`)·:492-494(`--profile =//prompt_overlay.md`), `compile_fragments` 의 `relative_to(release_root)` 검사 확인 | **유지** | +| overlay 는 조문·요율·기간 상수 없는 얇은 조각(버전 판단은 SG-04 위임) | `ensemble_v2.md:22`("special_law는 도메인 payload 플래그")·`:175`(SG-04 가 "적용 법률·특별법 후보, 시행일·경과규정 … version pinning" 흡수) 재확인. E-20·E-21 overlay:13 도 "SG-04로 pin"을 명시해 위임 구조가 이미 프롬프트에 내장돼 있음 | **유지** | + +**정정 판정(v.1 이 부정확했던 주장) 3건** + +1. **"3개 파일의 정규화 sha256 이 서로 달라야 한다"(v.1 계약 3) → 하드 계약 아님.** dedup(`duplicate_normalized_sha256`)의 스코프는 `compile_fragments()` **1회 호출 = 도메인 1개 컴파일 내부**다. 각 도메인은 SLP 를 1종만 선언하므로 한 컴파일에 `special_law_profile` 조각은 최대 1개이고, 서로 다른 도메인의 profile 파일끼리는 애초에 같은 dedup 집합에 들어가지 않는다. 실제 제약은 "같은 컴파일 안의 다른 조각(공통 계약·도메인 overlay 등)과 정규화 바이트가 같으면 탈락" 뿐이며 현실적으로 발생하지 않는다. 3개 파일 내용을 서로 다르게 쓰는 것은 여전히 옳지만(내용상 당연), 위반 시 조립이 깨지는 계약은 아니다 — **권장으로 강등**. + +2. **"카탈로그 `special_law_profiles/_registry_index.json` 을 함께 두어 `REFERENCE_CATALOG_NOT_PROVIDED` 경고 해소" → 파일 생성만으로는 해소되지 않는다.** `registry_validator` 의 카탈로그 주입 경로는 두 가지뿐이다: ① 도메인 인덱스의 `reference_catalogs`(또는 `reference_sets`) 맵 — 현재 없음, ② CLI `--reference-catalog KIND=PATH`. 파일을 만들어 두기만 하면 validator 는 그 존재를 모른다. 그리고 재실측 결과 **참조본 A0 는 카탈로그의 정본 위치를 이미 확정해 두었다**: `--reference-catalog profile=Default_Agent/platform/reference_catalogs/profile_ids.json`(YAML A0 호출부). 그 디렉터리에 실재하는 것은 `calculation_ids.json`·`domain_ids.json` 뿐이고 `profile_ids.json` 은 없다. 즉 카탈로그의 1차 신설 위치는 v.1 이 말한 `special_law_profiles/_registry_index.json` 이 아니라 **`platform/reference_catalogs/profile_ids.json`** 이며(형식은 `calculation_ids.json` 과 동형: `{"catalog_id": …, "entries": [{"profile_id": "SLP-…", …}]}` — validator 의 `_catalog_ids` 가 `profile_id` 키를 수집한다), `special_law_profiles/_registry_index.json` 은 별도 목적(자산 인덱스 + sha 봉인, 아래 종합 답변)으로 두는 이원 구조가 도메인 registry 와 동형이다. + +3. **"`runtime_manifest.json` 에 3~4건 등재하고 count 증분"(v.1 계약 5 포함) → 현행 장부 관례와 어긋나는 권고.** 재실측 결과 manifest 의 `domains/` 51건은 `module_role_projection.json` 25 + `structure_types.json` 25 + `_common/common_worker_contract.md` 1 이 전부로, **`domain_config.json`(26개)과 `seed_prompt_overlay.md`(26개)는 하나도 등재돼 있지 않다.** registry 자산의 무결성 봉인은 manifest 가 아니라 registry 인덱스가 담당한다(`config_sha256` 는 `load_registry` 가 로드 시 강제 대조, `prompt_overlay_sha256` 는 `registry_validator` 가 대조). 같은 성격의 자산인 SLP overlay 만 manifest 에 싣는 것은 장부 비대칭이며, 봉인의 동형 위치는 **SLP 자체 인덱스에 `prompt_overlay_sha256` 를 기재**하는 것이다. 또한 `runtime_artifact_count` 를 소비·대조하는 코드는 릴리스·YAML 어디에도 없다(grep 0건) — "길이와 일치해야 한다"는 실행 계약이 아니라 장부 자체 일관성 관례다. manifest 등재는 **선택 사항으로 강등**. + +**신규 발견(v.1 누락, 별건 플래그) 2건** + +- 참조본 A0 가 요구하는 자산 5개 중 **4개가 릴리스에 없다**: `platform/schemas/registry_index.schema.json`, `platform/schemas/domain_config.schema.json`, `platform/reference_catalogs/signal_ids.json`, `platform/reference_catalogs/profile_ids.json`. `load_json` 은 파일 부재 시 `FILE_NOT_FOUND` 로 raise 하므로 **참조본 A0 를 그대로 실행하면 SLP 문제에 도달하기 전에 죽는다.** 실행본 v.8 은 A1(프롬프트 조립)까지 도달했으므로 A0 요구가 참조본과 다르다는 뜻이다 — SLP 와 별개의 참조본↔릴리스 정합 결손으로 기록해 둔다. +- 도메인 인덱스의 `review_items` 는 `ASSEMBLY_INDEPENDENT_REVIEW_PENDING` 1건뿐, SLP 자산 미구축은 기록돼 있지 않다. `closure_gate.pass: true`·`missing_paths: []` 와 함께, **설계 장부가 이 결손을 인지하지 못한 채 닫혀 있었다**는 방증이다(전체 상태 표지는 `program_release_status: STAGE1_NOT_RELEASE_READY` 로 정직). + +### 1.2 개선 항목을 반영한 종합 답변 + +정공법(자산 신설) 결론과 3종 한정, raise 완화 기각, 얇은 조각 설계는 그대로 유지한다 — 기각 논거는 오히려 강해졌다(합성 강제 문구가 E-06 만이 아니라 E-20:13·E-21:13 까지 3개 도메인 전부의 프롬프트 본문에 있다). 다만 신설 범위를 v.1 의 "overlay 3개 + 카탈로그 1개 + manifest 등재"에서 다음 **4점 세트**로 정정한다. ① `Default_Agent/special_law_profiles/{SLP-PRODUCT-LIABILITY, SLP-IP, SLP-MEDIA-MEDIATION}/prompt_overlay.md` 3종 — 유일한 실행 차단 해소분이며, 조문·요율·기간 상수 없이 식별 단서·요건 슬롯·review code·SG-04 경계 선언만 담는 얇은 조각(UTF-8·LF·후행 개행 1, `special_law_profile` rank 40 은 정책에 기등록). ② `special_law_profiles/_registry_index.json` — profile_id·prompt_overlay_path·**prompt_overlay_sha256** 을 기재하는 자산 인덱스. sha 봉인의 동형 위치는 manifest 가 아니라 여기다(도메인 overlay 의 봉인이 도메인 인덱스에 있듯이). 스키마 주석이 전제한 파일이기도 하다. ③ `platform/reference_catalogs/profile_ids.json` — 참조본 A0 가 `--reference-catalog profile=` 로 이미 경로를 확정해 둔 검증용 ID 카탈로그(`calculation_ids.json` 과 동형, `profile_id` 키). 카탈로그는 파일만 만들면 발견되지 않고 validator 호출부 또는 인덱스 `reference_catalogs` 에 배선돼야 한다는 점이 v.1 에서 빠져 있었다. ④ `runtime_manifest.json` 등재는 **하지 않거나, 하려면 장부 정책을 먼저 정한다** — 현행 manifest 는 domain_config·seed_prompt_overlay 를 싣지 않으므로 SLP overlay 만 실으면 비대칭이 된다. "3개 파일 sha 상이" 는 계약이 아니라 권장(dedup 은 도메인별 컴파일 내부 스코프)으로 정정한다. 절단 문자열(`SLP-PRODUCT-LIABI`) 판단은 v.1 그대로 유지한다: 릴리스 내 절단형 0건, 스키마도 통과시키므로 원인은 실행 워크스페이스 `main.py` 생성부에서 찾아야 하고, 어느 쪽이든 자산 신설이 선행 조건이다. + +--- + +## 2. 질문 2 — ★ Insight 1·2번이 실행을 위해 수정해야 하는 문제인가 + +### 2.1 검증 내용 + +**Insight 1(인덱스 `[]` ↔ config 값 불일치) — v.1 의 "실행 비차단·위생 작업" 결론 유지, 근거 보강, 표현 1건 정정.** + +- 소비 경로 재확인: `load_registry` 는 엔트리를 `config`(domain_config.json 실물 로드본)와 `index_entry`(인덱스 원문)로 분리 저장하고(`registry_loader.py:111-117`), `prompt_compiler.py:228`·`domain_slice_compiler.py:482`·fanout 빌더(참조본 YAML:311) 는 전부 `config` 만 읽는다. `registry_validator` 의 `REFERENCE_FIELDS` 순회도 `config.get(field)` 만 본다. 인덱스의 `special_law_profiles: []` 를 읽는 코드는 **0곳** — 실행 비차단 확정. +- 해시 영향 재실측: 인덱스의 이 필드를 채워도 ① 각 `domain_config.json` 은 무변경이므로 `load_registry` 가 강제 대조하는 `config_sha256` 은 그대로 유효하고, ② 인덱스 파일 자체의 해시는 로드 시점에 `sha256_file` 로 **자기계산**될 뿐 기대값 장부가 어디에도 없다(`runtime_manifest.json` 에 `_registry_index.json` 미등재 실측). **v.1 의 "인덱스 자체 해시만 갱신되면 된다"는 표현은 정정한다 — 갱신할 해시 장부 자체가 존재하지 않으므로, 채움은 아무 해시도 깨뜨리지 않고 아무 갱신도 요구하지 않는다.** 유일한 유보: 실행본 A0 가 `--index-schema` 검증을 수행한다면 채운 값이 그 스키마와 정합해야 하는데, 참조본이 가리키는 `registry_index.schema.json` 실물이 없어 현재 확인 불가. +- 같은 유형의 더 큰 사례(신규): 인덱스가 E-06·E-20·E-21 에 선언한 `profile_schema_path`·`extension_schema_path`·`fixture_manifest_path`·`effect_projection_schema_path` 의 실물이 **12건 전부 MISSING** 이다(실측). 그러나 어떤 실행 경로도 이 4경로를 로드하지 않으므로 역시 비차단 — "인덱스 선언 ≠ 실물"은 SLP 필드 하나가 아니라 인덱스 전반의 상태이며, 위생 작업의 실제 범위는 Insight 1 이 지적한 것보다 넓다. + +**Insight 2(스키마 소문자 패턴 결함) — v.1 의 "이미 수정됨·주석만 잔재" 결론 유지, 실효성 보강.** + +- 패턴 재검증: `domain_slice.schema.v2.json:104` 의 현재 패턴 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|[a-z][a-z0-9_]{1,127})$` 에 3종 SLP ID 전부 매치(re 실측). 주석(:101)이 서술하는 "소문자 패턴이라 거부" 상태는 현재 파일에 존재하지 않는다. +- 실효성 보강(신규): 이 패턴은 장식이 아니라 **실행 경로에서 실제 소비된다** — 참조본 A2 가 `--slice-schema` 로 이 스키마를 `domain_slice_compiler` 에 전달하고(YAML:654·:686), `schema_subset_validator` 는 `pattern` 키워드를 지원하며 위반 시 `SCHEMA_PATTERN` 을 발행한다(:136-139). 즉 "패턴이 이미 고쳐져 있다"는 사실은 slice 검증 통과 여부를 실제로 좌우하는 확인이고, 수정 대상은 사실과 어긋난 `$comment` 문구(그리고 그 주석이 전제하는 `special_law_profiles/_registry_index.json` 의 실물화 — 질문 1 의 ② 신설로 자연 해소)뿐이다. + +### 2.2 개선 항목을 반영한 종합 답변 + +Insight 1·2번 모두 **stage 1 YAML 실행을 위해 수정해야 하는 문제가 아니라는 v.1 의 결론은 재검증에서 그대로 성립한다.** 1번의 인덱스 `special_law_profiles: []` 는 실행 경로 어디에서도 읽히지 않고(모든 소비자가 `config` 측을 읽는다), 채워 넣더라도 `config_sha256` 은 무변경·인덱스 자체는 기대값 해시 장부가 아예 없어 어떤 봉인도 깨지 않는다 — v.1 이 말한 "인덱스 해시 갱신"은 불필요한 것이 아니라 **개념적으로 존재하지 않는 절차**였다는 점만 정정한다. 아울러 위생 작업의 실제 범위는 이 필드 하나가 아니다: 같은 인덱스가 선언한 profile·extension·fixture·effect_projection 4경로의 실물이 세 도메인 12건 전부 부재한 상태이므로, 인덱스 정합화를 하려면 이 선언들까지 함께 실물화하거나 제거·보류 표기해야 장부가 다시 참말이 된다. 2번은 스키마 패턴이 이미 SLP 대문자 어휘를 받도록 수정돼 있고 그 패턴이 A2 의 `--slice-schema` 경유로 실제 slice 검증에 소비됨을 확인했으므로, 남은 작업은 실행과 무관한 문서 정확성 — 과거 결함 서술로 굳어 있는 `$comment` 를 현재형 사실("어휘는 SLP 카탈로그 소관, 패턴은 수용")로 고쳐 쓰는 것 — 뿐이다. 요컨대 에이전트가 stage 1 을 다시 실행 가능하게 만드는 유일한 수정은 질문 1 의 profile 자산 신설이고, Insight 1·2 는 그 작업에 얹어 처리하면 되는 장부·문서 위생이다. + +--- + +## 3. 총괄 — v.1 대비 변경 요약 + +| 항목 | v.1 | v.2 판정 | +|---|---|---| +| 자산 신설(정공법)·3종 한정·raise 유지 | 채택 | **유지** (합성 강제 문구 3개 도메인 전부로 논거 강화) | +| overlay 얇은 조각 설계(SG-04 위임) | 채택 | **유지** | +| 카탈로그 위치·배선 | `special_law_profiles/_registry_index.json` "두면 경고 해소" | **정정**: 1차 정본은 `platform/reference_catalogs/profile_ids.json`(A0 가 경로 기확정) + validator 배선 필요. SLP 자체 인덱스는 sha 봉인용으로 별도 | +| sha 봉인 위치 | `runtime_manifest.json` 등재 | **정정**: manifest 는 config·overlay 를 원래 싣지 않음(51건 실측). 봉인은 SLP 인덱스의 `prompt_overlay_sha256` 이 동형. manifest 등재·count 는 선택(외부 소비 0) | +| "3개 파일 sha 상이" 계약 | 하드 계약 | **정정**: dedup 은 도메인별 컴파일 내부 스코프 — 권장으로 강등 | +| Insight 1 실행 비차단 | 채택 | **유지** + "인덱스 해시 갱신" 표현 정정(갱신할 장부 없음) + 위생 범위 확장(선언 4경로 12건 MISSING) | +| Insight 2 실행 비차단 | 채택 | **유지** + 패턴의 실행 소비 확인(A2 `--slice-schema` + `SCHEMA_PATTERN`)으로 실효성 보강 | +| (신규) 참조본 A0 요구 자산 4종 부재 | — | **별건 플래그**: registry_index.schema·domain_config.schema·signal_ids·profile_ids 부재 — 참조본 그대로면 A0 가 `FILE_NOT_FOUND` 로 SLP 이전에 실패 | diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem_v.2_diff.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem_v.2_diff.md new file mode 100644 index 00000000..4ac0220d --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/assets_special_law_problem_v.2_diff.md @@ -0,0 +1,74 @@ +# special_law_profiles 그리고 insight 답변 분석 + +작성일 2026-08-21 · 검증 대상 `extension_research/assets_special_law_problem.md`(이하 "v.1") · 방법: v.1의 모든 사실 주장과 권고를 릴리스 실물(`Default_Agent/`)·런타임 코드(`stage1_runtime/`)·참조 런타임 YAML(`Claude_YAML/Stage_1_Registry_Runtime_v1.yml`)에 대해 명령 단위로 재실측 + +**검증 범위의 경계**: 실제 실패한 실행본 `stage_1_part_2_v.8.yml` 은 이 저장소에 없다(`/Users/jsahn/Works/작업결과/v2` 소재). 따라서 "런타임이 기대하는 경로" 류의 주장은 이 폴더의 참조본 YAML 근거이며, 실행본이 동일 규약을 쓴다는 것은 에러 문구(`special-law profile prompt path missing`)가 `prompt_compiler.py:232` 의 raise 문자열과 일치한다는 점으로 뒷받침된다. 이 경계는 v.1이 명시하지 않았던 것으로, 본 리포트에서 명시한다. + +--- + +## 1. 질문 1 — special_law_profiles 자산 처리 방향 + +### 1.1 검증 내용 + +**유지 판정(재실측으로 확인된 주장) 5건** + +| v.1 주장 | 재실측 결과 | 판정 | +|---|---|---| +| `special_law_profiles/` 실물 부재가 유일한 실행 차단 | `Default_Agent` 하위 해당 디렉터리·파일 0건, `prompt_compiler.py:231-232` 는 `profile_map` 에 키 없으면 무조건 raise | **유지** | +| 신설 범위는 선언된 3종 한정 | config 선언은 E-06·E-20·E-21 의 3종뿐. `SLP-AUTO-DAMAGE`·`SLP-STATE-COMPENSATION` 은 `domain_slice.schema.v2.json:101` **주석 문자열 안에만** 존재 | **유지** | +| raise 완화(2안) 기각 | v.1 은 E-06 근거만 들었으나, 재실측 결과 **3개 도메인 전부** 프롬프트 본문이 합성을 강제한다: E-06 overlay:5·:34("SLP-PRODUCT-LIABILITY 없이 종결하지 않는다"), E-20 overlay:13("**SLP-IP를 반드시 합성**하고 … SG-04로 pin한다"), E-21 overlay:13("**SLP-MEDIA-MEDIATION을 반드시 합성**…"). 완화 시 세 도메인 모두 자기 프롬프트가 금지한 산출물을 내게 된다 | **유지·강화** | +| 경로 계약 `special_law_profiles//prompt_overlay.md` + 릴리스 루트 하위 강제 | 참조본 YAML:473(`PROFILE_DIR`)·:492-494(`--profile =//prompt_overlay.md`), `compile_fragments` 의 `relative_to(release_root)` 검사 확인 | **유지** | +| overlay 는 조문·요율·기간 상수 없는 얇은 조각(버전 판단은 SG-04 위임) | `ensemble_v2.md:22`("special_law는 도메인 payload 플래그")·`:175`(SG-04 가 "적용 법률·특별법 후보, 시행일·경과규정 … version pinning" 흡수) 재확인. E-20·E-21 overlay:13 도 "SG-04로 pin"을 명시해 위임 구조가 이미 프롬프트에 내장돼 있음 | **유지** | + +**정정 판정(v.1 이 부정확했던 주장) 3건** + +1. **"3개 파일의 정규화 sha256 이 서로 달라야 한다"(v.1 계약 3) → 하드 계약 아님.** dedup(`duplicate_normalized_sha256`)의 스코프는 `compile_fragments()` **1회 호출 = 도메인 1개 컴파일 내부**다. 각 도메인은 SLP 를 1종만 선언하므로 한 컴파일에 `special_law_profile` 조각은 최대 1개이고, 서로 다른 도메인의 profile 파일끼리는 애초에 같은 dedup 집합에 들어가지 않는다. 실제 제약은 "같은 컴파일 안의 다른 조각(공통 계약·도메인 overlay 등)과 정규화 바이트가 같으면 탈락" 뿐이며 현실적으로 발생하지 않는다. 3개 파일 내용을 서로 다르게 쓰는 것은 여전히 옳지만(내용상 당연), 위반 시 조립이 깨지는 계약은 아니다 — **권장으로 강등**. + +2. **"카탈로그 `special_law_profiles/_registry_index.json` 을 함께 두어 `REFERENCE_CATALOG_NOT_PROVIDED` 경고 해소" → 파일 생성만으로는 해소되지 않는다.** `registry_validator` 의 카탈로그 주입 경로는 두 가지뿐이다: ① 도메인 인덱스의 `reference_catalogs`(또는 `reference_sets`) 맵 — 현재 없음, ② CLI `--reference-catalog KIND=PATH`. 파일을 만들어 두기만 하면 validator 는 그 존재를 모른다. 그리고 재실측 결과 **참조본 A0 는 카탈로그의 정본 위치를 이미 확정해 두었다**: `--reference-catalog profile=Default_Agent/platform/reference_catalogs/profile_ids.json`(YAML A0 호출부). 그 디렉터리에 실재하는 것은 `calculation_ids.json`·`domain_ids.json` 뿐이고 `profile_ids.json` 은 없다. 즉 카탈로그의 1차 신설 위치는 v.1 이 말한 `special_law_profiles/_registry_index.json` 이 아니라 **`platform/reference_catalogs/profile_ids.json`** 이며(형식은 `calculation_ids.json` 과 동형: `{"catalog_id": …, "entries": [{"profile_id": "SLP-…", …}]}` — validator 의 `_catalog_ids` 가 `profile_id` 키를 수집한다), `special_law_profiles/_registry_index.json` 은 별도 목적(자산 인덱스 + sha 봉인, 아래 종합 답변)으로 두는 이원 구조가 도메인 registry 와 동형이다. + +3. **"`runtime_manifest.json` 에 3~4건 등재하고 count 증분"(v.1 계약 5 포함) → 현행 장부 관례와 어긋나는 권고.** 재실측 결과 manifest 의 `domains/` 51건은 `module_role_projection.json` 25 + `structure_types.json` 25 + `_common/common_worker_contract.md` 1 이 전부로, **`domain_config.json`(26개)과 `seed_prompt_overlay.md`(26개)는 하나도 등재돼 있지 않다.** registry 자산의 무결성 봉인은 manifest 가 아니라 registry 인덱스가 담당한다(`config_sha256` 는 `load_registry` 가 로드 시 강제 대조, `prompt_overlay_sha256` 는 `registry_validator` 가 대조). 같은 성격의 자산인 SLP overlay 만 manifest 에 싣는 것은 장부 비대칭이며, 봉인의 동형 위치는 **SLP 자체 인덱스에 `prompt_overlay_sha256` 를 기재**하는 것이다. 또한 `runtime_artifact_count` 를 소비·대조하는 코드는 릴리스·YAML 어디에도 없다(grep 0건) — "길이와 일치해야 한다"는 실행 계약이 아니라 장부 자체 일관성 관례다. manifest 등재는 **선택 사항으로 강등**. + +**신규 발견(v.1 누락, 별건 플래그) 2건** + +- 참조본 A0 가 요구하는 자산 5개 중 **4개가 릴리스에 없다**: `platform/schemas/registry_index.schema.json`, `platform/schemas/domain_config.schema.json`, `platform/reference_catalogs/signal_ids.json`, `platform/reference_catalogs/profile_ids.json`. `load_json` 은 파일 부재 시 `FILE_NOT_FOUND` 로 raise 하므로 **참조본 A0 를 그대로 실행하면 SLP 문제에 도달하기 전에 죽는다.** 실행본 v.8 은 A1(프롬프트 조립)까지 도달했으므로 A0 요구가 참조본과 다르다는 뜻이다 — SLP 와 별개의 참조본↔릴리스 정합 결손으로 기록해 둔다. +- 도메인 인덱스의 `review_items` 는 `ASSEMBLY_INDEPENDENT_REVIEW_PENDING` 1건뿐, SLP 자산 미구축은 기록돼 있지 않다. `closure_gate.pass: true`·`missing_paths: []` 와 함께, **설계 장부가 이 결손을 인지하지 못한 채 닫혀 있었다**는 방증이다(전체 상태 표지는 `program_release_status: STAGE1_NOT_RELEASE_READY` 로 정직). + +### 1.2 개선 항목을 반영한 종합 답변 + +정공법(자산 신설) 결론과 3종 한정, raise 완화 기각, 얇은 조각 설계는 그대로 유지한다 — 기각 논거는 오히려 강해졌다(합성 강제 문구가 E-06 만이 아니라 E-20:13·E-21:13 까지 3개 도메인 전부의 프롬프트 본문에 있다). 다만 신설 범위를 v.1 의 "overlay 3개 + 카탈로그 1개 + manifest 등재"에서 다음 **4점 세트**로 정정한다. ① `Default_Agent/special_law_profiles/{SLP-PRODUCT-LIABILITY, SLP-IP, SLP-MEDIA-MEDIATION}/prompt_overlay.md` 3종 — 유일한 실행 차단 해소분이며, 조문·요율·기간 상수 없이 식별 단서·요건 슬롯·review code·SG-04 경계 선언만 담는 얇은 조각(UTF-8·LF·후행 개행 1, `special_law_profile` rank 40 은 정책에 기등록). ② `special_law_profiles/_registry_index.json` — profile_id·prompt_overlay_path·**prompt_overlay_sha256** 을 기재하는 자산 인덱스. sha 봉인의 동형 위치는 manifest 가 아니라 여기다(도메인 overlay 의 봉인이 도메인 인덱스에 있듯이). 스키마 주석이 전제한 파일이기도 하다. ③ `platform/reference_catalogs/profile_ids.json` — 참조본 A0 가 `--reference-catalog profile=` 로 이미 경로를 확정해 둔 검증용 ID 카탈로그(`calculation_ids.json` 과 동형, `profile_id` 키). 카탈로그는 파일만 만들면 발견되지 않고 validator 호출부 또는 인덱스 `reference_catalogs` 에 배선돼야 한다는 점이 v.1 에서 빠져 있었다. ④ `runtime_manifest.json` 등재는 **하지 않거나, 하려면 장부 정책을 먼저 정한다** — 현행 manifest 는 domain_config·seed_prompt_overlay 를 싣지 않으므로 SLP overlay 만 실으면 비대칭이 된다. "3개 파일 sha 상이" 는 계약이 아니라 권장(dedup 은 도메인별 컴파일 내부 스코프)으로 정정한다. 절단 문자열(`SLP-PRODUCT-LIABI`) 판단은 v.1 그대로 유지한다: 릴리스 내 절단형 0건, 스키마도 통과시키므로 원인은 실행 워크스페이스 `main.py` 생성부에서 찾아야 하고, 어느 쪽이든 자산 신설이 선행 조건이다. + +--- + +## 2. 질문 2 — ★ Insight 1·2번이 실행을 위해 수정해야 하는 문제인가 + +### 2.1 검증 내용 + +**Insight 1(인덱스 `[]` ↔ config 값 불일치) — v.1 의 "실행 비차단·위생 작업" 결론 유지, 근거 보강, 표현 1건 정정.** + +- 소비 경로 재확인: `load_registry` 는 엔트리를 `config`(domain_config.json 실물 로드본)와 `index_entry`(인덱스 원문)로 분리 저장하고(`registry_loader.py:111-117`), `prompt_compiler.py:228`·`domain_slice_compiler.py:482`·fanout 빌더(참조본 YAML:311) 는 전부 `config` 만 읽는다. `registry_validator` 의 `REFERENCE_FIELDS` 순회도 `config.get(field)` 만 본다. 인덱스의 `special_law_profiles: []` 를 읽는 코드는 **0곳** — 실행 비차단 확정. +- 해시 영향 재실측: 인덱스의 이 필드를 채워도 ① 각 `domain_config.json` 은 무변경이므로 `load_registry` 가 강제 대조하는 `config_sha256` 은 그대로 유효하고, ② 인덱스 파일 자체의 해시는 로드 시점에 `sha256_file` 로 **자기계산**될 뿐 기대값 장부가 어디에도 없다(`runtime_manifest.json` 에 `_registry_index.json` 미등재 실측). **v.1 의 "인덱스 자체 해시만 갱신되면 된다"는 표현은 정정한다 — 갱신할 해시 장부 자체가 존재하지 않으므로, 채움은 아무 해시도 깨뜨리지 않고 아무 갱신도 요구하지 않는다.** 유일한 유보: 실행본 A0 가 `--index-schema` 검증을 수행한다면 채운 값이 그 스키마와 정합해야 하는데, 참조본이 가리키는 `registry_index.schema.json` 실물이 없어 현재 확인 불가. +- 같은 유형의 더 큰 사례(신규): 인덱스가 E-06·E-20·E-21 에 선언한 `profile_schema_path`·`extension_schema_path`·`fixture_manifest_path`·`effect_projection_schema_path` 의 실물이 **12건 전부 MISSING** 이다(실측). 그러나 어떤 실행 경로도 이 4경로를 로드하지 않으므로 역시 비차단 — "인덱스 선언 ≠ 실물"은 SLP 필드 하나가 아니라 인덱스 전반의 상태이며, 위생 작업의 실제 범위는 Insight 1 이 지적한 것보다 넓다. + +**Insight 2(스키마 소문자 패턴 결함) — v.1 의 "이미 수정됨·주석만 잔재" 결론 유지, 실효성 보강.** + +- 패턴 재검증: `domain_slice.schema.v2.json:104` 의 현재 패턴 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|[a-z][a-z0-9_]{1,127})$` 에 3종 SLP ID 전부 매치(re 실측). 주석(:101)이 서술하는 "소문자 패턴이라 거부" 상태는 현재 파일에 존재하지 않는다. +- 실효성 보강(신규): 이 패턴은 장식이 아니라 **실행 경로에서 실제 소비된다** — 참조본 A2 가 `--slice-schema` 로 이 스키마를 `domain_slice_compiler` 에 전달하고(YAML:654·:686), `schema_subset_validator` 는 `pattern` 키워드를 지원하며 위반 시 `SCHEMA_PATTERN` 을 발행한다(:136-139). 즉 "패턴이 이미 고쳐져 있다"는 사실은 slice 검증 통과 여부를 실제로 좌우하는 확인이고, 수정 대상은 사실과 어긋난 `$comment` 문구(그리고 그 주석이 전제하는 `special_law_profiles/_registry_index.json` 의 실물화 — 질문 1 의 ② 신설로 자연 해소)뿐이다. + +### 2.2 개선 항목을 반영한 종합 답변 + +Insight 1·2번 모두 **stage 1 YAML 실행을 위해 수정해야 하는 문제가 아니라는 v.1 의 결론은 재검증에서 그대로 성립한다.** 1번의 인덱스 `special_law_profiles: []` 는 실행 경로 어디에서도 읽히지 않고(모든 소비자가 `config` 측을 읽는다), 채워 넣더라도 `config_sha256` 은 무변경·인덱스 자체는 기대값 해시 장부가 아예 없어 어떤 봉인도 깨지 않는다 — v.1 이 말한 "인덱스 해시 갱신"은 불필요한 것이 아니라 **개념적으로 존재하지 않는 절차**였다는 점만 정정한다. 아울러 위생 작업의 실제 범위는 이 필드 하나가 아니다: 같은 인덱스가 선언한 profile·extension·fixture·effect_projection 4경로의 실물이 세 도메인 12건 전부 부재한 상태이므로, 인덱스 정합화를 하려면 이 선언들까지 함께 실물화하거나 제거·보류 표기해야 장부가 다시 참말이 된다. 2번은 스키마 패턴이 이미 SLP 대문자 어휘를 받도록 수정돼 있고 그 패턴이 A2 의 `--slice-schema` 경유로 실제 slice 검증에 소비됨을 확인했으므로, 남은 작업은 실행과 무관한 문서 정확성 — 과거 결함 서술로 굳어 있는 `$comment` 를 현재형 사실("어휘는 SLP 카탈로그 소관, 패턴은 수용")로 고쳐 쓰는 것 — 뿐이다. 요컨대 에이전트가 stage 1 을 다시 실행 가능하게 만드는 유일한 수정은 질문 1 의 profile 자산 신설이고, Insight 1·2 는 그 작업에 얹어 처리하면 되는 장부·문서 위생이다. + +--- + +## 3. 총괄 — v.1 대비 변경 요약 + +| 항목 | v.1 | v.2 판정 | +|---|---|---| +| 자산 신설(정공법)·3종 한정·raise 유지 | 채택 | **유지** (합성 강제 문구 3개 도메인 전부로 논거 강화) | +| overlay 얇은 조각 설계(SG-04 위임) | 채택 | **유지** | +| 카탈로그 위치·배선 | `special_law_profiles/_registry_index.json` "두면 경고 해소" | **정정**: 1차 정본은 `platform/reference_catalogs/profile_ids.json`(A0 가 경로 기확정) + validator 배선 필요. SLP 자체 인덱스는 sha 봉인용으로 별도 | +| sha 봉인 위치 | `runtime_manifest.json` 등재 | **정정**: manifest 는 config·overlay 를 원래 싣지 않음(51건 실측). 봉인은 SLP 인덱스의 `prompt_overlay_sha256` 이 동형. manifest 등재·count 는 선택(외부 소비 0) | +| "3개 파일 sha 상이" 계약 | 하드 계약 | **정정**: dedup 은 도메인별 컴파일 내부 스코프 — 권장으로 강등 | +| Insight 1 실행 비차단 | 채택 | **유지** + "인덱스 해시 갱신" 표현 정정(갱신할 장부 없음) + 위생 범위 확장(선언 4경로 12건 MISSING) | +| Insight 2 실행 비차단 | 채택 | **유지** + 패턴의 실행 소비 확인(A2 `--slice-schema` + `SCHEMA_PATTERN`)으로 실효성 보강 | +| (신규) 참조본 A0 요구 자산 4종 부재 | — | **별건 플래그**: registry_index.schema·domain_config.schema·signal_ids·profile_ids 부재 — 참조본 그대로면 A0 가 `FILE_NOT_FOUND` 로 SLP 이전에 실패 | diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/remove_limit_on_assets.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/remove_limit_on_assets.md new file mode 100644 index 00000000..6170862a --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/remove_limit_on_assets.md @@ -0,0 +1,59 @@ +# Default_Agent 자산의 프롬프트 크기 limit 삭제 기록 + +작업일 2026-08-20 · 대상 `v.7/extension_research/Default_Agent/` · 구본 보관 `v.7/extension_research/assets_outdated/` + +## 1. 요청과 범위 + +`Default_Agent/` 자산(json·md·py·txt)에서 `utf8_bytes`·`unicode_scalars`에 걸린 limit만 삭제하고, 그 외 내용은 일절 손대지 않는다. 개정본은 제자리 overwrite, 구본은 `assets_outdated/`에 `_old`를 붙여 보존. + +전수 탐색 결과 limit 보유 파일은 **3개뿐**이었다. `.md` 29종에는 이 limit을 서술한 문서가 없었다(검색어: `utf8_bytes` `unicode_scalars` `96000` `24000` `size_guard` `상한` `최대` 등). + +| 파일 | 구본 sha(앞16) | 개정본 sha(앞16) | 행수 | +|---|---|---|---| +| `stage1_runtime/prompt_composition_policy.json` | `06dbc4de60ab3052` | `8d3f003ecc5e53eb` | 46 → 44 | +| `stage1_runtime/prompt_compiler.py` | `6922aa5274cea35a` | `d2f0256444f581d4` | 322 → 305 | +| `stage1_runtime/prompt_compiler.txt` | `6922aa5274cea35a` | `d2f0256444f581d4` | 322 → 305 | + +`.py`와 `.txt`는 개정 전후 모두 바이트 동일(동일 파일의 확장자 이본). + +## 2. 삭제 내역 + +**정책 JSON — 2행.** `deterministic_prompt_size_guard`에서 상한값 2개만 제거. + +``` +- "max_utf8_bytes": 96000, +- "max_unicode_scalars": 24000, +``` + +**컴파일러 py/txt — 17행.** 강제(enforcement) 경로 전부 + 매니페스트의 상한 반영 필드. + +- 상한 조회 3행 — `guard = policy.get("deterministic_prompt_size_guard", {})`, `maximum_bytes`, `maximum_scalars` +- 판정·차단 11행 — `exceeded = []`, 두 비교문, `raise RuntimeContractError(... PROMPT_SIZE_GUARD_EXCEEDED ...)` 블록 +- 매니페스트 3행 — `max_utf8_bytes`, `max_unicode_scalars`, `"pass": not exceeded` + +정책 JSON의 상한을 지워도 컴파일러가 `guard.get("max_utf8_bytes", 96000)` 형태로 **하드코딩 기본값을 갖고 있었기 때문에**, JSON만 고쳤다면 limit이 96,000/24,000으로 그대로 살아 있었을 것이다. 컴파일러 개정이 필수였던 이유다. + +## 3. 남긴 것과 근거 + +| 남긴 것 | 근거 | +|---|---| +| `scalar_count`, `estimate`, 매니페스트 `utf8_bytes`·`unicode_scalars`·`legacy_advisory_estimate` | **계측**이지 limit이 아니다. 관측성 유지 | +| `not_a_tokenizer`, `model_context_fit_not_proven` | 성격 고지 문구 | +| JSON `overflow_fail_code`, `fail_codes.prompt_size_guard_exceeded` | 실패코드 **레지스트리 항목**이지 수치 상한이 아니다. 현재 도달 불가(사문)지만 삭제는 "limit 외 무단절" 제약 위반 | + +`"pass": not exceeded`는 limit이 아니라 판정 결과지만, 삭제된 `exceeded`에만 의존해 `NameError` 없이 존치가 불가능했고 `True` 하드코딩은 내용 **추가**가 되므로 삭제가 유일한 준수 선택지였다. + +## 4. 검증 + +편집은 인덱스 사전 검증 후 행 삭제 방식으로만 수행(추가·수정 0). 자체 확인에 이어 독립 sub-agent 1회차 검증 **PASS, 결함 0** — 반복 불필요. + +- **삭제 전용 확인** — `git diff --numstat`이 세 파일 모두 `0` insertions(`0 17`, `0 17`, `0 2`), 변경(`c`) 헝크 0 +- **부수 피해 없음** — 구본에서 해당 행만 지운 결과가 개정본과 sha256 완전 일치(세 파일 모두) +- **유효성** — JSON 파싱 OK, `py_compile` OK, 삭제 변수(`maximum_bytes`·`maximum_scalars`·`exceeded`·`guard`) 잔존 참조 0 +- **기능 실증** — sub-agent가 scratchpad 사본에 한글 40,000자(≈120KB) 조각을 실제로 통과시켜 대조: 구본은 `PROMPT_SIZE_GUARD_EXCEEDED` 예외 발생, 개정본은 정상 완료(`utf8_bytes=120041`, `unicode_scalars=40041` 계측만 기록). limit이 실제로 더 이상 구속하지 않음을 실행으로 확인 +- **구본 무결성** — 백업 3본이 `git HEAD` 및 `runtime_manifest.json` 기록과 sha256 일치(3중 대조) + +## 5. 후속 확인이 필요한 사항 (이번 회차에서 손대지 않음) + +1. **`Default_Agent/runtime_manifest.json` 시효 만료.** 1644–1653행이 세 파일의 **개정 전** sha256을 기록하고 있어 현재 실물과 불일치한다. 무결성 재계산 검사를 돌리면 3건이 실패한다. 갱신은 limit 외 내용 변경이라 제약상 수행하지 않았다 — 매니페스트 재생성 여부는 결정이 필요하다. 현재 `Default_Agent/` 내에 이 파일을 **읽는** 코드는 없어 즉시 장애는 아니다. +2. **형제 배포 트리는 limit 보유 상태 유지.** `ver_8_yaml_candidates/Default_Agent_Stage_1/`, `P1-T2_to_T9_선결작업수행결과/00_배포트리/Default_Agent_Stage_1/`, `선행구축/Default_Agent_Stage_1/` 등에 같은 컴파일러 사본이 있고 limit이 살아 있다. 지정 경로가 `Default_Agent/`였으므로 범위 밖으로 두었다 — 배포 트리까지 일괄 적용할지 확인이 필요하다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/special_law_profile_update.md b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/special_law_profile_update.md new file mode 100644 index 00000000..504d8dce --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/special_law_profile_update.md @@ -0,0 +1,203 @@ +# special_law_profile 문제 해결 전략서 + +작성일 2026-08-21 · 근거 문서 `extension_research/assets_special_law_problem_v.2.md` §1.2(4점 세트) · 대상 릴리스 `v.7/extension_research/Default_Agent` + +## 0. 목적·범위·전제 + +`stage_1_part_2_v.8.yml` 실행을 막는 유일한 원인인 `special_law_profiles` 자산 부재를 v.2 §1.2의 확정 전략 — ① overlay 3종 신설, ② SLP 자산 인덱스 신설(sha 봉인), ③ `profile_ids.json` 카탈로그 신설·배선, ④ manifest 등재는 관례 정합 범위만 — 대로 해소하기 위한 실행 계획이다. `prompt_compiler.py:231-232` 의 raise 는 완화하지 않는다(3개 도메인 overlay 가 프롬프트 본문에서 profile 합성을 강제하는 법리 완결성 계약). 참조본 A0 요구 자산 4종 부재(registry_index.schema 등)는 **별건**으로 이 전략서 범위 밖이며, §2의 G 작업에서 실행 워크스페이스 확인 항목으로만 다룬다. + +**manifest 정책은 결정 사항이 아니라 관례에서 도출된다(실측)**: 현행 `runtime_manifest.json` 은 `platform/reference_catalogs/calculation_ids.json`·`domain_ids.json` 을 **등재**하고, `domain_config.json`·`seed_prompt_overlay.md`·`domains/_registry_index.json` 은 **미등재**다. 따라서 ③ `profile_ids.json` 은 등재(+1건, count 432→433), ① overlay 와 ② SLP 인덱스는 미등재가 동형이다. + +## 1. 작업 흐름도 — DAG (UML activity diagram 표기) + +`●` 시작 / `◉` 종료 / `━━` fork·join 동기화 바(바 아래 병렬, 바에서 합류) / `[ ]` activity / `◇` 결정(게이트) / `(선택)` 비필수 병렬 레인 + +``` + ● + │ + ▼ +[T0] 파라미터 고정 + │ (ID 3종·경로 규약·byte contract·rank 40 재확인) + │ + ━━━━━━━━━━━━━━━━━━━━━━ fork ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ + │ │ │ │ │ │ + ▼ ▼ ▼ ▼ ▼ ▼ +[A1] [A2] [A3] [C] [G] [H](선택) +overlay overlay overlay profile_ids 절단 원인 인덱스·주석 +SLP- SLP-IP SLP-MEDIA- .json 작성 조사(실행 위생(Insight +PRODUCT- MEDIATION +manifest 워크스페이스, 1·2 잔여) +LIABILITY 등재(+1) 읽기 전용) + │ │ │ │ │ │ + ━━━━ join(A1·A2·A3) ━━━━ │ │ │ + ▼ │ │ │ +[B] special_law_profiles/ │ │ │ + _registry_index.json │ │ │ + (3종 sha256 봉인) │ │ │ + │ │ │ │ + ━━━━━━━━━ join(B·C) ━━━━━━━━━━━━━━━━━ │ │ + ▼ │ │ + ━━━━━━━━━━━━━━━ fork ━━━━━━━━━━━━━━━ │ │ + │ │ │ │ │ + ▼ ▼ ▼ │ │ +[E1a] [E1b] [E1c] │ │ +E-06 컴파일 E-20 컴파일 E-21 컴파일 │ │ +스모크 스모크 스모크 │ │ + │ │ │ │ │ + ━━━━━━━━━ join(E1a·E1b·E1c) ━━━ │ │ + ▼ │ │ +[E2] registry_validator │ │ + 카탈로그 배선 스모크 │ │ + ▼ │ │ +[E3] sha 봉인 대조 │ │ + │ │ │ + ━━━━━━━━━━━━━ join(E3·G·H) ━━━━━━━━━━━━━━━━━━━━━━━━━ │ + ▼ (H는 미착수여도 통과) ━━━━━━ +[F] 기록·문서 반영 + │ + ▼ + ◇ 게이트: E1~E3 전건 PASS ∧ G 결론 확보 + │ yes + ▼ +[R] stage_1_part_2_v.8.yml 재실행 (실행 워크스페이스) + │ + ▼ + ◉ +``` + +**병렬/직렬 구분과 임계 경로** + +| 구분 | 작업 | 근거 | +|---|---|---| +| 병렬 1군 (T0 직후 동시 착수) | A1·A2·A3, C, G, H | 상호 입력 의존 없음. C는 ID 목록만 필요(overlay 내용 불요), G는 외부 워크스페이스 읽기 전용 | +| 직렬 | A(3건 전부)→B | B의 sha256 은 overlay **최종 바이트**에서 재계산해야 하므로 A 완료가 선행 | +| 직렬 | (B∧C)→E1→E2→E3 | E1 은 자산 실물 필요, E2 는 카탈로그 필요, E3 은 B 의 봉인값 필요 | +| 합류 | (E3∧G)→F→게이트→R | 절단 원인 미확인 상태로 재실행하면 같은 A1 단계에서 원인 불명 재실패 위험 | +| 임계 경로 | T0→A(최장 overlay)→B→E1→E2→E3→F→R | 법률 콘텐츠 작성(A)이 최장 작업. C·G·H 는 임계 경로 밖 | + +## 2. 작업별 세부 내역 + +### T0 — 파라미터 고정 (직렬 관문, 즉시 완료 가능) + +- **목적**: 신설 자산이 지켜야 할 계약값을 문서·코드 실측값으로 고정해 A~E 전 작업의 공통 전제로 삼는다. +- **세부 내역**: ① 대상 ID 3종 확정 — `SLP-PRODUCT-LIABILITY`(E-06), `SLP-IP`(E-20), `SLP-MEDIA-MEDIATION`(E-21). ② 경로 규약 — `Default_Agent/special_law_profiles//prompt_overlay.md`(참조본 YAML:473·:492 의 `--profile =//prompt_overlay.md` 조립식과 일치, 릴리스 루트 하위 필수). ③ byte contract — UTF-8·LF·후행 개행 정확히 1(`prompt_composition_policy.json` `byte_contract`). ④ 조각 카테고리 `special_law_profile` rank 40 기등록 확인 — **정책 파일 수정 없음**. ⑤ 스키마 패턴 `^(?:SLP-[A-Z][A-Z0-9-]{1,63}|…)$` 에 3종 매치 기확인 — **스키마 수정 없음**. +- **완료 기준**: 위 5항이 본 전략서에 기록됨(본 절로 충족). + +### A1·A2·A3 — profile overlay 3종 작성 (상호 병렬, 임계 경로) + +- **목적**: 유일한 실행 차단 해소분. 도메인 overlay 와 동일 톤의 "얇은 조각" 프롬프트 3건. +- **입력**: 각 도메인의 `seed_prompt_overlay.md`·`domain_config.json`(요건 슬롯·review code 어휘), `ensemble_v2.md` §4.2(SG-04 위임 규약). +- **공통 작성 지침(4절 구성)**: ① **식별 단서** — 이 특별법 체계가 적용 후보가 되는 원천(source-backed) 단서. ② **고유 요건 슬롯·증거 component** — 일반법 골격(각 도메인 overlay 소관)과 겹치지 않는 특별법 고유 쟁점의 슬롯화. ③ **발행할 review code** — 요건 미충족·경계 사안에서 침묵 대신 남길 표지. ④ **SG-04 경계 선언** — 적용 법률·버전·시행일·경과규정 판단은 operand 와 후보만 담아 SG-04 로 위임. **금지**: 법령 조문 번호·요율·기간·한도 상수 기재(도메인 overlay 의 "법정률·기간을 상수로 두지 않는다" 규율과 동일), 최종 결론(`final_conclusion_forbidden` 톤 유지 — 다만 마커 자체는 common contract 가 충족하므로 overlay 에 의무 문구는 아님). +- **개별 내역**: + - **A1 `SLP-PRODUCT-LIABILITY`**: 결함 3유형(제조·설계·표시) 후보 슬롯, 정상 사용·유통 당시 상태·대체원인 증거 component, 결함·인과 추정 구조는 operand 상태로만, 면책 항변 슬롯, E-06 config 의 `E06-N02`(단순 품질불만은 monitor) 경계와 정합. E-06 overlay:5·:34 의 "profile 없이 종결 금지" 규칙의 수신 측임을 명시. + - **A2 `SLP-IP`**: 권리 종류별(특허·저작권·상표·디자인·영업비밀) 특별법 후보 식별, 권리 유효·귀속·보호범위 대 피고 실시행위 대비 슬롯(E-20 config 의 scope/infringement review code 어휘 재사용), 손해액 산정 특칙은 계산 도메인(CE-01·CE-R1)·SG-04 위임. E-20 overlay:13 의 "반드시 합성 + SG-04 pin" 과 정합. + - **A3 `SLP-MEDIA-MEDIATION`**: 언론 보도·매체 유형별 구제수단(정정·반론·추후보도·손해배상) 후보 슬롯, 조정·중재 전치 여부는 절차 경계 review 로만 표지(X3 횡단 소관 침범 금지), 인격권·표현의 자유 형량은 결론 금지 준수. E-21 overlay:13 과 정합. +- **완료 기준(각 건)**: 4절 구성 충족, 금지 사항 0건, UTF-8·LF·후행 개행 1, 상수 검색(`grep -E '제[0-9]+조|[0-9]+%|[0-9]+년'`) 0건. +- **리스크**: 도메인 overlay 와 내용 중복이 크면 정규화 sha 동일로 같은 컴파일 내 dedup 탈락 가능(이론상) — 특별법 고유 층만 담으면 발생하지 않음. 크기 가드는 advisory 전환 상태라 차단 없음. + +### B — `special_law_profiles/_registry_index.json` 작성 (A 완료 후 직렬) + +- **목적**: SLP 자산 인덱스 + sha 봉인. 도메인 registry 의 `_registry_index.json` 과 동형 구조(봉인의 정위치는 manifest 가 아니라 자산군 자체 인덱스 — v.2 정정 ③). `domain_slice.schema.v2.json:101` 주석이 전제한 파일의 실물화이기도 하다. +- **세부 내역**: 엔트리 3건 — `profile_id`·`prompt_overlay_path`(인덱스 기준 상대경로)·`prompt_overlay_sha256`(**실물 파일 바이트의 sha256_file 값, 옮겨 적기 금지·재계산 원칙** — MEMORY 2026-08-20 교훈 1). `schema_version`·`catalog_id` 헤더 포함. 현재 이 인덱스를 읽는 코드는 없음(사람·후속 validator 용 장부) — validator 확장은 별건. +- **완료 기준**: JSON 파싱 OK, 3건 sha 를 독립 재계산으로 대조 일치(E3 에서 재확인). + +### C — `platform/reference_catalogs/profile_ids.json` 작성 (A 와 병렬) + +- **목적**: `registry_validator` 참조 무결성 검증용 ID 카탈로그. 참조본 A0 가 `--reference-catalog profile=<이 경로>` 로 이미 확정해 둔 정본 위치(v.2 정정 ②). +- **세부 내역**: `calculation_ids.json` 동형 — `{"catalog_id": "special_law_profiles", "entries": [{"profile_id": "SLP-…", "status": …}, ×3]}`. validator 의 `_catalog_ids` 가 `profile_id` 키를 수집하므로 이 키명이 계약. **배선 주의**: 파일 생성만으로는 발견되지 않는다 — 참조본 A0 는 기배선, 실행본 v.8 의 배선 여부는 G 의 확인 항목. (도메인 인덱스 `reference_catalogs` 맵 신설은 인덱스 수정을 수반하므로 채택하지 않음.) manifest 등재 +1건(카탈로그는 등재군 관례), `runtime_artifact_count` 432→433 — 값 치환은 앵커 기반 행 편집(MEMORY 2026-08-20 교훈 2). +- **완료 기준**: E2 에서 `special_law_profiles` 필드의 `REFERENCE_CATALOG_NOT_PROVIDED` 경고 소멸·`REFERENCE_NOT_FOUND` 0건. + +### G — 절단 문자열 원인 조사 (독립 병렬, 외부 워크스페이스, 읽기 전용) + +- **목적**: 에러의 `SLP-PRODUCT-LIABI`(17자) 절단 원인 규명. 릴리스 내 절단형 0건 기확인 — 원인은 실행 워크스페이스 측. +- **세부 내역**: `/Users/jsahn/Works/작업결과/v2` 의 Stage YAML 에서 ① `SLP-PRODUCT-LIABI` 문자열 검색, ② `main.py` 생성부의 `{{item_json}}` 직렬화·슬라이싱·길이 제한 검사, ③ v.8 A1 의 `--profile` 조립식이 T0 경로 규약과 일치하는지, ④ v.8 A0 에 profile 카탈로그 배선이 있는지(참조본과 달리 없을 수 있음 — 없으면 E2 는 릴리스 측 검증으로만 유효). **YAML 은 읽기만 하고 수정하지 않는다.** +- **완료 기준**: 절단 발생 지점 특정 또는 "워크스페이스 파일 교체로 소멸" 확인. 게이트(R) 선행 조건. + +### H — 장부·문서 위생 (선택 병렬, Insight 1·2 잔여) + +- **세부 내역**: ① 도메인 인덱스 E-06·E-20·E-21 엔트리의 `special_law_profiles: []` 를 config 값으로 채움 — 어떤 해시도 깨지 않음(config 무변경·인덱스는 기대값 장부 부재, v.2 §2.1 실측). ② `domain_slice.schema.v2.json:101` `$comment` 를 현재형 사실로 재작성(B 완료로 "실물 부재" 전제도 해소됨). ③ 인덱스 선언 4경로(profile_schema·extension_schema·fixture_manifest·effect_projection_schema) 12건 MISSING 은 실물화가 아니라 **기록만**(별건 결정 대상). +- **완료 기준**: 착수 시에만 적용. 미착수여도 게이트 통과에 영향 없음. + +### E1a·E1b·E1c — 컴파일 스모크 (B∧C 후, 상호 병렬) + +- **세부 내역(각 도메인)**: `prompt_compiler.main([--registry-index, --domain-id E-06|E-20|E-21, --policy, --common-contract, --profile =Default_Agent/special_law_profiles//prompt_overlay.md, --output-prompt/, --output-manifest/])` — 실행본 A1 의 호출 형태 재현. 산출 manifest 에서 ① rc 0, ② `special_law_profile` 카테고리 조각 존재·`source_domain_id` 일치, ③ `omitted` 에 profile 조각 없음, ④ `final_conclusion_forbidden` 마커 충족 확인. 회귀 대조로 SLP 비선언 도메인 1건(예: E-05) 컴파일 rc 0 확인. +- **완료 기준**: 3건 전부 4항 PASS. 산출물은 scratch 에만 기록. + +### E2 — validator 배선 스모크 (E1 후 직렬) + +- **세부 내역**: `registry_validator.main([--index, --reference-catalog profile=platform/reference_catalogs/profile_ids.json, --output/])`. 판정: `special_law_profiles` 필드의 closure 가 `catalog: "special_law_profiles"`… 아님 — `REFERENCE_FIELDS` 매핑상 kind 는 `special_law_profiles` 이므로 CLI 는 `--reference-catalog special_law_profiles=` 로 전달해야 한다(**참조본 A0 의 `profile=` kind 표기는 매핑 키와 불일치 — 실측 후 kind 문자열을 `special_law_profiles` 로 맞춰 검증하고, 이 불일치는 F 에서 별건 기록**). 기대: 해당 필드 경고 소멸, `REFERENCE_NOT_FOUND` 0, status 는 다른 결손(calculation·signal 카탈로그 미제공 경고 등)과 무관하게 errors 0 유지. +- **완료 기준**: E-06·E-20·E-21 의 `special_law_profiles` closure 가 `missing: []` 로 기록. + +### E3 — sha 봉인 대조 (E2 후 직렬) + +- **세부 내역**: B 인덱스의 3개 `prompt_overlay_sha256` 을 실물에서 전수 재계산 대조 + C 등재분 포함 manifest 전수 재계산(기대: 일치 156·불일치 0·부재 277 — 기존 155 에 profile_ids.json +1). MEMORY 2026-08-20 교훈 3(전수 재계산으로 상태 증명) 적용. +- **완료 기준**: 불일치 0. + +### F — 기록·문서 반영 (E3∧G∧H 합류 후) + +- **세부 내역**: ① `MEMORY.md` 압축 요약 추가. ② v.2 리포트에 후속 상태 부기(신설 자산 목록·검증 결과). ③ 별건 플래그 이관 기록 — 참조본 A0 자산 4종 부재, E2 에서 확인된 kind 표기(`profile=` vs `special_law_profiles=`) 불일치, H-③ 4경로 MISSING. 커밋은 사용자 지시 시에만. + +### R — 재실행 게이트 (◇ 통과 후, 실행 워크스페이스) + +- **선행 조건**: E1~E3 전건 PASS ∧ G 원인 결론. 워크스페이스 반영은 "신설 자산 복사" 방식(기존 4파일 교체 방식과 동일 절차)으로 하고, `stage_1_part_2_v.8.yml` 재실행으로 A1 통과를 확인한다. + +## 3. 완료 판정 총괄 (DoD) + +| # | 판정 항목 | 판정 방법 | +|---|---|---| +| 1 | overlay 3종 실재·계약 준수 | A 완료 기준 4항 | +| 2 | 조립 통과 | E1 3건 rc 0 + manifest 4항 | +| 3 | 참조 무결성 closure | E2 `missing: []` | +| 4 | 봉인 정합 | E3 불일치 0 (manifest 156/0/277) | +| 5 | 절단 원인 결론 | G 보고 | +| 6 | 장부 기록 | F 완료 (별건 3건 이관 포함) | +| 7 | 실행 복구 | R 에서 part 2 A1 통과 | + +--- + +# 4. 실행 결과 (2026-08-21 실행) + +전략서 §1 DAG 중 **T0 → A1·A2·A3 → B → C → E1 → E2 → E3 → F 를 실행 완료**했다. G 는 이 머신에서 수행 불가, H 는 보류, R 은 실행 워크스페이스 소관으로 미착수다. + +## 4.1 신설·변경 자산 + +| 자산 | 상태 | bytes | sha256 | +|---|---|---|---| +| `special_law_profiles/SLP-PRODUCT-LIABILITY/prompt_overlay.md` | 신설 | 5652 | `78281b81af2e61647024cbd72f9245d14bd3a6750413d16eb9aa093e6778d491` | +| `special_law_profiles/SLP-IP/prompt_overlay.md` | 신설 | 5720 | `7ed12d8183992733e79feb25622e6980642308d0705b8409f6a465394d3b7f13` | +| `special_law_profiles/SLP-MEDIA-MEDIATION/prompt_overlay.md` | 신설 | 5605 | `04ea71d164111934039640486ba113c22793c95168318c38974897d90ffe5a0e` | +| `special_law_profiles/_registry_index.json` | 신설(봉인 장부) | 1478 | `8491e715c47eaa7f0fe749a6356c8c654535f382aadcc08e59c8f3eb93420177` | +| `platform/reference_catalogs/profile_ids.json` | 신설(검증 카탈로그) | 363 | `84f3538229df73e2d7ca58eb32da3d9b38ab1037edb67d4ac43e31e51c20aea3` | +| `runtime_manifest.json` | 변경 | — | `profile_ids.json` 1건 등재, `runtime_artifact_count` 432→433 | + +**A 작성 시 반영한 실측 제약 1건(전략서 §2-A 에 없던 것)**: `routing/extension_payload_key_declarations.v1.json` 확인 결과 `special_law_profile_candidates` key 를 선언하는 도메인은 **E-06 뿐**이고, E-20·E-21 은 선언하지 않으며 세 도메인 모두 `additional_properties_allowed: false` 다. 따라서 세 overlay 에 같은 출력 key 를 지시하면 E-20·E-21 worker 출력이 스키마 위반이 된다. 각 overlay 의 §4 는 도메인별 선언 key 만 지목하도록 작성했다(E-06 → `special_law_profile_candidates`, E-20 → `ip_right_candidates` 등 7종, E-21 → `publication_candidates` 등 7종). review 발행 경로도 공통 계약이 정한 후보별 `review_items`·root `unknown_or_unrouted_reviews` 로 한정했다(`review_flags` 등 별도 key 지시 없음). + +`_registry_index.json`·overlay 3종은 manifest 에 등재하지 않았다 — 현행 manifest 가 `domain_config.json`·`seed_prompt_overlay.md`·`domains/_registry_index.json` 을 싣지 않는 관례와 동형이며, 봉인은 SLP 인덱스의 `prompt_overlay_sha256` 이 담당한다. 카탈로그(`profile_ids.json`)만 등재한 것은 `calculation_ids.json`·`domain_ids.json` 2/2 가 등재돼 있는 관례를 따른 것이다. + +## 4.2 검증 결과 (DoD 대조) + +| # | 판정 항목 | 결과 | +|---|---|---| +| 1 | overlay 계약 준수 | **PASS** — 법령 조문·요율·기간 상수 grep 0건, UTF-8·LF·후행 개행 1 (3건 전부) | +| 2 | 조립 통과(E1) | **PASS** — E-06·E-20·E-21 `rc=0`, `special_law_profile` 조각이 rank 40 위치에 1건씩 부착(`source_domain_id` 일치), `deduplicated_fragments` 빈 배열, `final_conclusion_forbidden` 마커 검증. 컴파일 크기 24,960 / 21,841 / 21,807 bytes. 회귀 대조로 SLP 비선언 도메인 E-05 도 `rc=0`(profile 조각 0건) | +| 3 | 참조 무결성(E2) | **PASS** — `--reference-catalog special_law_profiles=…` 주입 시 3개 도메인 closure `missing: []`, `REFERENCE_CATALOG_NOT_PROVIDED` 경고 54→51 감소, errors 0 | +| 4 | 봉인 정합(E3) | **PASS** — SLP 인덱스 3건 실물 재계산 대조 불일치 0, manifest 전수 재계산 **일치 156 · 불일치 0 · 파일없음 277**(직전 회차 155/0/277 에 카탈로그 1건 증가분과 일치) | +| 5 | 절단 원인(G) | **수행 불가** — `/Users/jsahn` 경로가 이 머신에 존재하지 않음(`No such file or directory`). 실행 워크스페이스 보유자만 확인 가능 | +| 6 | 장부 기록(F) | **완료** — 본 절, v.2 리포트 부기, `MEMORY.md` | +| 7 | 실행 복구(R) | **미착수** — 실행 워크스페이스 소관 | + +## 4.3 확정된 별건 3건 + +1. **validator 카탈로그 kind 표기 불일치(실측 확정)**: `registry_validator.REFERENCE_FIELDS` 는 `special_law_profiles` 필드를 kind `special_law_profiles` 로 매핑한다. 참조본 A0 의 `--reference-catalog profile=…` 표기로 실제 주입해 보면 경고가 54 로 유지되고 closure 가 `catalog: null / unverified` 로 남는다 — **즉 그 표기로는 카탈로그가 배선되지 않는다.** `special_law_profiles=` 로 고쳐야 검증이 성립한다. (같은 이유로 `calculation`·`signal` 표기도 각각 `computations`·`signals` 여야 하는지 별도 확인 필요 — `REFERENCE_FIELDS` 값은 `computations`·`signals` 다.) +2. **참조본 A0 요구 자산 잔여 부재 3종**: `platform/schemas/registry_index.schema.json`, `platform/schemas/domain_config.schema.json`, `platform/reference_catalogs/signal_ids.json`. (4종 중 `profile_ids.json` 은 이번 회차로 해소.) `load_json` 이 부재 시 `FILE_NOT_FOUND` 로 raise 하므로 참조본 A0 를 그대로 실행하면 여전히 A0 에서 정지한다. +3. **인덱스 선언 4경로 12건 MISSING**: E-06·E-20·E-21 의 `profile_schema_path`·`extension_schema_path`·`fixture_manifest_path`·`effect_projection_schema_path` 실물 부재. 실행 경로에서 로드되지 않아 비차단. + +## 4.4 H(장부·문서 위생) 보류 사유 + +- **H-① 인덱스 `special_law_profiles` 채움 보류**: 실행 워크스페이스의 `registry_index.schema.json` 을 아직 보지 못했다. 그 스키마가 slice 스키마와 같은 소문자 패턴 결함을 갖고 있으면, 현재 `[]` 로 통과하던 A0 인덱스 검증이 값을 채우는 순간 `REGISTRY_INDEX_SCHEMA_FAILED` 로 **새로 깨진다**. 실행 이득이 0인 반면 회귀 위험이 있으므로 그 스키마 확인 후로 미룬다. +- **H-② 스키마 `$comment` 수정 보류**: `domain_slice.schema.v2.json` 은 manifest 등재 자산이라 수정 시 sha 갱신이 따라붙고, 디버깅 진행 중 실행 워크스페이스 사본과 바이트 드리프트를 만든다. 실행 무관한 문서 수정이므로 H-① 과 함께 일괄 처리한다. + +## 4.5 다음 행동 (R, 사용자 소관) + +1. 신설 5개 파일(`special_law_profiles/` 4 + `platform/reference_catalogs/profile_ids.json`)과 갱신된 `runtime_manifest.json` 을 실행 워크스페이스 릴리스 트리에 반영한다. +2. `stage_1_part_2_v.8.yml` 을 재실행해 A1(프롬프트 조립) 통과를 확인한다. +3. 실패가 재현되면 에러의 profile ID 표기를 확인한다 — 절단형(`SLP-PRODUCT-LIABI`)이 그대로 나오면 §4.3-1 과 무관한 워크스페이스 `main.py` 생성부 문제이므로 G 를 그때 수행한다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/prompt_compiler.pre_special_law_wiring.py b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/prompt_compiler.pre_special_law_wiring.py new file mode 100644 index 00000000..e48c9b0c --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/prompt_compiler.pre_special_law_wiring.py @@ -0,0 +1,305 @@ +#!/usr/bin/env python3 +"""Deterministically compose common, domain, dependency, and special-law prompts.""" + +from __future__ import annotations + +import argparse +import json +import math +import sys +from dataclasses import dataclass +from pathlib import Path +from typing import Any + +from registry_loader import load_registry +from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json + + +@dataclass(frozen=True) +class FragmentSpec: + fragment_id: str + category: str + path: str + source_domain_id: str | None = None + + +def _priority(policy: dict[str, Any]) -> dict[str, int]: + out: dict[str, int] = {} + for index, item in enumerate(policy.get("fragment_priority", [])): + if isinstance(item, str): + out[item] = index * 10 + elif isinstance(item, dict) and item.get("category"): + out[str(item["category"])] = int(item.get("rank", index * 10)) + if not out: + raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty") + return out + + +def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str: + codes = policy.get("fail_codes") + return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback + + +def _normalize_fragment(path: str) -> tuple[str, bytes, str]: + source = Path(path) + try: + text = source.read_text(encoding="utf-8") + except FileNotFoundError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc + except UnicodeDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc + if source.suffix.lower() == ".json": + try: + value = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc + normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n" + else: + normalized = normalize_text(text) + data = normalized.encode("utf-8") + return normalized, data, sha256_bytes(data) + + +def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]: + priorities = _priority(policy) + required_categories = [str(item) for item in policy.get("required_categories", [])] + categories = {spec.category for spec in specs} + missing_categories = [item for item in required_categories if item not in categories] + if missing_categories: + raise RuntimeContractError( + _fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"), + "required prompt fragment category missing", + missing_categories, + ) + unknown = sorted(categories - set(priorities)) + if unknown: + raise RuntimeContractError( + _fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"), + "fragment category has no priority", + unknown, + ) + prepared: list[dict[str, Any]] = [] + by_id: dict[str, str] = {} + for spec in specs: + text, data, digest = _normalize_fragment(spec.path) + previous = by_id.get(spec.fragment_id) + if previous is not None and previous != digest: + raise RuntimeContractError( + _fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"), + f"same fragment_id has different normalized hashes: {spec.fragment_id}", + {"first": previous, "second": digest}, + ) + by_id[spec.fragment_id] = digest + prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest}) + prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path)) + seen_sha: set[str] = set() + kept: list[dict[str, Any]] = [] + omitted: list[dict[str, Any]] = [] + for item in prepared: + spec = item["spec"] + source_path = Path(spec.path).resolve() + release_root = Path(__file__).resolve().parents[1] + try: + logical_path = source_path.relative_to(release_root).as_posix() + except ValueError as exc: + raise RuntimeContractError( + "PROMPT_FRAGMENT_PATH_INVALID", + "prompt fragment must resolve under the release root", + source_path.name, + ) from exc + public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id} + if item["sha256"] in seen_sha: + omitted.append({**public, "reason": "duplicate_normalized_sha256"}) + continue + seen_sha.add(item["sha256"]) + kept.append(item) + separator = str(policy.get("byte_contract", {}).get("separator", "\n\n")) + compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n" + compiled_bytes = compiled.encode("utf-8") + required_markers = [str(item) for item in policy.get("required_markers", [])] + missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()] + if missing_markers: + raise RuntimeContractError( + _fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"), + "compiled prompt lacks required guard marker", + missing_markers, + ) + for conflict in policy.get("conflict_markers", []): + if isinstance(conflict, list) and len(conflict) > 1: + present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()] + if len(present) > 1: + raise RuntimeContractError( + _fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"), + "mutually exclusive prompt markers coexist", + present, + ) + scalar_count = len(compiled) + estimate = math.ceil(len(compiled_bytes) / 4) + manifest = { + "schema_version": "compiled_prompt_manifest.v1", + "compiled_prompt_sha256": sha256_bytes(compiled_bytes), + "compiled_prompt_bytes": len(compiled_bytes), + "deterministic_prompt_size_guard": { + "utf8_bytes": len(compiled_bytes), + "unicode_scalars": scalar_count, + "not_a_tokenizer": True, + "model_context_fit_not_proven": True, + "legacy_advisory_estimate": estimate, + }, + "normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator}, + "fragments": [ + { + "fragment_id": item["spec"].fragment_id, + "category": item["spec"].category, + "path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()), + "sha256": item["sha256"], + "hash_kind": "file_sha256", + "source_domain_id": item["spec"].source_domain_id, + } + for item in kept + ], + "deduplicated_fragments": omitted, + "required_markers_verified": required_markers, + } + return compiled, manifest + + +def _path_from_config(config_record: dict[str, Any], value: str) -> str: + config_path = config_record.get("config_path") + base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", ".")) + path = Path(value) + return str(path if path.is_absolute() else (base / path).resolve()) + + +def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]: + config = config_record["config"] + values: list[Any] = [] + if isinstance(config.get("prompt_fragments"), list): + values.extend(config["prompt_fragments"]) + for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"): + if isinstance(config.get(key), str): + values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"}) + index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {} + if not values and isinstance(index_entry.get("prompt_overlay_path"), str): + values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"}) + specs: list[FragmentSpec] = [] + for index, value in enumerate(values, 1): + if isinstance(value, str): + path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}" + elif isinstance(value, dict) and isinstance(value.get("path"), str): + path_value = value["path"] + fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}") + else: + continue + specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id)) + return specs + + +def collect_domain_fragments( + domain_id: str, + registry: dict[str, Any], + *, + common_contract_path: str, + profile_paths: dict[str, str] | None = None, + runtime_guard_path: str | None = None, + extra_specs: list[FragmentSpec] | None = None, +) -> list[FragmentSpec]: + entries = registry["entries"] + if domain_id not in entries: + raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}") + specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))] + visited: set[str] = set() + + def add_dependencies(current: str) -> None: + for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key): + if dependency not in entries: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}") + if dependency in visited: + continue + visited.add(dependency) + add_dependencies(dependency) + specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency)) + + add_dependencies(domain_id) + domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id) + if not domain_specs: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}") + specs.extend(domain_specs) + profiles = entries[domain_id]["config"].get("special_law_profiles", []) + profile_map = profile_paths or {} + for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key): + if profile not in profile_map: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}") + specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id)) + if runtime_guard_path: + specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve()))) + specs.extend(extra_specs or []) + return specs + + +def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]: + out: dict[str, str] = {} + for value in values: + if separator not in value: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value) + key, path = value.split(separator, 1) + out[key] = path + return out + + +def _parse_extra(values: list[str]) -> list[FragmentSpec]: + out: list[FragmentSpec] = [] + for value in values: + parts = value.split(":", 2) + if len(parts) != 3: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value) + out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve()))) + return out + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--registry-index", required=True) + parser.add_argument("--domain-id", required=True) + parser.add_argument("--policy", required=True) + parser.add_argument("--common-contract", required=True) + parser.add_argument("--runtime-guard") + parser.add_argument("--profile", action="append", default=[]) + parser.add_argument("--fragment", action="append", default=[]) + parser.add_argument("--output-prompt", required=True) + parser.add_argument("--output-manifest", required=True) + args = parser.parse_args(argv) + try: + registry = load_registry(args.registry_index) + specs = collect_domain_fragments( + args.domain_id, + registry, + common_contract_path=args.common_contract, + profile_paths=_parse_mapping(args.profile), + runtime_guard_path=args.runtime_guard, + extra_specs=_parse_extra(args.fragment), + ) + policy = load_json(args.policy) + compiled, manifest = compile_fragments(specs, policy) + prompt_path = Path(args.output_prompt) + prompt_path.parent.mkdir(parents=True, exist_ok=True) + prompt_path.write_text(compiled, encoding="utf-8", newline="") + from runtime_common import sha256_file + + manifest.update({ + "domain_id": args.domain_id, + "compiled_prompt_path": str(prompt_path), + "registry_index_sha256": registry["index_sha256"], + "composition_policy_path": args.policy, + "composition_policy_sha256": sha256_file(args.policy), + }) + write_json(args.output_manifest, manifest) + emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]}) + return 0 + except RuntimeContractError as exc: + emit_cli_result({"status": "FAILED", "error": exc.as_dict()}) + return 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/prompt_compiler.pre_special_law_wiring.txt b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/prompt_compiler.pre_special_law_wiring.txt new file mode 100644 index 00000000..e48c9b0c --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/prompt_compiler.pre_special_law_wiring.txt @@ -0,0 +1,305 @@ +#!/usr/bin/env python3 +"""Deterministically compose common, domain, dependency, and special-law prompts.""" + +from __future__ import annotations + +import argparse +import json +import math +import sys +from dataclasses import dataclass +from pathlib import Path +from typing import Any + +from registry_loader import load_registry +from runtime_common import RuntimeContractError, emit_cli_result, load_json, natural_key, normalize_text, sha256_bytes, write_json + + +@dataclass(frozen=True) +class FragmentSpec: + fragment_id: str + category: str + path: str + source_domain_id: str | None = None + + +def _priority(policy: dict[str, Any]) -> dict[str, int]: + out: dict[str, int] = {} + for index, item in enumerate(policy.get("fragment_priority", [])): + if isinstance(item, str): + out[item] = index * 10 + elif isinstance(item, dict) and item.get("category"): + out[str(item["category"])] = int(item.get("rank", index * 10)) + if not out: + raise RuntimeContractError("PROMPT_POLICY_INVALID", "fragment_priority cannot be empty") + return out + + +def _fail_code(policy: dict[str, Any], key: str, fallback: str) -> str: + codes = policy.get("fail_codes") + return str(codes.get(key, fallback)) if isinstance(codes, dict) else fallback + + +def _normalize_fragment(path: str) -> tuple[str, bytes, str]: + source = Path(path) + try: + text = source.read_text(encoding="utf-8") + except FileNotFoundError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment not found: {source}") from exc + except UnicodeDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"fragment is not UTF-8: {source}") from exc + if source.suffix.lower() == ".json": + try: + value = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeContractError("PROMPT_FRAGMENT_PATH_INVALID", f"JSON prompt fragment is invalid: {source.name}") from exc + normalized = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + "\n" + else: + normalized = normalize_text(text) + data = normalized.encode("utf-8") + return normalized, data, sha256_bytes(data) + + +def compile_fragments(specs: list[FragmentSpec], policy: dict[str, Any]) -> tuple[str, dict[str, Any]]: + priorities = _priority(policy) + required_categories = [str(item) for item in policy.get("required_categories", [])] + categories = {spec.category for spec in specs} + missing_categories = [item for item in required_categories if item not in categories] + if missing_categories: + raise RuntimeContractError( + _fail_code(policy, "required_fragment_missing", "PROMPT_REQUIRED_FRAGMENT_MISSING"), + "required prompt fragment category missing", + missing_categories, + ) + unknown = sorted(categories - set(priorities)) + if unknown: + raise RuntimeContractError( + _fail_code(policy, "unknown_priority_category", "PROMPT_PRIORITY_UNKNOWN"), + "fragment category has no priority", + unknown, + ) + prepared: list[dict[str, Any]] = [] + by_id: dict[str, str] = {} + for spec in specs: + text, data, digest = _normalize_fragment(spec.path) + previous = by_id.get(spec.fragment_id) + if previous is not None and previous != digest: + raise RuntimeContractError( + _fail_code(policy, "fragment_id_hash_conflict", "PROMPT_FRAGMENT_ID_HASH_CONFLICT"), + f"same fragment_id has different normalized hashes: {spec.fragment_id}", + {"first": previous, "second": digest}, + ) + by_id[spec.fragment_id] = digest + prepared.append({"spec": spec, "text": text, "bytes": data, "sha256": digest}) + prepared.sort(key=lambda item: (priorities[item["spec"].category], natural_key(item["spec"].fragment_id), item["spec"].path)) + seen_sha: set[str] = set() + kept: list[dict[str, Any]] = [] + omitted: list[dict[str, Any]] = [] + for item in prepared: + spec = item["spec"] + source_path = Path(spec.path).resolve() + release_root = Path(__file__).resolve().parents[1] + try: + logical_path = source_path.relative_to(release_root).as_posix() + except ValueError as exc: + raise RuntimeContractError( + "PROMPT_FRAGMENT_PATH_INVALID", + "prompt fragment must resolve under the release root", + source_path.name, + ) from exc + public = {"fragment_id": spec.fragment_id, "category": spec.category, "path": logical_path, "sha256": item["sha256"], "hash_kind": "file_sha256", "source_domain_id": spec.source_domain_id} + if item["sha256"] in seen_sha: + omitted.append({**public, "reason": "duplicate_normalized_sha256"}) + continue + seen_sha.add(item["sha256"]) + kept.append(item) + separator = str(policy.get("byte_contract", {}).get("separator", "\n\n")) + compiled = separator.join(item["text"].rstrip("\n") for item in kept).rstrip("\n") + "\n" + compiled_bytes = compiled.encode("utf-8") + required_markers = [str(item) for item in policy.get("required_markers", [])] + missing_markers = [marker for marker in required_markers if marker.casefold() not in compiled.casefold()] + if missing_markers: + raise RuntimeContractError( + _fail_code(policy, "required_guard_missing", "PROMPT_FINAL_CONCLUSION_GUARD_MISSING"), + "compiled prompt lacks required guard marker", + missing_markers, + ) + for conflict in policy.get("conflict_markers", []): + if isinstance(conflict, list) and len(conflict) > 1: + present = [marker for marker in conflict if str(marker).casefold() in compiled.casefold()] + if len(present) > 1: + raise RuntimeContractError( + _fail_code(policy, "exclusive_marker_conflict", "PROMPT_EXCLUSIVE_MARKER_CONFLICT"), + "mutually exclusive prompt markers coexist", + present, + ) + scalar_count = len(compiled) + estimate = math.ceil(len(compiled_bytes) / 4) + manifest = { + "schema_version": "compiled_prompt_manifest.v1", + "compiled_prompt_sha256": sha256_bytes(compiled_bytes), + "compiled_prompt_bytes": len(compiled_bytes), + "deterministic_prompt_size_guard": { + "utf8_bytes": len(compiled_bytes), + "unicode_scalars": scalar_count, + "not_a_tokenizer": True, + "model_context_fit_not_proven": True, + "legacy_advisory_estimate": estimate, + }, + "normalization": {"encoding": "UTF-8", "line_endings": "LF", "trailing_newline_count": 1, "separator": separator}, + "fragments": [ + { + "fragment_id": item["spec"].fragment_id, + "category": item["spec"].category, + "path": (Path(item["spec"].path).resolve().relative_to(Path(__file__).resolve().parents[1]).as_posix()), + "sha256": item["sha256"], + "hash_kind": "file_sha256", + "source_domain_id": item["spec"].source_domain_id, + } + for item in kept + ], + "deduplicated_fragments": omitted, + "required_markers_verified": required_markers, + } + return compiled, manifest + + +def _path_from_config(config_record: dict[str, Any], value: str) -> str: + config_path = config_record.get("config_path") + base = Path(config_path).parent if config_path else Path(config_record["index_entry"].get("base_path", ".")) + path = Path(value) + return str(path if path.is_absolute() else (base / path).resolve()) + + +def _config_prompt_specs(config_record: dict[str, Any], category: str, source_domain_id: str) -> list[FragmentSpec]: + config = config_record["config"] + values: list[Any] = [] + if isinstance(config.get("prompt_fragments"), list): + values.extend(config["prompt_fragments"]) + for key in ("prompt_overlay_ref", "seed_prompt_path", "prompt_path", "prompt_overlay_path"): + if isinstance(config.get(key), str): + values.append({"path": config[key], "fragment_id": f"{source_domain_id}:{key}"}) + index_entry = config_record.get("index_entry") if isinstance(config_record.get("index_entry"), dict) else {} + if not values and isinstance(index_entry.get("prompt_overlay_path"), str): + values.append({"path": index_entry["prompt_overlay_path"], "fragment_id": f"{source_domain_id}:index_prompt_overlay"}) + specs: list[FragmentSpec] = [] + for index, value in enumerate(values, 1): + if isinstance(value, str): + path_value, fragment_id = value, f"{source_domain_id}:prompt:{index:02d}" + elif isinstance(value, dict) and isinstance(value.get("path"), str): + path_value = value["path"] + fragment_id = str(value.get("fragment_id") or value.get("id") or f"{source_domain_id}:prompt:{index:02d}") + else: + continue + specs.append(FragmentSpec(fragment_id, category, _path_from_config(config_record, path_value), source_domain_id)) + return specs + + +def collect_domain_fragments( + domain_id: str, + registry: dict[str, Any], + *, + common_contract_path: str, + profile_paths: dict[str, str] | None = None, + runtime_guard_path: str | None = None, + extra_specs: list[FragmentSpec] | None = None, +) -> list[FragmentSpec]: + entries = registry["entries"] + if domain_id not in entries: + raise RuntimeContractError("DOMAIN_NOT_IN_REGISTRY", f"cannot compile prompt for unknown domain: {domain_id}") + specs = [FragmentSpec("common_worker_contract", "common_contract", str(Path(common_contract_path).resolve()))] + visited: set[str] = set() + + def add_dependencies(current: str) -> None: + for dependency in sorted([str(item) for item in entries[current]["config"].get("depends_on", [])], key=natural_key): + if dependency not in entries: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"dependency config missing: {dependency}") + if dependency in visited: + continue + visited.add(dependency) + add_dependencies(dependency) + specs.extend(_config_prompt_specs(entries[dependency], "common_dependency", dependency)) + + add_dependencies(domain_id) + domain_specs = _config_prompt_specs(entries[domain_id], "domain", domain_id) + if not domain_specs: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"domain has no prompt fragment: {domain_id}") + specs.extend(domain_specs) + profiles = entries[domain_id]["config"].get("special_law_profiles", []) + profile_map = profile_paths or {} + for profile in sorted([str(item.get("profile_id") or item.get("id")) if isinstance(item, dict) else str(item) for item in profiles], key=natural_key): + if profile not in profile_map: + raise RuntimeContractError("PROMPT_REQUIRED_FRAGMENT_MISSING", f"special-law profile prompt path missing: {profile}") + specs.append(FragmentSpec(f"special_law_profile:{profile}", "special_law_profile", str(Path(profile_map[profile]).resolve()), domain_id)) + if runtime_guard_path: + specs.append(FragmentSpec("runtime_guard", "runtime_guard", str(Path(runtime_guard_path).resolve()))) + specs.extend(extra_specs or []) + return specs + + +def _parse_mapping(values: list[str], separator: str = "=") -> dict[str, str]: + out: dict[str, str] = {} + for value in values: + if separator not in value: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", f"expected KEY{separator}PATH", value) + key, path = value.split(separator, 1) + out[key] = path + return out + + +def _parse_extra(values: list[str]) -> list[FragmentSpec]: + out: list[FragmentSpec] = [] + for value in values: + parts = value.split(":", 2) + if len(parts) != 3: + raise RuntimeContractError("PROMPT_ARGUMENT_INVALID", "fragment must be CATEGORY:ID:PATH", value) + out.append(FragmentSpec(parts[1], parts[0], str(Path(parts[2]).resolve()))) + return out + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--registry-index", required=True) + parser.add_argument("--domain-id", required=True) + parser.add_argument("--policy", required=True) + parser.add_argument("--common-contract", required=True) + parser.add_argument("--runtime-guard") + parser.add_argument("--profile", action="append", default=[]) + parser.add_argument("--fragment", action="append", default=[]) + parser.add_argument("--output-prompt", required=True) + parser.add_argument("--output-manifest", required=True) + args = parser.parse_args(argv) + try: + registry = load_registry(args.registry_index) + specs = collect_domain_fragments( + args.domain_id, + registry, + common_contract_path=args.common_contract, + profile_paths=_parse_mapping(args.profile), + runtime_guard_path=args.runtime_guard, + extra_specs=_parse_extra(args.fragment), + ) + policy = load_json(args.policy) + compiled, manifest = compile_fragments(specs, policy) + prompt_path = Path(args.output_prompt) + prompt_path.parent.mkdir(parents=True, exist_ok=True) + prompt_path.write_text(compiled, encoding="utf-8", newline="") + from runtime_common import sha256_file + + manifest.update({ + "domain_id": args.domain_id, + "compiled_prompt_path": str(prompt_path), + "registry_index_sha256": registry["index_sha256"], + "composition_policy_path": args.policy, + "composition_policy_sha256": sha256_file(args.policy), + }) + write_json(args.output_manifest, manifest) + emit_cli_result({"status": "PASS", "domain_id": args.domain_id, "compiled_prompt_path": str(prompt_path), "compiled_prompt_sha256": manifest["compiled_prompt_sha256"]}) + return 0 + except RuntimeContractError as exc: + emit_cli_result({"status": "FAILED", "error": exc.as_dict()}) + return 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/runtime_manifest.pre_special_law_wiring.json b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/runtime_manifest.pre_special_law_wiring.json new file mode 100644 index 00000000..efce5f2c --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/runtime_manifest.pre_special_law_wiring.json @@ -0,0 +1,1739 @@ +{ + "entries": [ + { + "path": "computations/CE-01/build_scenarios.py", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-01/build_scenarios.txt", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-01/calculate_preview.py", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-01/calculate_preview.txt", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-01/calculation_config.json", + "sha256": "3203d7794e560fadea0fdef3185d778c4e008035faa48073aa031bd1af1c4cf9" + }, + { + "path": "computations/CE-01/input.schema.json", + "sha256": "4b38fc245443c122964cd991b625e10dcd1c8197322fb4be6b67ce6d25dab2e3" + }, + { + "path": "computations/CE-01/legal_sources.json", + "sha256": "36b2c2b59f7d779045da04afe628993187db1f898c765caa6f6ef7124869e0e5" + }, + { + "path": "computations/CE-01/result.schema.json", + "sha256": "6b59729f53471e597b0f69c2dc2d9ad8699edbb14d7eb734f3bc1edac1acb157" + }, + { + "path": "computations/CE-01/rule_bindings.json", + "sha256": "c45cea648e1e31b3b6e3bc5a855f162146d95ffb64914cdc51321cb39bcc9086" + }, + { + "path": "computations/CE-01/runtime.py", + "sha256": "ae4c5dd04251be4fc5d95ac87f940a76c86efd174538de7e36fcad064d26c685" + }, + { + "path": "computations/CE-01/runtime.txt", + "sha256": "ae4c5dd04251be4fc5d95ac87f940a76c86efd174538de7e36fcad064d26c685" + }, + { + "path": "computations/CE-01/validate_operands.py", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-01/validate_operands.txt", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-02/build_scenarios.py", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-02/build_scenarios.txt", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-02/calculate_preview.py", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-02/calculate_preview.txt", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-02/calculation_config.json", + "sha256": "8c3ca126253713b8d0ab7b2fc7b642744a11da6e1289fb058cd26700376830d7" + }, + { + "path": "computations/CE-02/input.schema.json", + "sha256": "c793bf6ee703ff74949c3f6c6bd601043c353a5f2270d33cf1179c495d49f409" + }, + { + "path": "computations/CE-02/legal_sources.json", + "sha256": "516d1df3c30bca72ddca5ad553365b51428734d7d889737eced278df4b2a28d2" + }, + { + "path": "computations/CE-02/result.schema.json", + "sha256": "040e9b835292a22b7189f322a77326e884104fdd5eca1e6735ac8032b2d3e110" + }, + { + "path": "computations/CE-02/rule_bindings.json", + "sha256": "ef13267fe8a9d1465bc4e8d9cb17113e8bc8f5c67fa8d593057364e9b9ef5c41" + }, + { + "path": "computations/CE-02/runtime.py", + "sha256": "d3c137e3ddbab07e53fec3d8f0653e01c8b00c4db3c5cd0edf964aaa9cdae3c6" + }, + { + "path": "computations/CE-02/runtime.txt", + "sha256": "d3c137e3ddbab07e53fec3d8f0653e01c8b00c4db3c5cd0edf964aaa9cdae3c6" + }, + { + "path": "computations/CE-02/validate_operands.py", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-02/validate_operands.txt", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-03/build_scenarios.py", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-03/build_scenarios.txt", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-03/calculate_preview.py", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-03/calculate_preview.txt", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-03/calculation_config.json", + "sha256": "58c373598b08b6cc589684f6f6665882c9678f608eb95bc0f5b5b9f9a064e86a" + }, + { + "path": "computations/CE-03/input.schema.json", + "sha256": "11d587b3ca72fc6f09b13b4e032943d4111f9f07c13cbfd28b08ec292820f1f8" + }, + { + "path": "computations/CE-03/legal_sources.json", + "sha256": "fa5a8d24bf34e898637cb0928498e0930853a6e302c443725c98e7978ef7c866" + }, + { + "path": "computations/CE-03/result.schema.json", + "sha256": "584c593086729508bdc7dcf77f41165e087bee080b6f5eee4752e1a176f0d747" + }, + { + "path": "computations/CE-03/rule_bindings.json", + "sha256": "03903645939d2a24103c6e2ec38f406c58087e5be1c8e3ba2868beaf23a22884" + }, + { + "path": "computations/CE-03/runtime.py", + "sha256": "345259ad0e6642617465f5bece4070b7c869af4c95398992185166873851e7ce" + }, + { + "path": "computations/CE-03/runtime.txt", + "sha256": "345259ad0e6642617465f5bece4070b7c869af4c95398992185166873851e7ce" + }, + { + "path": "computations/CE-03/validate_operands.py", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-03/validate_operands.txt", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-04/build_scenarios.py", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-04/build_scenarios.txt", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-04/calculate_preview.py", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-04/calculate_preview.txt", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-04/calculation_config.json", + "sha256": "715362e80a35b578cd70a3ca6ea6a3081d777c4dfe65076097172d7090a43505" + }, + { + "path": "computations/CE-04/input.schema.json", + "sha256": "297f4f3909cdfff92b99af02d68eb05f3615138e7012e3480b18c50f4018955f" + }, + { + "path": "computations/CE-04/legal_sources.json", + "sha256": "abe56ff1f1c61cf033a4382c9f1522cc5c9ab396fb4afd2b1736388e2f35c1f0" + }, + { + "path": "computations/CE-04/result.schema.json", + "sha256": "acbd26ff73ff673ff77cb98a3cf888f8b0b11e0418ec7b794589dad78f544008" + }, + { + "path": "computations/CE-04/rule_bindings.json", + "sha256": "27da494fba541247ad6dd5608f82fcdf36cf6532d887790fc57a437bbafeb959" + }, + { + "path": "computations/CE-04/runtime.py", + "sha256": "db82038e271c68a4265412b052df9a142078621e798c812d7253849e5b3e73ac" + }, + { + "path": "computations/CE-04/runtime.txt", + "sha256": "db82038e271c68a4265412b052df9a142078621e798c812d7253849e5b3e73ac" + }, + { + "path": "computations/CE-04/validate_operands.py", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-04/validate_operands.txt", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-05/build_scenarios.py", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-05/build_scenarios.txt", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-05/calculate_preview.py", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-05/calculate_preview.txt", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-05/calculation_config.json", + "sha256": "b9fdd5c29cb3236860dc235c615cc3d9de76df1b83a913c539fdd1ed08525593" + }, + { + "path": "computations/CE-05/input.schema.json", + "sha256": "0a6a5685d8b371d21879d9471e4836170b556566748e59ab0bf8e92093551f65" + }, + { + "path": "computations/CE-05/legal_sources.json", + "sha256": "705f076994c0aac833fd693c262fffe9272654ae8e5119db9909229ab1f0526f" + }, + { + "path": "computations/CE-05/result.schema.json", + "sha256": "090a3d412d4dd3c73484a104566372be477d646cd489f32f2142a7bc80076966" + }, + { + "path": "computations/CE-05/rule_bindings.json", + "sha256": "7a7c80890e1f1dc7d31e791aedd6ff3cbf758808890b659817a6a9b24524dde1" + }, + { + "path": "computations/CE-05/runtime.py", + "sha256": "539c97da9a9e6efe5236a2f83de1b93ea0f29820abb8e9d85d75abad0a7fa066" + }, + { + "path": "computations/CE-05/runtime.txt", + "sha256": "539c97da9a9e6efe5236a2f83de1b93ea0f29820abb8e9d85d75abad0a7fa066" + }, + { + "path": "computations/CE-05/validate_operands.py", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-05/validate_operands.txt", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-06/build_scenarios.py", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-06/build_scenarios.txt", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-06/calculate_preview.py", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-06/calculate_preview.txt", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-06/calculation_config.json", + "sha256": "af6698a82611c6b433ab0325c85e023af8e03ed3243c0df93ad0d79a242e5453" + }, + { + "path": "computations/CE-06/input.schema.json", + "sha256": "e0d17b241b10a2100468502f9f603c58f286bf88334dd87a9092ddfd16d9c03f" + }, + { + "path": "computations/CE-06/legal_sources.json", + "sha256": "68226ef6ac76dfaf5c190bd18b2d14135c275aefaa9ce641661e864e1f0f47a4" + }, + { + "path": "computations/CE-06/result.schema.json", + "sha256": "fd75007c1efdadf38532ca8b472dce2141059a2a262e1afb6a4b112f8469607e" + }, + { + "path": "computations/CE-06/rule_bindings.json", + "sha256": "05e3929ff879676220e17d3692adda9e4291f8931d9a08dd55c0a89e6952211d" + }, + { + "path": "computations/CE-06/runtime.py", + "sha256": "ca4793e19e8584c02b8ca539f0a3ddf97e98bc3915c3f0f6c507a495b704fe22" + }, + { + "path": "computations/CE-06/runtime.txt", + "sha256": "ca4793e19e8584c02b8ca539f0a3ddf97e98bc3915c3f0f6c507a495b704fe22" + }, + { + "path": "computations/CE-06/validate_operands.py", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-06/validate_operands.txt", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-07/build_scenarios.py", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-07/build_scenarios.txt", + "sha256": "6ea1e4d6aaf320bcff54357f1e6cfed0b6db4cbdcf33dc5a3dd69a865ba0b4e4" + }, + { + "path": "computations/CE-07/calculate_preview.py", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-07/calculate_preview.txt", + "sha256": "d891d55e2b19838e882267aaab9f45414c0af308b0ea0453d87363757afc0fec" + }, + { + "path": "computations/CE-07/calculation_config.json", + "sha256": "9d8fbc73d1e258d6f07ea0752e5a41f1f3cab6ed263420bffe39bf78ee1a9732" + }, + { + "path": "computations/CE-07/input.schema.json", + "sha256": "5b746de91963c1acc7e91141a94c0e6a234220f40efaae019ff60d5aabda024b" + }, + { + "path": "computations/CE-07/legal_sources.json", + "sha256": "31b7e72fb91ece1cd8b7e910235951c23d2a9c3962c7eb247b0611a6e61a7634" + }, + { + "path": "computations/CE-07/result.schema.json", + "sha256": "0e68e5748eb7954457102c0a9fd145cab6e66dfd67295dab32aac322c4843c89" + }, + { + "path": "computations/CE-07/rule_bindings.json", + "sha256": "9c5dfbca715415fdd70006e3d9238293162d6cf24c89077c316fb5911657da24" + }, + { + "path": "computations/CE-07/runtime.py", + "sha256": "aa6b3064ca7a32cb06f8d86592ebc045b5a7c4761c65cce4ce810c0943754911" + }, + { + "path": "computations/CE-07/runtime.txt", + "sha256": "aa6b3064ca7a32cb06f8d86592ebc045b5a7c4761c65cce4ce810c0943754911" + }, + { + "path": "computations/CE-07/validate_operands.py", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-07/validate_operands.txt", + "sha256": "bb0e1f4a0229201889ec2ad3636a4a0d0f2940c7368fe28a6ae0da25ddbb6fe7" + }, + { + "path": "computations/CE-08/build_scenarios.py", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-08/build_scenarios.txt", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-08/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-08/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-08/calculation_config.json", + "sha256": "af76ca7365a2ff67cfcfe5a98bcdad381dd069d98a7a50c995bf8a4e022ea4f4" + }, + { + "path": "computations/CE-08/input.schema.json", + "sha256": "5173c015bb2d3d546cc6fb06f4c9a7228e13e02ab20c5930224de6ac835ea688" + }, + { + "path": "computations/CE-08/legal_sources.json", + "sha256": "a5914a15751ed248e2cec995fcbfe9a623b2f213b8dd1e02f5cf2f777b214d3f" + }, + { + "path": "computations/CE-08/result.schema.json", + "sha256": "87606d3cb13fe880eb73f3d43128d006c9d8956310de24fd48a1037a00613eba" + }, + { + "path": "computations/CE-08/rule_bindings.json", + "sha256": "12bf2326c39fa0d044a66b9f7bb9129221f2d168e8a02834893c8800f1672df0" + }, + { + "path": "computations/CE-08/validate_operands.py", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-08/validate_operands.txt", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-09/build_scenarios.py", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-09/build_scenarios.txt", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-09/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-09/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-09/calculation_config.json", + "sha256": "06e42c61676f1f0d2cd3bcfe25a4257f6b541f630652510568b66f979ff01294" + }, + { + "path": "computations/CE-09/input.schema.json", + "sha256": "aaff89593e8ad63f21d1cb08c22ca3cf604f8b45d991e5e0be6abbe4125773d0" + }, + { + "path": "computations/CE-09/legal_sources.json", + "sha256": "9153f64b23ec2770452c21e58da58626aa7d5adb55d2af33d15bc82834c02c2e" + }, + { + "path": "computations/CE-09/result.schema.json", + "sha256": "f13cfd422622a9b5f5b561ebe33b3ea3f873c5f7a116b551128ba54d6119093c" + }, + { + "path": "computations/CE-09/rule_bindings.json", + "sha256": "3eed1546d7d546eda76797a3e3f7f845bd54cb00213b6d266c9a370163d10401" + }, + { + "path": "computations/CE-09/validate_operands.py", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-09/validate_operands.txt", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-10/build_scenarios.py", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-10/build_scenarios.txt", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-10/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-10/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-10/calculation_config.json", + "sha256": "fb2cbb42cbee14327696d958b28cbc56ace3e658a603ce9ab4495a573bf30868" + }, + { + "path": "computations/CE-10/input.schema.json", + "sha256": "2614a0105da34d1cb92c6bea41939782e06911e6ade1084ab46089748daf93c0" + }, + { + "path": "computations/CE-10/legal_sources.json", + "sha256": "79405a0ef65ad1c68eed0a7858451e13a0611056b66f627402c7ab1eadbaa260" + }, + { + "path": "computations/CE-10/result.schema.json", + "sha256": "3f0f882ec80d917a9467d6ad103913689f19d76dd928be4b41c4e5148093145b" + }, + { + "path": "computations/CE-10/rule_bindings.json", + "sha256": "40c97d91bea995d115c87273275eb431713bc2765ea7282fe3ea4379c1ea42df" + }, + { + "path": "computations/CE-10/validate_operands.py", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-10/validate_operands.txt", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-11/build_scenarios.py", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-11/build_scenarios.txt", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-11/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-11/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-11/calculation_config.json", + "sha256": "c8227234e8be3c0070d6f227fba793fe07c32fa9f66b0448eb56090ccdf6a8b5" + }, + { + "path": "computations/CE-11/input.schema.json", + "sha256": "27abe619222ffba678fa5c79581c58018055cb4cc52822808fe9f67998e6d233" + }, + { + "path": "computations/CE-11/legal_sources.json", + "sha256": "d410993f1f1c539d9648f575b904e8cf44cc8a884a3de515f7dd9141c5fbf2b7" + }, + { + "path": "computations/CE-11/result.schema.json", + "sha256": "158d2cbe5512f321539e5ec15b3a7a0dd2bd7cee5e7696b5da614f0eb146c301" + }, + { + "path": "computations/CE-11/rule_bindings.json", + "sha256": "85633f8db119854985bd5977112f60b9a95ca372e1984435cc45b9c1906b4122" + }, + { + "path": "computations/CE-11/validate_operands.py", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-11/validate_operands.txt", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-12/build_scenarios.py", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-12/build_scenarios.txt", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-12/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-12/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-12/calculation_config.json", + "sha256": "37b3f54711741e7904470edd27edd73252941cfa9676795d7cf3d094b0e50e5f" + }, + { + "path": "computations/CE-12/input.schema.json", + "sha256": "dc0bf45b7b2cb025d8499db3018414d53e869b5f124d73b8a19568b6b6ea1da6" + }, + { + "path": "computations/CE-12/legal_sources.json", + "sha256": "c5f24a0d78643a318eaeb7d415542b7f229c9f883461e93716bbd5da67245670" + }, + { + "path": "computations/CE-12/result.schema.json", + "sha256": "bc23734a58e301ffe5fc14cdb5e856fd4ca19e797c2b744795f5ad27f7fe9a9e" + }, + { + "path": "computations/CE-12/rule_bindings.json", + "sha256": "e1eb7a94032bdd61f236f93fe40790cb7ba936b7516ed699ac039c4afb888c7d" + }, + { + "path": "computations/CE-12/validate_operands.py", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-12/validate_operands.txt", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-13/build_scenarios.py", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-13/build_scenarios.txt", + "sha256": "3e8a6566877a1ed0db0b2ecf44a9d0c2a155b298111a1febfa41e2c887da3b31" + }, + { + "path": "computations/CE-13/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-13/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-13/calculation_config.json", + "sha256": "0777b8d75002f503fe947cba412dd78c4f665974249f04d00ea510ba736e3669" + }, + { + "path": "computations/CE-13/input.schema.json", + "sha256": "c4fdb99ed30ad0527e337d3d5f93d783e3df4487b8ab976d0d522c8a8e9c633a" + }, + { + "path": "computations/CE-13/legal_sources.json", + "sha256": "9251de924bc65f1f1ccf112408333b774769fb4f8c578be5f46afa3dd250381a" + }, + { + "path": "computations/CE-13/result.schema.json", + "sha256": "9eba9125f2d9c740aad40674e46cbc999985e2f2845bb178f90d9b0ffc956c3d" + }, + { + "path": "computations/CE-13/rule_bindings.json", + "sha256": "4ffca8e7c9f7d6318ed476521a039455ca1be0c3c15e78633595dd27edf6ff70" + }, + { + "path": "computations/CE-13/validate_operands.py", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-13/validate_operands.txt", + "sha256": "7e59236c0a04581a63e89d33a0afa252b1fa2b4fcd8d82d8c6c6265b2161db8f" + }, + { + "path": "computations/CE-R1/build_scenarios.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R1/build_scenarios.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R1/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R1/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R1/calculation_config.json", + "sha256": "9f8b365ff6e2f530886bf6ec3d3cb0ad1bc0c4ab1f971b22789085e4c5d0d482" + }, + { + "path": "computations/CE-R1/input.schema.json", + "sha256": "128e2f1f9255d3a71db061aa8b7e0a63d0cbcdbf145f05077dd77e85fa0f92d7" + }, + { + "path": "computations/CE-R1/legal_sources.json", + "sha256": "e5b3fa59bb7c2abd0ab503f48dbca62323056a85dedb8ef87e1850b5e4e9bc1a" + }, + { + "path": "computations/CE-R1/result.schema.json", + "sha256": "8500fa2d64469705fb5b90fcd75ffdeea8a14ad343471436f29935bc9b306679" + }, + { + "path": "computations/CE-R1/rule_bindings.json", + "sha256": "6d91a467a0e8b70a119cc022ad9a0a7826a4c890dbb2a33b35de0e063c5c6e35" + }, + { + "path": "computations/CE-R1/validate_operands.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R1/validate_operands.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R2/build_scenarios.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R2/build_scenarios.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R2/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R2/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R2/calculation_config.json", + "sha256": "bd880eea154c8f7b14ee0fad1c30275eb94ab7a16261991e19c89a963ba17f07" + }, + { + "path": "computations/CE-R2/input.schema.json", + "sha256": "604862070ca05071d33035f056b06addf07eb95628093e5d78c8b969728dc44b" + }, + { + "path": "computations/CE-R2/legal_sources.json", + "sha256": "d7dfe4f134bfcbad54159d1dea0c58abcd3e39035766d1bcdeed3b26161974e0" + }, + { + "path": "computations/CE-R2/result.schema.json", + "sha256": "3b89d09e2166d31b48c413de038fb92f6e58d64713bf828ccb8a36c7ecc0f266" + }, + { + "path": "computations/CE-R2/rule_bindings.json", + "sha256": "22725fe7f52e3ab28dbb9b538e886115e7abb064c4b5e6d053832d58e2e3a45e" + }, + { + "path": "computations/CE-R2/validate_operands.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R2/validate_operands.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R3/build_scenarios.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R3/build_scenarios.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R3/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R3/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R3/calculation_config.json", + "sha256": "89c4cf1264ca3fb63624b1df6fd9bd3990d7c320d776fde5d88ad0f351c8880e" + }, + { + "path": "computations/CE-R3/input.schema.json", + "sha256": "eb8c3219a8e117ad8eb73ef974663638ab7f239a817d0462acc79f142dab2ba6" + }, + { + "path": "computations/CE-R3/legal_sources.json", + "sha256": "f5d6d3dab2c3654ed2a1186a67bfe165ee3a9cd41b2be58a88ac140ae53195a1" + }, + { + "path": "computations/CE-R3/result.schema.json", + "sha256": "0d5dd5a296da5f6afe469075c28d93e51d9c32de7ca882b575194a7b719a27bd" + }, + { + "path": "computations/CE-R3/rule_bindings.json", + "sha256": "1a7206f5f77e24bd200dca86286daaf80c00679744e37b31fbc429869f371582" + }, + { + "path": "computations/CE-R3/validate_operands.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R3/validate_operands.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R4/build_scenarios.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R4/build_scenarios.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R4/calculate_preview.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R4/calculate_preview.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R4/calculation_config.json", + "sha256": "aaa58513348b2f65a2da30753f94d4bd7ffec71568b59ed9607025006a970bef" + }, + { + "path": "computations/CE-R4/input.schema.json", + "sha256": "a9f1052844261ea32b0c40f2dda594681218fce6eead25d130fd3b1ab217028c" + }, + { + "path": "computations/CE-R4/legal_sources.json", + "sha256": "974ebfc8bd488d4105fb83637b1d3aed32ee631530ae1daaf7d653da104a21c6" + }, + { + "path": "computations/CE-R4/result.schema.json", + "sha256": "d71aa181f66e5b9bf96385e47b4302bc469f6517d09981b781938d658ff2cef2" + }, + { + "path": "computations/CE-R4/rule_bindings.json", + "sha256": "230cbc7c6785695dad0334a27d7da735b10b7198007b42134db4ba31be0acac7" + }, + { + "path": "computations/CE-R4/validate_operands.py", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/CE-R4/validate_operands.txt", + "sha256": "ba28e4c4f5a7cea416cd37a91b9ba68ca0e1c2f3156a13f69bee3dbf9ac58cfc" + }, + { + "path": "computations/_shared/numeric_guards.py", + "sha256": "498280d6f8080aafd5c1c46e70747514e4685a19f2a87f6bf254a6856b26d365" + }, + { + "path": "computations/_shared/numeric_guards.txt", + "sha256": "498280d6f8080aafd5c1c46e70747514e4685a19f2a87f6bf254a6856b26d365" + }, + { + "path": "computations/calculation_registry.json", + "sha256": "7ec7200acf2b6f31923edadf368e62784ea322b46862fd57e990b97d131b9da5" + }, + { + "path": "computations/calculation_registry_manifest.json", + "sha256": "2298d5b6fb8a1a977fce08079a78b18fefa675d32fc041640c083ce2cf8fe270" + }, + { + "path": "contracts/signals/s5_execution_contract.v2.json", + "sha256": "311ac8721ec297bc86e54c83f0866d61629bd99f15746f86bbbff32914d35e02" + }, + { + "path": "domains/E-00/module_role_projection.json", + "sha256": "278ecdd4ac6c004f9994d6357ef439d089456a5e65bed0f74f49170ba3315932" + }, + { + "path": "domains/E-00/structure_types.json", + "sha256": "ae4a94fdbb0e2a8e0273a259e38c28774a3f2f24985878286f18a9ae05cbabaa" + }, + { + "path": "domains/E-01/module_role_projection.json", + "sha256": "cb263ca339f7b749a51b57fa12d915fda3644f09e648d15a84202651a23a84fb" + }, + { + "path": "domains/E-01/structure_types.json", + "sha256": "f2c423a0346a347292b3ca33d3a9f33a82d540e83cbb5f79bb8dae7161c68b26" + }, + { + "path": "domains/E-02/module_role_projection.json", + "sha256": "e21465cf806fb080d75f89526f9c142010c392ee87454a5a19248b9a80eb7c28" + }, + { + "path": "domains/E-02/structure_types.json", + "sha256": "e470dd3c96f2f9b8bfa6c2c97b05259281005a84b6eea24c38d322abdf16ad8e" + }, + { + "path": "domains/E-03/module_role_projection.json", + "sha256": "f693146dfb3ba2ac5fa6a6221c6b2dc7e32084acb06823b1f46f423a56546704" + }, + { + "path": "domains/E-03/structure_types.json", + "sha256": "70c579a03c41226e6383fe7b4aa5e60f4b0f5a723db30f6292dd2a0bde4d17bb" + }, + { + "path": "domains/E-04/module_role_projection.json", + "sha256": "4c4bd473a12ad1dfcc8c9392503cd057017385a21287a89cb7ea147f3cf7d2d6" + }, + { + "path": "domains/E-04/structure_types.json", + "sha256": "e8b15e9e5a39a488bf9ef660d69c7bc2db832248eb33788ba2ae54a9b5200e73" + }, + { + "path": "domains/E-05/module_role_projection.json", + "sha256": "7fabfa3893a138dccb4e5bc797444f5f999c0e51ec6b0c8fbe0211eff3f03a6c" + }, + { + "path": "domains/E-05/structure_types.json", + "sha256": "22707fbce849991c099a08885cac929d1ba842aa896945f9921c3214e4c701d0" + }, + { + "path": "domains/E-06/module_role_projection.json", + "sha256": "805b73b2b15cb720b62cd83455de85a67bda434255b69d588eb98f2527f88b71" + }, + { + "path": "domains/E-06/structure_types.json", + "sha256": "a6520f72813caeaf72d62fd63eca10a1f46fd4fb3ffd97f97f222f37c53e8663" + }, + { + "path": "domains/E-07/module_role_projection.json", + "sha256": "9920ed415f9768beb517100e97fa7d2eb862ca90ca83f802ea460f8f6bbf2445" + }, + { + "path": "domains/E-07/structure_types.json", + "sha256": "fd4946e24cacc3bcc74c7fee5ce7fca99d3e5f2339fe3a9d698c54e2626a6c01" + }, + { + "path": "domains/E-08/module_role_projection.json", + "sha256": "4a8c284b5d3bbf8ca224e9f08398e6f563b00cec6de72f5128acc4cd8425eced" + }, + { + "path": "domains/E-08/structure_types.json", + "sha256": "a7ee6782c6541d0c6f6f6107a1d3615c2e7003a979741f2afc9d5d686a59e51d" + }, + { + "path": "domains/E-09/module_role_projection.json", + "sha256": "ab93656808a832cfd73a6d75ed28d75e50e6cc66c67d59d765fcb8bbf638205f" + }, + { + "path": "domains/E-09/structure_types.json", + "sha256": "c374b8b912679cc765b152c58ecd8320898f947b799e44ca509aa622669d7f4c" + }, + { + "path": "domains/E-10/module_role_projection.json", + "sha256": "80f760d9d8516cc81ebe503a2967418c43e11f8b8410231081111483dfa6477d" + }, + { + "path": "domains/E-10/structure_types.json", + "sha256": "07be97d40e422e14e570574c4e1da875307fe487c39e2ba9862ffd4b1623b748" + }, + { + "path": "domains/E-11/module_role_projection.json", + "sha256": "0a55ce91b1845d64d085710110aec22bfcea66967ab5fce3a4acf5179faf65f1" + }, + { + "path": "domains/E-11/structure_types.json", + "sha256": "6e565bab389f59866a0eafa8b7b7ff89f90ef2363ced3d9a4240100d3b3c3b90" + }, + { + "path": "domains/E-12/module_role_projection.json", + "sha256": "207ef7d4a1300c12aa54cb1defa1110a6a9d18850d0fc4e144baebe6ff2b3c27" + }, + { + "path": "domains/E-12/structure_types.json", + "sha256": "ef97f4b0ed05637165bfb211fc9be0cb3bcf0523e6b5a645c007e8dc2825bea9" + }, + { + "path": "domains/E-13/module_role_projection.json", + "sha256": "412df6800be576d7f789bc7fc2045091eb9c49fd346e442ba42d6f29300f7313" + }, + { + "path": "domains/E-13/structure_types.json", + "sha256": "cc18073ca443e98a34e139c08d11121ceb7d4ecd366ad3b3d5cb65ddb53e36f7" + }, + { + "path": "domains/E-14/module_role_projection.json", + "sha256": "cead5634615b6b417390b37a948a050ea89cba37485bb79617bd1296c06829a5" + }, + { + "path": "domains/E-14/structure_types.json", + "sha256": "5b0025a720a7b9f1ad7c74ce8d79f0969b282690abc0e3171e28dbf41933153b" + }, + { + "path": "domains/E-15/module_role_projection.json", + "sha256": "873f921bbbc86fc07a0c1ecfcbcacee67b1d9592b2efd7319290a10e8c1ba018" + }, + { + "path": "domains/E-15/structure_types.json", + "sha256": "784ceabd17c6de6d95a30f52dd5fd4036946fbedc556f2bad536225057a0f5fe" + }, + { + "path": "domains/E-16/module_role_projection.json", + "sha256": "24ad172b32d4942be6f4031c582f928d3c6feb6e442027e3ee02700519744b11" + }, + { + "path": "domains/E-16/structure_types.json", + "sha256": "b0bb30602ec424d35359c1a676ccca2bb2eda4a676997753da44d1dbbcc1f0b3" + }, + { + "path": "domains/E-17/module_role_projection.json", + "sha256": "498e2320ad04a907680b8588ee1b1b9ee7487e131bece96a601979d60796f5a4" + }, + { + "path": "domains/E-17/structure_types.json", + "sha256": "d306dbdf583453d51f643eb8398920efb1bb5bfa82b3a600caa79f45f790b2eb" + }, + { + "path": "domains/E-18/module_role_projection.json", + "sha256": "b0699a793cdc52e6f74e9df3254f403858e83080be9e402bf3c7dae3cd8ffd0e" + }, + { + "path": "domains/E-18/structure_types.json", + "sha256": "ce4eae3a54bfe7ac31f16b657cc0a891580e025f7bfecb1812c015e7fd0c2500" + }, + { + "path": "domains/E-19/module_role_projection.json", + "sha256": "00950a8c76d4367ccd6dc000476c810df33e49b07fce124545e350530b008b99" + }, + { + "path": "domains/E-19/structure_types.json", + "sha256": "58b9503558bcab3fb7ae76532659d54694208d813197e97bac148d50795feb22" + }, + { + "path": "domains/E-20/module_role_projection.json", + "sha256": "7aef9706b50161693005ee61fc576a1562929fdf20830a3187fdda2a94fae8e3" + }, + { + "path": "domains/E-20/structure_types.json", + "sha256": "6766546fd61d3f46f0d0b71bc3c41d5c63a525954b2568d8003c9f5e90234074" + }, + { + "path": "domains/E-21/module_role_projection.json", + "sha256": "de6b0f29ad8100744a1a4fe8393a44a8a921d604b5244f832256450cb547722a" + }, + { + "path": "domains/E-21/structure_types.json", + "sha256": "b027ff04f21cc4a86c8285e2cde6f8fd23e0ff3a3aeb3fa61a9bd10142814e36" + }, + { + "path": "domains/X1/module_role_projection.json", + "sha256": "1bceacd99f408384fec9547262017d8c857dbdb58ae906446c45ce8f1e322920" + }, + { + "path": "domains/X1/structure_types.json", + "sha256": "3344a27beb548c7d0b083b6e442a8f70ca89ce02d22614fda83203d3ff53e7e9" + }, + { + "path": "domains/X2/module_role_projection.json", + "sha256": "9282422d06035a78b592c110c197c1afa58e78294566eecd2666f0c61820cd5f" + }, + { + "path": "domains/X2/structure_types.json", + "sha256": "f51c0404bab1c5f4506012460ba02faf2cee6c958d49ee353095366defacd2ec" + }, + { + "path": "domains/X3/module_role_projection.json", + "sha256": "93d3709c7d01132b82f0445d18eab7e666ea8d7fe2afe9b308bf87af5477be10" + }, + { + "path": "domains/X3/structure_types.json", + "sha256": "91f7ef0a971b519bd7fea8fb2efc61770afeb83b3a6ee774f153aea28a9de913" + }, + { + "path": "domains/_common/common_worker_contract.md", + "sha256": "f133c54c5a9353ae14efc4d7a27d9bee18872b6727713a9a107778eef8990fb1" + }, + { + "path": "law_versions/_law_source_index.json", + "sha256": "b979f53d7062fb0594aeb193b759dff6b56b4797f58a0fd07598902a7d84c3fb" + }, + { + "path": "law_versions/calculator_runtime.py", + "sha256": "3d90b2af3a7bd1e743dfb654cc529b0044c867fd343ea9931c1bb1f3797e86ca" + }, + { + "path": "law_versions/calculator_runtime.txt", + "sha256": "3d90b2af3a7bd1e743dfb654cc529b0044c867fd343ea9931c1bb1f3797e86ca" + }, + { + "path": "law_versions/current_version_pin.json", + "sha256": "ab1d6cb8ad9b5257ec7e96b064ceb8395ff1709e042f478e713c0bdaf5459a0b" + }, + { + "path": "law_versions/forbidden_literal_dictionary.json", + "sha256": "7b0ca421eddc3b6a1913b62af2e419436eebe1899a09da203c32f47c65d738ca" + }, + { + "path": "law_versions/index.json", + "sha256": "44f746f4ec656d05d11f85cab7d39ba507a49a944bdcab2f46afce2a0b235b87" + }, + { + "path": "law_versions/law_source.schema.json", + "sha256": "70c786833aa7876f2e1babefea780f5b57fd1b69271e60e0673df76f18d7ed43" + }, + { + "path": "law_versions/law_version_registry.schema.json", + "sha256": "8c443ca46f50a60dedda3d7f52e4968ae26d82ebab2f7ac33a1685f6532fa250" + }, + { + "path": "law_versions/records/accelerated_litigation_interest.json", + "sha256": "3ffe86f04c6815955dedb6d377087116c96f25723a44adfcabc960cd13fd3d49" + }, + { + "path": "law_versions/records/ce13_attorney_fee_cost_inclusion.json", + "sha256": "9967d8ab167e17a41eb4d97618b4a4b21e31a4798c1d4c9d5f727613c1f1268c" + }, + { + "path": "law_versions/records/ce13_claim_value_rules.json", + "sha256": "43b22c0aa31ba80dd04b941c95f50084f07ae927b93f8a50eae4da3a080ed467" + }, + { + "path": "law_versions/records/ce13_stamp_fee_schedule.json", + "sha256": "f1f9899528050fee53b62e08edbb287b0910ed873f02695ecc259851e7aff241" + }, + { + "path": "law_versions/records/civil_commercial_statutory_interest.json", + "sha256": "7ae7dbba23221accd42c9d55270d09c9a0443a024eb3c78a07d3792aa1ea4b9a" + }, + { + "path": "law_versions/records/commercial_lease_major_versions.json", + "sha256": "108b5f23743aa6dff853c3c148881b6c3f1855284e6ef7d31c6feb720d3c52b7" + }, + { + "path": "law_versions/records/housing_lease_major_versions.json", + "sha256": "cd8bc45d6f19f7c43aae7bf3c61920129b03b23c07377bf1312008a80d8ee8bd" + }, + { + "path": "law_versions/records/interest_limitation_cap.json", + "sha256": "bd7a59ac226a8fc2168f1734b74c33a872c77a9e139fcc1dd3d7e7ea8d1f2d4a" + }, + { + "path": "law_versions/records/labor_delay_interest.json", + "sha256": "45a8992ea73c6ca519c61ed77f9c5189e555117efc63bc39648f561b396dcb86" + }, + { + "path": "law_versions/records/major_limitation_exclusion_periods.json", + "sha256": "64fbaa6e30a5a49d81a413a313b1c8d9f3b9f63c2077df3d7c28b58552d5eda3" + }, + { + "path": "law_versions/resolve_rule.py", + "sha256": "a96929c5960e431717fe4f2803fcabadc867d32430ed4b4213c8ce035be2efeb" + }, + { + "path": "law_versions/resolve_rule.txt", + "sha256": "a96929c5960e431717fe4f2803fcabadc867d32430ed4b4213c8ce035be2efeb" + }, + { + "path": "law_versions/rules/LR-ACCELERATED-JUDGMENT-INTEREST/2003-06-01.json", + "sha256": "20e168cfed8b02f7ac88e079812e138cc846c1d867e3fc417f87a7c146dc20a6" + }, + { + "path": "law_versions/rules/LR-ACCELERATED-JUDGMENT-INTEREST/2015-10-01.json", + "sha256": "758ccbe4422dd7d4e95d37ef34c673e48ddfb0dd7cf7ffd1257244041f18f4b3" + }, + { + "path": "law_versions/rules/LR-ACCELERATED-JUDGMENT-INTEREST/2019-06-01.json", + "sha256": "b2c0a43a1f96eb6f84779c7d0db296b800eeeebbea76a78ddf59488a7cbd73e3" + }, + { + "path": "law_versions/rules/LR-CE13-ATTORNEY-FEE-COST-INCLUSION/2018-04-01-current.json", + "sha256": "eb1424f3f619cafe1857f40ffb4e19243837b8686dfa8089960fb2a6453721f7" + }, + { + "path": "law_versions/rules/LR-CE13-CLAIM-VALUE-TABLE/2023-10-19-current.json", + "sha256": "bf9d3b9a11fb30240215d4c63108de958917a6d856c73eb5c063bcecd4e3688e" + }, + { + "path": "law_versions/rules/LR-CE13-FIRST-INSTANCE-STAMP/2023-10-19-current.json", + "sha256": "c39692cd45996fbeeda82e4380ae36efc509e06ee56cbc7f60ecbf7b2f12df47" + }, + { + "path": "law_versions/rules/LR-CIVIL-STATUTORY-INTEREST/civil-act-379-current.json", + "sha256": "89d87d5c2fd6fce98f9dd0eb1248ff3977a67d72ce8e8cb02567969da42ecbd5" + }, + { + "path": "law_versions/rules/LR-COMMERCIAL-LEASE-PANDEMIC/2020-09-29.json", + "sha256": "3af7daef66b399430f06797f3cc2d8c40abf54edcdea69348eaad45270bc571d" + }, + { + "path": "law_versions/rules/LR-COMMERCIAL-LEASE-PREMIUM/2015-05-13.json", + "sha256": "55aed9928928fb0a9d8f2351de7d82e8f2af59581067d9e8c5f8a3ffb607f266" + }, + { + "path": "law_versions/rules/LR-COMMERCIAL-LEASE-RENEWAL/2018-10-16.json", + "sha256": "0b2cf82d0bb1337e454f89e68055aac6051201f294f633b0ae06717adc289012" + }, + { + "path": "law_versions/rules/LR-COMMERCIAL-STATUTORY-INTEREST/commercial-act-54-current.json", + "sha256": "6922e051be8292c0f2330de4e8793137d9e43e2a25b47d83ed7eaa66f66fbd64" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-CANCELLATION-ACT/CivilAct-146-juridical-act.json", + "sha256": "a69b16e6169468c0e6e1e2cb3510c6fb17a3c90928529f9239ee4b11deba714f" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-CANCELLATION-RATIFICATION-POSSIBLE/CivilAct-146-ratification-possible.json", + "sha256": "c0deade3d1b4d18cd8d9737d915b8f7c8ca7ef176d2c5462651eb6983ffefbd0" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-CONTRACTOR-DEFECT/CivilAct-670.json", + "sha256": "b6533f2950549c41c65f9809bb43ae4990f21acb39f121a0081d97298609558a" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-CONTRACTOR-STRUCTURE-DESTRUCTION/CivilAct-671-destruction-damage.json", + "sha256": "479fe40fccfd01db22005adc50e3a7604520045a24e2b674e766dbdaaadf354f" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-CONTRACTOR-STRUCTURE-DURABLE/CivilAct-671-durable.json", + "sha256": "577fbc6c2c380f0763ac24e50853121b0bdeca87b319a0f2d10a4ba9ca94e620" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-CONTRACTOR-STRUCTURE-ORDINARY/CivilAct-671-ordinary.json", + "sha256": "8f35fe02575257a12018fae5b444c7406a80f89c53686adb6e24e50d255e0fa2" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-FRAUDULENT-ACT-ACT/CivilAct-406-2-act.json", + "sha256": "5fd1a86d29fa2a3511a996419415f2a91ed8e07554eecda21137ceb4e3c56f5f" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-FRAUDULENT-ACT-KNOWLEDGE/CivilAct-406-2-knowledge.json", + "sha256": "68fc66a41cf4070d1c24d91fb349043fdef0c6a1ae53385458a741969a32b360" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-INHERITANCE-RECOVERY-KNOWLEDGE/CivilAct-999-knowledge-current-line.json", + "sha256": "c9a9a44e37a0aa364c70c82354d0bf954b2ac403c6b7a1311de3d28fce7dc22f" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-INHERITANCE-RECOVERY-LONG/CivilAct-999-infringement-act.json", + "sha256": "2d0b9e77ef535918b9c0695290734a36b2605cbc3204670253535df84c3ef698" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-INHERITANCE-RECOVERY-LONG/CivilAct-999-inheritance-opening-pre-2002.json", + "sha256": "1e613701767916a0819be51cae8c23bde2b6cae769896b33f57736f11f271cf0" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-RESERVED-PORTION-KNOWLEDGE/CivilAct-1117-knowledge.json", + "sha256": "e0b511c2eb128edbaaa3cb04b7a86b469923d9e61c0c66a03be3480d4601aece" + }, + { + "path": "law_versions/rules/LR-EXCLUSION-RESERVED-PORTION-OPENING/CivilAct-1117-opening.json", + "sha256": "19700b5472b44b6905254b86d42901f0c878bc3eb1df9a9406390d366c5c24df" + }, + { + "path": "law_versions/rules/LR-HOUSING-LEASE-INFORMATION/2023-04-18.json", + "sha256": "eb72f4d629938f3ea173efa42459f01a0bdd98b0167cb9c6fa5bcd52909ed6c4" + }, + { + "path": "law_versions/rules/LR-HOUSING-LEASE-NOTICE-WINDOW/2020-12-10.json", + "sha256": "92c222923e86a47c1aaa605e3faba43a9359c1b2ccd29e4abea34618ebeef529" + }, + { + "path": "law_versions/rules/LR-HOUSING-LEASE-RENEWAL/2020-07-31.json", + "sha256": "a5808931cfbd30fda9e74dd398eae5ebc209e2d44aa4681816e6e83bb55b527a" + }, + { + "path": "law_versions/rules/LR-INTEREST-LIMITATION-CAP/2007-06-30.json", + "sha256": "7e408e9d54ce9645d9bbef0c94ef4e9afd7bb2e80bede91527920562cdf07e38" + }, + { + "path": "law_versions/rules/LR-INTEREST-LIMITATION-CAP/2014-07-15.json", + "sha256": "97a0e6bc47565fbd22f1abd88813cdff330656644d567ecaf2c2bcb1a540c702" + }, + { + "path": "law_versions/rules/LR-INTEREST-LIMITATION-CAP/2018-02-08.json", + "sha256": "ef0121da5eaa0755cc89f19f8514b20405403ae8f79abcf1f80d588d728d038a" + }, + { + "path": "law_versions/rules/LR-INTEREST-LIMITATION-CAP/2021-07-07.json", + "sha256": "d812a1183145084176b12d7359522d557667ddfcf62557d1922b4cf27eed577f" + }, + { + "path": "law_versions/rules/LR-LABOR-DELAY-INTEREST/2025-10-23.json", + "sha256": "58e66fb2850c1b93d8ed1469f6a8f256660911a9a09c38d044a241a3485a47bf" + }, + { + "path": "law_versions/rules/LR-LABOR-DELAY-INTEREST/pre-2025-10-23.json", + "sha256": "d6c1d43cc3e37764025dc15f17b50cf7516e17e3fa2f9639e9efa6e08dd939de" + }, + { + "path": "law_versions/rules/LR-LIMITATION-COMMERCIAL-CLAIM/CommercialAct-64.json", + "sha256": "db0970f4ace14329de4da7ba1dd8f8554bbbcab0a2e2f5ab3b4829b4b3522994" + }, + { + "path": "law_versions/rules/LR-LIMITATION-GENERAL-CLAIM/CivilAct-162-1.json", + "sha256": "2db0f112943535b6b47c8ec97872a212b27f99fa6607663e8e45b329ccf4964f" + }, + { + "path": "law_versions/rules/LR-LIMITATION-LABOR-WAGE-CLAIM/LaborStandardsAct-49-current-codification.json", + "sha256": "3518b0a00aac0633508bfcf9b2bdc49c6c3b3d35b0857e800e5522f324e42cb5" + }, + { + "path": "law_versions/rules/LR-LIMITATION-PROPERTY-RIGHT/CivilAct-162-2.json", + "sha256": "0cc9f43c4bc75dd1dc390d4c86446ae16e5a6e779742238562a3cf1e353d645b" + }, + { + "path": "law_versions/rules/LR-LIMITATION-SHORT-ONE-YEAR/CivilAct-164.json", + "sha256": "418ed935dc515c3d27270b2883991c173c9a6d5a4ebcb5c46e5345b11090fdd4" + }, + { + "path": "law_versions/rules/LR-LIMITATION-SHORT-THREE-YEAR/CivilAct-163.json", + "sha256": "50ec20bb0bf38788750b9bba1353349583f67dd71595f42764701d88da893000" + }, + { + "path": "law_versions/rules/LR-LIMITATION-TORT-ACT/CivilAct-766-2.json", + "sha256": "bd4c050039a94f623987800427b8e8c82d41fd59142c2ce0e993d90d8f6b3aab" + }, + { + "path": "law_versions/rules/LR-LIMITATION-TORT-KNOWLEDGE/CivilAct-766-1.json", + "sha256": "f272737f2f703ac7d21801d8a3b25a29e098ebdb2b3f98dd3852b539c55c2190" + }, + { + "path": "law_versions/source_status/source_occurrence_inventory.json", + "sha256": "2d2a2647dd0964eebfd9beadec753821e6ec158af8ab43b0561ae0ed88a2cb43" + }, + { + "path": "law_versions/source_status/source_repin_status.json", + "sha256": "f44e21988306748bfdde5d629d1e955f592d669c2af280eefb88f01c9ae74228" + }, + { + "path": "platform/contracts/determinism_scope_contract.json", + "sha256": "8734102d9e2e9b695a3b12e91e7ab1b7208e1c1403729c082d708ee2c4a70532" + }, + { + "path": "platform/reference_catalogs/calculation_ids.json", + "sha256": "fb14965b8315e01d47d694a10a674eda5db606ab9c7bbb6345de98495ba233b8" + }, + { + "path": "platform/reference_catalogs/domain_ids.json", + "sha256": "4740d98e6a31bd0e7d9f792189712a8af1d0a87ec0bbd181deba1e34edb08355" + }, + { + "path": "platform/reference_catalogs/profile_ids.json", + "sha256": "84f3538229df73e2d7ca58eb32da3d9b38ab1037edb67d4ac43e31e51c20aea3" + }, + { + "path": "platform/router/s3_sg01_router_config.json", + "sha256": "4913b119f5ee08f96b52e10bd1f8376326564f1ab4742909fd2d0f07ded4efc6" + }, + { + "path": "platform/router/s3_sg01_router_input.schema.json", + "sha256": "9895accee6271e7fb4637ee303be858cea4af7b461015db2d7f1c47dc019c7e2" + }, + { + "path": "platform/router/s3_sg01_router_output.schema.json", + "sha256": "4af85d52a60254149a3506cd5aa36cce1360e275b3c195d08307876379114b51" + }, + { + "path": "platform/schemas/client_goal_domain_profiles.schema.json", + "sha256": "ae2bfe0d754a09cbae16b2c15bf1518fc23f9e1bda8fa1f5f949606c8e42c010" + }, + { + "path": "platform/schemas/domain_fanout_plan.schema.json", + "sha256": "3b0948613a5996028b9c030a99f0b51d682f6035e019557756b1a15d43971113" + }, + { + "path": "platform/schemas/domain_seed_output.schema.v3.json", + "sha256": "992acf05dbccb34c65ead4e8c592f424e3b91672dc109cbd1bfa76a0a71a13c9" + }, + { + "path": "platform/schemas/domain_slice.schema.v2.json", + "sha256": "212a405088e7cf7ba2c65528a1c716938c946df7fe3bae3256b613051ed31aa3" + }, + { + "path": "platform/schemas/legal_effect_structures.schema.json", + "sha256": "e2879c7a6d7f163c0729e8b863c5ef79b1094482eace78e6b8e1c7d69ce201e3" + }, + { + "path": "platform/schemas/structure_seed_bundle.schema.json", + "sha256": "b7af9e422b6ac3876cffea39ec4f617eea76a631a57dfdfcd57d3785a83c667a" + }, + { + "path": "routing/activation_cue_digest.md", + "sha256": "7e98885a50e1880b5ee7e2f459ab47010f90387e5c463a484a152054c2321e97" + }, + { + "path": "routing/evidence_component_union.md", + "sha256": "ecd6e1877f476139012b79fdc04380c067b8e81af86707bcc5c2b311b3fabf79" + }, + { + "path": "routing/extension_payload_key_declarations.v1.json", + "sha256": "1684e13896bc9dfd0e7aa1a3b01354f58b2fa6ed00908bec6f861bce4b47235c" + }, + { + "path": "signals/_common/evidence_slot_status.schema.json", + "sha256": "292b03960b187cef668b8635a8d7539fde7c31f0c20d01af52c4f6ff8519d7b1" + }, + { + "path": "signals/_common/signal_item.schema.json", + "sha256": "de8695f98041c06cf50c0d8d2ebc31e7b3c518ca9d39a27da940438704c58bb1" + }, + { + "path": "signals/adapters/s3_domain_seed_adapter.py", + "sha256": "0aa58a1dd69bf1653d444477176303ff75ab84052a3162089160669b59d5373f" + }, + { + "path": "signals/adapters/s3_domain_seed_adapter.txt", + "sha256": "0aa58a1dd69bf1653d444477176303ff75ab84052a3162089160669b59d5373f" + }, + { + "path": "signals/adapters/s3_envelope_migration_adapter.py", + "sha256": "c87262a3a04b6da08e035bfb25abdc66b497381f67f12ca4c82494dde36700d6" + }, + { + "path": "signals/adapters/s3_envelope_migration_adapter.txt", + "sha256": "c87262a3a04b6da08e035bfb25abdc66b497381f67f12ca4c82494dde36700d6" + }, + { + "path": "signals/adapters/s4_calculation_adapter.py", + "sha256": "5fbec03c551cababd019f4022544553ee91f29fb1ba39f2113fad46af6c9315b" + }, + { + "path": "signals/adapters/s4_calculation_adapter.txt", + "sha256": "5fbec03c551cababd019f4022544553ee91f29fb1ba39f2113fad46af6c9315b" + }, + { + "path": "signals/adapters/sg01_activation_adapter.py", + "sha256": "b194eda84a781dda5c95999349b73cfd34a92573d3abae9aac655408663dd5ca" + }, + { + "path": "signals/adapters/sg01_activation_adapter.txt", + "sha256": "b194eda84a781dda5c95999349b73cfd34a92573d3abae9aac655408663dd5ca" + }, + { + "path": "signals/compiler/common.py", + "sha256": "3f5f0d133aac8f2086cd9beda4dee62309c57322d5005ba263880d0048c62ab5" + }, + { + "path": "signals/compiler/common.txt", + "sha256": "3f5f0d133aac8f2086cd9beda4dee62309c57322d5005ba263880d0048c62ab5" + }, + { + "path": "signals/compiler/projections.py", + "sha256": "b65c12ce7cb57dc77cfc5bac9ccd7fa55093f14a9432d309b1603da96fa259a7" + }, + { + "path": "signals/compiler/projections.txt", + "sha256": "b65c12ce7cb57dc77cfc5bac9ccd7fa55093f14a9432d309b1603da96fa259a7" + }, + { + "path": "signals/compiler/schema_validator.py", + "sha256": "374bb3eeb2b58aa51f2816a7950c846f658cc704b15fb9bdd260cad48eb1515b" + }, + { + "path": "signals/compiler/schema_validator.txt", + "sha256": "374bb3eeb2b58aa51f2816a7950c846f658cc704b15fb9bdd260cad48eb1515b" + }, + { + "path": "signals/compiler/signal_compiler.py", + "sha256": "b8d4026dcc302321f6580fe1f96d168722635203a8e85d10d1e5b2d5f2561222" + }, + { + "path": "signals/compiler/signal_compiler.txt", + "sha256": "b8d4026dcc302321f6580fe1f96d168722635203a8e85d10d1e5b2d5f2561222" + }, + { + "path": "signals/compiler/signal_gate.py", + "sha256": "8b16197c8a6fbc2806a31de009f0b17cf0fa0a20d745b49fbe21b88ff768a37c" + }, + { + "path": "signals/compiler/signal_gate.txt", + "sha256": "8b16197c8a6fbc2806a31de009f0b17cf0fa0a20d745b49fbe21b88ff768a37c" + }, + { + "path": "signals/compiler/transaction_writer.py", + "sha256": "33a1f61d23f76a95204b1c2ff324d39dd04fd88f3a80e5e82764c21dab6c52ac" + }, + { + "path": "signals/compiler/transaction_writer.txt", + "sha256": "33a1f61d23f76a95204b1c2ff324d39dd04fd88f3a80e5e82764c21dab6c52ac" + }, + { + "path": "signals/compiler/writer_boundary.py", + "sha256": "901a4d4636adf56c74c3c912af7dfbddebc92f33a62376346dfa6d9fda39d51a" + }, + { + "path": "signals/compiler/writer_boundary.txt", + "sha256": "901a4d4636adf56c74c3c912af7dfbddebc92f33a62376346dfa6d9fda39d51a" + }, + { + "path": "signals/contracts/admission_runtime_contract.json", + "sha256": "471915fdced832c15aed158e60b35503dd638dd6e1ec4be98b6e68ad8d8f12fc" + }, + { + "path": "signals/emitters/emit_sg02.py", + "sha256": "d988c900453b8518cdd53c67b5acbd120ee76563db7b20f57d67eaff90514a86" + }, + { + "path": "signals/emitters/emit_sg02.txt", + "sha256": "d988c900453b8518cdd53c67b5acbd120ee76563db7b20f57d67eaff90514a86" + }, + { + "path": "signals/emitters/emit_sg03.py", + "sha256": "63b880aea4e0b13307948e692f50fa2d144ccd0862cedffc861bed5a812c5cdf" + }, + { + "path": "signals/emitters/emit_sg03.txt", + "sha256": "63b880aea4e0b13307948e692f50fa2d144ccd0862cedffc861bed5a812c5cdf" + }, + { + "path": "signals/emitters/emit_sg04.py", + "sha256": "ea8799469f61f4b5a241751c91c6d2ac967529febfba1a88b543e807ecf334ae" + }, + { + "path": "signals/emitters/emit_sg04.txt", + "sha256": "ea8799469f61f4b5a241751c91c6d2ac967529febfba1a88b543e807ecf334ae" + }, + { + "path": "signals/emitters/emit_sg05.py", + "sha256": "d653ad284b0d90e6bca4af8010fd0ad47a74b4f0c76acc529b8b57c195823d40" + }, + { + "path": "signals/emitters/emit_sg05.txt", + "sha256": "d653ad284b0d90e6bca4af8010fd0ad47a74b4f0c76acc529b8b57c195823d40" + }, + { + "path": "signals/emitters/emit_sg06.py", + "sha256": "f7b2fd78672e5d8b16ac8c1221617ebc63949fbbdda44c28546a07772775658b" + }, + { + "path": "signals/emitters/emit_sg06.txt", + "sha256": "f7b2fd78672e5d8b16ac8c1221617ebc63949fbbdda44c28546a07772775658b" + }, + { + "path": "signals/emitters/emit_sg07.py", + "sha256": "adca1822d8b2cbded2875dde2226b0e669229febb9e02e187b0362478d423cba" + }, + { + "path": "signals/emitters/emit_sg07.txt", + "sha256": "adca1822d8b2cbded2875dde2226b0e669229febb9e02e187b0362478d423cba" + }, + { + "path": "signals/emitters/emit_sg08.py", + "sha256": "04878c1881a5ee96ef493ad7bf0d46f49a2a18a2e0c00dbf293d71929c7c83d4" + }, + { + "path": "signals/emitters/emit_sg08.txt", + "sha256": "04878c1881a5ee96ef493ad7bf0d46f49a2a18a2e0c00dbf293d71929c7c83d4" + }, + { + "path": "signals/emitters/emit_sg09.py", + "sha256": "c832e0b4bef81549fda1f0c11535256e80caa8933734fbfb3cd2f3cea4397b21" + }, + { + "path": "signals/emitters/emit_sg09.txt", + "sha256": "c832e0b4bef81549fda1f0c11535256e80caa8933734fbfb3cd2f3cea4397b21" + }, + { + "path": "signals/emitters/emit_sg10.py", + "sha256": "309b608ce02c1bb1cfe5eadddfa8824fd7a711bfd3d9b34a99a5b30e203674d8" + }, + { + "path": "signals/emitters/emit_sg10.txt", + "sha256": "309b608ce02c1bb1cfe5eadddfa8824fd7a711bfd3d9b34a99a5b30e203674d8" + }, + { + "path": "signals/emitters/emit_sg11.py", + "sha256": "ba5611d76827b2fde9b35458b675a0e364057dce5dedd697218032b6039d15d0" + }, + { + "path": "signals/emitters/emit_sg11.txt", + "sha256": "ba5611d76827b2fde9b35458b675a0e364057dce5dedd697218032b6039d15d0" + }, + { + "path": "signals/emitters/emit_sg12.py", + "sha256": "1ca89e3c375f045e298e2b489f1cc08aefbd8f8d592beef37da99cd841065431" + }, + { + "path": "signals/emitters/emit_sg12.txt", + "sha256": "1ca89e3c375f045e298e2b489f1cc08aefbd8f8d592beef37da99cd841065431" + }, + { + "path": "signals/emitters/emit_sg13.py", + "sha256": "2392ab26906c42f741402a9b867a2d57ad7af04fbe85700a65aeb08331095b04" + }, + { + "path": "signals/emitters/emit_sg13.txt", + "sha256": "2392ab26906c42f741402a9b867a2d57ad7af04fbe85700a65aeb08331095b04" + }, + { + "path": "signals/emitters/emitter_runtime.py", + "sha256": "160d717ef77fedfa76ab9b309894f6a711af320387a5cfe619ba166d8eb6e279" + }, + { + "path": "signals/emitters/emitter_runtime.txt", + "sha256": "160d717ef77fedfa76ab9b309894f6a711af320387a5cfe619ba166d8eb6e279" + }, + { + "path": "signals/schemas/asset_right_state_signals.schema.json", + "sha256": "fc34fbb3d33a284c3d57f3c278cbda8b3555ef26ee2f06b803fd2410ebce38b6" + }, + { + "path": "signals/schemas/calculation_requirements.schema.json", + "sha256": "7fdb5ef0f50d7af22ac417abc4022cd238f5ab0dc942866420616729a9e3571f" + }, + { + "path": "signals/schemas/defense_exception_signals.schema.json", + "sha256": "010148c15e60e4d112b142f80b1723c06e34ba22b3edefae9e4371f2353b053e" + }, + { + "path": "signals/schemas/domain_activation_manifest.schema.json", + "sha256": "013a6ebd230ebe46dda665af9f6c4448b267444b44e7b8f701f2fae80a2ee92a" + }, + { + "path": "signals/schemas/domain_signal_envelope.schema.v2.json", + "sha256": "1483d6c5f98083f59172feff9b7c15b44d3ed789db5b6172d0de05f06e9d3fbc" + }, + { + "path": "signals/schemas/evidence_proof_conflict_signals.schema.json", + "sha256": "c419f568e28c06c629bc715aff7b0737b77e9c4871c91d4fae8f6ecf04196390" + }, + { + "path": "signals/schemas/governing_law_version_signals.schema.json", + "sha256": "13a3f62f03356090d2cb24de2da0ba217928dfe8eb3c111d0f5e87c7df3119ee" + }, + { + "path": "signals/schemas/legal_effect_routes.schema.json", + "sha256": "c24cb740c370aa2477199a0225be8291164ef5c787601fd962a370c642cc3cc0" + }, + { + "path": "signals/schemas/legal_relation_lifecycle_signals.schema.json", + "sha256": "420613a5900c4360487b89b978efedde58f5ddc61644130e4b9e63ef8ab33d8b" + }, + { + "path": "signals/schemas/liability_causation_damage_signals.schema.json", + "sha256": "34102cb8eeda80773eb62a5ee61e3d714bf90424ed5350dcac4b7bf873a72c5a" + }, + { + "path": "signals/schemas/party_capacity_standing_signals.schema.json", + "sha256": "66de89ac53964166f6caabd50cbc03eb82dede0acf702d5e6d825c1d82ef81d0" + }, + { + "path": "signals/schemas/procedural_posture_relief_signals.schema.json", + "sha256": "fefb4317ad63088919b61777c71fe75ee6aa507b9f599dcf455d2af63dfc5e0d" + }, + { + "path": "signals/schemas/remedy_enforcement_signals.schema.json", + "sha256": "999e1969b983748f209e9b5239f7edd0ec43bc642d9ea8fd7edbf34f9ce653f3" + }, + { + "path": "signals/schemas/signal_manifest.schema.json", + "sha256": "5e72084780b82b29582c9ffcf48f3e4894d7c0b152e5ce8df394583c07dde681" + }, + { + "path": "signals/schemas/timeline_notice_condition_signals.schema.json", + "sha256": "99c66208524155cea6bbd5e24fd26998cc9b653c89b24b569c793e36f1623d35" + }, + { + "path": "signals/signal_registry.v2.json", + "sha256": "4392b40da458102f8dd11b40b40ae3f694b7b5911849b050e2e4118c569e5ab0" + }, + { + "path": "stage1_runtime/_asset_bundle.txt", + "sha256": "8237a5d2e19a3eab4b25d09e11142b5fdf41ee99949b2482eac6c92e34d7bb8d" + }, + { + "path": "stage1_runtime/activation_gate.py", + "sha256": "62c54e4165e876e688bbbdbfb3b089fcbdb3ce0ee0973af6b82b9cd9ce5d1a5b" + }, + { + "path": "stage1_runtime/activation_gate.txt", + "sha256": "62c54e4165e876e688bbbdbfb3b089fcbdb3ce0ee0973af6b82b9cd9ce5d1a5b" + }, + { + "path": "stage1_runtime/bo_surface_projection_policy.v1.json", + "sha256": "b3eba1d868ff543d761993396ada4910bd032a6762a8853403426655a3e9505a" + }, + { + "path": "stage1_runtime/domain_fanout_planner.py", + "sha256": "0ff8ec40b81cc7afb257195562a13a7020b8e9e043f4f95620f159ff611265d9" + }, + { + "path": "stage1_runtime/domain_fanout_planner.txt", + "sha256": "0ff8ec40b81cc7afb257195562a13a7020b8e9e043f4f95620f159ff611265d9" + }, + { + "path": "stage1_runtime/domain_slice_compiler.py", + "sha256": "c865eb281e201c5400d7b0d922b8d70f1df1d7de2001e66e3274fc5550a72526" + }, + { + "path": "stage1_runtime/domain_slice_compiler.txt", + "sha256": "c865eb281e201c5400d7b0d922b8d70f1df1d7de2001e66e3274fc5550a72526" + }, + { + "path": "stage1_runtime/prompt_compiler.py", + "sha256": "d2f0256444f581d41894c23b53f83e5bcc4c47f67e90e48464aa2eeace719f9a" + }, + { + "path": "stage1_runtime/prompt_compiler.txt", + "sha256": "d2f0256444f581d41894c23b53f83e5bcc4c47f67e90e48464aa2eeace719f9a" + }, + { + "path": "stage1_runtime/prompt_composition_policy.json", + "sha256": "8d3f003ecc5e53eb7112115126e325b77292a27b0a542a10c8cecfe6c6bbfdea" + }, + { + "path": "stage1_runtime/registry_loader.py", + "sha256": "b90e3a1ad6c26a1dca4d49635b8a97c8153d36147f2f1b1672429b2f7a0f02dc" + }, + { + "path": "stage1_runtime/registry_loader.txt", + "sha256": "b90e3a1ad6c26a1dca4d49635b8a97c8153d36147f2f1b1672429b2f7a0f02dc" + }, + { + "path": "stage1_runtime/registry_validator.py", + "sha256": "a8eff67f69433430efae9fa5292cba597e52dbb2e67d156ca18cb7055886538e" + }, + { + "path": "stage1_runtime/registry_validator.txt", + "sha256": "a8eff67f69433430efae9fa5292cba597e52dbb2e67d156ca18cb7055886538e" + }, + { + "path": "stage1_runtime/runtime_common.py", + "sha256": "5a76c488f15e54d7d143f2e18bbd3f104b3b32e5c13a84f0a98e276049bb9145" + }, + { + "path": "stage1_runtime/runtime_common.txt", + "sha256": "5a76c488f15e54d7d143f2e18bbd3f104b3b32e5c13a84f0a98e276049bb9145" + }, + { + "path": "stage1_runtime/schema_subset_validator.py", + "sha256": "18d470ff8dc775ce4b7da490e1f503913b821f7d95ad3d22ad0b8061d18aa377" + }, + { + "path": "stage1_runtime/schema_subset_validator.txt", + "sha256": "18d470ff8dc775ce4b7da490e1f503913b821f7d95ad3d22ad0b8061d18aa377" + }, + { + "path": "stage1_runtime/stage_a_context_builder.py", + "sha256": "68d984f9d565034c73721dc97c5aab0f9cdb8a8b46e1bf13eabd8c136ec4bcc0" + }, + { + "path": "stage1_runtime/stage_a_context_builder.txt", + "sha256": "68d984f9d565034c73721dc97c5aab0f9cdb8a8b46e1bf13eabd8c136ec4bcc0" + }, + { + "path": "stage1_runtime/structure_index_compiler.py", + "sha256": "c400adfb7c58e340a17e67aebcdc4a7cf92c35c422f418451c79ce7c531d201f" + }, + { + "path": "stage1_runtime/structure_index_compiler.txt", + "sha256": "c400adfb7c58e340a17e67aebcdc4a7cf92c35c422f418451c79ce7c531d201f" + }, + { + "path": "stage1_runtime/structure_index_policy.v1.json", + "sha256": "823a7e734a7bc9433966a5629a5a0b153a067c92d8570072fa14a5c42bb722bd" + }, + { + "path": "stage1_runtime/worker_output_validator.py", + "sha256": "e5b46f929082faf621731ffb8f4864ccb9bba461ab23caa56d1bcfe5b073fd9d" + }, + { + "path": "stage1_runtime/worker_output_validator.txt", + "sha256": "e5b46f929082faf621731ffb8f4864ccb9bba461ab23caa56d1bcfe5b073fd9d" + }, + { + "path": "tools/s3_router_v2.py", + "sha256": "bb3b163f8489e387befa1c4553ed125a12ea47b0f3a83013a5d2689daf32efd9" + }, + { + "path": "tools/s3_router_v2.txt", + "sha256": "bb3b163f8489e387befa1c4553ed125a12ea47b0f3a83013a5d2689daf32efd9" + }, + { + "path": "tools/s6_sg01_router.py", + "sha256": "7d958d73a1f04edc2ab3d89aa4c1b000d2bc8bb02b3fc22694b58a3021cd8be0" + }, + { + "path": "tools/s6_sg01_router.txt", + "sha256": "7d958d73a1f04edc2ab3d89aa4c1b000d2bc8bb02b3fc22694b58a3021cd8be0" + } + ], + "program_release_status": "STAGE1_NOT_RELEASE_READY", + "runtime_artifact_count": 433, + "schema_version": "stage1_runtime_manifest.v1" +} diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8.pre_special_law_wiring.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8.pre_special_law_wiring.yml new file mode 100644 index 00000000..257a6fb8 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8.pre_special_law_wiring.yml @@ -0,0 +1,3315 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, events_document, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-flash-lite' + llm_reasoning: high + llm_verbosity: medium + cache_control: + mode: auto + ttl: 15m + prompt: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "review_items": [], + "extensions": {} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + if worker_validator is not None: + report = worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, + schema=seed_schema, + expected_domain_id=domain_id, + slice_document=slice_doc) + for item in report.get("errors") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, + "detail": item}) + for item in report.get("warnings") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + for idx, cand in enumerate(cands): + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_21.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_21.yml new file mode 100644 index 00000000..154d6d27 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_21.yml @@ -0,0 +1,3349 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + profile_paths=profile_paths, extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, events_document, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-flash-lite' + llm_reasoning: high + llm_verbosity: medium + cache_control: + mode: auto + ttl: 15m + prompt: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "review_items": [], + "extensions": {} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + if worker_validator is not None: + report = worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, + schema=seed_schema, + expected_domain_id=domain_id, + slice_document=slice_doc) + for item in report.get("errors") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, + "detail": item}) + for item in report.get("warnings") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + for idx, cand in enumerate(cands): + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_21_11pm.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_21_11pm.yml new file mode 100644 index 00000000..1bdbd833 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_21_11pm.yml @@ -0,0 +1,3353 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + profile_paths=profile_paths, extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, events_document, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + # 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은 + # 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다 + # (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다. + "dynamic_fanout": plan_root.get("task_instances") or [], + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-flash-lite' + llm_reasoning: high + llm_verbosity: medium + cache_control: + mode: auto + ttl: 15m + prompt: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "review_items": [], + "extensions": {} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + if worker_validator is not None: + report = worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, + schema=seed_schema, + expected_domain_id=domain_id, + slice_document=slice_doc) + for item in report.get("errors") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, + "detail": item}) + for item in report.get("warnings") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + for idx, cand in enumerate(cands): + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_12am.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_12am.yml new file mode 100644 index 00000000..f4b005fa --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_12am.yml @@ -0,0 +1,3357 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + profile_paths=profile_paths, extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, events_document, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + # 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은 + # 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다 + # (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다. + "dynamic_fanout": plan_root.get("task_instances") or [], + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-flash-lite' + llm_reasoning: high + llm_verbosity: medium + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 15m + prompts: + - role: user + content: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "review_items": [], + "extensions": {} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + if worker_validator is not None: + report = worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, + schema=seed_schema, + expected_domain_id=domain_id, + slice_document=slice_doc) + for item in report.get("errors") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, + "detail": item}) + for item in report.get("warnings") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + for idx, cand in enumerate(cands): + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_15_35am.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_15_35am.yml new file mode 100644 index 00000000..eebdcfba --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_15_35am.yml @@ -0,0 +1,3418 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + profile_paths=profile_paths, extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + # P-2a — stage_a_context 를 slice 컴파일보다 먼저 만든다. slice 의 source_universe 가 + # 인용할 event 식별자의 정본이 event_candidate_map 이기 때문이다. 순서가 뒤였을 때 + # A0 는 원시 evidence_event_candidates 문서를 넘겼고, domain_slice_compiler._records 가 + # 그 문서-단위 items(30건)를 후보로 오인해 candidate_id 를 못 찾아 + # event:unidentified:NNNN 로 대체했다. 그 값은 source_universe_manifest 의 + # event_candidate_ids(EVT-...-NN, 84건)와 교집합이 0 이라, 워커가 규율을 지켜 + # slice 안의 것만 인용해도 R0 가 "source_refs outside Stage A universe" 로 차단했다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + # 맵을 통째로 넘기지 않는다 — domain_slice_compiler._records 는 dict 를 받으면 + # by_evidence_index(값이 id 리스트)에 먼저 걸려 빈 목록을 돌려준다. 후보 레코드 + # 목록으로 평탄화해 넘겨야 _record_id 가 candidate_id 를 찾아 EVT-...-NN 을 쓴다. + event_candidate_records = [ + row for row in ((stage_a.get("event_candidate_map") or {}).get("by_candidate_id") or {}).values() + if isinstance(row, dict)] + if not event_candidate_records: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_EVENT_CANDIDATE_MAP_EMPTY", + "message": "stage_a event_candidate_map.by_candidate_id 가 비었다. slice 의 event 식별자를 정본으로 실을 수 없다.", + }, ensure_ascii=False)) + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, event_candidate_records, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + # 계산은 P-2a 에서 이미 끝났다(slice 가 같은 식별자를 써야 하므로 앞당겼다). 여기서는 기록만 한다. + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + # 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은 + # 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다 + # (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다. + "dynamic_fanout": plan_root.get("task_instances") or [], + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-pro-preview' + llm_reasoning: high + llm_verbosity: low + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 15m + prompts: + - role: user + content: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "juristic_act_type": "<법률행위 유형 문자열 또는 null>", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "party_roles": [], + "time_facts": [], + "object_refs": [], + "amount_facts": [], + "review_items": [], + "extensions": {"domain_payload": {"action_summary": null, "action_type": null}} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + - 아래 자리들은 BO 호환면 투영(`bo_surface_projection_policy.v1`)이 읽는 1순위 출처다. + 비워 두면 BO.json 의 해당 칸이 폴백 값으로 채워지고 schema_field_fallback 검토가 발행된다. + slice 의 source_universe 안에 근거가 있으면 채운다. 근거가 없으면 비워 두고 사유를 남긴다 — + 추측으로 채우지 않는다. 사건종류 이름을 값으로 쓰지 않는다. + · `juristic_act_type` : 법률행위 유형 문자열 1개(없으면 null). -> JuristicAct.label + · `extensions.domain_payload.action_summary` : 이 후보가 무엇인지 한 문장. -> Action + · `extensions.domain_payload.action_type` : "법률행위(legal acts)" 또는 "사실행위(factual acts)". -> ActionType + · `legal_effect_candidates[]` : {"type_id": "<소문자_스네이크>", "source_refs": [], "registered": true|false}. -> Legal_Keywords + · `time_facts[]` : {"fact_type": "<소문자_스네이크>", "value": "<시점 문자열 또는 null>", "source_refs": []}. -> BehaviorTime · TimeText + · `object_refs[]` : 목적물 식별자 문자열. -> core_field_base.Object + · `amount_facts[]` : {"amount_type": "<소문자_스네이크>", "decimal_value": "<숫자 문자열 또는 null>", "currency": "KRW", "source_refs": []}. -> amount + · `party_roles[]` : {"role": "<소문자_스네이크>", "party_refs": []}. 투영 대상은 아니나 스키마 필드다. + - 아래 다섯 어휘는 스키마가 고정한 것이다. 다른 낱말을 쓰면 S0 신호 게이트가 경성으로 막는다. + R0 의 검증기는 스키마의 부분집합만 보므로 여기서 틀려도 그 단계에서는 걸리지 않는다. + · seed 최상위 `status` : READY | READY_WITH_REVIEW | NO_SUPPORT | BLOCKED | FAILED + · `evidence_slot_status[].status` : filled | partial | missing | conflicted + (요건 슬롯을 뒷받침하는 근거가 충분하면 filled, 일부만이면 partial, + 없으면 missing, 상충 근거가 함께 있으면 conflicted) + · `calculation_requests[].completeness` : ready | partial | blocked | deferred + · `review_items[].severity` 와 `unknown_or_unrouted_reviews[].severity` : info | review | hard_warning | block + · 세 후보 배열의 `source_kind` : meeting_clause | event_candidate | evidence | bo | fact + | signal | registry | law_version | calculation | other + - 아래 네 객체는 `additionalProperties: false` 다. 적힌 키 말고는 **한 개도** 넣지 않는다. + 필수 키를 빠뜨리거나 임의 키를 더하면 스키마 위반이다. + · `element_fact_candidates[]` · `opposing_fact_candidates[]` · `defense_candidates[]` : + {"source_id": "", "source_kind": "<위 어휘>", + "excerpt": "<선택: 근거 문구>", "payload": {}} + — 필수는 source_id · source_kind 둘이다. `slot_id` 나 `fact` 같은 키는 이 배열에 없다. + 슬롯 판정은 `evidence_slot_status[]` 가 맡는다. + · `evidence_slot_status[]` : + {"slot_id": "<요건 슬롯 id>", "status": "<위 어휘>", "source_refs": [], "review_code": null} + · `calculation_requests[]` : + {"calculation_domain": "", "completeness": "<위 어휘>", + "source_refs": [], "review_code": null} + · `review_items[]` : + {"review_code": "<대문자_스네이크>", "severity": "<위 어휘>", + "reason": "<왜 검토가 필요한지 한 문장>", "source_refs": []} + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + if worker_validator is not None: + report = worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, + schema=seed_schema, + expected_domain_id=domain_id, + slice_document=slice_doc) + for item in report.get("errors") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, + "detail": item}) + for item in report.get("warnings") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + for idx, cand in enumerate(cands): + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_1am.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_1am.yml new file mode 100644 index 00000000..8d885a9d --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_1am.yml @@ -0,0 +1,3374 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + profile_paths=profile_paths, extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, events_document, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + # 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은 + # 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다 + # (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다. + "dynamic_fanout": plan_root.get("task_instances") or [], + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-pro-preview' + llm_reasoning: high + llm_verbosity: low + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 15m + prompts: + - role: user + content: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "juristic_act_type": "<법률행위 유형 문자열 또는 null>", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "party_roles": [], + "time_facts": [], + "object_refs": [], + "amount_facts": [], + "review_items": [], + "extensions": {"domain_payload": {"action_summary": null, "action_type": null}} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + - 아래 자리들은 BO 호환면 투영(`bo_surface_projection_policy.v1`)이 읽는 1순위 출처다. + 비워 두면 BO.json 의 해당 칸이 폴백 값으로 채워지고 schema_field_fallback 검토가 발행된다. + slice 의 source_universe 안에 근거가 있으면 채운다. 근거가 없으면 비워 두고 사유를 남긴다 — + 추측으로 채우지 않는다. 사건종류 이름을 값으로 쓰지 않는다. + · `juristic_act_type` : 법률행위 유형 문자열 1개(없으면 null). -> JuristicAct.label + · `extensions.domain_payload.action_summary` : 이 후보가 무엇인지 한 문장. -> Action + · `extensions.domain_payload.action_type` : "법률행위(legal acts)" 또는 "사실행위(factual acts)". -> ActionType + · `legal_effect_candidates[]` : {"type_id": "<소문자_스네이크>", "source_refs": [], "registered": true|false}. -> Legal_Keywords + · `time_facts[]` : {"fact_type": "<소문자_스네이크>", "value": "<시점 문자열 또는 null>", "source_refs": []}. -> BehaviorTime · TimeText + · `object_refs[]` : 목적물 식별자 문자열. -> core_field_base.Object + · `amount_facts[]` : {"amount_type": "<소문자_스네이크>", "decimal_value": "<숫자 문자열 또는 null>", "currency": "KRW", "source_refs": []}. -> amount + · `party_roles[]` : {"role": "<소문자_스네이크>", "party_refs": []}. 투영 대상은 아니나 스키마 필드다. + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + if worker_validator is not None: + report = worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, + schema=seed_schema, + expected_domain_id=domain_id, + slice_document=slice_doc) + for item in report.get("errors") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, + "detail": item}) + for item in report.get("warnings") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + for idx, cand in enumerate(cands): + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_21h_30m.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_21h_30m.yml new file mode 100644 index 00000000..2cf42e9c --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_21h_30m.yml @@ -0,0 +1,3974 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + "Default_Agent/stage1_runtime/seed_admission_policy.v1.json", # R0 — 시드 수용 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + profile_paths=profile_paths, extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + # P-2a — stage_a_context 를 slice 컴파일보다 먼저 만든다. slice 의 source_universe 가 + # 인용할 event 식별자의 정본이 event_candidate_map 이기 때문이다. 순서가 뒤였을 때 + # A0 는 원시 evidence_event_candidates 문서를 넘겼고, domain_slice_compiler._records 가 + # 그 문서-단위 items(30건)를 후보로 오인해 candidate_id 를 못 찾아 + # event:unidentified:NNNN 로 대체했다. 그 값은 source_universe_manifest 의 + # event_candidate_ids(EVT-...-NN, 84건)와 교집합이 0 이라, 워커가 규율을 지켜 + # slice 안의 것만 인용해도 R0 가 "source_refs outside Stage A universe" 로 차단했다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + # 맵을 통째로 넘기지 않는다 — domain_slice_compiler._records 는 dict 를 받으면 + # by_evidence_index(값이 id 리스트)에 먼저 걸려 빈 목록을 돌려준다. 후보 레코드 + # 목록으로 평탄화해 넘겨야 _record_id 가 candidate_id 를 찾아 EVT-...-NN 을 쓴다. + event_candidate_records = [ + row for row in ((stage_a.get("event_candidate_map") or {}).get("by_candidate_id") or {}).values() + if isinstance(row, dict)] + if not event_candidate_records: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_EVENT_CANDIDATE_MAP_EMPTY", + "message": "stage_a event_candidate_map.by_candidate_id 가 비었다. slice 의 event 식별자를 정본으로 실을 수 없다.", + }, ensure_ascii=False)) + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, event_candidate_records, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + # 계산은 P-2a 에서 이미 끝났다(slice 가 같은 식별자를 써야 하므로 앞당겼다). 여기서는 기록만 한다. + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + # 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은 + # 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다 + # (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다. + "dynamic_fanout": plan_root.get("task_instances") or [], + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-pro-preview' + llm_reasoning: high + llm_verbosity: low + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 15m + prompts: + - role: user + content: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "juristic_act_type": "<법률행위 유형 문자열 또는 null>", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "party_roles": [], + "time_facts": [], + "object_refs": [], + "amount_facts": [], + "review_items": [], + "extensions": {"domain_payload": {"action_summary": null, "action_type": null}} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + - 아래 자리들은 BO 호환면 투영(`bo_surface_projection_policy.v1`)이 읽는 1순위 출처다. + 비워 두면 BO.json 의 해당 칸이 폴백 값으로 채워지고 schema_field_fallback 검토가 발행된다. + slice 의 source_universe 안에 근거가 있으면 채운다. 근거가 없으면 비워 두고 사유를 남긴다 — + 추측으로 채우지 않는다. 사건종류 이름을 값으로 쓰지 않는다. + · `juristic_act_type` : 법률행위 유형 문자열 1개(없으면 null). -> JuristicAct.label + · `extensions.domain_payload.action_summary` : 이 후보가 무엇인지 한 문장. -> Action + · `extensions.domain_payload.action_type` : "법률행위(legal acts)" 또는 "사실행위(factual acts)". -> ActionType + · `legal_effect_candidates[]` : {"type_id": "<소문자_스네이크>", "source_refs": [], "registered": true|false}. -> Legal_Keywords + · `time_facts[]` : {"fact_type": "<소문자_스네이크>", "value": "<시점 문자열 또는 null>", "source_refs": []}. -> BehaviorTime · TimeText + · `object_refs[]` : 목적물 식별자 문자열. -> core_field_base.Object + · `amount_facts[]` : {"amount_type": "<소문자_스네이크>", "decimal_value": "<숫자 문자열 또는 null>", "currency": "KRW", "source_refs": []}. -> amount + · `party_roles[]` : {"role": "<소문자_스네이크>", "party_refs": []}. 투영 대상은 아니나 스키마 필드다. + - 아래 다섯 어휘는 스키마가 고정한 것이다. 다른 낱말을 쓰면 S0 신호 게이트가 경성으로 막는다. + R0 의 검증기는 스키마의 부분집합만 보므로 여기서 틀려도 그 단계에서는 걸리지 않는다. + · seed 최상위 `status` : READY | READY_WITH_REVIEW | NO_SUPPORT | BLOCKED | FAILED + · `evidence_slot_status[].status` : filled | partial | missing | conflicted + (요건 슬롯을 뒷받침하는 근거가 충분하면 filled, 일부만이면 partial, + 없으면 missing, 상충 근거가 함께 있으면 conflicted) + · `calculation_requests[].completeness` : ready | partial | blocked | deferred + · `review_items[].severity` 와 `unknown_or_unrouted_reviews[].severity` : info | review | hard_warning | block + · 세 후보 배열의 `source_kind` : meeting_clause | event_candidate | evidence | bo | fact + | signal | registry | law_version | calculation | other + - 아래 네 객체는 `additionalProperties: false` 다. 적힌 키 말고는 **한 개도** 넣지 않는다. + 필수 키를 빠뜨리거나 임의 키를 더하면 스키마 위반이다. + · `element_fact_candidates[]` · `opposing_fact_candidates[]` · `defense_candidates[]` : + {"source_id": "", "source_kind": "<위 어휘>", + "excerpt": "<선택: 근거 문구>", "payload": {}} + — 필수는 source_id · source_kind 둘이다. `slot_id` 나 `fact` 같은 키는 이 배열에 없다. + 슬롯 판정은 `evidence_slot_status[]` 가 맡는다. + · `evidence_slot_status[]` : + {"slot_id": "<요건 슬롯 id>", "status": "<위 어휘>", "source_refs": [], "review_code": null} + · `calculation_requests[]` : + {"calculation_domain": "", "completeness": "<위 어휘>", + "source_refs": [], "review_code": null} + · `review_items[]` : + {"review_code": "<대문자_스네이크>", "severity": "<위 어휘>", + "reason": "<왜 검토가 필요한지 한 문장>", "source_refs": []} + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import copy + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + # R0-6 — 전체 스키마 검증은 신설이 아니다. worker_output_validator 가 이미 + # schema_subset_validator 로 전 계약을 검사하고 있었고(13도메인 1,151건 실측), + # R0 가 그 결과를 guard_warnings 로 강등하고 있었다. 이 정책은 그 결과에 처분을 준다. + SEED_ADMISSION_POLICY_PATH = "Default_Agent/stage1_runtime/seed_admission_policy.v1.json" + SEED_ADMISSION_SCHEMA_VERSION = "stage1_seed_admission_policy.v1" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if not isinstance(bo_type, str): + # 워커가 문자열이 아닌 값을 넣으면 집합 비교가 터진다(dict 면 unhashable). + # 수용 단계가 걸러야 하지만 투영은 마지막 방어선이라 여기서도 막는다 — + # 여기서 죽으면 회차 전체가 죽고, 원인이 워커 값이라는 것도 드러나지 않는다. + note("schema_field_fallback", "BOType", "non_string:%s" % type(bo_type).__name__) + bo_type = norm.get("bo_type_default") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ------------------------------------------------------------------ + # R0-6 — 시드 수용(admission). 검증 결과에 등급을 준다. + # 구조 위반은 막고(block), 어휘·형태 위반은 정규화·수확으로 살리고(repair), + # 살릴 수 없는 것은 후보 하나만 격리한다(quarantine). 회차 전체를 죽이지 않는다. + # 값을 지어내지 않는다 — 표에 없는 낱말과 universe 에 없는 참조는 복구 대상이 아니다. + # ------------------------------------------------------------------ + PATH_INDEX_RE = re.compile(r"\[(\d+)\]") + + def _norm_path(path: str) -> str: + """검증기 경로를 정책 키 모양으로 바꾼다. 인덱스는 [] 로 접는다.""" + body = path.replace("$.stage_b_domain_bo_seed_output.", "").replace("$.stage_b_domain_bo_seed_output", "") + return PATH_INDEX_RE.sub("[]", body).lstrip(".") + + def _path_tokens(path: str) -> list[Any]: + body = path.replace("$.stage_b_domain_bo_seed_output.", "").replace("$.stage_b_domain_bo_seed_output", "") + tokens: list[Any] = [] + for part in body.lstrip(".").split("."): + if not part: + continue + name = part.split("[", 1)[0] + if name: + tokens.append(name) + for hit in PATH_INDEX_RE.findall(part): + tokens.append(int(hit)) + return tokens + + def _resolve_parent(root: Any, tokens: list[Any]) -> tuple[Any, Any]: + """마지막 토큰의 부모 컨테이너와 그 토큰을 돌려준다. 못 찾으면 (None, None).""" + node = root + for tok in tokens[:-1]: + if isinstance(tok, int): + if not isinstance(node, list) or tok >= len(node): + return (None, None) + node = node[tok] + else: + if not isinstance(node, dict) or tok not in node: + return (None, None) + node = node[tok] + return (node, tokens[-1]) if tokens else (None, None) + + def _candidate_index(tokens: list[Any]) -> int | None: + for pos, tok in enumerate(tokens): + if tok == "bo_seed_candidates" and pos + 1 < len(tokens) and isinstance(tokens[pos + 1], int): + return tokens[pos + 1] + return None + + def _sink_for(seed_obj: dict[str, Any], tokens: list[Any], policy: dict[str, Any], norm: str): + """초과 정보를 담을 자리 두 곳을 돌려준다: (컨테이너 지역 payload, 후보 수준 보관 목록). + + 지역 payload 의 키는 **잎 이름 그대로** 쓴다 — 하류 어댑터가 + payload["status"] 처럼 짧은 이름으로 읽기 때문이다. 정규경로를 키로 쓰면 + 바이트는 남아도 아무도 읽지 못한다. + 후보 수준 보관은 dict 가 아니라 **덧붙이기 전용 목록**이다. 같은 정규경로의 + 두 번째 값이 첫 번째를 조용히 덮는 사고를 구조적으로 막는다. + """ + sink_cfg = _dict(policy.get("surplus_sink")) + local: dict[str, Any] | None = None + by_container = _dict(sink_cfg.get("by_container")) + container_key = norm.rsplit(".", 1)[0] if "." in norm else norm + local_name = by_container.get(container_key) + if local_name: + parent, _ = _resolve_parent(seed_obj, tokens) + if isinstance(parent, dict): + bucket = parent.get(local_name) + if not isinstance(bucket, dict): + bucket = {} + parent[local_name] = bucket + local = bucket + shelf: list[Any] | None = None + cand_idx = _candidate_index(tokens) + cands = seed_obj.get("bo_seed_candidates") + if cand_idx is not None and isinstance(cands, list) and cand_idx < len(cands) and isinstance(cands[cand_idx], dict): + node: Any = cands[cand_idx] + steps = str(sink_cfg.get("candidate_path") or "extensions.worker_surplus").split(".") + for step in steps[:-1]: + nxt = node.get(step) + if not isinstance(nxt, dict): + nxt = {} + node[step] = nxt + node = nxt + leaf = steps[-1] + if not isinstance(node.get(leaf), list): + node[leaf] = [] + shelf = node[leaf] + return (local, shelf) + + def _universe_kind(value: str, universe: dict[str, set]) -> str | None: + if value in universe["source_evidence_indexes"]: + return "evidence" + if value in universe["source_event_candidate_ids"]: + return "event_candidate" + if value in universe["source_meeting_clause_ids"]: + return "meeting_clause" + return None + + def _derive_required(seed_obj, tokens, norm, rule, universe, ctx): + """선언된 규칙으로만 채운다. 규칙이 없거나 실패하면 (False, None).""" + parent, key = _resolve_parent(seed_obj, tokens) + if not isinstance(parent, dict): + return (False, None) + how = str(rule.get("rule") or "") + if how == "constant": + value = rule.get("value") + return (True, copy.deepcopy(value)) if "value" in rule else (False, None) + if how == "from_context": + ckey = str(rule.get("key") or "") + return (True, copy.deepcopy(ctx[ckey])) if ckey in ctx else (False, None) + if how == "first_universe_member_of_sibling_array": + known = (universe["source_evidence_indexes"] | universe["source_event_candidate_ids"] + | universe["source_meeting_clause_ids"]) + for sib in rule.get("sibling_candidates") or []: + raw = parent.get(sib) + values = raw if isinstance(raw, list) else ([raw] if isinstance(raw, str) else []) + for item in values: + if isinstance(item, str) and item in known: + return (True, item) + return (False, None) + if how == "universe_kind_of_sibling": + sib = parent.get(str(rule.get("sibling") or "")) + kind = _universe_kind(sib, universe) if isinstance(sib, str) else None + if not kind: + # 형제가 아직 채워지지 않았을 수 있다. 같은 원소의 참조 배열에서 직접 읽는다. + for name in rule.get("fallback_sibling_arrays") or []: + raw = parent.get(name) + values = raw if isinstance(raw, list) else ([raw] if isinstance(raw, str) else []) + for item in values: + if isinstance(item, str): + kind = _universe_kind(item, universe) + if kind: + break + if kind: + break + return (True, kind) if kind else (False, None) + if how == "rename_sibling": + # 워커가 같은 뜻을 다른 이름으로 적었을 때 이름만 바로잡는다. 값은 그대로 옮긴다. + # 옮긴 뒤 원래 키를 지운다 — 남겨 두면 다음 패스에서 초과 속성으로 다시 걸린다. + for sib in rule.get("sibling_candidates") or []: + if sib in parent and parent.get(sib) not in (None, ""): + return (True, parent.pop(sib)) + return (False, None) + if how == "sibling_matching_pattern": + pat = re.compile(str(rule.get("pattern") or "^$")) + for sib in rule.get("sibling_candidates") or []: + value = parent.get(sib) + if isinstance(value, str) and pat.fullmatch(value): + return (True, value) + for sib in rule.get("fallback_rename_sibling") or []: + if sib in parent and parent.get(sib) not in (None, ""): + return (True, parent.pop(sib)) + return (False, None) + return (False, None) + + def _promote_alias(parent, key, norm, policy, record): + """수확 직전에 한 번 더 본다 — 이 키가 선언된 자리의 다른 이름일 뿐인가. + + 그렇다면 자유 공간으로 밀어 넣지 않고 제 자리로 올린다. 워커가 excerpt 를 + value·content 로 부르는 표류가 실측됐고, 그 값은 하류가 실제로 읽는 칸이다. + """ + if not isinstance(parent, dict) or not isinstance(key, str): + return False + container = norm.rsplit(".", 1)[0] if "." in norm else "" + for target_path, rule in _dict(policy.get("required_derivation")).items(): + if str(rule.get("rule") or "") != "rename_sibling": + continue + if target_path.rsplit(".", 1)[0] != container: + continue + target = target_path.rsplit(".", 1)[-1] + if target in parent and parent.get(target) not in (None, ""): + continue + if key in (rule.get("sibling_candidates") or []): + parent[target] = parent.pop(key) + record["received"] = _clip(parent[target]) + record["applied"] = "promoted_to:%s" % target + record["rule_id"] = "alias_promotion:%s" % target_path + return True + return False + + def _harvest(seed_obj, tokens, policy, norm, record): + """규약 밖 값을 버리지 않고 자유 공간으로 옮긴다. 옮긴 사실을 기록한다.""" + parent, key = _resolve_parent(seed_obj, tokens) + if parent is None: + return False + if _promote_alias(parent, key, norm, policy, record): + return True + local, shelf = _sink_for(seed_obj, tokens, policy, norm) + if isinstance(parent, list) and isinstance(key, int): + if key >= len(parent): + return False + moved = parent.pop(key) + leaf = None + elif isinstance(parent, dict): + if key not in parent: + return False + moved = parent.pop(key) + leaf = key + else: + return False + placed = [] + # 1) 컨테이너 지역 payload — 하류가 읽는 짧은 이름으로. 이미 있으면 덮지 않는다. + if isinstance(local, dict) and isinstance(leaf, str) and leaf not in local: + local[leaf] = moved + placed.append("payload.%s" % leaf) + # 2) 후보 수준 보관 — 덧붙이기 전용이라 어떤 값도 덮이지 않는다. 원래 경로를 함께 남긴다. + if isinstance(shelf, list): + # D11 — norm 은 인덱스를 접으므로 같은 컨테이너의 두 값이 구별되지 않는다. + # 원본 경로를 함께 남겨야 부모 원소와의 결합(예: role ↔ party_refs)을 복원할 수 있다. + shelf.append({"path": norm, "source_path": record.get("path"), "value": moved}) + placed.append("worker_surplus[]") + if not placed: + # 보관할 자리가 없으면 지우지 않는다. 되돌려 놓고 실패로 돌려주면 + # 이 위반은 잔여로 남아 정책이 정한 처분(기본 review)으로 간다. + # 뿌리 수준 배열(unknown_or_unrouted_reviews 등)이 여기 해당한다 — + # 후보에 매이지 않아 보관처가 없는데, 그렇다고 사건 자료를 버릴 수는 없다. + if isinstance(parent, list) and isinstance(key, int): + parent.insert(key, moved) + elif isinstance(parent, dict) and isinstance(key, str): + parent[key] = moved + return False + record["received"] = _clip(moved) + record["applied"] = "+".join(placed) + # 기계적 키 이동과, 실질 서술을 담은 채 **객체째** 밀려난 것은 검토 무게가 다르다. + # 후자를 같은 등급에 섞으면 1,200건 속 몇 건을 사람이 찾아내야 한다. + # 판정은 dict 로 좁힌다 — 스칼라 한 개의 자리 이동(당사자 이름·slot_id 문자열)은 + # 원소가 사라진 것이 아니라 키가 옮겨진 것이라 무게가 다르다. + if isinstance(moved, dict): + for _k in ("excerpt", "value", "fact", "content", "reason", "statement"): + _v = moved.get(_k) + if isinstance(_v, str) and _v.strip(): + record["content_bearing"] = True + break + return True + + def _evict(seed_obj, tokens, policy, norm, record, quarantined=None): + """복구 불가한 객체 하나를 배열에서 들어내 보관한다. 후보 자체는 삭제하지 않는다. + + D4·D8 — 후보를 pop 하면 ① 예산에 잡히지 않아 소리 없이 사라지고 + ② 뒤 후보의 인덱스가 밀려 이미 기록한 격리 표시가 다른 후보를 가리킨다. + 후보 수준이면 삭제 대신 격리로 돌린다. + """ + trimmed = list(tokens) + while trimmed and not isinstance(trimmed[-1], int): + trimmed.pop() + if not trimmed: + return False + if len(trimmed) == 2 and trimmed[0] == "bo_seed_candidates": + if quarantined is None: + return False + quarantined.add(trimmed[1]) + record["applied"] = "quarantined_candidate" + return True + return _harvest(seed_obj, trimmed, policy, _norm_path("$.stage_b_domain_bo_seed_output." + norm), record) + + def _clip(value: Any, limit: int = 200) -> Any: + try: + text = json.dumps(value, ensure_ascii=False) + except Exception: + text = str(value) + return text if len(text) <= limit else text[:limit] + "…" + + def admit_seed(seed_obj, domain_id, schema, slice_doc, universe, policy, records, plan_row=None, validator=None): + """검증 -> 처분 -> 복구를 수렴할 때까지 돌리고, 남은 것은 후보 격리로 넘긴다. + + 돌려주는 것: (수용된 seed, 격리된 후보 인덱스 집합, 경성 중단 사유 목록) + """ + # D1 — 검증기는 인자로 받는다. main() 지역 이름을 전역처럼 읽으면 NameError 로 즉사한다. + if validator is None: + return (seed_obj, set(), []) + + def _raw_validate(obj): + return validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": obj}, schema=schema, + expected_domain_id=domain_id, slice_document=slice_doc) + + def _validate(obj): + """검증기 자체가 터질 수 있다(예: bo_type 이 dict 면 unhashable). + + D9 — 예외를 코드로 바꾸기만 하면 부족하다. 그 오류의 path 는 뿌리라 + 후보를 지목하지 못하고, 처분이 뿌리 잔여(review)로 강등돼 문제 후보가 + 그대로 투영으로 흘러 project_to_bo_surface 에서 다시 죽는다. + 그래서 **어느 후보가 터뜨렸는지 후보 단위로 좁혀** path 에 인덱스를 실어 준다. + """ + try: + return _raw_validate(obj) + except Exception as exc: + detail = "%s: %s" % (type(exc).__name__, str(exc)[:160]) + errors = [] + cands = _list(obj.get("bo_seed_candidates")) + for idx in range(len(cands)): + probe = dict(obj) + probe["bo_seed_candidates"] = [cands[idx]] + try: + _raw_validate(probe) + except Exception: + errors.append({"code": "VALIDATOR_CRASHED", + "path": "$.stage_b_domain_bo_seed_output.bo_seed_candidates[%d]" % idx, + "message": detail}) + if not errors: + # 후보를 좁히지 못했다. 뿌리 문제이므로 명시적으로 끊는다 — + # review 로 강등해 투영에서 죽게 두는 것이 최악이다. + errors.append({"code": "VALIDATOR_CRASHED_ROOT", + "path": "$.stage_b_domain_bo_seed_output", "message": detail}) + return {"errors": errors, "warnings": []} + disposition = _dict(policy.get("disposition_by_code")) + unknown_disp = str(policy.get("unknown_code_disposition") or "quarantine") + synonyms = _dict(policy.get("enum_synonyms")) + case_fold = bool(policy.get("enum_case_fold")) + enum_unmapped = str(policy.get("enum_unmapped_disposition") or "repair_evict") + derivations = _dict(policy.get("required_derivation")) + derive_failed = str(policy.get("derivation_failed_disposition") or "repair_evict") + order = [str(v) for v in (policy.get("repair_order") or [])] + max_passes = int(policy.get("max_repair_passes") or 3) + blocks: list[dict[str, Any]] = [] + quarantined: set[int] = set() + + for _pass in range(max_passes): + cands = _list(seed_obj.get("bo_seed_candidates")) + # task_instance_id 는 지어내지 않는다 — A0 의 fan-out 계획 행이 든 값을 쓴다. + ctx = {"domain_id": domain_id, "seed_count": len(cands), + "emitted_seed_ids": [str(_dict(c).get("seed_id") or "") for c in cands]} + _plan_tid = _dict(plan_row).get("task_instance_id") + if isinstance(_plan_tid, str) and _plan_tid: + ctx["task_instance_id"] = _plan_tid + # r0_membership_key 는 R0 가 자기 관측으로 결정적으로 만든다(워커가 알 수 없는 값이다). + ctx["r0_membership_key"] = "%s:%s" % ( + domain_id, hashlib.sha256(json.dumps(ctx["emitted_seed_ids"], ensure_ascii=False, + sort_keys=True).encode("utf-8")).hexdigest()[:32]) + report = _validate(seed_obj) + errors = [e for e in (report.get("errors") or []) if isinstance(e, dict)] + if not errors: + break + buckets: dict[str, list[dict[str, Any]]] = {} + for err in errors: + disp = str(disposition.get(str(err.get("code"))) or unknown_disp) + buckets.setdefault(disp, []).append(err) + for err in buckets.get("block") or []: + blocks.append({"code": err.get("code"), "path": err.get("path"), "message": err.get("message")}) + if blocks: + return (seed_obj, quarantined, blocks) + for err in buckets.get("review") or []: + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "review", "received": None, "applied": None, + "rule_id": "disposition:review"}) + progressed = False + for action in order: + # 원소를 들어내는 처분(harvest·evict)만 내림차순으로 돈다 — 앞 인덱스를 먼저 + # 지우면 뒤 경로가 밀리기 때문이다. 값을 채우는 처분(normalize·derive)은 + # 오름차순이어야 한다: source_kind 는 형제 source_id 를 읽으므로 순서가 뒤집히면 + # 형제가 아직 없어 파생이 실패하고, 그 실패가 요건사실 원소를 통째로 들어낸다. + _removes = action in ("repair_harvest", "repair_evict") + for err in sorted(buckets.get(action) or [], + key=lambda e: PATH_INDEX_RE.sub( + lambda m: "[%05d]" % int(m.group(1)), str(e.get("path") or "")), + reverse=_removes): + path = str(err.get("path") or "") + tokens = _path_tokens(path) + if not tokens: + continue + norm = _norm_path(path) + rec = {"domain_id": domain_id, "code": err.get("code"), "path": path, + "action": action, "received": None, "applied": None, "rule_id": None} + done = False + if action == "repair_normalize": + parent, key = _resolve_parent(seed_obj, tokens) + table = _dict(synonyms.get(norm)) + if isinstance(parent, dict) and key in parent and table: + raw = parent.get(key) + probe = raw.casefold() if (case_fold and isinstance(raw, str)) else raw + mapped = table.get(probe) if isinstance(probe, str) else None + if mapped is not None: + rec["received"], rec["applied"] = _clip(raw), mapped + rec["rule_id"] = "enum_synonyms:%s" % norm + parent[key] = mapped + done = True + if not done and enum_unmapped == "repair_evict": + rec["rule_id"] = "enum_unmapped:%s" % norm + done = _evict(seed_obj, tokens, policy, norm, rec, quarantined) + elif action == "repair_derive": + rule = _dict(derivations.get(norm)) + if rule: + ok, value = _derive_required(seed_obj, tokens, norm, rule, universe, ctx) + if ok: + parent, key = _resolve_parent(seed_obj, tokens) + if isinstance(parent, dict): + rec["received"], rec["applied"] = None, _clip(value) + rec["rule_id"] = "required_derivation:%s" % norm + parent[key] = value + done = True + if not done and derive_failed == "repair_evict": + rec["rule_id"] = "derivation_failed:%s" % norm + done = _evict(seed_obj, tokens, policy, norm, rec, quarantined) + elif action == "repair_harvest": + rec["rule_id"] = "surplus_sink:%s" % norm + done = _harvest(seed_obj, tokens, policy, norm, rec) + elif action == "repair_evict": + rec["rule_id"] = "evict:%s" % norm + done = _evict(seed_obj, tokens, policy, norm, rec, quarantined) + elif action == "quarantine": + idx = _candidate_index(tokens) + if idx is not None: + quarantined.add(idx) + rec["rule_id"] = "quarantine:%s" % norm + done = True + if done: + cand_idx = _candidate_index(tokens) + if cand_idx is not None: + cand = _list(seed_obj.get("bo_seed_candidates")) + if cand_idx < len(cand): + rec["candidate_ref"] = _dict(cand[cand_idx]).get("candidate_ref") or _dict(cand[cand_idx]).get("seed_id") + records.append(rec) + progressed = True + if progressed: + break + if not progressed: + break + + # 수렴하지 않고 남은 위반은 후보 단위로 격리한다. 회차는 계속된다. + report = _validate(seed_obj) + for err in (report.get("errors") or []): + if not isinstance(err, dict): + continue + # review 로 수용하기로 선언된 코드는 남아 있는 것이 정상이다. 다시 격리하지 않는다. + if str(disposition.get(str(err.get("code"))) or "") == "review": + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "review", "received": None, "applied": None, + "rule_id": "disposition:review(residual)"}) + continue + tokens = _path_tokens(str(err.get("path") or "")) + idx = _candidate_index(tokens) + if idx is None: + root_disp = str(policy.get("residual_root_disposition") or "review") + if root_disp == "block": + blocks.append({"code": err.get("code"), "path": err.get("path"), + "message": err.get("message"), "note": "root-level residue"}) + else: + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "review", "received": None, "applied": None, + "rule_id": "residual_root"}) + else: + quarantined.add(idx) + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "quarantine", "received": None, "applied": None, + "rule_id": "residual_after_repair"}) + return (seed_obj, quarantined, blocks) + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + admission_policy = _dict(json.loads(_verify_asset(SEED_ADMISSION_POLICY_PATH))) + if admission_policy.get("schema_version") != SEED_ADMISSION_SCHEMA_VERSION: + raise RuntimeError("PART2_SEED_ADMISSION_POLICY_INVALID") + admission_enforcing = str(admission_policy.get("enforcement") or "enforce") == "enforce" + admission_records: list[dict[str, Any]] = [] + declared_candidate_total = 0 + admission_blocks: list[dict[str, Any]] = [] + quarantined_by_domain: dict[str, set] = {} + admitted_total = 0 + quarantined_total = 0 + + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R0-6 — 검증 결과를 경고로 흘리지 않고 등급대로 처분한다. + # shadow 모드에서는 계산·기록만 하고 seed 를 바꾸지 않는다(첫 전환 회차용). + # D13 — 회계 누적은 검증기 유무와 무관하다. 분기 안에 두면 검증기 반입이 + # 실패한 회차에서 declared=0 · admitted=N 이 되어 헛경보(lost 음수)가 난다. + declared_candidate_total += len(_list(seed_obj.get("bo_seed_candidates"))) + if worker_validator is not None: + for item in (worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, schema=seed_schema, + expected_domain_id=domain_id, slice_document=slice_doc).get("warnings") or []): + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + work_obj = seed_obj if admission_enforcing else copy.deepcopy(seed_obj) + admitted_obj, quarantined_idx, blocks = admit_seed( + work_obj, domain_id, seed_schema, slice_doc, universe, + admission_policy, admission_records, PLAN_ROWS.get(domain_id) or {}, + worker_validator) + if blocks: + admission_blocks.extend({"domain_id": domain_id, **b} for b in blocks) + if admission_enforcing: + seed_obj = admitted_obj + quarantined_by_domain[domain_id] = quarantined_idx + else: + quarantined_by_domain[domain_id] = set() + for b in blocks: + guard_warnings.append({"code": "SEED_ADMISSION_SHADOW_BLOCK", + "domain_id": domain_id, "detail": b}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + skip_idx = quarantined_by_domain.get(domain_id) or set() + for idx, cand in enumerate(cands): + if idx in skip_idx: + quarantined_total += 1 + continue + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # R0-6b — 수용 결산. 격리는 후보 단위이고, 회차 중단은 예산을 넘을 때만이다. + admitted_total = sum(len(v) for v in projected_candidates.values()) + # D4 — 수용 단계에서 후보가 사라지는 경로는 격리 하나뿐이어야 한다. + # 원본 후보 수와 (수용 + 격리)가 맞지 않으면 소리 없이 없어진 것이 있다는 뜻이다. + if admission_enforcing and declared_candidate_total != admitted_total + quarantined_total: + guard_warnings.append({"code": "SEED_ADMISSION_ACCOUNTING_DRIFT", + "declared": declared_candidate_total, "admitted": admitted_total, + "quarantined": quarantined_total, + "lost": declared_candidate_total - admitted_total - quarantined_total}) + if admission_blocks and admission_enforcing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SEED_ADMISSION_BLOCKED", + "message": "시드 뿌리 계약이 깨졌다. 복구 대상이 아니다.", + "blocks": admission_blocks[:20], "block_count": len(admission_blocks), + }, ensure_ascii=False)) + budget = _dict(admission_policy.get("quarantine_budget")) + seen_total = admitted_total + quarantined_total + ratio = (quarantined_total / seen_total) if seen_total else 0.0 + max_ratio = budget.get("max_quarantined_candidate_ratio") + min_admitted = budget.get("min_admitted_candidates_run") + if admission_enforcing and isinstance(max_ratio, (int, float)) and seen_total and ratio > float(max_ratio): + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SEED_ADMISSION_BUDGET_EXCEEDED", + "message": "격리 비율이 한계를 넘었다. 모델 잡음이 아니라 계약 파손으로 본다.", + "quarantined": quarantined_total, "seen": seen_total, + "ratio": round(ratio, 4), "max_ratio": max_ratio, + }, ensure_ascii=False)) + if admission_enforcing and isinstance(min_admitted, int) and admitted_total < min_admitted: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SEED_ADMISSION_EMPTY", + "message": "수용된 후보가 없다.", "admitted": admitted_total, + }, ensure_ascii=False)) + _adm_review = _dict(admission_policy.get("review")) + _adm_counter = 0 + for rec in admission_records: + _adm_counter += 1 + _is_q = rec.get("action") == "quarantine" + _is_c = bool(rec.get("content_bearing")) and not _is_q + review_handoff_items.append({ + "review_id": "R0:admission:%03d" % _adm_counter, + "source_domain": rec.get("domain_id"), + "severity": str((_adm_review.get("quarantine_severity") if _is_q + else _adm_review.get("evicted_with_content_severity") if _is_c + else _adm_review.get("severity")) or "SOFT_WARNING"), + "issue_type": str((_adm_review.get("quarantine_issue_type") if _is_q + else _adm_review.get("evicted_with_content_issue_type") if _is_c + else _adm_review.get("issue_type")) or "seed_admission_repair"), + "source_review_code": rec.get("code"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": str(_adm_review.get("downstream_owner") or "Stage2"), + "template_note": "수용 단계 처분: %s · 경로 %s · 규칙 %s · 원값 %s -> 적용 %s (후보 %s)" % ( + rec.get("action"), rec.get("path"), rec.get("rule_id"), + rec.get("received"), rec.get("applied"), rec.get("candidate_ref")), + }) + guard_warnings.append({"code": "SEED_ADMISSION_SUMMARY", + "enforcement": admission_policy.get("enforcement"), + "repairs": sum(1 for r in admission_records if str(r.get("action") or "").startswith("repair")), + "reviews": sum(1 for r in admission_records if r.get("action") == "review"), + "quarantined_candidates": quarantined_total, + "admitted_candidates": admitted_total, + "quarantine_ratio": round(ratio, 4)}) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if not isinstance(bo_type, str): + # 워커가 문자열이 아닌 값을 넣으면 집합 비교 자체가 터진다(unhashable). + # 수용 단계가 걸러야 하지만, 투영은 마지막 방어선이라 여기서도 막는다. + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", + "received": type(bo_type).__name__, + "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_9_45am.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_9_45am.yml new file mode 100644 index 00000000..5b3bbfe4 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/outdated/stage_1_part_2_v.8_Aug_22_9_45am.yml @@ -0,0 +1,3393 @@ +# ============================================================================= +# Liti-agent Stage 1 Part 2 v.8 — BO 컴파일 (동적 도메인 fan-out) · 통합 실행본 +# +# 정본 근거 +# 전체 DAG : stage_1_update_strategy.md §1 Part 2 블록 · §3 세부 워크플로우 +# 개정 전략 : stage_1_part_2_optimal_update_strategy_v.2.md +# 선결작업 : part_2_우선작업_report.md · part_2_우선작업_2_report.md +# 패치 : part_2_prerequisite/patches/C1_C2_registry_component_ids.patch.md +# +# 구성 — task 6종 +# [개정] Task_C_BO_A0_context_and_domain_slice_compiler 상수 3덩어리 -> 모듈 5개 호출 +# [신설] Task_C_B_domain_worker_* 정적 worker 5개 대체 템플릿 +# [개정] Task_C_BO_R0_seed_reducer_and_exception_planner fan-out 기대집합 + validator + C-2 +# [무변경] Task_C_BO_R1_exception_adjudicator v3 원문 바이트 동일 +# [개정] Task_C_BO_F0_final_bo_compiler_gate_writer BOType 어휘 registry 합집합 +# [개정] Task_C_BO_S0_signal_bundle_writer 인라인 모듈 -> 조립본 모듈 반입 +# +# 삭제 — Task_C_BO_Stage_B_B1~B5 다섯 (v3 1502~2826행, 1,325행) +# §6.7 규율대로 즉시 삭제하지 않는다. 템플릿으로 승계 5도메인을 돌려 같은 BO 가 나오는 +# 것을 확인한 뒤(Q-4) 삭제한다(Q-5). 이 파일은 그 확인이 끝난 상태를 전제한다. +# +# 확정 계약 (stage_1_update_strategy.md §0.3) +# slice runtime/domain_slices/.json task_c_bo_stage_b_domain_slice.v2 +# worker 산출 runtime/domain_seed_outputs/.json task_c_bo_stage_b_domain_bo_seed.v3 +# fan-out fanout/domain_fanout_plan.json domain_fanout_plan.v1 +# worker 이름 Task_C_B_domain_worker_* · 인스턴스 DOMAIN-<도메인ID> +# 실행 인자 --asset-root · --execution-root · --logical-root +# 구 slice/seed 경로(stage1_tmp/task_c_bo/domain_slices|domain_seed_outputs)는 쓰지 않는다 +# (legacy_paths_forbidden). stage_a_context·source_universe_manifest(P-1 복귀)와 +# postb_* 3종은 stage1_tmp/task_c_bo/ 를 정본 경로로 유지한다. +# +# 모듈 반입 — Part 1 D0 규약 R-1~R-5 승계 +# .txt 미러를 read_raw 로 읽고 runtime_manifest.json 의 sha256 과 대조한 뒤 +# /tmp/s1/_rt 에 .py 로 기록하고 sys.path 에 넣는다. 미러는 정본 .py 옆에 있다. +# +# 이 파일은 스테이지 하나다. 스테이지 선언 1벌 · task_procedure 1벌 · tasks 1벌. +# 들여쓰기는 Part 2 v3 관례(Stages 2 · tasks 4 · task_name 4)를 유지한다. +# ============================================================================= +--- +Agent: + name: Liti-agent_Civil_Suit_Plaintiff_Stage_1_Part_2 + description: 민사소송 원고 송무 초지능 AI변호사 - Stage 1 Part 2 + version: v.2 + Stages: + - name: stage1_BO_시그널_생성 + description: BO 생성, 시그널 생성 + llm_provider: openai + llm_model: gpt-4o-2024-08-06 + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + description: Get the content of local documents + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + description: Run scripts of programming languages + headers: + Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM= + tasks: + - task_name: Task_C_BO_A0_context_and_domain_slice_compiler + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx" + network: "agent-network" + timeout: 300 + code: | + #!/usr/bin/env python3 + # Task_C_BO_A0_context_and_domain_slice_compiler (v4) + # 1) 자산 반입 -> 2) 봉인 검증 -> 3) registry 로드·검증 -> + # 4) 프롬프트 조립 -> 5) slice 컴파일 -> 6) fan-out 계획 -> 7) 기록 + # 도메인 상수를 두지 않는다. 라우팅 판정은 모듈 안에서만 일어난다. + import contextlib + import datetime + import hashlib + import io + import itertools + import json + import os + import posixpath + import pathlib + import sys + import unicodedata + + import httpx + + # ------------------------------------------------------------------ + # localdocs 보일러플레이트 (SKILL.md 5장 / 5.2장) + # clientInfo 에 {{__user_hash__}} / {{__workspace_hash__}} 를 반드시 넣는다. + # 빠지면 localdocs 가 루트 경로를 보므로 사용자 파일을 찾지 못한다. + # Task_A0_domain_screener_02.yml 의 검증 완료본을 그대로 복사했다. + # ------------------------------------------------------------------ + TASK_NAME = "Task_C_BO_A0_context_and_domain_slice_compiler" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=120) + MSG_ID_COUNTER = itertools.count(10) + + + def next_msg_id(): + return next(MSG_ID_COUNTER) + + + def _init(): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-03-26", "capabilities": {}, + "clientInfo": {"name": TASK_NAME, "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}"}} + }, headers=MCP_HEADERS) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post(LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS).raise_for_status() + + + def _parse_mcp(text): + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + return None + try: + return json.loads(text) + except Exception: + return None + + + def _call(name, args, mid): + r = CLIENT.post(LOCALDOCS_URL, json={ + "jsonrpc": "2.0", "id": mid, "method": "tools/call", + "params": {"name": name, "arguments": args} + }, headers=MCP_HEADERS) + r.raise_for_status() + p = _parse_mcp(r.text) + if not p or "result" not in p: + raise RuntimeError("MCP_CALL_FAILED:%s" % name) + return p + + + def read_raw(name): + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # registry_index_sha256 과 screening_sha256 은 원문 바이트의 해시여야 + # 하므로 재직렬화를 절대 허용하지 않는다. + p = _call("read_docs", {"doc_names": [name]}, next_msg_id()) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + def read_json(name): + raw = read_raw(name) + s = raw.strip() + if s.startswith("```"): + for part in s.split("```"): + part = part.strip() + if part.startswith("json"): + part = part[4:].strip() + if part.startswith("{") or part.startswith("["): + s = part + break + try: + return json.loads(s) + except json.JSONDecodeError: + obj, _ = json.JSONDecoder().raw_decode(s) + return obj + + + def write_doc(path, content): + _call("write_file", {"path": path, "content": content, "overwrite": True}, + next_msg_id()) + # ------------------------------------------------------------------ + # 실행 뿌리 세 개 — D-5 §2.4 0-c-2 확정값. 모듈에는 argv 로만 넘긴다. + # ------------------------------------------------------------------ + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + + WORK = pathlib.Path(EXECUTION_ROOT) + RT = WORK / "_rt" + + # 미러는 정본 .py 옆에 놓인다. 이름이 아니라 논리 경로로 지목한다. + MODULE_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "registry_validator": "Default_Agent/stage1_runtime/registry_validator.txt", + "prompt_compiler": "Default_Agent/stage1_runtime/prompt_compiler.txt", + "domain_slice_compiler": "Default_Agent/stage1_runtime/domain_slice_compiler.txt", + "domain_fanout_planner": "Default_Agent/stage1_runtime/domain_fanout_planner.txt", + "stage_a_context_builder": "Default_Agent/stage1_runtime/stage_a_context_builder.txt", + } + MODULES = ["runtime_common", "schema_subset_validator", "registry_loader", + "registry_validator", "prompt_compiler", "domain_slice_compiler", + "domain_fanout_planner", "stage_a_context_builder"] + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + + # 사건 산출물. Part 1 이 낸 것만 읽는다. + HANDOFF = "quality_gates/stage1_part1_soft_gate_handoff.json" + ACTIVATION_MANIFEST = "routing/domain_activation_manifest.json" + SCREENING = "routing/domain_screening.json" + EVIDENCE = "evidence_indexed.json" + EVENTS = "evidence_event_candidates.json" + # meeting_clause_ids 는 evidence·event 문서에 없다. 실측으로 확인했다 + # (구 매니페스트 41개 중 두 문서에서 발견되는 것 0개). 원문을 읽어야 나온다. + MEETING = "client_meeting.md" + # R-3 — Part 1 screener 03 이 낸 어휘 사전. 여덟 갈래 중 여섯을 E|O|V|D|R| 줄로 담는다. + # 네 번째 digest 생성기를 만들지 않는다 — 이미 있는 것을 프롬프트 조각으로 붙인다. + VOCABULARY = "routing/candidate_profile_vocabulary.md" + + # 정적 자산. + REGISTRY_INDEX = "Default_Agent/domains/_registry_index.json" + COMMON_CONTRACT = "Default_Agent/domains/_common/common_worker_contract.md" + POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" + SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" + FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" + + # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. + # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 + # 여기 넣지 않는다. 그 경계는 의도한 것이다. + PART2_REQUIRED_ASSETS = ( + SLICE_SCHEMA, + FANOUT_SCHEMA, + COMMON_CONTRACT, + POLICY, + "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json", # R0 + "Default_Agent/signals/signal_registry.v2.json", # S0 + "Default_Agent/contracts/signals/s5_execution_contract.v2.json", # S0 + "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 + "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 + "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + ) + RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" + # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. + OVERLAY_ERROR_CODES = ("PROMPT_OVERLAY_HASH_MISMATCH", "PROMPT_OVERLAY_NOT_FOUND", + "PROMPT_OVERLAY_PATH_INVALID", "PROMPT_OVERLAY_REFERENCE_DIVERGENCE") + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + + # 산출 경로 — 새 계약만 쓴다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + PROMPT_DIR = "runtime/compiled_prompts" + SEED_DIR = "runtime/domain_seed_outputs" + FANOUT_PATH = "fanout/domain_fanout_plan.json" + # P-1 — v4 개정에서 구 slice 경로를 걷어내며 이 둘의 접두까지 벗겼던 것을 되돌린다. + # 이 둘은 slice 가 아니며 R0·F0·S0 가 여기서 읽는다(v3 1104·1105행). + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + SOURCE_MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + RECEIPT_PATH = "validation_assets/routing/stage_receipt.json" + + ERRORS = [] + WARNINGS = [] + + + def warn(code, message): + WARNINGS.append({"code": code, "message": message}) + + + def sha_text(text): + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + + def utc_now(): + # stage_a_context 의 created_at_utc 전용이다. + # 조립 프롬프트 해시에는 들어가지 않으므로 결정성(판정 2)에 영향이 없다. + return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + + def canonical(value): + return json.dumps(value, ensure_ascii=False, sort_keys=True, + separators=(",", ":")) + "\n" + + + def stage_text(logical_name, body): + # 논리 이름을 그대로 실행 뿌리 아래 상대경로로 쓴다. + target = WORK / logical_name + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(body.encode("utf-8")) + return len(body.encode("utf-8")) + + + # ------------------------------------------------------------------ + # 0) 배포 완전성 — F-2. 모듈 반입보다 앞이고 봉인 검증보다도 앞이다. + # 봉인은 사건 산출물의 문제이고 이것은 조립본의 문제라 원인이 다르다. + # 첫 실패에서 멈추지 않고 전부 모은다 — 배포는 한 번에 고쳐야 한다. + # ------------------------------------------------------------------ + def assert_deployment(): + """조립본이 Part 2 개정 델타를 한 벌로 받았는지 본다. 읽기만 한다.""" + manifest = json.loads(read_raw(RUNTIME_MANIFEST)) + if manifest.get("schema_version") != RUNTIME_MANIFEST_SCHEMA: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "schema_version", + "expected": RUNTIME_MANIFEST_SCHEMA, "actual": manifest.get("schema_version"), + }, ensure_ascii=False)) + rows = [row for row in (manifest.get("entries") or []) if isinstance(row, dict)] + paths = [row.get("path") for row in rows] + duplicates = sorted({p for p in paths if paths.count(p) > 1}) + if manifest.get("runtime_artifact_count") != len(rows) or duplicates: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_RUNTIME_MANIFEST_INVALID", "detail": "count_or_duplicate", + "declared_count": manifest.get("runtime_artifact_count"), "actual_count": len(rows), + "duplicate_paths": duplicates, + }, ensure_ascii=False)) + expected = {row["path"]: row["sha256"] for row in rows} + unregistered, mismatch, unreadable = [], [], [] + for logical in sorted(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))): + rel = logical[len(ASSET_ROOT):] + want = expected.get(rel) + if want is None: + unregistered.append(rel) + try: + body = read_raw(logical) + except Exception: + unreadable.append(rel) + continue + if want is not None and want != sha_text(body): + mismatch.append({"path": rel, "expected": want, "actual": sha_text(body)}) + if unregistered or mismatch or unreadable: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", + "message": "조립본에 Part 2 개정 자산이 한 벌로 반영되지 않았다.", + "unregistered": unregistered, "hash_mismatch": mismatch, "unreadable": unreadable, + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + return manifest + + + # ------------------------------------------------------------------ + # 1) 모듈 반입 — R-1~R-5. 해시가 어긋나면 실행하지 않는다. + # ------------------------------------------------------------------ + def materialize_modules(): + RT.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] + for row in manifest_doc.get("entries") or []} + staged = [] + for name in MODULES: + logical = MODULE_MIRRORS[name] + raw = read_raw(logical).encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (RT / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(RT) not in sys.path: + sys.path.insert(0, str(RT)) + return staged + + + # ------------------------------------------------------------------ + # 2) 봉인 검증 — 세 해시는 read_raw 원문에서 계산한다 (§6.1-3). + # ------------------------------------------------------------------ + def verify_seal(handoff, screening_raw, manifest_raw, index_raw): + root = handoff.get("stage1_part1_soft_gate_handoff", handoff) + guard = root.get("digest_guard") or {} + pairs = [("screening_sha256", sha_text(screening_raw)), + ("activation_manifest_sha256", sha_text(manifest_raw)), + ("registry_index_sha256", sha_text(index_raw))] + for key, actual in pairs: + declared = guard.get(key) + if declared is None: + raise RuntimeError("SEAL_KEY_MISSING:%s" % key) + if declared != actual: + raise RuntimeError("SEAL_FAILED:%s" % key) + return {key: value for key, value in pairs} + + + # ------------------------------------------------------------------ + # 3) C-1 — evidence authority map 에 registry_component_ids 통과 (P0 판정 B) + # component_keys(문서 추출 구조 이름)와 계층이 다르므로 섞지 않는다. + # ------------------------------------------------------------------ + def evidence_authority_map(evidence_document): + root = evidence_document.get("evidence_indexed", evidence_document) + items = root.get("items") if isinstance(root, dict) else evidence_document + out = {} + for item in items if isinstance(items, list) else []: + if not isinstance(item, dict): + continue + index = item.get("evidence_index") or item.get("evidence_index_proposed") + if not isinstance(index, str) or not index: + continue + out[index] = { + "evidence_index": index, + "doc_uid": item.get("doc_uid"), + "doc_type": item.get("doc_type"), + "source_pointer": item.get("source_pointer") or {}, + "registry_component_ids": [ + str(value) for value in (item.get("registry_component_ids") or []) + if isinstance(value, str) and value + ], + } + return out + + + # ------------------------------------------------------------------ + # 4) 본체 + # ------------------------------------------------------------------ + def main(): + # F-2 — 게이트가 먼저다. 반입도 봉인도 그 뒤다. + gate_manifest = assert_deployment() + staged_modules = materialize_modules() + import registry_loader + import registry_validator + import prompt_compiler + import domain_slice_compiler + import domain_fanout_planner + import stage_a_context_builder + + handoff = read_json(HANDOFF) + screening_raw = read_raw(SCREENING) + manifest_raw = read_raw(ACTIVATION_MANIFEST) + index_raw = read_raw(REGISTRY_INDEX) + seal = verify_seal(handoff, screening_raw, manifest_raw, index_raw) + + stage_text(REGISTRY_INDEX, index_raw) + index_doc = json.loads(index_raw) + index = index_doc.get("domain_registry_index", index_doc) + for entry in index.get("entries") or []: + config_path = entry.get("config_path") + if not isinstance(config_path, str) or not config_path: + raise RuntimeError("REGISTRY_CONFIG_PATH_MISSING:%s" % entry.get("domain_id")) + logical = unicodedata.normalize("NFC", "Default_Agent/domains/" + config_path + if not config_path.startswith("Default_Agent/") + else config_path) + config_text = read_raw(logical) + stage_text(logical, config_text) + # 프롬프트 조각도 함께 반입한다. prompt_compiler 가 도메인별 + # seed_prompt_overlay 를 읽으므로 config 만 실으면 fragment not found 로 멈춘다. + # 파일 이름을 짓지 않는다 — config 가 선언한 prompt_overlay_ref 를 따라간다. + overlay_ref = json.loads(config_text).get("prompt_overlay_ref") + if isinstance(overlay_ref, str) and overlay_ref: + overlay_logical = unicodedata.normalize( + "NFC", overlay_ref if overlay_ref.startswith("Default_Agent/") + else posixpath.join(posixpath.dirname(logical), overlay_ref)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) + stage_text(POLICY, read_raw(POLICY)) + + slice_schema = json.loads(read_raw(SLICE_SCHEMA)) + fanout_schema = json.loads(read_raw(FANOUT_SCHEMA)) + + # F-3 — 입력 능력 검사. 장부(F-2)가 아니라 의미를 본다. + # 매니페스트와 스키마를 함께 옛 판본으로 되돌리면 장부는 자기들끼리 맞아 통과한다. + # 그 자리에서 유일하게 남는 검사가 이것이다. + _sb = (slice_schema.get("properties") or {}).get("stage_b_domain_slice") or {} + _props = _sb.get("properties") or {} + _missing = [k for k in ("domain_declarations",) if k not in _props] + if "hash_kind" not in ((_props.get("compiled_prompt") or {}).get("properties") or {}): + _missing.append("compiled_prompt.hash_kind") + if _missing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_SCHEMA_STALE", + "message": "슬라이스 스키마가 컴파일러가 내는 키를 선언하지 않는다. 조립본의 스키마가 개정 전 판본이다.", + "path": SLICE_SCHEMA, "missing_declarations": _missing, + "remedy": "domain_slice.schema.v2.json 을 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + os.chdir(EXECUTION_ROOT) + registry = registry_loader.load_registry(REGISTRY_INDEX) + validation = registry_validator.validate_registry(REGISTRY_INDEX) + # F-4d — overlay 계열 네 코드만 경성으로 올린다. validate_registry 전체를 올리면 + # 지금 통과 중인 다른 review 항목까지 막힌다. 부분 복사에서 흔한 것은 훼손이 아니라 + # 누락이고, 누락은 PROMPT_OVERLAY_NOT_FOUND 로 나온다. + _ovl = [e for e in (validation.get("errors") or []) + if isinstance(e, dict) and e.get("code") in OVERLAY_ERROR_CODES] + if _ovl: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "detail": "prompt_overlay", + "codes": sorted({str(e.get("code")) for e in _ovl}), + "domains": sorted({str(e.get("domain_id")) for e in _ovl}), + "remedy": PART2_DEPLOYMENT_REMEDY, + }, ensure_ascii=False)) + if validation.get("status") not in ("PASS", "READY", "OK"): + warn("REGISTRY_VALIDATION_NOT_PASS", str(validation.get("status"))) + + manifest = json.loads(manifest_raw) + manifest_root = manifest.get("domain_activation_manifest", manifest) + # D-3 넓은 정의 — execution_eligible 이 유일한 결정 필드다. + eligible = sorted({row.get("domain_id") + for row in manifest_root.get("domain_entries") or [] + if row.get("execution_eligible") is True}) + expected_runnable = sorted(set(manifest_root.get("expected_runnable_domain_ids") or [])) + if eligible != expected_runnable: + raise RuntimeError("A0_EXPECTED_RUNNABLE_SET_MISMATCH") + + evidence_document = json.loads(read_raw(EVIDENCE)) + events_document = json.loads(read_raw(EVENTS)) + authority = evidence_authority_map(evidence_document) + + # 프롬프트 조립 — rank 10 -> 20 -> 30 -> 40 -> 50. 상한 96,000 B / 24,000 자. + policy = json.loads(read_raw(POLICY)) + # R-3 — 어휘 사전을 실행 뿌리에 실어 조각으로 붙인다. 부재는 경고로 남기고 진행한다 + # (Part 1 이 아직 그 파일을 내지 않은 배포에서도 조립은 되어야 한다). + vocabulary_specs = [] + try: + vocabulary_text = read_raw(VOCABULARY) + stage_text(VOCABULARY, vocabulary_text) + vocabulary_specs = [prompt_compiler.FragmentSpec( + fragment_id="candidate_profile_vocabulary", + category="common_dependency", + path=str(pathlib.Path(EXECUTION_ROOT) / VOCABULARY))] + except Exception as exc: + warn("VOCABULARY_FRAGMENT_ABSENT", "%s: %s" % (VOCABULARY, exc)) + prompt_manifests = {} + for domain_id in expected_runnable: + specs = prompt_compiler.collect_domain_fragments( + domain_id, registry, common_contract_path=COMMON_CONTRACT, + profile_paths=profile_paths, extra_specs=vocabulary_specs) + text, manifest_row = prompt_compiler.compile_fragments(specs, policy) + rel = "%s/%s.md" % (PROMPT_DIR, domain_id) + stage_text(rel, text) + write_doc(rel, text) + row = dict(manifest_row) + row["compiled_prompt_path"] = rel + row["compiled_prompt_sha256"] = sha_text(text) + row.setdefault("composition_policy_sha256", sha_text(read_raw(POLICY))) + row["_manifest_dir"] = EXECUTION_ROOT + prompt_manifests[domain_id] = row + + # R-2 — Part 1 screener 02 가 CALC_NOT_IN_BINDINGS 로 이미 검증해 낸 + # requested_calculation_domains 를 통과시킨다. 새 registry 를 적재하지 않는다. + # 봉인용 원문 바이트(screening_raw)는 손대지 않고 파싱만 따로 한다. + # 파싱 실패와 계약 위반을 갈라 둔다. try 로 함께 감싸면 계약 위반이 경고로 + # 강등되어 조용히 통과한다 — 애초에 고치려던 것이 그 조용함이다. + screening_calc = {} + try: + screening_doc = json.loads(screening_raw) + except Exception as exc: + screening_doc = None + warn("SCREENING_CALC_PARSE_SKIPPED", str(exc)) + if screening_doc is not None: + # 루트 래핑을 벗긴다. Part 1 은 {"domain_screening": {...}} 로 쓰고 + # 스키마가 그 키를 required 로 못박는다. 벗기지 않으면 candidates 가 + # 늘 None 이 되어 예외도 없이 아무 일도 일어나지 않는다. + screening_root = screening_doc.get("domain_screening", screening_doc) \ + if isinstance(screening_doc, dict) else None + if not isinstance(screening_root, dict): + raise RuntimeError("SCREENING_ROOT_INVALID") + candidate_rows = screening_root.get("candidates") + if not isinstance(candidate_rows, list) or not candidate_rows: + raise RuntimeError("SCREENING_CANDIDATES_EMPTY") + for row in candidate_rows: + if not isinstance(row, dict): + raise RuntimeError("SCREENING_CANDIDATE_INVALID") + domain_id = row.get("domain_id") + codes = [str(v) for v in (row.get("requested_calculation_domains") or []) + if isinstance(v, str) and v] + if isinstance(domain_id, str) and domain_id and codes: + screening_calc[domain_id] = sorted(set(codes)) + + # P-2a — stage_a_context 를 slice 컴파일보다 먼저 만든다. slice 의 source_universe 가 + # 인용할 event 식별자의 정본이 event_candidate_map 이기 때문이다. 순서가 뒤였을 때 + # A0 는 원시 evidence_event_candidates 문서를 넘겼고, domain_slice_compiler._records 가 + # 그 문서-단위 items(30건)를 후보로 오인해 candidate_id 를 못 찾아 + # event:unidentified:NNNN 로 대체했다. 그 값은 source_universe_manifest 의 + # event_candidate_ids(EVT-...-NN, 84건)와 교집합이 0 이라, 워커가 규율을 지켜 + # slice 안의 것만 인용해도 R0 가 "source_refs outside Stage A universe" 로 차단했다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + # 맵을 통째로 넘기지 않는다 — domain_slice_compiler._records 는 dict 를 받으면 + # by_evidence_index(값이 id 리스트)에 먼저 걸려 빈 목록을 돌려준다. 후보 레코드 + # 목록으로 평탄화해 넘겨야 _record_id 가 candidate_id 를 찾아 EVT-...-NN 을 쓴다. + event_candidate_records = [ + row for row in ((stage_a.get("event_candidate_map") or {}).get("by_candidate_id") or {}).values() + if isinstance(row, dict)] + if not event_candidate_records: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_EVENT_CANDIDATE_MAP_EMPTY", + "message": "stage_a event_candidate_map.by_candidate_id 가 비었다. slice 의 event 식별자를 정본으로 실을 수 없다.", + }, ensure_ascii=False)) + try: + result = domain_slice_compiler.compile_domain_slices( + manifest, registry, evidence_document, event_candidate_records, + manifest_sha256=seal["activation_manifest_sha256"], + evidence_sha256=sha_text(read_raw(EVIDENCE)), + events_sha256=sha_text(read_raw(EVENTS)), + slice_schema=slice_schema, + compiled_prompt_manifests=prompt_manifests, + expected_output_dir=SEED_DIR, + screening_calculation_domains=screening_calc) + except TypeError as exc: + # F-3b 앞단 — 옛 컴파일러는 screening_calculation_domains 를 받지 않는다. 그대로 두면 + # 배포 원인을 말하지 않는 TypeError 로 끝난다. 이름을 붙여 같은 코드로 내보낸다. + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 새 인자를 받지 않는다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "detail": "signature_mismatch", "signature_error": str(exc)[:200], + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + # F-3b — 출력 능력 검사. 기록 루프 앞이다. 여기서 멈추면 슬라이스가 한 벌도 나가지 않는다. + # 새 스키마가 domain_declarations 를 required 로 올리지 않으므로(P0 판정 C) 옛 컴파일러의 + # 산출도 스키마 검증은 26/26 통과한다. 장부가 볼 수 없는 그 자리를 이 검사가 막는다. + _bad = [] + for _did, _obj in sorted((result.get("slices") or {}).items()): + _root = (_obj or {}).get(SLICE_ROOT_KEY) or _obj or {} + if ("domain_declarations" not in _root + or "hash_kind" not in (_root.get("compiled_prompt") or {})): + _bad.append(_did) + if _bad: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_COMPILER_STALE", + "message": "컴파일러가 registry 선언 블록을 싣지 않았다. 조립본의 domain_slice_compiler 가 개정 전 판본이다.", + "domains_without_declarations": _bad, + "remedy": "domain_slice_compiler.py|.txt 를 후보 오버레이 판본으로 덮어쓴다.", + }, ensure_ascii=False)) + + slice_hashes = {} + for domain_id, slice_obj in (result.get("slices") or {}).items(): + rel = "%s/%s.json" % (SLICE_DIR, domain_id) + text = canonical(slice_obj) + stage_text(rel, text) + write_doc(rel, text) + slice_hashes[domain_id] = sha_text(text) + + plan = domain_fanout_planner.build_fanout_plan( + manifest, registry, + activation_manifest_sha256=seal["activation_manifest_sha256"], + slice_dir=SLICE_DIR, seed_output_dir=SEED_DIR, + fanout_schema=fanout_schema) + plan_root = plan.get("domain_fanout_plan", plan) + + # barrier 기대 집합은 expected_runnable_domain_ids 다. active_domain_ids 가 아니다. + planned = sorted({row.get("domain_id") + for row in plan_root.get("task_instances") or []}) + if planned != expected_runnable: + raise RuntimeError("A0_FANOUT_SET_MISMATCH") + + write_doc(FANOUT_PATH, canonical(plan)) + + # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. + # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. + # 계산은 P-2a 에서 이미 끝났다(slice 가 같은 식별자를 써야 하므로 앞당겼다). 여기서는 기록만 한다. + write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) + write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) + # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. + # 스키마를 안 넘겨도 같은 모양의 성공이 돌아오므로 산출물만으로는 "통과"와 + # "안 함"을 가를 수 없다. 그래서 넘긴 사실과 대상 수를 여기에 적어 둔다. + write_doc(RECEIPT_PATH, canonical({ + "schema_version": "stage1_stage_receipt.v2", + "stage": "P2-A0", + "loader_mode": "registry_modules", + "worker_mode": "template_fanout", + "activation_source": "sg01_manifest", + "slice_sha256_by_domain": slice_hashes, + "schema_injection": { + "slice_schema_path": SLICE_SCHEMA, + "slice_schema_sha256": sha_text(read_raw(SLICE_SCHEMA)), + "slice_schema_argument": "slice_schema", + "fanout_schema_path": FANOUT_SCHEMA, + "fanout_schema_sha256": sha_text(read_raw(FANOUT_SCHEMA)), + "fanout_schema_argument": "fanout_schema", + "validated_slice_count": len(slice_hashes), + "validated_fanout_instance_count": len(plan_root.get("task_instances") or []), + "domain_declarations_projected": sorted( + (result.get("slices") or {}).keys()), + "screening_calculation_domains": screening_calc, + "vocabulary_fragment_injected": bool(vocabulary_specs), + "keyword_support_checker": "validation_assets/routing/_check_schema_keyword_support.py", + "note": "넘김이 곧 검증은 아니다. 대상 수가 0 이면 검증도 0 회다.", + }, + "deployment_gate": { + "checked_count": len(set(list(MODULE_MIRRORS.values()) + list(PART2_REQUIRED_ASSETS))), + "runtime_manifest_sha256": sha_text(read_raw(RUNTIME_MANIFEST)), + "runtime_artifact_count": gate_manifest.get("runtime_artifact_count"), + "schema_capability_checked": ["domain_declarations", "compiled_prompt.hash_kind"], + "compiler_output_checked": True, + "overlay_codes_enforced": list(OVERLAY_ERROR_CODES), + "note": "무엇을 봤는지 적는다. 상수 PASS 는 증거가 아니다.", + }, + "created_by": TASK_NAME, + })) + + return {"status": "READY", "written": True, + "modules": staged_modules, + "expected_runnable_domain_ids": expected_runnable, + "slice_count": len(slice_hashes), + "fanout_instance_count": len(plan_root.get("task_instances") or []), + # 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은 + # 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다 + # (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다. + "dynamic_fanout": plan_root.get("task_instances") or [], + "digest_guard": seal, + "errors": ERRORS, "warnings": WARNINGS} + + + _sink = io.StringIO() + with contextlib.redirect_stdout(_sink): + _init() + RESULT = main() + print(json.dumps(RESULT, ensure_ascii=False)) + + - task_name: Task_C_B_domain_worker_* + max_concurrency: 8 + preflight_files: + - "{{item.compiled_prompt_path}}" + - "{{item.slice_path}}" + llm_provider: google + llm_model: 'gemini-3.1-pro-preview' + llm_reasoning: high + llm_verbosity: low + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 15m + prompts: + - role: user + content: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + + + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + + + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + + + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + + + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "juristic_act_type": "<법률행위 유형 문자열 또는 null>", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "party_roles": [], + "time_facts": [], + "object_refs": [], + "amount_facts": [], + "review_items": [], + "extensions": {"domain_payload": {"action_summary": null, "action_type": null}} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true + } + } + } + + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + - 아래 자리들은 BO 호환면 투영(`bo_surface_projection_policy.v1`)이 읽는 1순위 출처다. + 비워 두면 BO.json 의 해당 칸이 폴백 값으로 채워지고 schema_field_fallback 검토가 발행된다. + slice 의 source_universe 안에 근거가 있으면 채운다. 근거가 없으면 비워 두고 사유를 남긴다 — + 추측으로 채우지 않는다. 사건종류 이름을 값으로 쓰지 않는다. + · `juristic_act_type` : 법률행위 유형 문자열 1개(없으면 null). -> JuristicAct.label + · `extensions.domain_payload.action_summary` : 이 후보가 무엇인지 한 문장. -> Action + · `extensions.domain_payload.action_type` : "법률행위(legal acts)" 또는 "사실행위(factual acts)". -> ActionType + · `legal_effect_candidates[]` : {"type_id": "<소문자_스네이크>", "source_refs": [], "registered": true|false}. -> Legal_Keywords + · `time_facts[]` : {"fact_type": "<소문자_스네이크>", "value": "<시점 문자열 또는 null>", "source_refs": []}. -> BehaviorTime · TimeText + · `object_refs[]` : 목적물 식별자 문자열. -> core_field_base.Object + · `amount_facts[]` : {"amount_type": "<소문자_스네이크>", "decimal_value": "<숫자 문자열 또는 null>", "currency": "KRW", "source_refs": []}. -> amount + · `party_roles[]` : {"role": "<소문자_스네이크>", "party_refs": []}. 투영 대상은 아니나 스키마 필드다. + + + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + + + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + + use_tools: + - localdocs + - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_R0_seed_reducer_and_exception_planner (v3) + # publisher + domain_join + PostB_1 통합 결정적 reducer. + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # v4 — {{prev.Task_C_BO_Stage_B_B*}} 다섯을 걷어냈다. + # worker 산출은 wildcard fan-out 인스턴스가 파일로 남기므로 경로로 읽는다. + # v4 — seed 목록은 상수가 아니라 A0 의 fan-out 계획이 정한다. + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SLICE_DIR = "runtime/domain_slices" + SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + SLICE_ROOT_KEY = "stage_b_domain_slice" + # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + EXECUTION_ROOT = "/tmp/s1_r0" + VALIDATOR_MIRRORS = { + "runtime_common": "Default_Agent/stage1_runtime/runtime_common.txt", + "schema_subset_validator": "Default_Agent/stage1_runtime/schema_subset_validator.txt", + "registry_loader": "Default_Agent/stage1_runtime/registry_loader.txt", + "worker_output_validator": "Default_Agent/stage1_runtime/worker_output_validator.txt", + } + + def _seed_docs_from_plan(plan): + # v5 — 경로만이 아니라 계획 행 전체를 보관한다. validate_seed_object 가 + # slice_sha256 · compiled_prompt_sha256 기대값을 이 행에서 대조한다(R0-6). + root = plan.get("domain_fanout_plan", plan) + out = {} + rows = {} + for row in root.get("task_instances") or []: + domain_id = row.get("domain_id") + path = row.get("expected_output_path") + if isinstance(domain_id, str) and isinstance(path, str) and domain_id and path: + out[domain_id] = path + rows[domain_id] = row + if not out: + raise RuntimeError("R0_FANOUT_PLAN_EMPTY") + return out, rows + # v4 — 계획이 정하는 두 목록. 상수가 아니므로 비워 두고 main 에서 내용만 채운다. + # 재바인딩하지 않고 갱신만 하므로 아래 도우미들이 같은 객체를 본다. + SEED_DOCS: dict[str, str] = {} + PLAN_ROWS: dict[str, dict[str, Any]] = {} + DOMAIN_ORDER: list[str] = [] + + # DOMAIN_ORDER 는 fan-out 계획의 등재 순서를 그대로 쓴다. 상수 순서를 두지 않는다. + def _domain_order(seed_docs): + return list(seed_docs.keys()) + # v5 — 구 이름 표(DOMAIN_LABELS)와 _domain_label 을 걷어냈다. 유일 소비처가 되쓰기 + # (R0-5 에서 삭제)의 transport_metadata 였다. 이로써 R0 에 구 명세서(B1~B5) 이름 의존이 없다. + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + MANIFEST_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + # R0-1 — BO 투영 정책. 투영 규칙의 정본은 코드가 아니라 이 선언 자산이다. + BO_PROJECTION_POLICY = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + # v3 계약이 정본이다. 도메인 ID 는 registry 값(E-00 · EC-00 · X1 …)이고 구 이름도 아직 들어올 수 있으므로 + # 접두사는 도메인에 묶지 않고 형식만 본다 — 도메인 일치는 validate_candidate 의 prefix 검사가 맡는다. + CANDIDATE_REF_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_.-]{0,63}:[0-9]{3}$") + REVIEW_ISSUE_ENUM = { + "missing_source", "source_conflict", "cross_domain_merge_needed", + "amount_or_date_uncertain", "legal_effect_uncertain", "review_required", + "legal_theory_required", "near_duplicate_kept_separate", + "meeting_only_evidence_gap", "schema_field_fallback", "prior_link_ambiguous", + } + DOWNSTREAM_OWNER_ENUM = {"publisher", "domain_join", "C0", "C1", "C2", "C3", "C5", "D", "E", "Stage2"} + # v5 — ALLOWED_SEED_KEYS(v2 화이트리스트)를 걷어냈다. v3 후보 18필드와의 교집합이 + # extensions 하나뿐이라 워커 산출을 통째로 버리던 자리다(C-1). 원장 payload 의 + # 키 집합은 project_to_bo_surface 의 반환문이 유일한 정의다. + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-r0-seed-reducer-and-exception-planner", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 미러 해시 대조의 전제다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + def _assert_mirror_consistent(logical: str) -> None: + """F-4b — 부재는 통과(materialize_validator 의 기존 관용 유지). 미등재·불일치만 막는다. + + 예외 종류를 바꿔 try 를 뚫는 우회(SystemExit 등)는 쓰지 않는다. 그것은 __main__ 가드의 + stdout 출력과 예행 하네스의 단계 기록까지 건너뛴다. 판정을 try 밖으로 옮기는 것이 답이다. + """ + try: + body = read_raw(logical) + except Exception: + return + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + + + def materialize_validator() -> list[str]: + """worker_output_validator 와 그 의존 셋을 반입한다. 실패는 경고로 남기고 진행한다. + + 이 검증은 덧붙이는 층이다 — 반입이 안 되는 배포에서도 R0 본체는 돌아야 한다. + """ + import hashlib + import os + import pathlib + rt = pathlib.Path(EXECUTION_ROOT) / "_rt" + rt.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + staged: list[str] = [] + for name, logical in VALIDATOR_MIRRORS.items(): + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != hashlib.sha256(raw).hexdigest(): + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + (rt / (name + ".py")).write_bytes(raw) + staged.append(name) + if str(rt) not in sys.path: + sys.path.insert(0, str(rt)) + return staged + + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + def clip(value: Any, limit: int = 120) -> str: + text = " ".join(str(value or "").split()) + return text if len(text) <= limit else text[:limit].rstrip() + "..." + + # v4 — parse_llm_json 을 걷어냈다. worker 가 {{item.expected_output_path}} 에 + # strict JSON 파일을 직접 쓰므로 LLM 원문 관용 파싱 경로가 없어졌다. + # salvage_notes 는 산출 스키마에 남지만 v4 에서는 항상 빈 목록이다 — 구제할 원문이 없다. + # v5 — DOMAIN_PAYLOAD_CANON · _canon_payload 를 걷어냈다. canon 키가 구 이름(B1~B4)뿐이라 + # registry ID 26종 전부에서 no-op 였다(사문 코드). 확장 payload 는 워커 발행 형태 그대로 둔다. + def ensure_candidate_ref(cand: dict[str, Any], domain_id: str, idx: int) -> dict[str, Any]: + """v3 워커 출력에는 candidate_ref 가 없다 — seed 스키마가 additionalProperties: false 로 봉인돼 + 워커가 실을 수 없는 필드다. v3 이 주는 순번에서 R0 내부 식별자를 결정적으로 만든다. + 이미 실려 있으면(구 판본 산출) 그대로 둔다.""" + ref = cand.get("candidate_ref") + if isinstance(ref, str) and ref: + return cand + out = dict(cand) + out["candidate_ref"] = "%s:%03d" % (domain_id, idx + 1) + return out + + def project_to_bo_surface(cand: dict[str, Any], domain_id: str, universe: dict[str, set[str]], + policy: dict[str, Any], allowed_bo_types: set[str], + reviews: list[dict[str, Any]]) -> dict[str, Any]: + """v3 후보를 BO 호환면으로 투영한다. 값의 정본은 registry 이고 규칙은 정책 파일이 선언한다. + + 전임자 둘(expand_candidate + _seed_payload)은 v2 키를 기본값으로 깔고 v2 화이트리스트로 + 걸렀다. v3 후보를 넣으면 워커가 실은 값이 extensions 하나만 남았고, 그 결과 중복 판정 키 + 여덟 성분이 전부 비어 사건 전체가 한 버킷으로 접혔다(C-1·C-2). 여기서는 v3 필드에서 + 끌어오고, registry 가 말해 주지 않는 칸은 채우지 않고 reviews 에 올린다. + 반환 키 집합은 입력과 무관하게 고정이다 — 이 반환문이 원장 payload 키 집합의 유일한 정의다. + """ + ref = str(cand.get("candidate_ref")) + + def note(issue_type: str, field: str, source: str) -> None: + reviews.append({"issue_type": issue_type, "candidate_ref": ref, + "field": field, "source": source}) + + refs = _strings(cand.get("source_refs")) + evidence = sorted(set(refs) & universe["source_evidence_indexes"]) + events = sorted(set(refs) & universe["source_event_candidate_ids"]) + clauses = sorted(set(refs) & universe["source_meeting_clause_ids"]) + + norm = _dict(policy.get("f0_normalization")) + bo_type = cand.get("bo_type") + if allowed_bo_types and bo_type not in allowed_bo_types: + note("legal_effect_uncertain", "BOType", "bo_type") + + ext = dict(_dict(cand.get("extensions"))) + if not isinstance(ext.get("domain_payload"), dict): + ext["domain_payload"] = {} + domain_payload = _dict(ext.get("domain_payload")) + + action_type = domain_payload.get("action_type") + if not (isinstance(action_type, str) and action_type in set(_strings(norm.get("action_type_enum")))): + # registry 근거가 없는 칸이다. 기본값은 선언이며 추정이 아니다 — 반드시 검토로 올린다. + action_type = norm.get("action_type_default") + note("schema_field_fallback", "ActionType", "policy_default") + + effect_type_ids = sorted({str(e.get("type_id")).strip() + for e in _list(cand.get("legal_effect_candidates")) + if isinstance(e, dict) and str(e.get("type_id") or "").strip()}) + action_summary = domain_payload.get("action_summary") + if isinstance(action_summary, str) and action_summary.strip(): + action = action_summary.strip() + elif effect_type_ids: + # 값은 registry token 이지 서술문이 아니다. Stage 2 는 review_handoff 의 action_source 를 함께 읽는다. + action = "%s:%s" % (bo_type, effect_type_ids[0]) + note("schema_field_fallback", "Action", "legal_effect_type_id") + else: + action = str(bo_type) + note("schema_field_fallback", "Action", "bo_type") + + time_facts = [t for t in _list(cand.get("time_facts")) if isinstance(t, dict)] + behavior_time = None + time_text = None + if time_facts: + pick = sorted(time_facts, key=lambda t: (str(t.get("fact_type") or ""), str(t.get("value") or "")))[0] + behavior_time = pick.get("value") + time_text = pick.get("value") + distinct_times = {str(t.get("value") or "").strip() for t in time_facts if str(t.get("value") or "").strip()} + if len(distinct_times) > 1: + note("amount_or_date_uncertain", "core_field_base.BehaviorTime", "time_facts") + + object_refs = sorted(_strings(cand.get("object_refs"))) + + amount_facts = [a for a in _list(cand.get("amount_facts")) if isinstance(a, dict)] + amount = None + if amount_facts: + pick = sorted(amount_facts, key=lambda a: (str(a.get("amount_type") or ""), str(a.get("decimal_value") or "")))[0] + # v3 amount_facts 는 {amount_type, decimal_value, currency, source_refs} 닫힌 스키마다 — + # value_text 필드가 없으므로 정책 규칙대로 decimal_value 원문을 그대로 쓴다. + amount = {"value_text": pick.get("decimal_value"), + "numeric_value": pick.get("decimal_value"), + "currency": pick.get("currency")} + distinct_amounts = {str(a.get("decimal_value") or "").strip() for a in amount_facts if str(a.get("decimal_value") or "").strip()} + if len(distinct_amounts) > 1: + note("amount_or_date_uncertain", "amount", "amount_facts") + + return { + "candidate_ref": ref, + "source_domain": domain_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": _normalize_juristic(cand.get("juristic_act_type")), + "Action": action, + "Reason": None, + "PriorAct": None, + "ReasonRefs": [], + "Legal_Keywords": effect_type_ids, + "core_field_base": {"BehaviorTime": behavior_time, "TimeText": time_text, + "Object": object_refs[0] if object_refs else None, + "StatementType": bo_type}, + "amount": amount, + "source_evidence_indexes": evidence, + "provenance": {"source_event_candidate_ids": events, + "source_meeting_clause_ids": clauses, + "source_evidence_indexes": evidence, + "source_domain": domain_id}, + "downstream_seed_refs": {}, + "extensions": ext, + "registry_component_ids": _strings(cand.get("registry_component_ids")), + } + + def expand_review_item(item: Any, domain_id: str, idx: int) -> dict[str, Any]: + """v3 검토 항목(후보별 review_items · 루트 unknown_or_unrouted_reviews)을 handoff 항목형으로 사상한다. + + 구판은 v3 루트 13키에 없는 domain_review_queue 를 읽었다 — 봉인(additionalProperties: false)이 + 워커에게 발행을 금지한 키라 워커 검토가 한 건도 도달하지 못했다(C-6). 사상 규칙은 정책 + review_item_projection 이 선언한다. 원본 코드는 덧붙이기 키 source_review_code 로 보존한다. + """ + src = _dict(item) + raw_type = str(src.get("unresolved_type") or "").strip() + raw_code = str(src.get("review_code") or "").strip() + severity = src.get("severity") if src.get("severity") in ("SOFT_WARNING", "HARD_WARNING") else "SOFT_WARNING" + return { + "review_id": str(src.get("review_id") or f"{domain_id}:review:{idx:03d}"), + "issue_type": raw_type if raw_type in REVIEW_ISSUE_ENUM else "review_required", + "severity": severity, + # 원본 review_code(v3 필수 키)를 잃지 않는다 — 정책 additive_keys 의 목적이 그것이다. + "source_review_code": raw_code or raw_type or None, + "reason": str(src.get("reason") or "").strip(), + "source_refs": _strings(src.get("source_refs")), + "recommended_downstream_owner": src.get("recommended_downstream_owner") or "Stage2", + } + + # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. + + 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), + 신선도는 워커가 실을 수 없는 transport_metadata.slice_guard 를 읽는 죽은 검사였다. + 신선도의 제 필드는 v3 루트의 slice_sha256 · compiled_prompt_sha256 이고(둘 다 required + — 워커가 반드시 echo 한다), 기대값은 fan-out 계획 행이 든다. + """ + if seed_obj.get("schema_version") != SEED_SCHEMA_VERSION: + raise ValueError(f"{domain_id}: seed schema_version mismatch") + if seed_obj.get("domain_id") != domain_id: + raise ValueError(f"{domain_id}: seed domain_id mismatch") + status = seed_obj.get("status") + if status in ("BLOCKED", "FAILED"): + # 워커 실패 신호다. fail-open 은 활성화 판정의 원칙이고, 실패의 침묵 흡수는 금지 원칙이 막는다. + raise ValueError(f"{domain_id}: worker reported {status}") + if status == "NO_SUPPORT": + # 적법한 "실을 것 없음". 후보가 있으면 상태·내용 모순이다. + if _list(seed_obj.get("bo_seed_candidates")): + raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") + elif status not in ("READY", "READY_WITH_REVIEW"): + raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + for key in ("slice_sha256", "compiled_prompt_sha256"): + want = plan_row.get(key) + if isinstance(want, str) and want: + if seed_obj.get(key) != want: + raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + else: + warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) + + def validate_candidate(cand: dict[str, Any], domain_id: str, idx: int) -> None: + prefix = domain_id + ref = cand.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref invalid") + if not ref.startswith(prefix + ":"): + raise ValueError(f"{domain_id}.bo_seed_candidates[{idx}].candidate_ref prefix mismatch") + for forbidden in ("BO_ID", "id", "Evidence", "EvidenceTitles"): + if forbidden in cand: + raise ValueError(f"{domain_id}.{ref}: final field {forbidden} is prohibited") + if cand.get("Reason") is not None: + raise ValueError(f"{domain_id}.{ref}: Reason must be null/absent") + if cand.get("PriorAct") is not None: + raise ValueError(f"{domain_id}.{ref}: PriorAct must be null/absent") + if cand.get("ReasonRefs") not in ([], None): + raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- + def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: + provenance = _dict(seed.get("provenance")) + return { + "source_evidence_indexes": _strings(seed.get("source_evidence_indexes") or provenance.get("source_evidence_indexes")), + "source_event_candidate_ids": _strings(provenance.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(provenance.get("source_meeting_clause_ids")), + } + + def _sort_key(seed: dict[str, Any]) -> dict[str, Any]: + core = _dict(seed.get("core_field_base")) + domain = seed.get("source_domain") + juristic = _dict(seed.get("JuristicAct")) + return { + "BehaviorTime": core.get("BehaviorTime"), + "domain_order": DOMAIN_ORDER.index(domain) if domain in DOMAIN_ORDER else len(DOMAIN_ORDER), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicActLabel": juristic.get("label"), + "Action": seed.get("Action"), + "candidate_ref": seed.get("candidate_ref"), + } + + def _duplicate_key(seed: dict[str, Any]) -> tuple[Any, ...]: + core = _dict(seed.get("core_field_base")) + juristic = _dict(seed.get("JuristicAct")) + refs = _source_refs(seed) + return ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + juristic.get("label"), + str(seed.get("Action") or "").strip(), + str(core.get("BehaviorTime") or "").strip(), + str(core.get("Object") or "").strip(), + ) + + def _normalize_juristic(value: Any) -> Any: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _compact_exception_payload(seeds: list[dict[str, Any]]) -> list[dict[str, Any]]: + compact = [] + for seed in seeds: + compact.append({ + "candidate_ref": seed.get("candidate_ref"), + "source_domain": seed.get("source_domain"), + "BOType": seed.get("BOType"), + "ActionType": seed.get("ActionType"), + "JuristicAct": seed.get("JuristicAct"), + "Action": seed.get("Action"), + "core_field_base": seed.get("core_field_base"), + "amount": seed.get("amount"), + "source_refs": _source_refs(seed), + }) + return compact + + def main() -> None: + _init() + # F-4a — 자기 정적 입력. try 밖이어야 한다. 안에 넣으면 아래 except Exception 이 + # 삼켜 WORKER_VALIDATOR_UNAVAILABLE 경고로 강등되고 R0 이 계속 돈다. + _seed_schema_body = _verify_asset(SEED_SCHEMA_PATH) + # R0-1 — 투영 정책 반입 (F-4a 와 같은 규율: try 밖 경성). 정책이 없거나 낡았는데 + # 조용히 옛 규칙으로 도는 것이 이번 결손(v2 잔재)의 재발 경로다. + projection_policy = _dict(json.loads(_verify_asset(BO_PROJECTION_POLICY))) + if projection_policy.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("PART2_PROJECTION_POLICY_INVALID") + # F-4b — 미러 넷의 무결성. 부재는 통과시키고 미등재·불일치만 막는다. + for _mirror in VALIDATOR_MIRRORS.values(): + _assert_mirror_consistent(_mirror) + salvage_notes: list[dict[str, Any]] = [] + guard_warnings: list[dict[str, Any]] = [] + # v4 — seed 목록과 그 순서는 A0 의 fan-out 계획이 정한다. 이 파일은 목록을 만들지 않는다. + worker_validator = None + seed_schema = None + try: + materialize_validator() + import worker_output_validator as worker_validator + seed_schema = json.loads(_seed_schema_body) + except Exception as exc: + guard_warnings.append({"code": "WORKER_VALIDATOR_UNAVAILABLE", "message": str(exc)[:200]}) + worker_validator = None + _docs, _rows = _seed_docs_from_plan(_dict(read_json_doc(FANOUT_PLAN_PATH))) + SEED_DOCS.update(_docs) + PLAN_ROWS.update(_rows) + DOMAIN_ORDER.extend(_domain_order(SEED_DOCS)) + stage_a_outer = read_json_doc(STAGE_A_PATH) + stage_a = _dict(_dict(stage_a_outer).get("stage_a_context") or stage_a_outer) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + raise RuntimeError("Stage A context must be READY task_c_bo_stage_a_context.v1") + manifest = _dict(read_json_doc(MANIFEST_PATH)) + universe = { + "source_event_candidate_ids": set(_strings(manifest.get("event_candidate_ids"))), + "source_evidence_indexes": set(_strings(manifest.get("evidence_index_set"))), + "source_meeting_clause_ids": set(_strings(manifest.get("meeting_clause_ids"))), + } + if not universe["source_evidence_indexes"]: + raise RuntimeError("source universe manifest has no evidence indexes") + + # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — + # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + seed_objects: dict[str, dict[str, Any]] = {} + projected_candidates: dict[str, list[dict[str, Any]]] = {} + review_handoff_items: list[dict[str, Any]] = [] + allowed_bo_types_by_domain: dict[str, set[str]] = {} + projection_review_counter = 0 + for domain_id in DOMAIN_ORDER: + # v4 — worker 가 {{item.expected_output_path}} 에 자기 seed 를 직접 쓴다. + # {{prev}} 원문 관용 파싱이 아니라 계획이 정한 경로에서 읽는다. + outer = _dict(read_json_doc(SEED_DOCS[domain_id])) + seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) + if not seed_obj: + raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") + validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 + # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} + allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) + allowed_bo_types_by_domain[domain_id] = allowed_bo_types + # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + if worker_validator is not None: + report = worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, + schema=seed_schema, + expected_domain_id=domain_id, + slice_document=slice_doc) + for item in report.get("errors") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, + "detail": item}) + for item in report.get("warnings") or []: + guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, + "detail": item}) + cands = _list(seed_obj.get("bo_seed_candidates")) + projected: list[dict[str, Any]] = [] + projection_reviews: list[dict[str, Any]] = [] + for idx, cand in enumerate(cands): + if not isinstance(cand, dict): + raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") + cand = ensure_candidate_ref(cand, domain_id, idx) + validate_candidate(cand, domain_id, idx) + # membership 검사 — v3 의 평평한 source_refs 를 universe 와 대조 (hard BLOCK). + # 이쪽을 보지 않으면 membership 게이트가 v3 산출에서는 통과만 하는 빈 검사가 된다. + known_sources = (universe["source_event_candidate_ids"] | universe["source_meeting_clause_ids"] + | universe["source_evidence_indexes"]) + ref_bad = [v for v in _strings(cand.get("source_refs")) if v not in known_sources] + if ref_bad: + raise RuntimeError(f"BLOCK: {domain_id}.{cand.get('candidate_ref')}: source_refs outside Stage A universe: {ref_bad}") + projected.append(project_to_bo_surface(cand, domain_id, universe, projection_policy, + allowed_bo_types, projection_reviews)) + # R0-5 — 되쓰기 없음. seed_objects 는 워커 원본 그대로다(S0 와 signal adapter 가 + # v3 적합 원본을 읽는다). 투영본은 projected_candidates 가 따로 든다(R0-2 배선). + seed_objects[domain_id] = seed_obj + projected_candidates[domain_id] = projected + # R0-4 — v3 검토 채널: 후보별 review_items + 루트 unknown_or_unrouted_reviews. + # list(...) 복사는 워커 원본 목록을 제자리 변형하지 않기 위한 것이다. + worker_reviews = list(_list(seed_obj.get("unknown_or_unrouted_reviews"))) + for cand in _list(seed_obj.get("bo_seed_candidates")): + worker_reviews.extend(_list(_dict(cand).get("review_items"))) + if seed_obj.get("status") == "NO_SUPPORT": + worker_reviews.append({"review_id": f"{domain_id}:status:NO_SUPPORT", + "review_code": "NO_SUPPORT", + "unresolved_type": "review_required", + "severity": "SOFT_WARNING", + "reason": "worker reported NO_SUPPORT (nothing to carry for this domain)"}) + for idx, item in enumerate(worker_reviews, start=1): + mapped = expand_review_item(item, domain_id, idx) + refs = set(mapped.get("source_refs") or []) + review_handoff_items.append({ + "review_id": mapped["review_id"], + "source_domain": domain_id, + "severity": mapped["severity"], + "issue_type": mapped["issue_type"], + "source_review_code": mapped.get("source_review_code"), + "source_event_candidate_ids": sorted(refs & universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(refs & universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(refs & universe["source_meeting_clause_ids"]), + "downstream_owner": mapped["recommended_downstream_owner"] if mapped.get("recommended_downstream_owner") in DOWNSTREAM_OWNER_ENUM else "Stage2", + "template_note": mapped.get("reason") or "후속 단계에서 해당 review 항목의 증거와 법률상 의미를 재검토한다.", + }) + for note_item in projection_reviews: + projection_review_counter += 1 + entry = { + "review_id": "R0:projection:%03d" % projection_review_counter, + "source_domain": domain_id, + "severity": "SOFT_WARNING", + "issue_type": note_item["issue_type"], + "source_review_code": note_item.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "투영 규칙이 채우지 못했거나 기본값을 적용한 칸이다: %s ← %s (%s)" % ( + note_item.get("field"), note_item.get("source"), note_item.get("candidate_ref")), + } + if note_item.get("field") == "Action": + entry["action_source"] = note_item.get("source") + review_handoff_items.append(entry) + + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). + # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. + input_candidate_total = 0 + seeds: list[dict[str, Any]] = [] + for domain_id in DOMAIN_ORDER: + projected = projected_candidates[domain_id] + input_candidate_total += len(projected) + seeds.extend(projected) + if not seeds: + raise RuntimeError("no seed candidate from Stage B workers") + + seen_refs: set[str] = set() + ledger_candidates: list[dict[str, Any]] = [] + deterministic_decisions: list[dict[str, Any]] = [] + exceptions: list[dict[str, Any]] = [] + duplicate_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + policy_review_counter = 0 + + def add_policy_review(domain_id: str, issue_type: str, refs: dict[str, list[str]], severity: str = "SOFT_WARNING") -> None: + nonlocal policy_review_counter + policy_review_counter += 1 + review_handoff_items.append({ + "review_id": f"R0:policy:{policy_review_counter:03d}", + "source_domain": domain_id, + "severity": severity, + "issue_type": issue_type, + "source_event_candidate_ids": refs.get("source_event_candidate_ids", []), + "source_evidence_indexes": refs.get("source_evidence_indexes", []), + "source_meeting_clause_ids": refs.get("source_meeting_clause_ids", []), + "downstream_owner": "Stage2", + "template_note": "결정적 defer 정책에 의해 보존된 검토 항목이다.", + }) + + for seed in seeds: + ref = seed.get("candidate_ref") + if not isinstance(ref, str) or not CANDIDATE_REF_RE.fullmatch(ref): + raise RuntimeError(f"invalid candidate_ref: {ref!r}") + if ref in seen_refs: + raise RuntimeError(f"duplicate candidate_ref: {ref}") + seen_refs.add(ref) + refs = _source_refs(seed) + # hard BLOCK: universe 밖 ref는 soft-strip 이후에도 남아있으면 안 된다 (방어적 재검사) + for key, values in refs.items(): + allowed = universe.get(key, set()) + outside = [v for v in values if allowed and v not in allowed] + if outside: + raise RuntimeError(f"{ref}: {key} outside Stage A universe: {outside}") + flags: list[str] = [] + # 결정적 defer 정책 (개선전략서 X-2): + # C-2 (P0 판정 B) — 증거 단계에서 붙은 registry 구성요소는 증거 유래 근거다. + # _source_refs 가 seed 루트와 provenance 를 모두 보는 관례를 그대로 따른다. + registry_components = [ + str(value) + for value in (seed.get("registry_component_ids") + or _dict(seed.get("provenance")).get("registry_component_ids") + or []) + if isinstance(value, str) and value + ] + if not refs["source_evidence_indexes"] and not registry_components: + flags.append("meeting_only_evidence_gap") + add_policy_review(seed.get("source_domain"), "meeting_only_evidence_gap", refs) + domain_allowed = allowed_bo_types_by_domain.get(str(seed.get("source_domain"))) or set() + if (domain_allowed and seed.get("BOType") not in domain_allowed) or not seed.get("ActionType") or not ( + seed.get("Action") or _dict(_dict(seed.get("extensions")).get("domain_payload")).get("action_summary") + ): + flags.append("schema_field_fallback") + add_policy_review(seed.get("source_domain"), "schema_field_fallback", refs) + link_candidates = _strings(_dict(seed.get("downstream_seed_refs")).get("prior_candidate_refs")) + if len(link_candidates) > 1: + flags.append("prior_link_ambiguous") + add_policy_review(seed.get("source_domain"), "prior_link_ambiguous", refs) + duplicate_buckets.setdefault(_duplicate_key(seed), []).append(seed) + ledger_candidates.append({ + "candidate_ref": ref, + "source_domain": seed.get("source_domain"), + "seed_payload": seed, + "source_refs": refs, + "deterministic_sort_key": _sort_key(seed), + "flags": flags, + }) + + # exact duplicate: provenance union 무손실이므로 canonical merge (v2 규칙 계승) + for bucket in duplicate_buckets.values(): + if len(bucket) <= 1: + continue + canonical = bucket[0].get("candidate_ref") + duplicates = [item.get("candidate_ref") for item in bucket[1:]] + deterministic_decisions.append({ + "decision_type": "EXACT_DUPLICATE_MERGE", + "canonical_candidate_ref": canonical, + "duplicate_candidate_refs": duplicates, + "basis": "exact duplicate deterministic rule (provenance-lossless union)", + }) + + # near duplicate: KEEP_SEPARATE + cluster id + review (LLM 금지 — defer 정책) + near_buckets: dict[tuple[Any, ...], list[dict[str, Any]]] = {} + for seed in seeds: + refs = _source_refs(seed) + key = ( + tuple(sorted(refs["source_evidence_indexes"])), + tuple(sorted(refs["source_event_candidate_ids"])), + seed.get("BOType"), + seed.get("ActionType"), + ) + near_buckets.setdefault(key, []).append(seed) + near_cluster_count = 0 + pack_field_conflicts: list[dict[str, Any]] = [] + for bucket in near_buckets.values(): + if len(bucket) <= 1 or len({_duplicate_key(s) for s in bucket}) <= 1: + continue + near_cluster_count += 1 + cluster_id = f"near-dup-{near_cluster_count:03d}" + cluster_refs = [str(s.get("candidate_ref")) for s in bucket] + for item in ledger_candidates: + if item["candidate_ref"] in cluster_refs: + item.setdefault("near_dup_cluster_id", cluster_id) + if "near_duplicate_kept_separate" not in item["flags"]: + item["flags"].append("near_duplicate_kept_separate") + add_policy_review(bucket[0].get("source_domain"), "near_duplicate_kept_separate", + {"source_evidence_indexes": _source_refs(bucket[0])["source_evidence_indexes"], + "source_event_candidate_ids": _source_refs(bucket[0])["source_event_candidate_ids"], + "source_meeting_clause_ids": []}) + # non-deferrable 판정(X-3 4중 조건): 같은 near cluster에서 BehaviorTime 또는 amount가 + # 서로 다른 non-null 값으로 충돌하면 writer가 단일 값을 고를 수 없으므로 pack에 수록 + times = {str(_dict(s.get("core_field_base")).get("BehaviorTime")) for s in bucket if _dict(s.get("core_field_base")).get("BehaviorTime")} + amounts = set() + for s in bucket: + av = s.get("amount") + if isinstance(av, dict) and av.get("value_text"): + amounts.add(str(av.get("value_text"))) + elif isinstance(av, str) and av.strip(): + amounts.add(av.strip()) + if len(times) > 1 or len(amounts) > 1: + pack_field_conflicts.append({ + "exception_id": f"EX-FIELD-{len(pack_field_conflicts) + 1:03d}", + "exception_type": "field_conflict", + "candidate_refs": cluster_refs, + "reason": "same-source candidates carry conflicting BehaviorTime/amount values", + "conflicting_values": {"BehaviorTime": sorted(times), "amount": sorted(amounts)}, + "compact_candidate_payload": _compact_exception_payload(bucket), + "allowed_decisions": ["KEEP_SEPARATE", "MERGE", "SPLIT", "DROP", "BLOCK_REVIEW"], + "escalation_flag": True, + }) + + exceptions.extend(pack_field_conflicts) + has_exceptions = bool(exceptions) + + # 3) conservation invariant (write 전) + merged_absorbed = sum(len(_strings(d.get("duplicate_candidate_refs"))) for d in deterministic_decisions) + if len(ledger_candidates) != input_candidate_total: + raise RuntimeError(f"ledger candidate count {len(ledger_candidates)} != input candidates {input_candidate_total}") + if len(seen_refs) != input_candidate_total: + raise RuntimeError("candidate_ref conservation failed") + + ledger = { + "postb_seed_ledger": { + "schema_version": "task_c_bo_postb_seed_ledger.v1", + "status": "READY", + "source_stage_a_created_at_utc": stage_a.get("created_at_utc"), + "input_digests_sha256": stage_a.get("input_digests_sha256"), + "source_universe": { + "source_event_candidate_ids": sorted(universe["source_event_candidate_ids"]), + "source_evidence_indexes": sorted(universe["source_evidence_indexes"]), + "source_meeting_clause_ids": sorted(universe["source_meeting_clause_ids"]), + }, + "stage_b_source_contract": { + "schema_version": "task_c_bo_stage_b_bo_seed_universe.compat_from_r0.v1", + "status": "READY", + "compatibility_source": "r0_seed_reducer.direct_worker_outputs", + }, + "ledger_candidates": sorted(ledger_candidates, key=lambda item: ( + item["deterministic_sort_key"].get("BehaviorTime") is None, + item["deterministic_sort_key"].get("BehaviorTime") or "", + item["deterministic_sort_key"].get("domain_order", 99), + item["deterministic_sort_key"].get("BOType") or "", + item["deterministic_sort_key"].get("ActionType") or "", + item["deterministic_sort_key"].get("JuristicActLabel") or "", + item["deterministic_sort_key"].get("Action") or "", + item["deterministic_sort_key"].get("candidate_ref") or "", + )), + "deterministic_decisions": deterministic_decisions, + "exception_pack": { + "has_exceptions": has_exceptions, + "clusters": [], + "field_conflicts": pack_field_conflicts, + "link_ambiguities": [], + "schema_risks": [], + }, + "audit_trace": { + "removed_or_sidecar_fields": [], + "source_membership_policy": "outside-universe source ref => hard BLOCK (defer 정책 §7)", + "normalization_notes": salvage_notes + guard_warnings, + }, + } + } + write_doc(LEDGER_PATH, json.dumps(ledger, ensure_ascii=False, indent=2)) + + pack = { + "schema_version": "stage1_part2_exception_pack.v1", + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "exceptions": exceptions, + "budget": {"max_candidates_per_exception": 8, "max_payload_chars_per_candidate": 2000}, + } + write_doc(PACK_PATH, json.dumps(pack, ensure_ascii=False, indent=2)) + + handoff = { + "schema_version": "stage1_part2_review_handoff.v1", + "status": "PENDING_FINALIZE", + "review_items": review_handoff_items, + } + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "READY", + "message": "R0 seed reducer 완료: ledger/exception pack/review handoff 생성", + "ledger_path": LEDGER_PATH, + "exception_pack_path": PACK_PATH, + "review_handoff_path": REVIEW_HANDOFF_PATH, + "has_exceptions": has_exceptions, + "exception_count": len(exceptions), + "candidate_counts": { + "input": input_candidate_total, + "ledger": len(ledger_candidates), + "exact_duplicate_absorbed": merged_absorbed, + "near_dup_clusters": near_cluster_count, + }, + "review_item_count": len(review_handoff_items), + "salvage_count": len(salvage_notes), + }, ensure_ascii=False)) + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "R0 seed reducer 실패: downstream 진행 금지", + "reason": str(exc)}, ensure_ascii=False)) + raise + + - task_name: Task_C_BO_R1_exception_adjudicator + llm_provider: google + llm_model: gemini-3.1-flash-lite + llm_reasoning: low + llm_verbosity: low + max_iterations: 1 + use_tools: + - localdocs + cache_control: + mode: auto + ttl: 20m + preflight: true + preflight_files: + - quality_gates/stage1_part2_exception_pack.json + prompts: + - role: user + content: |- + + + - You are executing one LLM sub-task inside Stage 1 of a Korean civil-litigation complaint-generation pipeline. + - The current task's static block, role overlay, assigned inputs, output schema, and writer boundary control. + - This common prefix cannot expand the current task's input set, output set, legal domain, validation authority, writer authority, or reasoning depth. + - If any common rule appears broader than the current task, apply only the narrower current-task version. + - Stage 1 prepares verified structured artifacts. Do not draft complaint prose or final counsel-level conclusions unless the current task explicitly authorizes a validation or gate conclusion. + + + + - Use only assigned files, provided context inputs, prior outputs, and allowed tools. + - Do not import facts, law, procedural history, parties, dates, amounts, IDs, document contents, or source meanings from memory, outside knowledge, or unassigned files. + - Treat prior outputs as authority only to the extent the current task names them or provides them as context. + - If a value is unsupported, missing, conflicting, stale, or out of scope, use only the current schema's allowed null, empty, unknown, warning, blocked, or needs_review path. + + + + - Preserve exact source identifiers required by the current schema. + - Maintain separation among raw fact, inferred fact, legal signal, evidence support, fact support, validation issue, and final gate decision when the current schema distinguishes them. + - Do not upgrade meeting-only or indirect material into direct proof. + - Do not silently resolve material conflicts. If the current schema has a conflict or uncertainty field, use it; otherwise stay within the task's allowed warning or review path. + + + + - Follow required JSON shape, key names, enum values, ordering, file names, and status strings exactly. + - Do not add arbitrary keys, prose, markdown fences, alternative files, unauthorized repair, or explanatory material outside allowed fields. + - Create, mutate, normalize, merge, or finalize IDs only when the current task explicitly authorizes it. + - Write final files only when the current task is the authorized writer. Validators and guards report issues in their own authorized schema and do not silently repair unless instructed. + + + + - Prefer the current prompt and schema, assigned structured upstream artifacts, compact indexes, ledgers, manifests, bundles, and gates. + - Read raw evidence or meeting text only when the current task requires direct provenance, ambiguity resolution, or a schema-required value missing from structured artifacts. + - For map or projection tasks, process only the assigned item, domain, or batch. Reducers aggregate only the inputs assigned to them. + - Do not restate, summarize, cite, or copy this common prefix in any output. + + + + - Return only the requested structured artifact, concise allowed rationale fields, validation notes, or status object. + - Keep chain-of-thought private. + - Stop when the current schema is complete and safe. + + + + + + TASK_NAME: Task_C_BO_R1_exception_adjudicator + STAGE: PostB conditional exception adjudicator (Part 1 v3 GB 패턴) + MISSION: 결정적 reducer(R0)가 non-deferrable로 판정한 compact exception만 판정한다. 병합·최종 파일 작성·사실 창작은 하지 않는다. + + + + - 유일한 입력은 preflight로 제공된 `quality_gates/stage1_part2_exception_pack.json`이다. + - Stage A context, seed ledger 전문, raw evidence, meeting 원문을 읽거나 요청하지 않는다. + - pack에 없는 exception_id·candidate_ref·bh# id를 창작하지 않는다. + - BO.json, ledger, review handoff, signal 파일을 작성하지 않는다. + - 출력 파일은 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json` 하나뿐이다. + + + + 1. preflight로 제공된 exception pack의 `has_exceptions`를 확인한다. + 2. `has_exceptions == false`이면: `write_file(overwrite=true)`로 아래 no-exception 객체를 `stage1_tmp/task_c_bo/postb_adjudication_decisions.json`에 저장하고 `"NO_EXCEPTIONS"`만 출력한 뒤 즉시 종료한다(terminate). 다른 어떤 파일도 읽지 않는다. + {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", "status": "READY_NO_EXCEPTIONS", "exception_count": 0, "canonical_decisions": [], "field_decisions": [], "link_decisions": [], "semantic_gate_decisions": [], "blocked_review_items": []}} + 3. `has_exceptions == true`이면: 각 exception을 `compact_candidate_payload`만으로 판정한다. 추가 read는 금지된다. + 4. 판정 규칙: + - canonical decision: KEEP_SEPARATE, MERGE, SPLIT, DROP, BLOCK_REVIEW 중 하나. MERGE는 `input_candidate_refs`와 `merge_target_ref`를 명시한다. + - field conflict: 제공된 conflicting_values 중 하나를 `selected_value`로 선택하거나 BLOCK_REVIEW. 임의 값 창작 금지. + - link ambiguity: exception에 나열된 candidate ref 중 선택, NO_LINK, 또는 BLOCK_REVIEW. + - semantic risk: PASS, WARNING, BLOCK_REVIEW. + - compact payload로 확정할 수 없으면 반드시 `blocked_review_items`에 넣는다(확신 없는 확정 금지 — 인간 검토 라우팅). + 5. `write_file(overwrite=true)`로 결과를 저장한다. root는 `postb_exception_adjudication`이며 schema_version은 `task_c_bo_postb_exception_adjudication.v1`, status는 `READY`, `exception_count`는 판정한 exception 수다. 모든 decision은 pack의 `exception_id`를 인용한다. + 6. `"R1 예외 판정 완료 (decisions=<건수>)"`만 출력하고 작업을 끝낸다(terminate). + + + + - Stage 1은 법률효과·청구원인을 확정하지 않는다. 두 값을 모두 보존하거나 Stage 2로 defer할 수 있는 사안은 이미 R0가 결정적으로 처리했으므로, 여기 도달한 항목은 final writer가 단일 값을 선택해야만 진행되는 사안이다. + - 같은 source에 근거한 상충 값(BehaviorTime·amount)은: 원문 근거가 더 구체적인 쪽(payload의 core_field_base·amount 기재가 더 완전한 후보)을 선택하고, 우열을 가릴 수 없으면 BLOCK_REVIEW. + - KEEP_SEPARATE가 provenance를 보존하는 기본값이다. MERGE는 provenance 합집합이 무손실일 때만 선택한다. + - DROP은 어떤 경우에도 source 유일 후보에 적용하지 않는다. + + + + - exception pack 부재·파싱 불가: 즉시 중단하고 채팅으로만 보고한다. decisions 파일은 쓰지 않는다. + - tool 오류: 1회만 재시도. 재실패 시 `FAILED: `만 보고하고 종료한다. + + + - task_name: Task_C_BO_F0_final_bo_compiler_gate_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 240 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_F0_final_bo_compiler_gate_writer (v3) + # PostB_3(final compiler) + PostB_4(final gate/writer) 통합. 입력은 파일 계약(ledger/decisions/stage_a). + # Spec: Part_2_Improvement_Strategy_Claude_v1.md §9 (bh# 규칙 N-6, Reason/PriorAct 정책 R-5) + from __future__ import annotations + import hashlib + import itertools + import json + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + STAGE_A_PATH = "stage1_tmp/task_c_bo/stage_a_context.json" + LEDGER_PATH = "stage1_tmp/task_c_bo/postb_seed_ledger.json" + DECISIONS_PATH = "stage1_tmp/task_c_bo/postb_adjudication_decisions.json" + REVIEW_HANDOFF_PATH = "quality_gates/stage1_part2_review_handoff.json" + EXCEPTION_PACK_PATH = "quality_gates/stage1_part2_exception_pack.json" + BUNDLE_COMPACT_PATH = "stage1_tmp/task_c_bo/postb_compiled_bundle_compact.json" + TARGET_NAME = "BO.json" + + # v4 신설 — BOType 어휘와 확장 payload 선언의 정본은 registry 다. 코드에 어휘를 두지 않는다. + # registry 를 런타임에 적재하지 않는다. 그러려면 index 1 + domain_config 26 + extension schema 26 + # 을 읽어야 하고 그것은 읽기 53회다. 값이 사건마다 달라지지 않으므로 배포 시점에 한 번 + # 접어 둔 자산 하나만 읽는다. 생성기는 routing/_build_extension_payload_declarations.py 다. + EXTENSION_DECLARATIONS_PATH = "Default_Agent/routing/extension_payload_key_declarations.v1.json" + RUNTIME_MANIFEST_PATH = "Default_Agent/runtime_manifest.json" + DECLARATIONS_SCHEMA_VERSION = "stage1_extension_payload_key_declarations.v1" + BO_TYPE_SOURCE = "registry_union" + UNDECLARED_KEY_REVIEW_CODE = "EXTENSION_PAYLOAD_KEY_UNDECLARED" + # F0-2 — BO 투영 정책 (정규화 기본값의 정본). R0 와 같은 자산을 읽는다. + BO_PROJECTION_POLICY_PATH = "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json" + BO_PROJECTION_SCHEMA_VERSION = "stage1_bo_surface_projection_policy.v1" + + BH_ID_RE = re.compile(r"^bh[1-9][0-9]*$") + ACTION_TYPE_ENUM = { + "법률행위(legal acts)", + "준법률행위(quasi-legal acts)", + "사실행위(factual acts)", + "위법행위(unlawful acts)", + "소송행위(litigation acts)", + } + ALLOWED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", "amount", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", "extensions", + } + REQUIRED_TOP_LEVEL = { + "BO_ID", "id", "BOType", "ActionType", "JuristicAct", "Action", "Reason", + "PriorAct", "ReasonRefs", "Legal_Keywords", "core_field_base", + "EvidenceTitles", "Evidence", "source_evidence_indexes", "provenance", + "downstream_seed_refs", + } + CORE_KEYS = [ + "Performer", "PerformerType", "Action_proposal", "Subject", "Object", + "BehaviorTime", "TimeText", "TimePrecision", "StatementType", "Perspective", + ] + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-f0-final-bo-compiler-gate-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + def read_raw(name: str) -> str: + # 재직렬화 없이 원문 그대로 돌려준다. 해시 대조의 전제다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + # ---------- v4 신설: registry 선언 조회 ---------- + def _load_declarations() -> dict[str, Any]: + # 어휘의 정본이므로 훼손되면 BOType 검증이 조용히 넓어진다. + # 원문 바이트의 sha256 을 runtime_manifest 와 대조한 뒤에만 쓴다(D0 반입 규약과 같은 규율). + body = read_raw(EXTENSION_DECLARATIONS_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + EXTENSION_DECLARATIONS_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_EXTENSION_DECLARATIONS_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != DECLARATIONS_SCHEMA_VERSION: + raise RuntimeError("F0_EXTENSION_DECLARATIONS_SCHEMA_MISMATCH") + if not doc.get("bo_types_union"): + raise RuntimeError("F0_REGISTRY_BO_TYPES_EMPTY") + if not doc.get("declared_key_union"): + raise RuntimeError("F0_EXTENSION_DECLARED_KEYS_EMPTY") + return doc + + def _load_projection_policy() -> dict[str, Any]: + # F0-2 — 정규화 기본값·어휘의 정본. _load_declarations 와 같은 규율로 sha256 대조 후에만 쓴다. + body = read_raw(BO_PROJECTION_POLICY_PATH) + manifest = json.loads(read_raw(RUNTIME_MANIFEST_PATH)) + want = {row["path"]: row["sha256"] for row in manifest.get("entries") or []}.get( + BO_PROJECTION_POLICY_PATH[len("Default_Agent/"):]) + if want is None: + raise RuntimeError("F0_PROJECTION_POLICY_UNREGISTERED") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise RuntimeError("F0_PROJECTION_POLICY_HASH_MISMATCH") + doc = json.loads(body) + if not isinstance(doc, dict) or doc.get("schema_version") != BO_PROJECTION_SCHEMA_VERSION: + raise RuntimeError("F0_PROJECTION_POLICY_SCHEMA_MISMATCH") + if not isinstance(doc.get("f0_normalization"), dict): + raise RuntimeError("F0_PROJECTION_POLICY_NORMALIZATION_MISSING") + return doc + + def _declared_bo_types(declarations: dict[str, Any]) -> set[str]: + return {str(v) for v in declarations.get("bo_types_union") or [] if isinstance(v, str) and v} + + def _resolve_domain_ids(declarations: dict[str, Any], source_domain: Any) -> list[str]: + """source_domain 은 도메인 ID 이거나 구 이름(B1~B5)이다. 구 이름은 별칭표 primary 로 옮긴다.""" + name = str(source_domain or "").strip() + if not name: + return [] + known = {str(row.get("domain_id")) for row in declarations.get("domains") or []} + if name in known: + return [name] + targets = _dict(declarations.get("legacy_alias_targets")).get(name) + return [str(v) for v in targets or [] if str(v) in known] + + def _declared_keys_for(declarations: dict[str, Any], domain_ids: list[str]) -> set[str]: + """도메인을 특정하지 못하면 전체 합집합을 상대로 한다. 좁히지 못한 것을 위반으로 세지 않는다.""" + if not domain_ids: + return {str(v) for v in declarations.get("declared_key_union") or []} + wanted = set(domain_ids) + out: set[str] = set() + for row in declarations.get("domains") or []: + if str(row.get("domain_id")) in wanted: + out.update(str(v) for v in row.get("declared_keys") or []) + return out + + def _extension_key_reviews(bo_items: list[dict[str, Any]], declarations: dict[str, Any]) -> list[dict[str, Any]]: + """확장 payload 키를 registry 선언과 대조한다. 선언 밖 키는 review 로 남기고 값은 지우지 않는다.""" + reviews: list[dict[str, Any]] = [] + for item in bo_items: + payload = _dict(_dict(item.get("extensions")).get("domain_payload")) + if not payload: + continue + source_domain = _dict(item.get("provenance")).get("source_domain") + domain_ids = _resolve_domain_ids(declarations, source_domain) + undeclared = sorted(set(payload) - _declared_keys_for(declarations, domain_ids)) + if undeclared: + reviews.append({ + "bo_id": item.get("BO_ID"), + "source_domain": source_domain, + "resolved_domain_ids": domain_ids, + "resolution": "registry_domain_ids" if domain_ids else "declared_key_union_fallback", + "undeclared_keys": undeclared, + "review_code": UNDECLARED_KEY_REVIEW_CODE, + }) + return reviews + + def _dict(value: Any) -> dict[str, Any]: + return value if isinstance(value, dict) else {} + + def _list(value: Any) -> list[Any]: + return value if isinstance(value, list) else [] + + def _strings(value: Any) -> list[str]: + out: list[str] = [] + for item in _list(value): + if item is None: + continue + text = str(item).strip() + if text and text not in out: + out.append(text) + return out + + # ---------- PostB_3 이식 ---------- + def _field_decision_map(adj: dict[str, Any]) -> dict[tuple[str, str], Any]: + out: dict[tuple[str, str], Any] = {} + for item in _list(adj.get("field_decisions")): + if isinstance(item, dict) and item.get("candidate_ref") and item.get("field") and item.get("selected_value") != "BLOCK_REVIEW": + out[(str(item["candidate_ref"]), str(item["field"]))] = item.get("selected_value") + return out + + def _link_decision_map(adj: dict[str, Any]) -> dict[str, dict[str, Any]]: + out: dict[str, dict[str, Any]] = {} + for item in _list(adj.get("link_decisions")): + if isinstance(item, dict) and item.get("candidate_ref"): + out[str(item["candidate_ref"])] = item + return out + + def _decision_sets(ledger: dict[str, Any], adj: dict[str, Any], blockers: list[Any]) -> tuple[set[str], dict[str, str]]: + dropped: set[str] = set() + merge_into: dict[str, str] = {} + for decision in _list(ledger.get("deterministic_decisions")): + if not isinstance(decision, dict) or decision.get("decision_type") != "EXACT_DUPLICATE_MERGE": + continue + canonical = decision.get("canonical_candidate_ref") + for dup in _strings(decision.get("duplicate_candidate_refs")): + if canonical: + merge_into[dup] = str(canonical) + dropped.add(dup) + for decision in _list(adj.get("canonical_decisions")): + if not isinstance(decision, dict): + continue + kind = decision.get("decision") + refs = _strings(decision.get("input_candidate_refs")) + if kind == "DROP": + dropped.update(_strings(decision.get("drop_candidate_refs")) or refs) + elif kind == "MERGE": + target = decision.get("merge_target_ref") or (refs[0] if refs else None) + if target: + for ref in refs: + if ref != target: + merge_into[ref] = str(target) + dropped.add(ref) + elif kind == "BLOCK_REVIEW": + blockers.append(decision) + return dropped, merge_into + + def _sort_tuple(item: dict[str, Any]) -> tuple[Any, ...]: + key = _dict(item.get("deterministic_sort_key")) + return ( + key.get("BehaviorTime") is None, + key.get("BehaviorTime") or "", + key.get("domain_order", 99), + key.get("BOType") or "", + key.get("ActionType") or "", + key.get("JuristicActLabel") or "", + key.get("Action") or "", + key.get("candidate_ref") or "", + ) + + def _juristic(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + label = value.get("label") + return {"label": str(label).strip()} if label not in (None, "") else {"label": None} + text = str(value).strip() + return {"label": text} if text else None + + def _core(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("core_field_base")) + out = {key: src.get(key) for key in CORE_KEYS} + if out.get("Action_proposal") is None and seed.get("Action"): + out["Action_proposal"] = seed.get("Action") + if out.get("StatementType") is None: + out["StatementType"] = seed.get("BOType") + if out.get("Perspective") is None: + out["Perspective"] = "plaintiff" + return out + + def _amount(value: Any) -> dict[str, Any] | None: + if value is None: + return None + if isinstance(value, dict): + return { + "value_text": value.get("value_text") or value.get("text"), + "numeric_value": value.get("numeric_value"), + "currency": value.get("currency"), + } + text = str(value).strip() + return {"value_text": text, "numeric_value": None, "currency": None} if text else None + + def _evidence_item(index: str, source: Any, gaps: list[Any], bo_id: str) -> dict[str, Any]: + obj = _dict(source) + title = obj.get("source_title") or obj.get("title") or obj.get("evidence_title") or obj.get("document_title") or index + relevant = obj.get("relevant_content") or obj.get("excerpt") or obj.get("summary") or obj.get("content") + if relevant in (None, ""): + gaps.append({"BO_ID": bo_id, "evidence_index": index, "gap": "missing_relevant_content"}) + relevant = None + return { + "evidence_index": index, + "source_title": str(title), + "priority_class": obj.get("priority_class") or obj.get("priority") or None, + "relevant_content": relevant, + "authentication_status": obj.get("authentication_status") or obj.get("auth_status") or None, + "corroboration": obj.get("corroboration") or None, + "selection_basis": "source_evidence_indexes membership", + } + + def _downstream_refs(seed: dict[str, Any]) -> dict[str, Any]: + src = _dict(seed.get("downstream_seed_refs")) + return { + "claim_group_seed_refs_proposed": _strings(src.get("claim_group_seed_refs_proposed") or src.get("claim_group_seed_refs")), + "canonical_theory_graph_seed_ref_proposed": src.get("canonical_theory_graph_seed_ref_proposed") or src.get("canonical_theory_graph_seed_ref"), + "legal_effect_structure_seed_ref_proposed": src.get("legal_effect_structure_seed_ref_proposed") or src.get("legal_effect_structure_seed_ref"), + } + + def _keywords(seed: dict[str, Any], juristic: dict[str, Any] | None) -> list[str]: + out = _strings(seed.get("Legal_Keywords")) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + out.extend(_strings(domain_payload.get("legal_effect_tags"))) + if isinstance(juristic, dict) and juristic.get("label"): + out.append(str(juristic["label"])) + deduped: list[str] = [] + for item in out: + if item not in deduped: + deduped.append(item) + return deduped + + def _evidence_map_from_stage_a(stage_a: dict[str, Any]) -> dict[str, Any]: + evidence_map = _dict(stage_a.get("evidence_authority_map")) + by_index = _dict(evidence_map.get("by_evidence_index")) + if by_index: + return by_index + out: dict[str, Any] = {} + for item in _list(evidence_map.get("items")): + if isinstance(item, dict): + idx = item.get("evidence_index") or item.get("index") or item.get("id") + if idx is not None: + out[str(idx)] = item + return out + + def _stage_a_universe(stage_a: dict[str, Any]) -> dict[str, set[str]]: + event_map = _dict(stage_a.get("event_candidate_map")) + evidence_map = _dict(stage_a.get("evidence_authority_map")) + meeting_map = _dict(stage_a.get("meeting_clause_map")) + event_ids = set(_strings(event_map.get("candidate_id_set"))) + evidence_ids = set(_strings(evidence_map.get("evidence_index_set"))) + meeting_ids = set(_strings(meeting_map.get("clause_order"))) + event_ids.update(str(k) for k in _dict(event_map.get("by_event_candidate_id")).keys()) + evidence_ids.update(str(k) for k in _dict(evidence_map.get("by_evidence_index")).keys()) + meeting_ids.update(str(k) for k in _dict(meeting_map.get("by_clause_id")).keys()) + return { + "source_event_candidate_ids": event_ids, + "source_evidence_indexes": evidence_ids, + "source_meeting_clause_ids": meeting_ids, + } + + # ---------- PostB_4 이식: 게이트 ---------- + def _add_gate(gates: list[dict[str, Any]], key: str, passed: bool, detail: str) -> None: + gates.append({"gate_key": key, "status": "PASS" if passed else "FAILED", "detail": detail}) + + def _validate_item(item: Any, idx: int, ids: set[str], universe: dict[str, set[str]], + bo_types: set[str]) -> list[str]: + errors: list[str] = [] + if not isinstance(item, dict): + return [f"item {idx} must be object"] + extra = sorted(set(item.keys()) - ALLOWED_TOP_LEVEL) + missing = sorted(REQUIRED_TOP_LEVEL - set(item.keys())) + if extra: + errors.append(f"{item.get('BO_ID', idx)} additional fields: {extra}") + if missing: + errors.append(f"{item.get('BO_ID', idx)} missing fields: {missing}") + bo_id = item.get("BO_ID") + expected = f"bh{idx}" + if bo_id != expected or item.get("id") != bo_id or not isinstance(bo_id, str) or not BH_ID_RE.fullmatch(bo_id): + errors.append(f"BO_ID/id sequence mismatch: expected {expected}") + # v4 — 어휘의 정본은 registry 합집합이다. 코드에 {"event","state"} 를 두지 않는다. + if item.get("BOType") not in bo_types: + errors.append(f"{bo_id}.BOType invalid") + if item.get("ActionType") not in ACTION_TYPE_ENUM: + errors.append(f"{bo_id}.ActionType invalid") + juristic = item.get("JuristicAct") + if juristic is not None and (not isinstance(juristic, dict) or set(juristic.keys()) != {"label"}): + errors.append(f"{bo_id}.JuristicAct invalid") + for key in ("Action", "Reason"): + if not isinstance(item.get(key), str) or not item.get(key).strip(): + errors.append(f"{bo_id}.{key} must be non-empty string") + prior = item.get("PriorAct") + if prior is not None and prior not in ids: + errors.append(f"{bo_id}.PriorAct references missing BO_ID") + for ref in _list(item.get("ReasonRefs")): + if ref not in ids: + errors.append(f"{bo_id}.ReasonRefs references missing BO_ID {ref}") + core = item.get("core_field_base") + if not isinstance(core, dict) or set(core.keys()) != set(CORE_KEYS): + errors.append(f"{bo_id}.core_field_base keys invalid") + amount = item.get("amount") + if amount is not None and (not isinstance(amount, dict) or set(amount.keys()) - {"value_text", "numeric_value", "currency"}): + errors.append(f"{bo_id}.amount invalid") + evidence = _list(item.get("Evidence")) + evidence_indexes = _strings(item.get("source_evidence_indexes")) + evidence_index_set: set[str] = set() + titles: list[str] = [] + for ev in evidence: + if not isinstance(ev, dict): + errors.append(f"{bo_id}.Evidence item must be object") + continue + required_ev = {"evidence_index", "source_title", "priority_class", "relevant_content", "authentication_status", "corroboration", "selection_basis"} + if set(ev.keys()) != required_ev: + errors.append(f"{bo_id}.Evidence item keys invalid") + if isinstance(ev.get("evidence_index"), str): + evidence_index_set.add(ev["evidence_index"]) + if isinstance(ev.get("source_title"), str) and ev.get("source_title") not in titles: + titles.append(ev["source_title"]) + if item.get("EvidenceTitles") != titles: + errors.append(f"{bo_id}.EvidenceTitles mismatch") + if set(evidence_indexes) != evidence_index_set: + errors.append(f"{bo_id}.source_evidence_indexes must equal Evidence[].evidence_index") + if universe["source_evidence_indexes"] and not set(evidence_indexes).issubset(universe["source_evidence_indexes"]): + errors.append(f"{bo_id}.source_evidence_indexes outside Stage A universe") + provenance = item.get("provenance") + if not isinstance(provenance, dict) or set(provenance.keys()) != {"source_event_candidate_ids", "source_meeting_clause_ids", "source_domain"}: + errors.append(f"{bo_id}.provenance invalid") + else: + if universe["source_event_candidate_ids"] and not set(_strings(provenance.get("source_event_candidate_ids"))).issubset(universe["source_event_candidate_ids"]): + errors.append(f"{bo_id}.provenance.source_event_candidate_ids outside Stage A universe") + if universe["source_meeting_clause_ids"] and not set(_strings(provenance.get("source_meeting_clause_ids"))).issubset(universe["source_meeting_clause_ids"]): + errors.append(f"{bo_id}.provenance.source_meeting_clause_ids outside Stage A universe") + downstream = item.get("downstream_seed_refs") + if not isinstance(downstream, dict) or set(downstream.keys()) != { + "claim_group_seed_refs_proposed", "canonical_theory_graph_seed_ref_proposed", "legal_effect_structure_seed_ref_proposed", + }: + errors.append(f"{bo_id}.downstream_seed_refs invalid") + extensions = item.get("extensions", {"domain_payload": {}}) + if extensions is not None and (not isinstance(extensions, dict) or set(extensions.keys()) - {"domain_payload"} or not isinstance(extensions.get("domain_payload", {}), dict)): + errors.append(f"{bo_id}.extensions invalid") + return errors + + def _fail(message: str, gates: list[dict[str, Any]], reasons: list[str]) -> None: + print(json.dumps({ + "status": "FAILED", + "message": message, + "write_target": TARGET_NAME, + "gate_results": gates, + "failure_reasons": reasons[:40], + }, ensure_ascii=False)) + sys.exit(1) + + def main() -> None: + _init() + gates: list[dict[str, Any]] = [] + # v4 — registry 선언을 한 번 읽는다. BOType 어휘와 확장 payload 선언이 여기서 나온다. + declarations = _load_declarations() + f0_norm = _dict(_load_projection_policy().get("f0_normalization")) + bo_types = _declared_bo_types(declarations) + stage_a = _dict(_dict(read_json_doc(STAGE_A_PATH)).get("stage_a_context") or read_json_doc(STAGE_A_PATH)) + if stage_a.get("schema_version") != "task_c_bo_stage_a_context.v1" or stage_a.get("status") != "READY": + _fail("Stage A freshness guard failed", gates, ["stage_a not READY"]) + universe = _stage_a_universe(stage_a) + evidence_map = _evidence_map_from_stage_a(stage_a) + ledger = _dict(_dict(read_json_doc(LEDGER_PATH)).get("postb_seed_ledger")) + if ledger.get("schema_version") != "task_c_bo_postb_seed_ledger.v1" or ledger.get("status") != "READY": + _fail("R0 seed ledger not READY", gates, [str(ledger.get("status"))]) + # P-11 — R1 산출은 조건부다. R1 은 예외가 없어도 no-exception 객체를 반드시 쓰므로 + # 파일 부재는 "예외 없음"이 아니라 "R1 이 돌지 않았거나 실패했다"를 뜻한다. + # 종전의 무조건 fallback 은 그 둘을 가르지 못하고 판정을 조용히 삼켰다. + # 예외 팩의 exception_count 가 필수 여부를 정한다. + try: + pack = _dict(read_json_doc(EXCEPTION_PACK_PATH)) + except Exception: + pack = {} + pack_root = _dict(pack.get("postb_exception_pack") or pack) + declared_exceptions = pack_root.get("exception_count") + if not isinstance(declared_exceptions, int): + declared_exceptions = len(_list(pack_root.get("exceptions"))) + r1_state = "READ" + try: + adj_doc = read_json_doc(DECISIONS_PATH) + except Exception as exc: + if declared_exceptions > 0: + _fail("R1 adjudication decisions required but unreadable", gates, + ["exception_count=%d" % declared_exceptions, str(exc)]) + r1_state = "R1_SKIPPED" + adj_doc = {"postb_exception_adjudication": {"schema_version": "task_c_bo_postb_exception_adjudication.v1", + "status": "READY_NO_EXCEPTIONS", "exception_count": 0, + "canonical_decisions": [], "field_decisions": [], + "link_decisions": [], "semantic_gate_decisions": [], + "blocked_review_items": []}} + _add_gate(gates, "r1_decision_presence", True, + "exception_count=%d state=%s" % (declared_exceptions, r1_state)) + adj = _dict(_dict(adj_doc).get("postb_exception_adjudication")) + if adj.get("schema_version") != "task_c_bo_postb_exception_adjudication.v1": + _fail("R1 adjudication schema mismatch", gates, [str(adj.get("schema_version"))]) + if adj.get("status") not in {"READY", "READY_NO_EXCEPTIONS"}: + _fail("R1 adjudication status invalid", gates, [str(adj.get("status"))]) + blocked = _list(adj.get("blocked_review_items")) + block_decisions: list[Any] = [] + field_decisions = _field_decision_map(adj) + link_decisions = _link_decision_map(adj) + dropped, merge_into = _decision_sets(ledger, adj, block_decisions) + if blocked or block_decisions: + _fail("R1 returned BLOCK_REVIEW items: 인간 검토 필요", gates, + [json.dumps(x, ensure_ascii=False)[:200] for x in (blocked + block_decisions)]) + + candidates = [item for item in _list(ledger.get("ledger_candidates")) if isinstance(item, dict)] + survivors = [item for item in candidates if item.get("candidate_ref") not in dropped] + survivors.sort(key=_sort_tuple) + if not survivors: + _fail("no surviving BO candidates after decisions", gates, []) + + candidate_ref_to_bo_id: dict[str, str] = {} + for idx, item in enumerate(survivors, start=1): + candidate_ref_to_bo_id[str(item["candidate_ref"])] = f"bh{idx}" + for source_ref, target_ref in merge_into.items(): + if target_ref in candidate_ref_to_bo_id: + candidate_ref_to_bo_id[source_ref] = candidate_ref_to_bo_id[target_ref] + + bo_items: list[dict[str, Any]] = [] + normalization_notes: list[dict[str, Any]] = [] + evidence_gaps: list[Any] = [] + prior_link_notes: list[dict[str, Any]] = [] + + for idx, ledger_item in enumerate(survivors, start=1): + seed = _dict(ledger_item.get("seed_payload")) + candidate_ref = str(ledger_item.get("candidate_ref")) + bo_id = f"bh{idx}" + bo_type = field_decisions.get((candidate_ref, "BOType"), seed.get("BOType")) + action_type = field_decisions.get((candidate_ref, "ActionType"), seed.get("ActionType")) + juristic = _juristic(field_decisions.get((candidate_ref, "JuristicAct.label"), seed.get("JuristicAct"))) + domain_payload = _dict(_dict(seed.get("extensions")).get("domain_payload")) + action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) + # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 + # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if bo_type not in bo_types: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") + if action_type not in ACTION_TYPE_ENUM: + normalization_notes.append({"candidate_ref": candidate_ref, "field": "ActionType", "received": action_type, "fallback": f0_norm.get("action_type_default")}) + action_type = f0_norm.get("action_type_default") + if not isinstance(action, str) or not action.strip(): + normalization_notes.append({"candidate_ref": candidate_ref, "field": "Action", "fallback": "source-backed BO"}) + action = str(f0_norm.get("action_default_template") or " source-backed BO").replace("", candidate_ref) + + refs = _dict(ledger_item.get("source_refs")) + source_evidence_indexes = _strings(refs.get("source_evidence_indexes")) + evidence_items = [_evidence_item(eidx, evidence_map.get(eidx), evidence_gaps, bo_id) for eidx in source_evidence_indexes] + evidence_titles: list[str] = [] + for ev in evidence_items: + title = ev["source_title"] + if title not in evidence_titles: + evidence_titles.append(title) + + link = link_decisions.get(candidate_ref, {}) + reason_ref_candidates = _strings(link.get("reason_refs_candidate_refs")) + prior_candidate = link.get("prior_candidate_ref") + if prior_candidate == "NO_LINK": + prior_candidate = None + explicit_refs = _dict(seed.get("downstream_seed_refs")) + if not reason_ref_candidates: + reason_ref_candidates = _strings(explicit_refs.get("reason_refs_candidate_refs")) + if not prior_candidate: + prior_list = _strings(explicit_refs.get("prior_candidate_refs")) + if len(prior_list) == 1: + prior_candidate = prior_list[0] + elif len(prior_list) > 1: + # 결정적 defer 정책 (R-5): PriorAct 불명은 null 유지 + review note (blocker 아님) + prior_candidate = None + prior_link_notes.append({"candidate_ref": candidate_ref, "prior_candidates": prior_list, + "policy": "prior_link_ambiguous_kept_null"}) + reason_refs = [candidate_ref_to_bo_id[ref] for ref in reason_ref_candidates if ref in candidate_ref_to_bo_id and candidate_ref_to_bo_id[ref] != bo_id] + if prior_candidate and prior_candidate in candidate_ref_to_bo_id: + prior_act = candidate_ref_to_bo_id[prior_candidate] + elif reason_refs: + prior_act = reason_refs[0] + else: + prior_act = None + reason = "ReasonRefs에 기재된 선행 BO와 source evidence/event chain으로 연결됨" if reason_refs else "source evidence 및 event candidate에 의해 독립적으로 확인되는 BO" + + bo_items.append({ + "BO_ID": bo_id, + "id": bo_id, + "BOType": bo_type, + "ActionType": action_type, + "JuristicAct": juristic, + "Action": str(action).strip(), + "Reason": reason, + "PriorAct": prior_act, + "ReasonRefs": reason_refs, + "Legal_Keywords": _keywords(seed, juristic), + "core_field_base": _core(seed), + "amount": _amount(seed.get("amount")), + "EvidenceTitles": evidence_titles, + "Evidence": evidence_items, + "source_evidence_indexes": source_evidence_indexes, + "provenance": { + "source_event_candidate_ids": _strings(refs.get("source_event_candidate_ids")), + "source_meeting_clause_ids": _strings(refs.get("source_meeting_clause_ids")), + "source_domain": seed.get("source_domain"), + }, + "downstream_seed_refs": _downstream_refs(seed), + "extensions": {"domain_payload": domain_payload}, + }) + + # ---------- conservation + 게이트 (PostB_4 이식) ---------- + total_ledger = len(candidates) + absorbed = len(dropped) + _add_gate(gates, "candidate_conservation", len(bo_items) + absorbed == total_ledger, + f"BO {len(bo_items)} + absorbed {absorbed} == ledger {total_ledger}") + _add_gate(gates, "bo_items_array_non_empty", len(bo_items) > 0, "bo_items must be non-empty array") + ids = {item["BO_ID"] for item in bo_items} + errors: list[str] = [] + for idx, item in enumerate(bo_items, start=1): + errors.extend(_validate_item(item, idx, ids, universe, bo_types)) + _add_gate(gates, "bo_schema_and_reference_validation", not errors, "BO_JSON_Schema target validation") + # v4 신설 — 확장 payload 키를 registry 선언과 대조한다. + # 실패로 세지 않는다. 선언 밖 키는 review 로 남기고 값은 그대로 둔다. + extension_key_reviews = _extension_key_reviews(bo_items, declarations) + _add_gate(gates, "extension_payload_key_declaration_check", True, + f"bo_type_source={BO_TYPE_SOURCE} bo_types={len(bo_types)} " + f"declared_keys={len(declarations.get('declared_key_union') or [])} " + f"undeclared_records={len(extension_key_reviews)}") + if any(g["status"] != "PASS" for g in gates) or errors: + _fail("pre-write gate failed", gates, errors) + + payload = json.dumps(bo_items, ensure_ascii=False, indent=2) + "\n" + write_doc(TARGET_NAME, payload) + reread = read_json_doc(TARGET_NAME) + _add_gate(gates, "post_write_json_parse", isinstance(reread, list) and len(reread) == len(bo_items), "BO.json reread JSON parse") + if not isinstance(reread, list) or len(reread) != len(bo_items): + _fail("post-write verification failed", gates, ["reread mismatch"]) + + write_doc(BUNDLE_COMPACT_PATH, json.dumps({ + "schema_version": "task_c_bo_postb_compiled_bundle_compact.v1", + "status": "READY", + "candidate_ref_to_bo_id": candidate_ref_to_bo_id, + "bo_item_count": len(bo_items), + "normalization_notes": normalization_notes, + "evidence_gap_items": evidence_gaps, + "prior_link_notes": prior_link_notes, + "extension_key_reviews": extension_key_reviews, + }, ensure_ascii=False, indent=2)) + + # review handoff 최종 status 갱신 + try: + handoff = _dict(read_json_doc(REVIEW_HANDOFF_PATH)) + except Exception: + handoff = {"schema_version": "stage1_part2_review_handoff.v1", "review_items": []} + handoff["status"] = "FINALIZED" + handoff["bo_item_count"] = len(bo_items) + for review in extension_key_reviews: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:extension_key:{review['bo_id']}", + "source_domain": review["source_domain"], + "severity": "SOFT_WARNING", + "issue_type": UNDECLARED_KEY_REVIEW_CODE, + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "확장 payload 에 registry 선언 밖 키가 있다: " + + ", ".join(review["undeclared_keys"][:12]), + }) + for note in prior_link_notes: + handoff.setdefault("review_items", []).append({ + "review_id": f"F0:prior:{note['candidate_ref']}", + "source_domain": None, + "severity": "SOFT_WARNING", + "issue_type": "prior_link_ambiguous", + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "선행행위 후보가 복수여서 PriorAct를 null로 보존하였다.", + }) + # F0-2 — 정규화는 노트로 끝내지 않고 handoff 에도 올린다. 침묵하는 폴백과 + # 선언된 기본값의 차이는 관측 가능성이다 (M-f 관측점). + for note in normalization_notes: + handoff.setdefault("review_items", []).append({ + "review_id": "F0:normalization:%s:%s" % (note.get("candidate_ref"), note.get("field")), + "source_domain": str(note.get("candidate_ref") or "").split(":")[0] or None, + "severity": "SOFT_WARNING", + "issue_type": "schema_field_fallback", + "source_review_code": note.get("field"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": "Stage2", + "template_note": "F0 정규화가 적용된 칸이다. 값의 출처와 타당성을 재검토한다.", + }) + write_doc(REVIEW_HANDOFF_PATH, json.dumps(handoff, ensure_ascii=False, indent=2)) + + print(json.dumps({ + "status": "PASS", + "message": f"BO.json 작성 완료 (BO {len(bo_items)}건)", + "write_target": TARGET_NAME, + "bo_item_count": len(bo_items), + "absorbed_by_merge": absorbed, + "gate_results": gates, + "bundle_compact_path": BUNDLE_COMPACT_PATH, + "bo_type_source": BO_TYPE_SOURCE, + "registry_version": declarations.get("generated_from", {}).get("registry_version"), + "extension_key_review_count": len(extension_key_reviews), + "r1_decision_state": r1_state, + "declared_exception_count": declared_exceptions, + }, ensure_ascii=False)) + + if __name__ == "__main__": + main() + + - task_name: Task_C_BO_S0_signal_bundle_writer + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + # Task_C_BO_S0_signal_bundle_writer (v4) + # 정본 signal 거래 1건을 기록한다. 생성기·사영기·기록기는 조립본 모듈이며 여기서 만들지 않는다. + # Spec: stage_1_part_2_optimal_update_strategy_v.2.md §6.5 + from __future__ import annotations + import contextlib + import hashlib + import io + import itertools + import json + import pathlib + import posixpath + import re + import sys + from typing import Any + import httpx + + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_HEADERS = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + CLIENT = httpx.Client(timeout=60) + MSG_ID = itertools.count(10) + JSON_DECODER = json.JSONDecoder() + + # ---- 실행 뿌리 셋 — D-5 §2.4 0-c-2 확정값 ---- + ASSET_ROOT = "Default_Agent/" + EXECUTION_ROOT = "/tmp/s1" + LOGICAL_ROOT = "/tmp/s1" + WORK = pathlib.Path(EXECUTION_ROOT) + # 모듈의 디렉터리 산술을 그대로 재현한다(머리말 '반입 배치' 참조). + # SIGNALS_ROOT.parents[1] == ANCHOR 이므로 계약은 ANCHOR/contracts 아래다. + ANCHOR = WORK / "_sig" + SIGNALS_ROOT = ANCHOR / "pkg" / "signals" + CONTRACT_DIR = ANCHOR / "contracts" + OUTPUT_DIR = WORK / "_signal_out" + + # ---- 반입 대상 ---- + RUNTIME_MANIFEST = "Default_Agent/runtime_manifest.json" + COMPILER_MODULES = ["common", "projections", "schema_validator", "signal_compiler", + "signal_gate", "transaction_writer", "writer_boundary"] + ADAPTER_MODULES = ["s3_domain_seed_adapter", "s3_envelope_migration_adapter", + "s4_calculation_adapter", "sg01_activation_adapter"] + EMITTER_MODULES = ["emitter_runtime"] + ["emit_sg%02d" % n for n in range(2, 14)] + SIGNAL_REGISTRY = "Default_Agent/signals/signal_registry.v2.json" + EXECUTION_CONTRACT = "Default_Agent/contracts/signals/s5_execution_contract.v2.json" + + # ---- 사건 입력 ---- + # v4 — 구 경로·정적 이름을 걷어냈다. seed 는 fan-out 계획의 expected_output_path 로 읽는다. + ACTIVATION_MANIFEST_PATH = "routing/domain_activation_manifest.json" + FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" + SOURCE_UNIVERSE_PATH = "stage1_tmp/task_c_bo/source_universe_manifest.json" + BO_PATH = "BO.json" + SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" + # R-5 — 도메인이 선언한 방출 signal 집합. A0 가 슬라이스에 실어 둔 것을 읽는다. + # registry 를 여기서 다시 적재하지 않는다. + SLICE_DIR = "runtime/domain_slices" + SLICE_ROOT_KEY = "stage_b_domain_slice" + DECLARED_EMISSION_REVIEW_CODE = "SIGNAL_EMISSION_NOT_DECLARED" + + # ---- 산출 ---- + SIGNAL_OUTPUT_PREFIX = "signals/" + TRANSACTION_ID_RE = r"^S5TX-[a-f0-9]{20}$" + CANONICAL_WRITER_MODULE = "compiler/transaction_writer.py" + COMPATIBILITY_ROOT_ALIASES = { + "compatibility_views/actio_case_signals.json": "actio_case_signals.json", + "compatibility_views/case_liability_signals.json": "case_liability_signals.json", + "compatibility_views/legal_effect_signals.json": "legal_effect_signals.json", + } + # 각 호환 뷰가 어느 정본 signal 의 사영인지. projections.py 의 서명이 정본이다. + COMPATIBILITY_VIEW_SOURCES = { + "compatibility_views/actio_case_signals.json": [], + "compatibility_views/case_liability_signals.json": ["SG-05", "SG-08"], + "compatibility_views/legal_effect_signals.json": ["SG-13"], + } + SIGNAL_FILE_BY_CODE = { + "SG-05": "legal_relation_lifecycle_signals.json", + "SG-08": "liability_causation_damage_signals.json", + "SG-13": "legal_effect_routes.json", + } + COMPATIBILITY_EMPTY_REVIEW_CODE = "COMPATIBILITY_VIEW_EMPTY" + + + def _mid() -> int: + return next(MSG_ID) + + def _init() -> None: + r = CLIENT.post( + LOCALDOCS_URL, + json={ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-03-26", + "capabilities": {}, + "clientInfo": { + "name": "task-c-bo-s0-signal-bundle-writer", + "version": "1.0", + "user_id": "{{__user_hash__}}", + "workspace_id": "{{__workspace_hash__}}", + }, + }, + }, + headers=MCP_HEADERS, + ) + r.raise_for_status() + sid = r.headers.get("mcp-session-id") + if sid: + MCP_HEADERS["mcp-session-id"] = sid + CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "method": "notifications/initialized"}, + headers=MCP_HEADERS, + ).raise_for_status() + + def _parse_mcp(text: str) -> Any: + for line in text.strip().split("\n"): + if line.startswith("data: "): + try: + return json.loads(line[6:]) + except Exception: + pass + try: + return json.loads(text) + except Exception: + return None + + def _call(name: str, args: dict[str, Any]) -> Any: + r = CLIENT.post( + LOCALDOCS_URL, + json={"jsonrpc": "2.0", "id": _mid(), "method": "tools/call", + "params": {"name": name, "arguments": args}}, + headers=MCP_HEADERS, + ) + r.raise_for_status() + payload = _parse_mcp(r.text) + if not payload or "result" not in payload: + raise RuntimeError(f"MCP {name} failed") + return payload + + def read_json_doc(path: str) -> Any: + result = _call("read_docs", {"doc_names": [path]}) + content = result["result"].get("content", []) + text = content[0].get("text", "") if content else "" + try: + parsed = json.loads(text) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Invalid read_docs envelope for {path}") from exc + if isinstance(parsed, dict) and "results" in parsed: + rows = parsed.get("results") or [] + if not rows or rows[0].get("error"): + raise RuntimeError(f"read_docs failed for {path}: {rows[:1]}") + inner = rows[0].get("content") or "" + if isinstance(inner, str): + try: + return json.loads(inner) + except json.JSONDecodeError: + obj, end = JSON_DECODER.raw_decode(inner.strip()) + if inner.strip()[end:].strip(): + raise RuntimeError(f"{path}: non-json tail") + return obj + return inner + return parsed + + def write_doc(path: str, content: str) -> None: + _call("write_file", {"path": path, "content": content, "overwrite": True}) + + + def read_raw(name: str) -> str: + # 파일 내용을 재직렬화 없이 원문 그대로 돌려준다. + # 모듈 미러의 sha256 은 원문 바이트의 해시여야 하므로 재직렬화를 허용하지 않는다. + # Task_D0_domain_activation_gate.yml 의 검증 완료본을 _call 서명만 맞춰 옮겼다. + p = _call("read_docs", {"doc_names": [name]}) + text = (p["result"].get("content") or [{}])[0].get("text", "") + if not text: + raise RuntimeError("EMPTY_RESPONSE:%s" % name) + try: + outer = json.loads(text) + except Exception: + return text + if isinstance(outer, str): + return outer + if isinstance(outer, dict) and "results" in outer: + r0 = (outer.get("results") or [{}])[0] + inner = r0.get("content") + if inner is None: + inner = r0.get("text") + if not isinstance(inner, str) or not inner: + raise RuntimeError("EMPTY_CONTENT:%s" % name) + return inner + raise RuntimeError("UNEXPECTED_READ_DOCS_ENVELOPE:%s" % name) + + + # ------------------------------------------------------------------ + # F-4 — 자기 입력의 배포 무결성. A0 게이트가 앞에서 걸러 주지만 이 task 는 + # 개별 재실행이 가능하다(예행 하네스가 실제로 task 단위로 돌린다). + # 매니페스트는 task 당 한 번만 읽는다. 호출마다 읽으면 S0 의 $ref 폐포 walk 에서만 19 회다. + # ------------------------------------------------------------------ + PART2_DEPLOYMENT_REMEDY = ( + "배포 원본 extension_research/Default_Agent 트리를 서버 Default_Agent/ 에 통째로 배포하고 " + "runtime_manifest.json 의 등재·sha256 과 대조해 결손 경로를 복원한다. 부분 복사는 허용되지 않는다.") + _MANIFEST_MAP = None + + + def _manifest_map(): + global _MANIFEST_MAP + if _MANIFEST_MAP is None: + _MANIFEST_MAP = {row["path"]: row["sha256"] + for row in (json.loads(read_raw(RUNTIME_MANIFEST)).get("entries") or []) + if isinstance(row, dict)} + return _MANIFEST_MAP + + + def _deployment_error(logical, detail, extra=None): + payload = {"reason_code": "PART2_ASSET_DEPLOYMENT_INCOMPLETE", "path": logical, + "detail": detail, "remedy": PART2_DEPLOYMENT_REMEDY} + if extra: + payload.update(extra) + return RuntimeError(json.dumps(payload, ensure_ascii=False)) + + + def _verify_asset(logical: str) -> str: + """원문을 돌려준다. 반환값을 소비해야 읽기 횟수가 늘지 않는다.""" + body = read_raw(logical) + want = _manifest_map().get(logical[len("Default_Agent/"):]) + if want is None: + raise _deployment_error(logical, "unregistered") + if want != hashlib.sha256(body.encode("utf-8")).hexdigest(): + raise _deployment_error(logical, "hash_mismatch", {"expected": want}) + return body + + + # ------------------------------------------------------------------ + # 1) 자산 반입 — D0 반입 규약 R-1~R-5 를 그대로 따른다. + # + # 디렉터리 산술을 흉내내야 하는 이유(실측). + # signals/compiler/*.py 는 _SIGNALS_ROOT = Path(__file__).resolve().parents[1] + # 로 signals 뿌리를 잡고, 실행 계약을 _SIGNALS_ROOT.parents[1]/contracts/ + # s5_execution_contract.v2.json 에서 읽는다. 즉 계약은 signals 의 조부모 아래다. + # 조립본은 계약을 Default_Agent/contracts/signals/ 에 두므로 그 산술이 조립본 + # 배치로는 풀리지 않는다. 반입 시에는 우리가 배치를 정하므로 모듈이 기대하는 + # 산술을 그대로 재현한다 — signals 를 /pkg/signals 에 두고 계약을 + # /contracts 에 둔다. 모듈 원문은 한 글자도 고치지 않는다. + # ------------------------------------------------------------------ + def _relative_refs(node: Any) -> list[str]: + """상대 파일 $ref 만 모은다. 로컬 포인터(#/...)는 검증기가 스스로 푼다.""" + out: list[str] = [] + if isinstance(node, dict): + ref = node.get("$ref") + if isinstance(ref, str) and ref and not ref.startswith("#"): + out.append(ref.split("#", 1)[0]) + for value in node.values(): + out.extend(_relative_refs(value)) + elif isinstance(node, list): + for value in node: + out.extend(_relative_refs(value)) + return [item for item in out if item] + + + def _stage_bytes(target, body: str) -> int: + raw = body.encode("utf-8") + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + return len(raw) + + + def materialize() -> dict[str, Any]: + SIGNALS_ROOT.mkdir(parents=True, exist_ok=True) + CONTRACT_DIR.mkdir(parents=True, exist_ok=True) + OUTPUT_DIR.mkdir(parents=True, exist_ok=True) + manifest_doc = json.loads(read_raw(RUNTIME_MANIFEST)) + expected = {row["path"]: row["sha256"] for row in manifest_doc.get("entries") or []} + + staged: dict[str, Any] = {"modules": [], "schemas": [], "unregistered": []} + for sub, names in (("compiler", COMPILER_MODULES), + ("adapters", ADAPTER_MODULES), + ("emitters", EMITTER_MODULES)): + for name in names: + logical = "%ssignals/%s/%s.txt" % (ASSET_ROOT, sub, name) + body = read_raw(logical) + raw = body.encode("utf-8") + want = expected.get(logical[len(ASSET_ROOT):]) + got = hashlib.sha256(raw).hexdigest() + if want is None: + raise RuntimeError("MODULE_MIRROR_UNREGISTERED:%s" % logical) + if want != got: + raise RuntimeError("MODULE_MIRROR_HASH_MISMATCH:%s" % logical) + _stage_bytes(SIGNALS_ROOT / sub / (name + ".py"), body) + staged["modules"].append("%s/%s" % (sub, name)) + + _stage_bytes(SIGNALS_ROOT / "signal_registry.v2.json", _verify_asset(SIGNAL_REGISTRY)) + _stage_bytes(CONTRACT_DIR / "s5_execution_contract.v2.json", _verify_asset(EXECUTION_CONTRACT)) + + # 스키마 목록을 이 코드가 만들지 않는다. registry 가 선언한 참조에서 출발해 + # 상대 파일 $ref 를 따라간다. _common/ 아래 조각도 그렇게 저절로 딸려 온다. + registry = json.loads((SIGNALS_ROOT / "signal_registry.v2.json").read_text(encoding="utf-8")) + pending = list(dict.fromkeys( + [str(row["schema"]) for row in registry["entries"] if row.get("schema")] + + [str(registry["domain_envelope"]), str(registry["manifest_schema"])])) + seen: set[str] = set() + while pending: + rel = posixpath.normpath(pending.pop(0)) + if rel in seen or rel.startswith(".."): + continue + seen.add(rel) + # F-4c — 이 한 줄이 폐포가 끌어오는 signal 스키마 전부를 덮는다. + # 목록을 상수로 굳히지 않는다 — registry 가 바뀌면 조용히 어긋난다. + body = _verify_asset("%ssignals/%s" % (ASSET_ROOT, rel)) + _stage_bytes(SIGNALS_ROOT / rel, body) + staged["schemas"].append(rel) + for child in _relative_refs(json.loads(body)): + pending.append(posixpath.join(posixpath.dirname(rel), child)) + + sys.path.insert(0, str(SIGNALS_ROOT)) + staged["signals_root"] = str(SIGNALS_ROOT) + staged["module_count"] = len(staged["modules"]) + staged["schema_count"] = len(staged["schemas"]) + return staged + + + # ------------------------------------------------------------------ + # 2) 입력 조립 — 정적 어휘를 두지 않는다. 계획서와 매니페스트가 목록을 정한다. + # ------------------------------------------------------------------ + def build_inputs() -> tuple[dict[str, Any], dict[str, Any]]: + activation = read_json_doc(ACTIVATION_MANIFEST_PATH) + if not isinstance(activation, dict) or not isinstance( + activation.get("domain_activation_manifest"), dict): + raise RuntimeError("SG01_INPUT_REQUIRED: Part 1 activation gate output is required") + + plan = read_json_doc(FANOUT_PLAN_PATH) + plan_root = plan.get("domain_fanout_plan") if isinstance(plan, dict) else None + plan_root = plan_root if isinstance(plan_root, dict) else (plan if isinstance(plan, dict) else {}) + instances = [x for x in (plan_root.get("task_instances") or []) if isinstance(x, dict)] + if not instances: + raise RuntimeError("S0_FANOUT_PLAN_EMPTY") + + seeds: dict[str, Any] = {} + seed_paths: list[str] = [] + declared_emissions: dict[str, list[str]] = {} + for instance in instances: + path = instance.get("expected_output_path") + domain_id = str(instance.get("domain_id") or "") + if not isinstance(path, str) or not path or not domain_id: + raise RuntimeError("S0_FANOUT_INSTANCE_INVALID:%s" % json.dumps(instance, ensure_ascii=False)[:120]) + document = read_json_doc(path) + root = document.get("stage_b_domain_bo_seed_output") if isinstance(document, dict) else None + if not isinstance(root, dict): + raise RuntimeError("S0_SEED_ROOT_MISSING:%s" % path) + if root.get("schema_version") != SEED_SCHEMA_VERSION: + raise RuntimeError("S3_SEED_SCHEMA_VERSION_MISMATCH:%s" % path) + if root.get("domain_id") != domain_id: + raise RuntimeError("S0_SEED_DOMAIN_MISMATCH:%s" % path) + seeds[domain_id] = document + seed_paths.append(path) + # R-5 — 같은 도메인의 슬라이스에서 emits_signals 선언을 읽는다. 부재는 조용히 넘긴다. + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_root = slice_doc.get(SLICE_ROOT_KEY) if isinstance(slice_doc, dict) else None + slice_root = slice_root if isinstance(slice_root, dict) else (slice_doc if isinstance(slice_doc, dict) else {}) + declarations = slice_root.get("domain_declarations") + codes = [str(v) for v in ((declarations or {}).get("emits_signals") or []) + if isinstance(v, str) and v] + if codes: + declared_emissions[domain_id] = sorted(set(codes)) + except Exception: + pass + + universe_doc = read_json_doc(SOURCE_UNIVERSE_PATH) + universe_doc = universe_doc if isinstance(universe_doc, dict) else {} + bo_items = read_json_doc(BO_PATH) + bo_ids = sorted({str(x.get("BO_ID")) for x in bo_items + if isinstance(x, dict) and x.get("BO_ID")}) if isinstance(bo_items, list) else [] + evidence_ids = sorted({str(v) for v in (universe_doc.get("evidence_index_set") or [])}) + event_ids = sorted({str(v) for v in (universe_doc.get("event_candidate_ids") or [])}) + meeting_ids = sorted({str(v) for v in (universe_doc.get("meeting_clause_ids") or [])}) + if not (evidence_ids or event_ids or meeting_ids): + raise RuntimeError("S0_SOURCE_UNIVERSE_EMPTY:%s" % SOURCE_UNIVERSE_PATH) + + # fact_ids 와 law_version_ids 는 이 매니페스트가 선언하지 않는다. + # 비워 둔다. 레코드가 그 종류를 실으면 게이트의 source_membership 이 잡는다. + # 조용히 통과시키지 않는 쪽이 맞다. + source_universe = { + "bo_ids": bo_ids, + "fact_ids": [], + "evidence_ids": evidence_ids, + "meeting_clause_ids": meeting_ids, + "law_version_ids": [], + "event_ids": event_ids, + "all_source_refs": sorted(set(bo_ids) | set(evidence_ids) | set(event_ids) | set(meeting_ids)), + "unrouted_evidence_count": int(len( + activation["domain_activation_manifest"].get("unrouted_material") or [])), + } + inputs = { + "declared_emissions": declared_emissions, + "domain_activation_manifest": activation, + "domain_seed_outputs": seeds, + "source_universe": source_universe, + # v4 — 구 signal 원문을 넣지 않는다. 세 호환 뷰는 정본 signal 의 사영일 뿐이다. + "legacy_signals": {}, + "signal_candidates": {}, + } + receipt = { + "seed_count": len(seeds), + "declared_emission_domains": sorted(declared_emissions), + "seed_paths": seed_paths, + "bo_id_count": len(bo_ids), + "evidence_count": len(evidence_ids), + "event_count": len(event_ids), + "meeting_count": len(meeting_ids), + "fact_ids_declared": False, + "law_version_ids_declared": False, + } + return inputs, receipt + + + # ------------------------------------------------------------------ + # 3) 생성기 12 · 사영기 3 · 단일 기록기 호출 + # 호출 본문은 이 한 함수뿐이다. 생성기와 사영기는 순수 함수이며 파일을 쓰지 않는다. + # 실행기 안에서 파일을 쓰는 것은 compiler/transaction_writer.py 하나다 — + # signal_gate 의 canonical_writer_uniqueness 가 그것을 강제한다. + # ------------------------------------------------------------------ + def compile_and_validate(inputs: dict[str, Any]) -> tuple[dict[str, Any], dict[str, Any]]: + from compiler.signal_compiler import compile_transaction + from compiler.signal_gate import validate_output + + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + manifest = compile_transaction(inputs, OUTPUT_DIR) + gate = validate_output(inputs, OUTPUT_DIR, SIGNALS_ROOT) + if not re.fullmatch(TRANSACTION_ID_RE, str(manifest.get("transaction_id") or "")): + raise RuntimeError("S0_TRANSACTION_ID_PATTERN:%s" % manifest.get("transaction_id")) + if gate.get("canonical_writer_modules") != [CANONICAL_WRITER_MODULE]: + raise RuntimeError("S0_CANONICAL_WRITER_NOT_UNIQUE:%s" + % json.dumps(gate.get("canonical_writer_modules"), ensure_ascii=False)) + if gate.get("status") != "PASS": + raise RuntimeError("S0_SIGNAL_GATE_FAILED:%s" + % json.dumps(gate.get("errors")[:8], ensure_ascii=False)) + return manifest, gate + + + # ------------------------------------------------------------------ + # 4) 반출 — 거래가 낸 바이트를 그대로 옮긴다. 재직렬화하지 않는다. + # ------------------------------------------------------------------ + def publish(manifest: dict[str, Any]) -> dict[str, Any]: + written: list[dict[str, Any]] = [] + local: dict[str, bytes] = {} + for path in sorted(OUTPUT_DIR.rglob("*.json")): + rel = path.relative_to(OUTPUT_DIR).as_posix() + raw = path.read_bytes() + local[rel] = raw + write_doc(SIGNAL_OUTPUT_PREFIX + rel, raw.decode("utf-8")) + written.append({"path": SIGNAL_OUTPUT_PREFIX + rel, + "sha256": hashlib.sha256(raw).hexdigest(), "bytes": len(raw)}) + + # 구 이름 세 개는 Part 3·4 가 읽는 최대 호환면이다. 같은 바이트를 그대로 한 벌 더 놓는다. + # 두 번째 생산자가 아니라 운반이다 — 내용은 거래가 낸 것과 바이트 동일하다. + aliases: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + raw = local.get(canonical_rel) + if raw is None: + raise RuntimeError("S0_COMPATIBILITY_VIEW_MISSING:%s" % canonical_rel) + write_doc(alias, raw.decode("utf-8")) + aliases.append({"alias": alias, "canonical": SIGNAL_OUTPUT_PREFIX + canonical_rel, + "sha256": hashlib.sha256(raw).hexdigest()}) + + # 기록 후 재읽기. 봉인 파일 하나를 원문 바이트로 되읽어 해시를 대조한다. + reread = read_raw(SIGNAL_OUTPUT_PREFIX + "signal_manifest.json").encode("utf-8") + if hashlib.sha256(reread).hexdigest() != hashlib.sha256(local["signal_manifest.json"]).hexdigest(): + raise RuntimeError("S0_POST_WRITE_MANIFEST_HASH_MISMATCH") + return {"written": written, "compatibility_root_aliases": aliases, + "file_count": len(written)} + + + def emission_notices(manifest: dict[str, Any], + declared_emissions: dict[str, list[str]]) -> list[dict[str, Any]]: + """도메인이 선언한 emits_signals 와 기록이 실린 정본 signal 을 대조한다. + + 실패로 세지 않는다. 선언은 registry 의 것이고 실제 방출은 사건 재료에 달려 있어 + 선언보다 적게 나오는 것은 정상이다. 반대로 **선언 밖에서 기록이 나오면** 어휘 밖의 + 산출이므로 지목한다 — 137종 일반성은 그 어휘 안에서 성립해야 한다. + """ + if not declared_emissions: + return [] + union: set[str] = set() + for codes in declared_emissions.values(): + union.update(codes) + by_path = {row["path"]: row for row in manifest.get("files") or []} + emitted: set[str] = set() + for code, filename in SIGNAL_FILE_BY_CODE.items(): + if (by_path.get(filename) or {}).get("record_count"): + emitted.add(code) + undeclared = sorted(code for code in emitted if code not in union) + if not undeclared: + return [] + return [{ + "review_code": DECLARED_EMISSION_REVIEW_CODE, + "undeclared_signals": undeclared, + "declared_union": sorted(union), + "declared_by_domain": {k: v for k, v in sorted(declared_emissions.items())}, + "note": "선언 밖 signal 에 기록이 실렸다. registry 의 emits_signals 를 넓히거나 산출을 좁힌다.", + }] + + + def compatibility_notices(manifest: dict[str, Any]) -> list[dict[str, Any]]: + """호환 뷰가 비었는데 정본 signal 에는 기록이 있으면 조용히 넘기지 않고 지목한다. + + v3 은 세 파일을 BO.json 에서 직접 만들었고, v4 는 정본 signal 의 사영으로 만든다. + 사영 대상은 compatibility_key/compatibility_route 를 단 기록뿐이며 그 표식은 + 구 signal 원문에서만 붙는다. 따라서 구 원문을 넣지 않는 v4 에서는 뷰가 빌 수 있다. + Part 3·4 는 signal_manifest.downstream_read_sets 가 선언한 정본 집합으로 옮겨야 한다. + 그 이관은 Part 3·4 개정의 몫이므로 여기서는 사실만 남긴다. + """ + by_path = {row["path"]: row for row in manifest.get("files") or []} + notices: list[dict[str, Any]] = [] + for canonical_rel, alias in COMPATIBILITY_ROOT_ALIASES.items(): + view = by_path.get(canonical_rel) or {} + if view.get("state") != "empty": + continue + sources = COMPATIBILITY_VIEW_SOURCES[canonical_rel] + populated = sorted(code for code in sources + if (by_path.get(SIGNAL_FILE_BY_CODE.get(code, "")) or {}).get("record_count")) + if populated: + notices.append({ + "review_code": COMPATIBILITY_EMPTY_REVIEW_CODE, + "alias": alias, + "canonical_view": canonical_rel, + "populated_canonical_signals": populated, + "downstream_read_sets": manifest.get("downstream_read_sets"), + "note": "구 이름 파일이 비었다. Part 3·4 는 정본 signal 집합으로 읽어야 한다.", + }) + return notices + + + def main() -> None: + _init() + staged = materialize() + inputs, input_receipt = build_inputs() + manifest, gate = compile_and_validate(inputs) + published = publish(manifest) + notices = compatibility_notices(manifest) + notices.extend(emission_notices(manifest, inputs.get("declared_emissions") or {})) + + print(json.dumps({ + "status": "READY_WITH_REVIEW" if notices else "READY", + "message": "정본 signal 거래 1건 기록 완료 (파일 %d종)" % published["file_count"], + "schema_version": "stage1_canonical_signal_writer.v1", + "transaction_id": manifest.get("transaction_id"), + "manifest_status": manifest.get("status"), + "signal_manifest_path": SIGNAL_OUTPUT_PREFIX + "signal_manifest.json", + "module_import": { + "module_count": staged["module_count"], + "schema_count": staged["schema_count"], + "hash_source": RUNTIME_MANIFEST, + "signals_root": staged["signals_root"], + }, + "inputs": input_receipt, + "gate": { + "status": gate.get("status"), + "error_count": gate.get("error_count"), + "canonical_writer_modules": gate.get("canonical_writer_modules"), + "source_membership_pass": gate.get("source_membership_pass"), + "domain_source_membership_pass": gate.get("domain_source_membership_pass"), + "meeting_only_promotion_pass": gate.get("meeting_only_promotion_pass"), + "negative_conflict_preservation_pass": gate.get("negative_conflict_preservation_pass"), + "compatibility_projection_pass": gate.get("compatibility_projection_pass"), + "manifest_hash_pass": gate.get("manifest_hash_pass"), + "forbidden_conclusion_key_pass": gate.get("forbidden_conclusion_key_pass"), + }, + "published": published, + "active_domains": manifest.get("active_domains"), + "unrouted_counts": manifest.get("unrouted_counts"), + "compatibility_notices": notices, + }, ensure_ascii=False)) + + + if __name__ == "__main__": + try: + main() + except Exception as exc: + print(json.dumps({"status": "FAILED", "message": "S0 signal bundle writer 실패", + "reason": str(exc)}, ensure_ascii=False)) + raise + + task_procedure: + # A0 가 fan-out 계획을 낸 뒤에야 worker 인스턴스가 생긴다. 그래서 직렬이다. + IN: + nexts: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + wait_until: [] + + Task_C_BO_A0_context_and_domain_slice_compiler: + nexts: ["Task_C_B_domain_worker_*"] + wait_until: ["IN"] + + # 활성 도메인 병렬 x M. 인스턴스는 domain_fanout_plan.task_instances[] 가 만든다. + Task_C_B_domain_worker_*: + nexts: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + wait_until: ["Task_C_BO_A0_context_and_domain_slice_compiler"] + + # barrier — 정확 일치. 누락·초과·중복 모두 실패다. + Task_C_BO_R0_seed_reducer_and_exception_planner: + nexts: ["Task_C_BO_R1_exception_adjudicator"] + wait_until: ["all Task_C_B_domain_worker_*"] + + # 조건부. 예외 pack 이 비면 통과만 한다. + Task_C_BO_R1_exception_adjudicator: + nexts: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + wait_until: ["Task_C_BO_R0_seed_reducer_and_exception_planner"] + + Task_C_BO_F0_final_bo_compiler_gate_writer: + nexts: ["Task_C_BO_S0_signal_bundle_writer"] + wait_until: ["Task_C_BO_R1_exception_adjudicator"] + + Task_C_BO_S0_signal_bundle_writer: + nexts: ["OUT"] + wait_until: ["Task_C_BO_F0_final_bo_compiler_gate_writer"] + + OUT: + nexts: [] + wait_until: ["Task_C_BO_S0_signal_bundle_writer"] + + prevs: [] + nexts: [] diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/stage_1_part_2_v.8.yml b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/stage_1_part_2_v.8.yml index 257a6fb8..17910f80 100644 --- a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/stage_1_part_2_v.8.yml +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/ver_8_yaml_candidates/stage_1_part_2_v.8.yml @@ -238,6 +238,7 @@ Agent: POLICY = "Default_Agent/stage1_runtime/prompt_composition_policy.json" SLICE_SCHEMA = "Default_Agent/platform/schemas/domain_slice.schema.v2.json" FANOUT_SCHEMA = "Default_Agent/platform/schemas/domain_fanout_plan.schema.json" + SPECIAL_LAW_INDEX = "Default_Agent/special_law_profiles/_registry_index.json" # F-2 — Part 2 가 조립본에서 읽는 정적 자산 중 경로가 고정된 것. 이 목록이 곧 배포 요구 선언이다. # S0 의 signal 스키마 폐포 17종과 미러 24종은 런타임에 계산되거나 S0 가 이미 경성으로 대조하므로 @@ -253,6 +254,7 @@ Agent: "Default_Agent/routing/extension_payload_key_declarations.v1.json", # F0 "Default_Agent/stage1_runtime/worker_output_validator.txt", # R0 전용 미러 "Default_Agent/stage1_runtime/bo_surface_projection_policy.v1.json", # R0·F0 — BO 투영 정책 + "Default_Agent/stage1_runtime/seed_admission_policy.v1.json", # R0 — 시드 수용 정책 ) RUNTIME_MANIFEST_SCHEMA = "stage1_runtime_manifest.v1" # registry_validator 는 overlay 오류를 모으기만 한다. 네 코드는 배포 문제이므로 경성으로 올린다. @@ -464,6 +466,39 @@ Agent: stage_text(overlay_logical, read_raw(overlay_logical)) except Exception as exc: warn("PROMPT_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + # 특별법 profile 조각 반입 — prompt_compiler.collect_domain_fragments 는 + # domain_config.special_law_profiles 가 선언한 profile_id 를 profile_paths 인자에서 + # 찾는다. 그 인자를 넘기지 않으면 파일이 배포돼 있어도 + # PROMPT_REQUIRED_FRAGMENT_MISSING 으로 멈춘다(디스크를 보지 않는 검사다). + # 파일 이름을 짓지 않는다 — profile registry 가 선언한 prompt_overlay_path 를 따라간다. + profile_paths = {} + try: + slp_index_raw = read_raw(SPECIAL_LAW_INDEX) + except Exception as exc: + warn("SPECIAL_LAW_INDEX_ABSENT", "%s: %s" % (SPECIAL_LAW_INDEX, exc)) + else: + stage_text(SPECIAL_LAW_INDEX, slp_index_raw) + slp_doc = json.loads(slp_index_raw) + slp_index = slp_doc.get("special_law_profile_registry_index", slp_doc) + slp_base = posixpath.dirname(SPECIAL_LAW_INDEX) + for entry in slp_index.get("entries") or []: + profile_id = entry.get("profile_id") + overlay_path = entry.get("prompt_overlay_path") + if not isinstance(profile_id, str) or not profile_id: + continue + if not isinstance(overlay_path, str) or not overlay_path: + warn("SPECIAL_LAW_OVERLAY_PATH_MISSING", str(profile_id)) + continue + overlay_logical = unicodedata.normalize( + "NFC", overlay_path if overlay_path.startswith("Default_Agent/") + else posixpath.join(slp_base, overlay_path)) + try: + stage_text(overlay_logical, read_raw(overlay_logical)) + except Exception as exc: + warn("SPECIAL_LAW_OVERLAY_STAGE_SKIPPED", "%s: %s" % (overlay_logical, exc)) + continue + profile_paths[profile_id] = overlay_logical + stage_text(COMMON_CONTRACT, read_raw(COMMON_CONTRACT)) stage_text(POLICY, read_raw(POLICY)) @@ -536,7 +571,7 @@ Agent: for domain_id in expected_runnable: specs = prompt_compiler.collect_domain_fragments( domain_id, registry, common_contract_path=COMMON_CONTRACT, - extra_specs=vocabulary_specs) + profile_paths=profile_paths, extra_specs=vocabulary_specs) text, manifest_row = prompt_compiler.compile_fragments(specs, policy) rel = "%s/%s.md" % (PROMPT_DIR, domain_id) stage_text(rel, text) @@ -579,9 +614,48 @@ Agent: if isinstance(domain_id, str) and domain_id and codes: screening_calc[domain_id] = sorted(set(codes)) + # P-2a — stage_a_context 를 slice 컴파일보다 먼저 만든다. slice 의 source_universe 가 + # 인용할 event 식별자의 정본이 event_candidate_map 이기 때문이다. 순서가 뒤였을 때 + # A0 는 원시 evidence_event_candidates 문서를 넘겼고, domain_slice_compiler._records 가 + # 그 문서-단위 items(30건)를 후보로 오인해 candidate_id 를 못 찾아 + # event:unidentified:NNNN 로 대체했다. 그 값은 source_universe_manifest 의 + # event_candidate_ids(EVT-...-NN, 84건)와 교집합이 0 이라, 워커가 규율을 지켜 + # slice 안의 것만 인용해도 R0 가 "source_refs outside Stage A universe" 로 차단했다. + meeting_raw = read_raw(MEETING) + created_at_utc = utc_now() + input_digests = { + MEETING: sha_text(meeting_raw), + EVIDENCE: sha_text(read_raw(EVIDENCE)), + EVENTS: sha_text(read_raw(EVENTS)), + SCREENING: seal["screening_sha256"], + ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], + REGISTRY_INDEX: seal["registry_index_sha256"], + } + stage_a = stage_a_context_builder.build_stage_a_context( + meeting_text=meeting_raw, + evidence_obj=evidence_document, + event_obj=events_document, + input_digests_sha256=input_digests, + created_at_utc=created_at_utc, + digest_guard=seal, + expected_runnable_domain_ids=expected_runnable) + source_manifest = stage_a_context_builder.build_source_universe_manifest( + stage_a, input_digests_sha256=input_digests, + registry_index_sha256=seal["registry_index_sha256"]) + # 맵을 통째로 넘기지 않는다 — domain_slice_compiler._records 는 dict 를 받으면 + # by_evidence_index(값이 id 리스트)에 먼저 걸려 빈 목록을 돌려준다. 후보 레코드 + # 목록으로 평탄화해 넘겨야 _record_id 가 candidate_id 를 찾아 EVT-...-NN 을 쓴다. + event_candidate_records = [ + row for row in ((stage_a.get("event_candidate_map") or {}).get("by_candidate_id") or {}).values() + if isinstance(row, dict)] + if not event_candidate_records: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_EVENT_CANDIDATE_MAP_EMPTY", + "message": "stage_a event_candidate_map.by_candidate_id 가 비었다. slice 의 event 식별자를 정본으로 실을 수 없다.", + }, ensure_ascii=False)) try: result = domain_slice_compiler.compile_domain_slices( - manifest, registry, evidence_document, events_document, + manifest, registry, evidence_document, event_candidate_records, manifest_sha256=seal["activation_manifest_sha256"], evidence_sha256=sha_text(read_raw(EVIDENCE)), events_sha256=sha_text(read_raw(EVENTS)), @@ -641,27 +715,7 @@ Agent: # P-2 — stage_a_context 와 원천 우주 매니페스트는 R0·F0·S0 의 소비 계약이다. # v3 의 세 builder 를 그대로 이식한 모듈이 만든다. 여기서 모양을 짓지 않는다. - meeting_raw = read_raw(MEETING) - created_at_utc = utc_now() - input_digests = { - MEETING: sha_text(meeting_raw), - EVIDENCE: sha_text(read_raw(EVIDENCE)), - EVENTS: sha_text(read_raw(EVENTS)), - SCREENING: seal["screening_sha256"], - ACTIVATION_MANIFEST: seal["activation_manifest_sha256"], - REGISTRY_INDEX: seal["registry_index_sha256"], - } - stage_a = stage_a_context_builder.build_stage_a_context( - meeting_text=meeting_raw, - evidence_obj=evidence_document, - event_obj=events_document, - input_digests_sha256=input_digests, - created_at_utc=created_at_utc, - digest_guard=seal, - expected_runnable_domain_ids=expected_runnable) - source_manifest = stage_a_context_builder.build_source_universe_manifest( - stage_a, input_digests_sha256=input_digests, - registry_index_sha256=seal["registry_index_sha256"]) + # 계산은 P-2a 에서 이미 끝났다(slice 가 같은 식별자를 써야 하므로 앞당겼다). 여기서는 기록만 한다. write_doc(STAGE_A_PATH, canonical({"stage_a_context": stage_a})) write_doc(SOURCE_MANIFEST_PATH, canonical(source_manifest)) # P-13 — 판정 7. compile_domain_slices 의 반환에는 검증 수행 여부 필드가 없다. @@ -707,6 +761,10 @@ Agent: "expected_runnable_domain_ids": expected_runnable, "slice_count": len(slice_hashes), "fanout_instance_count": len(plan_root.get("task_instances") or []), + # 오케스트레이터는 계획 파일을 읽지 않는다. wildcard fan-out 은 + # 반환 JSON 최상위 dynamic_fanout 리스트로만 확장된다 + # (agent.py _extract_fanout_items). 항목은 planner 가 이미 만든 것을 그대로 넘긴다. + "dynamic_fanout": plan_root.get("task_instances") or [], "digest_guard": seal, "errors": ERRORS, "warnings": WARNINGS} @@ -723,128 +781,182 @@ Agent: - "{{item.compiled_prompt_path}}" - "{{item.slice_path}}" llm_provider: google - llm_model: 'gemini-3.1-flash-lite' + llm_model: 'gemini-3.1-pro-preview' llm_reasoning: high - llm_verbosity: medium + llm_verbosity: low + use_tools: + - localdocs cache_control: mode: auto ttl: 15m - prompt: |- - - Task_C_B_domain_worker - - You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial - litigation counsel. Your role in this task is a **per-domain BO seed worker** within - the Stage 1 Part 2 dynamic fan-out. - - - 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 - `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. - 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 - 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. - - + prompts: + - role: user + content: |- + + Task_C_B_domain_worker + + You are an MCP-enabled LLM agent assisting plaintiff-side Korean civil/commercial + litigation counsel. Your role in this task is a **per-domain BO seed worker** within + the Stage 1 Part 2 dynamic fan-out. + + + 본 task 는 오케스트레이터가 runtime parameter 로 주입한 단일 도메인 + `{{item.domain_id}}` 하나만 처리한다. 다른 도메인의 사실을 자기 산출에 넣지 않는다. + 읽어야 할 것은 두 파일뿐이다 — 조립 프롬프트 `{{item.compiled_prompt_path}}` 와 + 도메인 slice `{{item.slice_path}}`. 프롬프트를 다시 조립하지 않는다. + + - - - - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) - - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) - - - - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) - - - - 다른 도메인의 slice 나 seed 를 읽지 않는다. - - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. - - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. - - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. - - slice 의 source_universe 밖 출처를 인용하지 않는다. - - + + + - `{{item.compiled_prompt_path}}` (조립 프롬프트. preflight 로 이미 실려 있다) + - `{{item.slice_path}}` (도메인 slice. 최상위 키 stage_b_domain_slice) + + + - `{{item.expected_output_path}}` (본 인스턴스의 seed 파일 1개만) + + + - 다른 도메인의 slice 나 seed 를 읽지 않는다. + - 프롬프트를 재조립하지 않는다. 조각을 다시 이어 붙이지 않는다. + - 최종 청구권을 고르지 않는다. 최종 요건충족을 판단하지 않는다. + - BO 식별자를 확정하지 않는다. BO_ID · Evidence · EvidenceTitles 키를 쓰지 않는다. + - slice 의 source_universe 밖 출처를 인용하지 않는다. + + - - - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) - → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. - - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. - - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. - + + - 조립 프롬프트는 rank 10(공통 계약) → 20(의존 공통층) → 30(도메인 overlay) + → 40(특별법 overlay) → 50(실행 가드) 순으로 이미 합성되어 있다. + - 그 본문이 이 task 의 실질 지시다. 본 래퍼는 입출력 계약만 규정한다. + - 프롬프트와 slice 가 어긋나 보이면 임의로 고르지 말고 review_items 에 남긴다. + - - - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. - - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 - component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). - - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. - 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. - - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. - + + - 모든 근거는 slice 의 `source_universe[*].source_id` 안에 있어야 한다. + - 증거 구성요소 이름은 `Default_Agent/routing/evidence_component_union.md` 의 + component_id 만 쓴다. 목록에 없는 이름을 만들지 않는다(P0 판정 A·B). + - 인용한 component_id 는 각 후보의 `registry_component_ids` 배열에 싣는다. + 그 배열이 비어 있지 않은 후보는 R0 에서 증거 유래로 인정된다. + - 붙일 근거가 slice 안에서 직접 읽히지 않으면 비워 두고 review 로 남긴다. + - - 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 - `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 - `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. + + 최상위는 `stage_b_domain_bo_seed_output` 한 키다. 스키마는 + `Default_Agent/platform/schemas/domain_seed_output.schema.v3.json` 이며 + `schema_version` 은 `task_c_bo_stage_b_domain_bo_seed.v3` 로 고정이다. - { - "stage_b_domain_bo_seed_output": { - "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", - "status": "READY", - "task_instance_id": "{{item.task_instance_id}}", - "domain_id": "{{item.domain_id}}", - "registry_version": "", - "registry_index_sha256": "", - "domain_config_sha256": "", - "slice_sha256": "{{item.slice_sha256}}", - "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", - "bo_seed_candidates": [ - { - "seed_id": "<도메인슬러그-001 꼴>", - "bo_type": "", - "source_refs": [], - "registry_component_ids": [], - "element_fact_candidates": [], - "opposing_fact_candidates": [], - "defense_candidates": [], - "evidence_slot_status": [], - "calculation_requests": [], - "dependency_refs": [], - "legal_effect_candidates": [], - "review_items": [], - "extensions": {} + { + "stage_b_domain_bo_seed_output": { + "schema_version": "task_c_bo_stage_b_domain_bo_seed.v3", + "status": "READY", + "task_instance_id": "{{item.task_instance_id}}", + "domain_id": "{{item.domain_id}}", + "registry_version": "", + "registry_index_sha256": "", + "domain_config_sha256": "", + "slice_sha256": "{{item.slice_sha256}}", + "compiled_prompt_sha256": "{{item.compiled_prompt_sha256}}", + "bo_seed_candidates": [ + { + "seed_id": "<도메인슬러그-001 꼴>", + "bo_type": "", + "juristic_act_type": "<법률행위 유형 문자열 또는 null>", + "source_refs": [], + "registry_component_ids": [], + "element_fact_candidates": [], + "opposing_fact_candidates": [], + "defense_candidates": [], + "evidence_slot_status": [], + "calculation_requests": [], + "dependency_refs": [], + "legal_effect_candidates": [], + "party_roles": [], + "time_facts": [], + "object_refs": [], + "amount_facts": [], + "review_items": [], + "extensions": {"domain_payload": {"action_summary": null, "action_type": null}} + } + ], + "unknown_or_unrouted_reviews": [], + "completion_receipt": {}, + "contract_guards": { + "final_conclusion_forbidden": true, + "unknown_values_require_review": true, + "source_membership_required": true, + "strict_json_output": true } - ], - "unknown_or_unrouted_reviews": [], - "completion_receipt": {}, - "contract_guards": { - "final_conclusion_forbidden": true, - "unknown_values_require_review": true, - "source_membership_required": true, - "strict_json_output": true } } - } - 추가 제약 - - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates - · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. - - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. - - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. - 억지로 채우는 것이 비워 두는 것보다 나쁘다. - + 추가 제약 + - 다섯 배열(element_fact_candidates · opposing_fact_candidates · defense_candidates + · calculation_requests · dependency_refs)의 이름은 스키마가 정한 것이다. 바꾸지 않는다. + - `dependency_refs` 는 연결만 남긴다. 의존 도메인의 결론을 복사하지 않는다. + - 후보를 만들 수 없으면 빈 배열로 두고 review_items 에 사유를 남긴다. + 억지로 채우는 것이 비워 두는 것보다 나쁘다. + - 아래 자리들은 BO 호환면 투영(`bo_surface_projection_policy.v1`)이 읽는 1순위 출처다. + 비워 두면 BO.json 의 해당 칸이 폴백 값으로 채워지고 schema_field_fallback 검토가 발행된다. + slice 의 source_universe 안에 근거가 있으면 채운다. 근거가 없으면 비워 두고 사유를 남긴다 — + 추측으로 채우지 않는다. 사건종류 이름을 값으로 쓰지 않는다. + · `juristic_act_type` : 법률행위 유형 문자열 1개(없으면 null). -> JuristicAct.label + · `extensions.domain_payload.action_summary` : 이 후보가 무엇인지 한 문장. -> Action + · `extensions.domain_payload.action_type` : "법률행위(legal acts)" 또는 "사실행위(factual acts)". -> ActionType + · `legal_effect_candidates[]` : {"type_id": "<소문자_스네이크>", "source_refs": [], "registered": true|false}. -> Legal_Keywords + · `time_facts[]` : {"fact_type": "<소문자_스네이크>", "value": "<시점 문자열 또는 null>", "source_refs": []}. -> BehaviorTime · TimeText + · `object_refs[]` : 목적물 식별자 문자열. -> core_field_base.Object + · `amount_facts[]` : {"amount_type": "<소문자_스네이크>", "decimal_value": "<숫자 문자열 또는 null>", "currency": "KRW", "source_refs": []}. -> amount + · `party_roles[]` : {"role": "<소문자_스네이크>", "party_refs": []}. 투영 대상은 아니나 스키마 필드다. + - `slice_sha256` 와 `compiled_prompt_sha256` 은 골격이 준 값을 **글자 그대로** 옮기는 자리다. + 슬라이스 본문을 뒤져 비슷한 이름을 찾지 않는다. 슬라이스에서 읽어 오는 sha 는 + `registry_index_sha256` 과 `domain_config_sha256` 둘뿐이다. + 특히 `activation.activation_manifest_sha256` 은 전 도메인이 같은 값이라 혼동하기 쉽다 — + 그것을 집으면 R0 가 전사 실수로 기록하고 이 산출에 검토 표시를 남긴다. + 다른 도메인의 sha 를 집으면 "남의 자료로 작업했다"로 판정되어 회차가 중단된다. + 그 둘 어디에도 해당하지 않는 값(어디서 왔는지 설명되지 않는 sha)을 적으면 + 더 무겁게 다뤄져 이 후보가 격리된다. 값을 지어내는 것이 가장 나쁘다. + - 아래 다섯 어휘는 스키마가 고정한 것이다. 다른 낱말을 쓰면 S0 신호 게이트가 경성으로 막는다. + R0 의 검증기는 스키마의 부분집합만 보므로 여기서 틀려도 그 단계에서는 걸리지 않는다. + · seed 최상위 `status` : READY | READY_WITH_REVIEW | NO_SUPPORT | BLOCKED | FAILED + · `evidence_slot_status[].status` : filled | partial | missing | conflicted + (요건 슬롯을 뒷받침하는 근거가 충분하면 filled, 일부만이면 partial, + 없으면 missing, 상충 근거가 함께 있으면 conflicted) + · `calculation_requests[].completeness` : ready | partial | blocked | deferred + · `review_items[].severity` 와 `unknown_or_unrouted_reviews[].severity` : info | review | hard_warning | block + · 세 후보 배열의 `source_kind` : meeting_clause | event_candidate | evidence | bo | fact + | signal | registry | law_version | calculation | other + - 아래 네 객체는 `additionalProperties: false` 다. 적힌 키 말고는 **한 개도** 넣지 않는다. + 필수 키를 빠뜨리거나 임의 키를 더하면 스키마 위반이다. + · `element_fact_candidates[]` · `opposing_fact_candidates[]` · `defense_candidates[]` : + {"source_id": "", "source_kind": "<위 어휘>", + "excerpt": "<선택: 근거 문구>", "payload": {}} + — 필수는 source_id · source_kind 둘이다. `slot_id` 나 `fact` 같은 키는 이 배열에 없다. + 슬롯 판정은 `evidence_slot_status[]` 가 맡는다. + · `evidence_slot_status[]` : + {"slot_id": "<요건 슬롯 id>", "status": "<위 어휘>", "source_refs": [], "review_code": null} + · `calculation_requests[]` : + {"calculation_domain": "", "completeness": "<위 어휘>", + "source_refs": [], "review_code": null} + · `review_items[]` : + {"review_code": "<대문자_스네이크>", "severity": "<위 어휘>", + "reason": "<왜 검토가 필요한지 한 문장>", "source_refs": []} + - - - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. - - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. - - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. - - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. - - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. - - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. - + + - 최상위가 `stage_b_domain_bo_seed_output` 한 키인지 점검한다. + - `domain_id` 와 `task_instance_id` 가 주입값과 정확히 같은지 점검한다. + - 모든 `source_refs` 원소가 slice 의 source_universe 안에 있는지 점검한다. + - `bo_type` 이 slice 의 allowed_legal_effect_bo_types 안에 있는지 점검한다. + - `registry_component_ids` 원소가 합집합 목록 안에 있는지 점검한다. + - 금지 키(BO_ID · Evidence · EvidenceTitles · final_*)가 없는지 점검한다. + - - - 자기 도메인 밖으로 나가지 않는다. - - 프롬프트를 다시 만들지 않는다. - - 결론을 내리지 않는다. 후보만 남긴다. - - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. - + + - 자기 도메인 밖으로 나가지 않는다. + - 프롬프트를 다시 만들지 않는다. + - 결론을 내리지 않는다. 후보만 남긴다. + - `write_file(overwrite=true)` 로 `{{item.expected_output_path}}` 하나만 쓴다. + use_tools: - localdocs - task_name: Task_C_BO_R0_seed_reducer_and_exception_planner @@ -861,6 +973,7 @@ Agent: # publisher + domain_join + PostB_1 통합 결정적 reducer. # Spec: Part_2_Improvement_Strategy_Claude_v1.md §7 (defer policy = 개선전략서 X-2, pack 조건 = X-3) from __future__ import annotations + import copy import hashlib import itertools import json @@ -881,6 +994,11 @@ Agent: FANOUT_PLAN_PATH = "fanout/domain_fanout_plan.json" SLICE_DIR = "runtime/domain_slices" SEED_SCHEMA_PATH = "Default_Agent/platform/schemas/domain_seed_output.schema.v3.json" + # R0-6 — 전체 스키마 검증은 신설이 아니다. worker_output_validator 가 이미 + # schema_subset_validator 로 전 계약을 검사하고 있었고(13도메인 1,151건 실측), + # R0 가 그 결과를 guard_warnings 로 강등하고 있었다. 이 정책은 그 결과에 처분을 준다. + SEED_ADMISSION_POLICY_PATH = "Default_Agent/stage1_runtime/seed_admission_policy.v1.json" + SEED_ADMISSION_SCHEMA_VERSION = "stage1_seed_admission_policy.v1" SEED_SCHEMA_VERSION = "task_c_bo_stage_b_domain_bo_seed.v3" SLICE_ROOT_KEY = "stage_b_domain_slice" # R-4 — 머리말이 약속한 worker_output_validator 를 실제로 부른다. 반입은 D0 규약 R-1~R-5. @@ -1195,6 +1313,12 @@ Agent: norm = _dict(policy.get("f0_normalization")) bo_type = cand.get("bo_type") + if not isinstance(bo_type, str): + # 워커가 문자열이 아닌 값을 넣으면 집합 비교가 터진다(dict 면 unhashable). + # 수용 단계가 걸러야 하지만 투영은 마지막 방어선이라 여기서도 막는다 — + # 여기서 죽으면 회차 전체가 죽고, 원인이 워커 값이라는 것도 드러나지 않는다. + note("schema_field_fallback", "BOType", "non_string:%s" % type(bo_type).__name__) + bo_type = norm.get("bo_type_default") if allowed_bo_types and bo_type not in allowed_bo_types: note("legal_effect_uncertain", "BOType", "bo_type") @@ -1297,7 +1421,7 @@ Agent: } # ---------- 워커 출력 수용 검증 (v3 계약 정본 · 정책 status_policy · 계획 해시 대조) ---------- - def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]]) -> None: + def validate_seed_object(seed_obj: dict[str, Any], domain_id: str, plan_row: dict[str, Any], warnings: list[dict[str, Any]], echo_mismatches: list[dict[str, Any]] | None = None) -> None: """v3 seed 루트를 검증한다. status 는 v3 enum 5종을 정책 status_policy 로 가른다. 구판은 READY 계열 2종만 허용해 계약상 적법한 NO_SUPPORT 가 R0 전체를 중단시켰고(C-11), @@ -1319,11 +1443,20 @@ Agent: raise ValueError(f"{domain_id}: NO_SUPPORT with non-empty bo_seed_candidates") elif status not in ("READY", "READY_WITH_REVIEW"): raise ValueError(f"{domain_id}: seed status outside v3 enum: {status!r}") + # R0-7 — 신선도를 워커의 자기 신고로 판정하지 않는다. + # echo 된 sha 는 모델이 옮겨 적은 값이라 전사 실수가 곧 회차 사망이 된다. + # 실측: E-00 이 slice_sha256 자리에 슬라이스 안의 activation_manifest_sha256 을 집었고, + # 그 값은 14개 슬라이스 전부에 같은 값으로 들어 있어 어느 도메인에서든 재발할 수 있다. + # 진짜 신선도는 모델을 거치지 않는 실물 두 값(계획서 sha ↔ R0 가 직접 읽은 문서 sha)이 + # 말한다. 그 대조는 호출부가 하고(PART2_SLICE_STALE), 여기서는 echo 불일치를 + # 수용 정책 관할로 넘길 수 있도록 모아만 둔다. + _echo = echo_mismatches if echo_mismatches is not None else [] for key in ("slice_sha256", "compiled_prompt_sha256"): want = plan_row.get(key) if isinstance(want, str) and want: if seed_obj.get(key) != want: - raise ValueError(f"{domain_id}: stale seed output: {key} mismatch") + _echo.append({"domain_id": domain_id, "key": key, + "expected": want, "received": seed_obj.get(key)}) else: warnings.append({"domain_id": domain_id, "warning": f"fanout plan carries no {key} expectation"}) @@ -1344,6 +1477,450 @@ Agent: if cand.get("ReasonRefs") not in ([], None): raise ValueError(f"{domain_id}.{ref}: ReasonRefs must be []/absent") + # ------------------------------------------------------------------ + # R0-6 — 시드 수용(admission). 검증 결과에 등급을 준다. + # 구조 위반은 막고(block), 어휘·형태 위반은 정규화·수확으로 살리고(repair), + # 살릴 수 없는 것은 후보 하나만 격리한다(quarantine). 회차 전체를 죽이지 않는다. + # 값을 지어내지 않는다 — 표에 없는 낱말과 universe 에 없는 참조는 복구 대상이 아니다. + # ------------------------------------------------------------------ + PATH_INDEX_RE = re.compile(r"\[(\d+)\]") + + def _norm_path(path: str) -> str: + """검증기 경로를 정책 키 모양으로 바꾼다. 인덱스는 [] 로 접는다.""" + body = path.replace("$.stage_b_domain_bo_seed_output.", "").replace("$.stage_b_domain_bo_seed_output", "") + return PATH_INDEX_RE.sub("[]", body).lstrip(".") + + def _path_tokens(path: str) -> list[Any]: + body = path.replace("$.stage_b_domain_bo_seed_output.", "").replace("$.stage_b_domain_bo_seed_output", "") + tokens: list[Any] = [] + for part in body.lstrip(".").split("."): + if not part: + continue + name = part.split("[", 1)[0] + if name: + tokens.append(name) + for hit in PATH_INDEX_RE.findall(part): + tokens.append(int(hit)) + return tokens + + def _resolve_parent(root: Any, tokens: list[Any]) -> tuple[Any, Any]: + """마지막 토큰의 부모 컨테이너와 그 토큰을 돌려준다. 못 찾으면 (None, None).""" + node = root + for tok in tokens[:-1]: + if isinstance(tok, int): + if not isinstance(node, list) or tok >= len(node): + return (None, None) + node = node[tok] + else: + if not isinstance(node, dict) or tok not in node: + return (None, None) + node = node[tok] + return (node, tokens[-1]) if tokens else (None, None) + + def _candidate_index(tokens: list[Any]) -> int | None: + for pos, tok in enumerate(tokens): + if tok == "bo_seed_candidates" and pos + 1 < len(tokens) and isinstance(tokens[pos + 1], int): + return tokens[pos + 1] + return None + + def _sink_for(seed_obj: dict[str, Any], tokens: list[Any], policy: dict[str, Any], norm: str): + """초과 정보를 담을 자리 두 곳을 돌려준다: (컨테이너 지역 payload, 후보 수준 보관 목록). + + 지역 payload 의 키는 **잎 이름 그대로** 쓴다 — 하류 어댑터가 + payload["status"] 처럼 짧은 이름으로 읽기 때문이다. 정규경로를 키로 쓰면 + 바이트는 남아도 아무도 읽지 못한다. + 후보 수준 보관은 dict 가 아니라 **덧붙이기 전용 목록**이다. 같은 정규경로의 + 두 번째 값이 첫 번째를 조용히 덮는 사고를 구조적으로 막는다. + """ + sink_cfg = _dict(policy.get("surplus_sink")) + local: dict[str, Any] | None = None + by_container = _dict(sink_cfg.get("by_container")) + container_key = norm.rsplit(".", 1)[0] if "." in norm else norm + local_name = by_container.get(container_key) + if local_name: + parent, _ = _resolve_parent(seed_obj, tokens) + if isinstance(parent, dict): + bucket = parent.get(local_name) + if not isinstance(bucket, dict): + bucket = {} + parent[local_name] = bucket + local = bucket + shelf: list[Any] | None = None + cand_idx = _candidate_index(tokens) + cands = seed_obj.get("bo_seed_candidates") + if cand_idx is not None and isinstance(cands, list) and cand_idx < len(cands) and isinstance(cands[cand_idx], dict): + node: Any = cands[cand_idx] + steps = str(sink_cfg.get("candidate_path") or "extensions.worker_surplus").split(".") + for step in steps[:-1]: + nxt = node.get(step) + if not isinstance(nxt, dict): + nxt = {} + node[step] = nxt + node = nxt + leaf = steps[-1] + if not isinstance(node.get(leaf), list): + node[leaf] = [] + shelf = node[leaf] + return (local, shelf) + + def _universe_kind(value: str, universe: dict[str, set]) -> str | None: + if value in universe["source_evidence_indexes"]: + return "evidence" + if value in universe["source_event_candidate_ids"]: + return "event_candidate" + if value in universe["source_meeting_clause_ids"]: + return "meeting_clause" + return None + + def _derive_required(seed_obj, tokens, norm, rule, universe, ctx): + """선언된 규칙으로만 채운다. 규칙이 없거나 실패하면 (False, None).""" + parent, key = _resolve_parent(seed_obj, tokens) + if not isinstance(parent, dict): + return (False, None) + how = str(rule.get("rule") or "") + if how == "constant": + value = rule.get("value") + return (True, copy.deepcopy(value)) if "value" in rule else (False, None) + if how == "from_context": + ckey = str(rule.get("key") or "") + return (True, copy.deepcopy(ctx[ckey])) if ckey in ctx else (False, None) + if how == "first_universe_member_of_sibling_array": + known = (universe["source_evidence_indexes"] | universe["source_event_candidate_ids"] + | universe["source_meeting_clause_ids"]) + for sib in rule.get("sibling_candidates") or []: + raw = parent.get(sib) + values = raw if isinstance(raw, list) else ([raw] if isinstance(raw, str) else []) + for item in values: + if isinstance(item, str) and item in known: + return (True, item) + return (False, None) + if how == "universe_kind_of_sibling": + sib = parent.get(str(rule.get("sibling") or "")) + kind = _universe_kind(sib, universe) if isinstance(sib, str) else None + if not kind: + # 형제가 아직 채워지지 않았을 수 있다. 같은 원소의 참조 배열에서 직접 읽는다. + for name in rule.get("fallback_sibling_arrays") or []: + raw = parent.get(name) + values = raw if isinstance(raw, list) else ([raw] if isinstance(raw, str) else []) + for item in values: + if isinstance(item, str): + kind = _universe_kind(item, universe) + if kind: + break + if kind: + break + return (True, kind) if kind else (False, None) + if how == "rename_sibling": + # 워커가 같은 뜻을 다른 이름으로 적었을 때 이름만 바로잡는다. 값은 그대로 옮긴다. + # 옮긴 뒤 원래 키를 지운다 — 남겨 두면 다음 패스에서 초과 속성으로 다시 걸린다. + for sib in rule.get("sibling_candidates") or []: + if sib in parent and parent.get(sib) not in (None, ""): + return (True, parent.pop(sib)) + return (False, None) + if how == "sibling_matching_pattern": + pat = re.compile(str(rule.get("pattern") or "^$")) + for sib in rule.get("sibling_candidates") or []: + value = parent.get(sib) + if isinstance(value, str) and pat.fullmatch(value): + return (True, value) + for sib in rule.get("fallback_rename_sibling") or []: + if sib in parent and parent.get(sib) not in (None, ""): + return (True, parent.pop(sib)) + return (False, None) + return (False, None) + + def _promote_alias(parent, key, norm, policy, record): + """수확 직전에 한 번 더 본다 — 이 키가 선언된 자리의 다른 이름일 뿐인가. + + 그렇다면 자유 공간으로 밀어 넣지 않고 제 자리로 올린다. 워커가 excerpt 를 + value·content 로 부르는 표류가 실측됐고, 그 값은 하류가 실제로 읽는 칸이다. + """ + if not isinstance(parent, dict) or not isinstance(key, str): + return False + container = norm.rsplit(".", 1)[0] if "." in norm else "" + for target_path, rule in _dict(policy.get("required_derivation")).items(): + if str(rule.get("rule") or "") != "rename_sibling": + continue + if target_path.rsplit(".", 1)[0] != container: + continue + target = target_path.rsplit(".", 1)[-1] + if target in parent and parent.get(target) not in (None, ""): + continue + if key in (rule.get("sibling_candidates") or []): + parent[target] = parent.pop(key) + record["received"] = _clip(parent[target]) + record["applied"] = "promoted_to:%s" % target + record["rule_id"] = "alias_promotion:%s" % target_path + return True + return False + + def _harvest(seed_obj, tokens, policy, norm, record): + """규약 밖 값을 버리지 않고 자유 공간으로 옮긴다. 옮긴 사실을 기록한다.""" + parent, key = _resolve_parent(seed_obj, tokens) + if parent is None: + return False + if _promote_alias(parent, key, norm, policy, record): + return True + local, shelf = _sink_for(seed_obj, tokens, policy, norm) + if isinstance(parent, list) and isinstance(key, int): + if key >= len(parent): + return False + moved = parent.pop(key) + leaf = None + elif isinstance(parent, dict): + if key not in parent: + return False + moved = parent.pop(key) + leaf = key + else: + return False + placed = [] + # 1) 컨테이너 지역 payload — 하류가 읽는 짧은 이름으로. 이미 있으면 덮지 않는다. + if isinstance(local, dict) and isinstance(leaf, str) and leaf not in local: + local[leaf] = moved + placed.append("payload.%s" % leaf) + # 2) 후보 수준 보관 — 덧붙이기 전용이라 어떤 값도 덮이지 않는다. 원래 경로를 함께 남긴다. + if isinstance(shelf, list): + # D11 — norm 은 인덱스를 접으므로 같은 컨테이너의 두 값이 구별되지 않는다. + # 원본 경로를 함께 남겨야 부모 원소와의 결합(예: role ↔ party_refs)을 복원할 수 있다. + shelf.append({"path": norm, "source_path": record.get("path"), "value": moved}) + placed.append("worker_surplus[]") + if not placed: + # 보관할 자리가 없으면 지우지 않는다. 되돌려 놓고 실패로 돌려주면 + # 이 위반은 잔여로 남아 정책이 정한 처분(기본 review)으로 간다. + # 뿌리 수준 배열(unknown_or_unrouted_reviews 등)이 여기 해당한다 — + # 후보에 매이지 않아 보관처가 없는데, 그렇다고 사건 자료를 버릴 수는 없다. + if isinstance(parent, list) and isinstance(key, int): + parent.insert(key, moved) + elif isinstance(parent, dict) and isinstance(key, str): + parent[key] = moved + return False + record["received"] = _clip(moved) + record["applied"] = "+".join(placed) + # 기계적 키 이동과, 실질 서술을 담은 채 **객체째** 밀려난 것은 검토 무게가 다르다. + # 후자를 같은 등급에 섞으면 1,200건 속 몇 건을 사람이 찾아내야 한다. + # 판정은 dict 로 좁힌다 — 스칼라 한 개의 자리 이동(당사자 이름·slot_id 문자열)은 + # 원소가 사라진 것이 아니라 키가 옮겨진 것이라 무게가 다르다. + if isinstance(moved, dict): + for _k in ("excerpt", "value", "fact", "content", "reason", "statement"): + _v = moved.get(_k) + if isinstance(_v, str) and _v.strip(): + record["content_bearing"] = True + break + return True + + def _evict(seed_obj, tokens, policy, norm, record, quarantined=None): + """복구 불가한 객체 하나를 배열에서 들어내 보관한다. 후보 자체는 삭제하지 않는다. + + D4·D8 — 후보를 pop 하면 ① 예산에 잡히지 않아 소리 없이 사라지고 + ② 뒤 후보의 인덱스가 밀려 이미 기록한 격리 표시가 다른 후보를 가리킨다. + 후보 수준이면 삭제 대신 격리로 돌린다. + """ + trimmed = list(tokens) + while trimmed and not isinstance(trimmed[-1], int): + trimmed.pop() + if not trimmed: + return False + if len(trimmed) == 2 and trimmed[0] == "bo_seed_candidates": + if quarantined is None: + return False + quarantined.add(trimmed[1]) + record["applied"] = "quarantined_candidate" + return True + return _harvest(seed_obj, trimmed, policy, _norm_path("$.stage_b_domain_bo_seed_output." + norm), record) + + def _clip(value: Any, limit: int = 200) -> Any: + try: + text = json.dumps(value, ensure_ascii=False) + except Exception: + text = str(value) + return text if len(text) <= limit else text[:limit] + "…" + + def admit_seed(seed_obj, domain_id, schema, slice_doc, universe, policy, records, plan_row=None, validator=None): + """검증 -> 처분 -> 복구를 수렴할 때까지 돌리고, 남은 것은 후보 격리로 넘긴다. + + 돌려주는 것: (수용된 seed, 격리된 후보 인덱스 집합, 경성 중단 사유 목록) + """ + # D1 — 검증기는 인자로 받는다. main() 지역 이름을 전역처럼 읽으면 NameError 로 즉사한다. + if validator is None: + return (seed_obj, set(), []) + + def _raw_validate(obj): + return validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": obj}, schema=schema, + expected_domain_id=domain_id, slice_document=slice_doc) + + def _validate(obj): + """검증기 자체가 터질 수 있다(예: bo_type 이 dict 면 unhashable). + + D9 — 예외를 코드로 바꾸기만 하면 부족하다. 그 오류의 path 는 뿌리라 + 후보를 지목하지 못하고, 처분이 뿌리 잔여(review)로 강등돼 문제 후보가 + 그대로 투영으로 흘러 project_to_bo_surface 에서 다시 죽는다. + 그래서 **어느 후보가 터뜨렸는지 후보 단위로 좁혀** path 에 인덱스를 실어 준다. + """ + try: + return _raw_validate(obj) + except Exception as exc: + detail = "%s: %s" % (type(exc).__name__, str(exc)[:160]) + errors = [] + cands = _list(obj.get("bo_seed_candidates")) + for idx in range(len(cands)): + probe = dict(obj) + probe["bo_seed_candidates"] = [cands[idx]] + try: + _raw_validate(probe) + except Exception: + errors.append({"code": "VALIDATOR_CRASHED", + "path": "$.stage_b_domain_bo_seed_output.bo_seed_candidates[%d]" % idx, + "message": detail}) + if not errors: + # 후보를 좁히지 못했다. 뿌리 문제이므로 명시적으로 끊는다 — + # review 로 강등해 투영에서 죽게 두는 것이 최악이다. + errors.append({"code": "VALIDATOR_CRASHED_ROOT", + "path": "$.stage_b_domain_bo_seed_output", "message": detail}) + return {"errors": errors, "warnings": []} + disposition = _dict(policy.get("disposition_by_code")) + unknown_disp = str(policy.get("unknown_code_disposition") or "quarantine") + synonyms = _dict(policy.get("enum_synonyms")) + case_fold = bool(policy.get("enum_case_fold")) + enum_unmapped = str(policy.get("enum_unmapped_disposition") or "repair_evict") + derivations = _dict(policy.get("required_derivation")) + derive_failed = str(policy.get("derivation_failed_disposition") or "repair_evict") + order = [str(v) for v in (policy.get("repair_order") or [])] + max_passes = int(policy.get("max_repair_passes") or 3) + blocks: list[dict[str, Any]] = [] + quarantined: set[int] = set() + + for _pass in range(max_passes): + cands = _list(seed_obj.get("bo_seed_candidates")) + # task_instance_id 는 지어내지 않는다 — A0 의 fan-out 계획 행이 든 값을 쓴다. + ctx = {"domain_id": domain_id, "seed_count": len(cands), + "emitted_seed_ids": [str(_dict(c).get("seed_id") or "") for c in cands]} + _plan_tid = _dict(plan_row).get("task_instance_id") + if isinstance(_plan_tid, str) and _plan_tid: + ctx["task_instance_id"] = _plan_tid + # r0_membership_key 는 R0 가 자기 관측으로 결정적으로 만든다(워커가 알 수 없는 값이다). + ctx["r0_membership_key"] = "%s:%s" % ( + domain_id, hashlib.sha256(json.dumps(ctx["emitted_seed_ids"], ensure_ascii=False, + sort_keys=True).encode("utf-8")).hexdigest()[:32]) + report = _validate(seed_obj) + errors = [e for e in (report.get("errors") or []) if isinstance(e, dict)] + if not errors: + break + buckets: dict[str, list[dict[str, Any]]] = {} + for err in errors: + disp = str(disposition.get(str(err.get("code"))) or unknown_disp) + buckets.setdefault(disp, []).append(err) + for err in buckets.get("block") or []: + blocks.append({"code": err.get("code"), "path": err.get("path"), "message": err.get("message")}) + if blocks: + return (seed_obj, quarantined, blocks) + for err in buckets.get("review") or []: + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "review", "received": None, "applied": None, + "rule_id": "disposition:review"}) + progressed = False + for action in order: + # 원소를 들어내는 처분(harvest·evict)만 내림차순으로 돈다 — 앞 인덱스를 먼저 + # 지우면 뒤 경로가 밀리기 때문이다. 값을 채우는 처분(normalize·derive)은 + # 오름차순이어야 한다: source_kind 는 형제 source_id 를 읽으므로 순서가 뒤집히면 + # 형제가 아직 없어 파생이 실패하고, 그 실패가 요건사실 원소를 통째로 들어낸다. + _removes = action in ("repair_harvest", "repair_evict") + for err in sorted(buckets.get(action) or [], + key=lambda e: PATH_INDEX_RE.sub( + lambda m: "[%05d]" % int(m.group(1)), str(e.get("path") or "")), + reverse=_removes): + path = str(err.get("path") or "") + tokens = _path_tokens(path) + if not tokens: + continue + norm = _norm_path(path) + rec = {"domain_id": domain_id, "code": err.get("code"), "path": path, + "action": action, "received": None, "applied": None, "rule_id": None} + done = False + if action == "repair_normalize": + parent, key = _resolve_parent(seed_obj, tokens) + table = _dict(synonyms.get(norm)) + if isinstance(parent, dict) and key in parent and table: + raw = parent.get(key) + probe = raw.casefold() if (case_fold and isinstance(raw, str)) else raw + mapped = table.get(probe) if isinstance(probe, str) else None + if mapped is not None: + rec["received"], rec["applied"] = _clip(raw), mapped + rec["rule_id"] = "enum_synonyms:%s" % norm + parent[key] = mapped + done = True + if not done and enum_unmapped == "repair_evict": + rec["rule_id"] = "enum_unmapped:%s" % norm + done = _evict(seed_obj, tokens, policy, norm, rec, quarantined) + elif action == "repair_derive": + rule = _dict(derivations.get(norm)) + if rule: + ok, value = _derive_required(seed_obj, tokens, norm, rule, universe, ctx) + if ok: + parent, key = _resolve_parent(seed_obj, tokens) + if isinstance(parent, dict): + rec["received"], rec["applied"] = None, _clip(value) + rec["rule_id"] = "required_derivation:%s" % norm + parent[key] = value + done = True + if not done and derive_failed == "repair_evict": + rec["rule_id"] = "derivation_failed:%s" % norm + done = _evict(seed_obj, tokens, policy, norm, rec, quarantined) + elif action == "repair_harvest": + rec["rule_id"] = "surplus_sink:%s" % norm + done = _harvest(seed_obj, tokens, policy, norm, rec) + elif action == "repair_evict": + rec["rule_id"] = "evict:%s" % norm + done = _evict(seed_obj, tokens, policy, norm, rec, quarantined) + elif action == "quarantine": + idx = _candidate_index(tokens) + if idx is not None: + quarantined.add(idx) + rec["rule_id"] = "quarantine:%s" % norm + done = True + if done: + cand_idx = _candidate_index(tokens) + if cand_idx is not None: + cand = _list(seed_obj.get("bo_seed_candidates")) + if cand_idx < len(cand): + rec["candidate_ref"] = _dict(cand[cand_idx]).get("candidate_ref") or _dict(cand[cand_idx]).get("seed_id") + records.append(rec) + progressed = True + if progressed: + break + if not progressed: + break + + # 수렴하지 않고 남은 위반은 후보 단위로 격리한다. 회차는 계속된다. + report = _validate(seed_obj) + for err in (report.get("errors") or []): + if not isinstance(err, dict): + continue + # review 로 수용하기로 선언된 코드는 남아 있는 것이 정상이다. 다시 격리하지 않는다. + if str(disposition.get(str(err.get("code"))) or "") == "review": + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "review", "received": None, "applied": None, + "rule_id": "disposition:review(residual)"}) + continue + tokens = _path_tokens(str(err.get("path") or "")) + idx = _candidate_index(tokens) + if idx is None: + root_disp = str(policy.get("residual_root_disposition") or "review") + if root_disp == "block": + blocks.append({"code": err.get("code"), "path": err.get("path"), + "message": err.get("message"), "note": "root-level residue"}) + else: + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "review", "received": None, "applied": None, + "rule_id": "residual_root"}) + else: + quarantined.add(idx) + records.append({"domain_id": domain_id, "code": err.get("code"), "path": err.get("path"), + "action": "quarantine", "received": None, "applied": None, + "rule_id": "residual_after_repair"}) + return (seed_obj, quarantined, blocks) + # ---------- PostB_1 이식: sort key / duplicate keys / schema risk ---------- def _source_refs(seed: dict[str, Any]) -> dict[str, list[str]]: provenance = _dict(seed.get("provenance")) @@ -1451,6 +2028,17 @@ Agent: # 1) 워커 출력 수용: 검증 -> 투영. 워커 seed 파일은 손대지 않는다 — # 선언표(stage1_part_interface.v1)가 기록자를 워커 하나로 정했다(R0-5). + admission_policy = _dict(json.loads(_verify_asset(SEED_ADMISSION_POLICY_PATH))) + if admission_policy.get("schema_version") != SEED_ADMISSION_SCHEMA_VERSION: + raise RuntimeError("PART2_SEED_ADMISSION_POLICY_INVALID") + admission_enforcing = str(admission_policy.get("enforcement") or "enforce") == "enforce" + admission_records: list[dict[str, Any]] = [] + declared_candidate_total = 0 + admission_blocks: list[dict[str, Any]] = [] + quarantined_by_domain: dict[str, set] = {} + admitted_total = 0 + quarantined_total = 0 + seed_objects: dict[str, dict[str, Any]] = {} projected_candidates: dict[str, list[dict[str, Any]]] = {} review_handoff_items: list[dict[str, Any]] = [] @@ -1463,33 +2051,124 @@ Agent: seed_obj = _dict(outer.get("stage_b_domain_bo_seed_output")) if not seed_obj: raise RuntimeError(f"{domain_id}: stage_b_domain_bo_seed_output missing") - validate_seed_object(seed_obj, domain_id, PLAN_ROWS.get(domain_id) or {}, guard_warnings) + _plan_row = PLAN_ROWS.get(domain_id) or {} + _echo_mismatches: list[dict[str, Any]] = [] + validate_seed_object(seed_obj, domain_id, _plan_row, guard_warnings, _echo_mismatches) # 슬라이스는 검증기 유무와 무관하게 읽는다 — worker_output_validator 와 # BOType 허용 어휘(allowed_legal_effect_bo_types, registry 유래)가 이 값을 쓴다. + # 원문도 함께 든다 — 신선도 판정의 실물 근거가 이 바이트열의 해시다. + # D3 — 원문 읽기가 실패해도 문서 자체는 기존 경로로 다시 시도한다. read_raw 는 봉투가 + # 다르면 raise 하므로, 그것 때문에 slice_doc 까지 잃으면 allowed_legal_effect_bo_types 가 + # 조용히 빈 집합이 된다. + _slice_raw = None + slice_doc = None try: - slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + _slice_raw = read_raw("%s/%s.json" % (SLICE_DIR, domain_id)) + slice_doc = json.loads(_slice_raw) except Exception: - slice_doc = None + _slice_raw = None + try: + slice_doc = read_json_doc("%s/%s.json" % (SLICE_DIR, domain_id)) + except Exception: + slice_doc = None + # R0-7a — 신선도의 실물 근거. 계획서가 든 sha 와 R0 가 지금 읽은 문서의 sha 를 맞춘다. + # 둘 다 모델을 거치지 않으므로 여기서 어긋나면 정말로 낡았거나 배포가 어긋난 것이다. + _want_slice = _plan_row.get("slice_sha256") + if isinstance(_want_slice, str) and _want_slice and _slice_raw is not None: + # sha_text 는 A0 블록 전용이다. R0 는 이 블록의 관용대로 인라인으로 센다. + _actual_slice = hashlib.sha256(_slice_raw.encode("utf-8")).hexdigest() + if _actual_slice != _want_slice: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SLICE_STALE", + "message": "fan-out 계획이 든 slice sha 와 실제 slice 문서가 다르다. 시드가 아니라 슬라이스가 어긋났다.", + "domain_id": domain_id, "expected": _want_slice, "actual": _actual_slice, + }, ensure_ascii=False)) + # R0-7b — echo 불일치의 등급은 정책이 정한다. 다만 두 실패는 성격이 다르므로 코드를 가른다. + # · 다른 슬라이스의 sha 를 echo 했다 = 워커가 **엉뚱한 슬라이스로 작업**했다는 뜻이다. + # 실물 대조로는 원리상 못 잡는다(그 도메인 슬라이스는 멀쩡하다). 중대 위반이다. + # · 그 밖(예: activation_manifest_sha256 을 집음)은 전사 실수다. 검토로 족하다. + # D2 — 슬라이스를 못 읽어 실물 대조를 못 한 도메인은 신선도 방어가 0 이 된다. + # 그 경우에는 echo 를 검토로 강등하지 않는다. + # 판정은 slice 에 한정하지 않는다 — 다른 도메인의 compiled_prompt sha 를 집은 것도 + # 그 도메인 자료를 읽었다는 뜻이라 같은 무게다. 두 필드를 합쳐 "남의 자료" 집합을 만든다. + _foreign_material_shas = set() + _own_material_shas = {str(_plan_row.get(_mk)) for _mk in ("slice_sha256", "compiled_prompt_sha256") + if _plan_row.get(_mk)} + for _d2, _r in PLAN_ROWS.items(): + if _d2 == domain_id or not isinstance(_r, dict): + continue + for _mk in ("slice_sha256", "compiled_prompt_sha256"): + if _r.get(_mk): + _foreign_material_shas.add(str(_r.get(_mk))) + # D7 — 두 도메인의 자료가 우연히 같아지면(예: 빈 도메인 둘) 정상 산출이 남의 것으로 오판된다. + # 자기 값을 빼서 원천 차단한다. 현 데이터에는 중복이 없지만 데이터의 성질에 기대지 않는다. + _foreign_material_shas -= _own_material_shas + # D8 — 우리가 양성으로 아는 유일한 오집합은 activation_manifest_sha256 이다(값으로 식별된다). + # 그것만 전사 실수로 보고, 설명되지 않는 값은 미상으로 두어 더 무겁게 다룬다. + _known_slip_sha = str(_dict(_dict(_dict(slice_doc).get(SLICE_ROOT_KEY)).get("activation")).get( + "activation_manifest_sha256") or "") + _disp_map = _dict(admission_policy.get("disposition_by_code")) + for _em in _echo_mismatches: + _rcv = str(_em.get("received")) + if _rcv in _foreign_material_shas: + _code = "SEED_ECHO_FOREIGN_SLICE" + elif _slice_raw is None: + # 원문이 없으면 실물 대조가 불가능하다. 이 상태에서는 known slip 이라도 강등하지 않는다 — + # activation_manifest_sha256 은 전 도메인이 공유하는 값이라 어느 슬라이스를 읽었는지에 + # 대해 아무것도 말해 주지 않는다. 그것을 근거로 가볍게 볼 수 없다. + _code = "SEED_ECHO_SHA_UNVERIFIABLE" + elif _known_slip_sha and _rcv == _known_slip_sha: + _code = "SEED_ECHO_SHA_MISMATCH" + else: + _code = "SEED_ECHO_SHA_UNEXPLAINED" + _disp = str(_disp_map.get(_code) or ("block" if _code == "SEED_ECHO_FOREIGN_SLICE" + else "review" if _code == "SEED_ECHO_SHA_MISMATCH" + else "quarantine")) + if _disp == "block": + raise RuntimeError(json.dumps(dict( + {"reason_code": "PART2_%s" % _code}, **_em), ensure_ascii=False)) + admission_records.append({ + "domain_id": domain_id, "code": _code, + "path": "$.stage_b_domain_bo_seed_output.%s" % _em["key"], + "action": _disp, "received": _em.get("received"), "applied": None, + "rule_id": "echo_sha_not_trusted:%s" % _em["key"]}) slice_root = _dict(_dict(slice_doc).get(SLICE_ROOT_KEY)) if isinstance(slice_doc, dict) else {} allowed_bo_types = set(_strings(slice_root.get("allowed_legal_effect_bo_types"))) allowed_bo_types_by_domain[domain_id] = allowed_bo_types - # R-4 — 스키마와 슬라이스를 실제로 넘긴다. 넘기지 않으면 검증이 조용히 건너뛰어진다. + # R0-6 — 검증 결과를 경고로 흘리지 않고 등급대로 처분한다. + # shadow 모드에서는 계산·기록만 하고 seed 를 바꾸지 않는다(첫 전환 회차용). + # D13 — 회계 누적은 검증기 유무와 무관하다. 분기 안에 두면 검증기 반입이 + # 실패한 회차에서 declared=0 · admitted=N 이 되어 헛경보(lost 음수)가 난다. + declared_candidate_total += len(_list(seed_obj.get("bo_seed_candidates"))) if worker_validator is not None: - report = worker_validator.validate_worker_output( - {"stage_b_domain_bo_seed_output": seed_obj}, - schema=seed_schema, - expected_domain_id=domain_id, - slice_document=slice_doc) - for item in report.get("errors") or []: - guard_warnings.append({"code": "WORKER_OUTPUT_ERROR", "domain_id": domain_id, - "detail": item}) - for item in report.get("warnings") or []: + for item in (worker_validator.validate_worker_output( + {"stage_b_domain_bo_seed_output": seed_obj}, schema=seed_schema, + expected_domain_id=domain_id, slice_document=slice_doc).get("warnings") or []): guard_warnings.append({"code": "WORKER_OUTPUT_REVIEW", "domain_id": domain_id, "detail": item}) + work_obj = seed_obj if admission_enforcing else copy.deepcopy(seed_obj) + admitted_obj, quarantined_idx, blocks = admit_seed( + work_obj, domain_id, seed_schema, slice_doc, universe, + admission_policy, admission_records, PLAN_ROWS.get(domain_id) or {}, + worker_validator) + if blocks: + admission_blocks.extend({"domain_id": domain_id, **b} for b in blocks) + if admission_enforcing: + seed_obj = admitted_obj + quarantined_by_domain[domain_id] = quarantined_idx + else: + quarantined_by_domain[domain_id] = set() + for b in blocks: + guard_warnings.append({"code": "SEED_ADMISSION_SHADOW_BLOCK", + "domain_id": domain_id, "detail": b}) cands = _list(seed_obj.get("bo_seed_candidates")) projected: list[dict[str, Any]] = [] projection_reviews: list[dict[str, Any]] = [] + skip_idx = quarantined_by_domain.get(domain_id) or set() for idx, cand in enumerate(cands): + if idx in skip_idx: + quarantined_total += 1 + continue if not isinstance(cand, dict): raise RuntimeError(f"{domain_id}.bo_seed_candidates[{idx}] must be object") cand = ensure_candidate_ref(cand, domain_id, idx) @@ -1552,6 +2231,70 @@ Agent: entry["action_source"] = note_item.get("source") review_handoff_items.append(entry) + # R0-6b — 수용 결산. 격리는 후보 단위이고, 회차 중단은 예산을 넘을 때만이다. + admitted_total = sum(len(v) for v in projected_candidates.values()) + # D4 — 수용 단계에서 후보가 사라지는 경로는 격리 하나뿐이어야 한다. + # 원본 후보 수와 (수용 + 격리)가 맞지 않으면 소리 없이 없어진 것이 있다는 뜻이다. + if admission_enforcing and declared_candidate_total != admitted_total + quarantined_total: + guard_warnings.append({"code": "SEED_ADMISSION_ACCOUNTING_DRIFT", + "declared": declared_candidate_total, "admitted": admitted_total, + "quarantined": quarantined_total, + "lost": declared_candidate_total - admitted_total - quarantined_total}) + if admission_blocks and admission_enforcing: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SEED_ADMISSION_BLOCKED", + "message": "시드 뿌리 계약이 깨졌다. 복구 대상이 아니다.", + "blocks": admission_blocks[:20], "block_count": len(admission_blocks), + }, ensure_ascii=False)) + budget = _dict(admission_policy.get("quarantine_budget")) + seen_total = admitted_total + quarantined_total + ratio = (quarantined_total / seen_total) if seen_total else 0.0 + max_ratio = budget.get("max_quarantined_candidate_ratio") + min_admitted = budget.get("min_admitted_candidates_run") + if admission_enforcing and isinstance(max_ratio, (int, float)) and seen_total and ratio > float(max_ratio): + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SEED_ADMISSION_BUDGET_EXCEEDED", + "message": "격리 비율이 한계를 넘었다. 모델 잡음이 아니라 계약 파손으로 본다.", + "quarantined": quarantined_total, "seen": seen_total, + "ratio": round(ratio, 4), "max_ratio": max_ratio, + }, ensure_ascii=False)) + if admission_enforcing and isinstance(min_admitted, int) and admitted_total < min_admitted: + raise RuntimeError(json.dumps({ + "reason_code": "PART2_SEED_ADMISSION_EMPTY", + "message": "수용된 후보가 없다.", "admitted": admitted_total, + }, ensure_ascii=False)) + _adm_review = _dict(admission_policy.get("review")) + _adm_counter = 0 + for rec in admission_records: + _adm_counter += 1 + _is_q = rec.get("action") == "quarantine" + _is_c = bool(rec.get("content_bearing")) and not _is_q + review_handoff_items.append({ + "review_id": "R0:admission:%03d" % _adm_counter, + "source_domain": rec.get("domain_id"), + "severity": str((_adm_review.get("quarantine_severity") if _is_q + else _adm_review.get("evicted_with_content_severity") if _is_c + else _adm_review.get("severity")) or "SOFT_WARNING"), + "issue_type": str((_adm_review.get("quarantine_issue_type") if _is_q + else _adm_review.get("evicted_with_content_issue_type") if _is_c + else _adm_review.get("issue_type")) or "seed_admission_repair"), + "source_review_code": rec.get("code"), + "source_event_candidate_ids": [], + "source_evidence_indexes": [], + "source_meeting_clause_ids": [], + "downstream_owner": str(_adm_review.get("downstream_owner") or "Stage2"), + "template_note": "수용 단계 처분: %s · 경로 %s · 규칙 %s · 원값 %s -> 적용 %s (후보 %s)" % ( + rec.get("action"), rec.get("path"), rec.get("rule_id"), + rec.get("received"), rec.get("applied"), rec.get("candidate_ref")), + }) + guard_warnings.append({"code": "SEED_ADMISSION_SUMMARY", + "enforcement": admission_policy.get("enforcement"), + "repairs": sum(1 for r in admission_records if str(r.get("action") or "").startswith("repair")), + "reviews": sum(1 for r in admission_records if r.get("action") == "review"), + "quarantined_candidates": quarantined_total, + "admitted_candidates": admitted_total, + "quarantine_ratio": round(ratio, 4)}) + # 2) ledger 구성 — 원장은 워커 원본이 아니라 투영본을 읽는다 (R0-2 배선). # 워커 원본에는 candidate_ref 가 없으므로(봉인 스키마) 원본을 넣으면 아래 검사에서 즉사한다. input_candidate_total = 0 @@ -2507,6 +3250,13 @@ Agent: action = field_decisions.get((candidate_ref, "Action"), seed.get("Action") or domain_payload.get("action_summary") or _dict(seed.get("core_field_base")).get("Action_proposal")) # F0-1 — 어휘의 정본은 registry 합집합(bo_types)이다. {"event","state"} 하드코딩은 # claim 등 여덟 도메인의 선언값을 침묵 덮어쓰던 자리다(C-5). 기본값은 정책 f0_normalization 이 선언한다. + if not isinstance(bo_type, str): + # 워커가 문자열이 아닌 값을 넣으면 집합 비교 자체가 터진다(unhashable). + # 수용 단계가 걸러야 하지만, 투영은 마지막 방어선이라 여기서도 막는다. + normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", + "received": type(bo_type).__name__, + "fallback": f0_norm.get("bo_type_default")}) + bo_type = f0_norm.get("bo_type_default") if bo_type not in bo_types: normalization_notes.append({"candidate_ref": candidate_ref, "field": "BOType", "received": bo_type, "fallback": f0_norm.get("bo_type_default")}) bo_type = f0_norm.get("bo_type_default") diff --git a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/확장_최적워크플로우_연구_프롬프트.txt b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/확장_최적워크플로우_연구_프롬프트.txt index f29020cf..218b0501 100644 --- a/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/확장_최적워크플로우_연구_프롬프트.txt +++ b/Case_02_Comparison_Research/YAML_Prompts/1. Stage_1/v.7/extension_research/확장_최적워크플로우_연구_프롬프트.txt @@ -2282,6 +2282,14 @@ stage_1_2_3_assets_distribution_location.md`의 `§8. 특기사항 (판독 중 ================================== +┌────────────────────────────────────┐ +│ Stage 1 -- Part 4 │ +│ YAML Docs Update Strategy │ +└────────────────────────────────────┘ +┌────────────────────────────────────────────┐ +│ Stage 1 -- Part 1|2||3 │ +│ yaml 실행 체크하면서 동시 진행 │ +└────────────────────────────────────────────┘