diff --git a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00.yml b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00.yml index df39cc88..593248b1 100644 --- a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00.yml +++ b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00.yml @@ -1,6502 +1,3057 @@ Agent: - name: Stage_2_S2_00_v2 - version: "1.2.0" - description: >- - Stage 1 Part 1-4의 고정 입력을 검증·보존·정규화하고 claim-neutral - cluster slice와 S2_10 structured-context handoff plan을 생성하는 - S2_00 deterministic ingress Agent. + name: Stage_2_S2_00_v3 + version: 3.0.0 + description: Stage 1 사건·배포 root와 {{prev.###}} 결과물 참조를 직접 받아 C00 검증, C05 원본 보존·정규화, C10 claim-neutral cluster, + C15 묶음·status-last 발행을 단일 비 LLM task로 수행한다. metadata: workflow_id: S2_00 execution_class: NON-LLM-DETERMINISTIC execution_authority: MCP_CODE_EXECUTOR_INLINE - release_ref: Default_Agent/Stage_2_Clean/manifest/stage2_release.json - workflow_contract_ref: Default_Agent/Stage_2_Clean/workflows/S2_00_stage1_ingress_normalize_and_bundle_compile.yml - deployment_binding_ref: Default_Agent/Stage_2_Clean/deployment/stage2_code_executor_binding.yml - implementation_status: S2_00_PREPARE_OFFLINE_VERIFIED_BACKEND_ARGUMENT_GATE_AND_SERIALIZATION_UNVERIFIED_LIVE_ADMISSION_PENDING + implementation_status: IMPLEMENTED_OFFLINE_VERIFIED_LIVE_NOT_RUN + algorithm_version: s2_00_direct_ingress/3.0.0 + input_contract: + stage1_run_root_ref: '{{prev.stage1_run_root_ref}}' + stage1_deployment_root_ref: '{{prev.stage1_deployment_root_ref}}' + stage1_result_reference: '{{prev.###}}; ### = Stage 1 결과물 파일' + source_contract_authority: parameters.code::SOURCE_POLICY + output_contract: + root: stage2_runs/from-stage1//s2_00/ + states: + - READY + - READY_WITH_ISSUES + - BLOCKED + normal_artifact_count: 5 + status_last: ingress/ingress_status.json + publication_semantics: STATUS_LAST_LOGICAL_COMMIT; SAME_ROOT_CONCURRENT_WRITERS_UNVERIFIED + execution_admission: DEV_FIXTURE_RELEASE + standalone_contract: true Stages: - - name: S2_00 - description: >- - 첫 Code Executor task에서 고정 request를 준비하고, 다음 ingress - task의 한 호출 안에서 C00, C05, C10, C15를 순차 실행한다. - localdocs binary IO와 status-last 논리 배리어로 결과를 발행한다. - prevs: [] - nexts: [] - tools: - mcpServers: - localdocs: - type: streamable-http - url: http://mcp-localdocs:8012/mcp - code-executor: - type: streamable-http - url: https://code-executor.mcp.eroomai.com/mcp - tasks: - - task_name: Task_S2_00_prepare_request - description: >- - 네 외부 문자열 인자 request_id, attempt_id, stage1_run_root_ref, - stage1_deployment_root_ref를 검증하고 고정 localdocs request 경로에 - 저장·read-back한다. Backend 인자 결속 및 성공 게이트는 미검증이며 - 인자가 결속되지 않은 직접 실행은 실패로 종료한다. - mcp: code-executor - tool_name: run_code - parameters: - language: python - requirements: "httpx==0.28.1" - network: agent-network - timeout: 300 - code: |- - #!/usr/bin/env python3 - """Prepare the fixed S2_00 control request from four explicit caller values. + - name: S2_00 + description: 두 root 직접 전달과 Stage 1 결과물 참조 사용. 별도 준비 task 없이 한 run_code에서 원본·배포 검증, 의미 보존, cluster/bundle 구성, + 자체 출력 검증 및 마지막 status 발행. + prevs: [] + nexts: [] + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + tasks: + - task_name: Task_S2_00_deterministic_ingress + description: backend Python 환경이 제공하는 {{prev.###}} 원본 결과물을 재사용한다. 필요한 배포·원본 검증만 수행하고 자체 5개 정상 산출물 또는 차단 진단을 + 발행한다. 기존 DEV release의 실제 사건 발행 금지를 유지한다. + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx==0.28.1 + network: agent-network + timeout: 300 + code: | + #!/usr/bin/env python3 + """S2_00 direct Stage 1 ingress; one deterministic Code Executor task. - The backend argument-delivery contract is not bound. Direct execution fails - closed; callers must invoke ``prepare_request`` with the four named values. - """ + Stage 1 results are supplied through the backend's prev file references. + This module owns its contract; no downstream Agent or output schema is loaded. + MCP transport follows the required Code Executor notebook and SKILL guide. + """ + from __future__ import annotations + import base64 + import binascii + from collections import Counter, defaultdict + import contextlib + from dataclasses import dataclass + import hashlib + import io + import itertools + import json + import math + import os + from pathlib import Path, PurePosixPath + import posixpath + import re + import stat + import sys + import tempfile + import unicodedata + from typing import Any, Callable, Iterable, Mapping, MutableMapping, Sequence - from __future__ import annotations - - import base64 - import hashlib - import json - from pathlib import PurePosixPath - import re - import sys - import unicodedata - from typing import Any, Mapping + ALGORITHM_VERSION = "s2_00_direct_ingress/3.0.0" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_PROTOCOL_VERSION = "2025-03-26" + INLINE_CLIENT_NAME = "liti-stage2-s2-00-direct" + INLINE_CLIENT_VERSION = "3.0.0" + INLINE_USER_HASH = r"""{{__user_hash__}}""" + INLINE_WORKSPACE_HASH = r"""{{__workspace_hash__}}""" + RAW_RUN_ROOT = r"""{{prev.stage1_run_root_ref}}""" + RAW_DEPLOYMENT_ROOT = r"""{{prev.stage1_deployment_root_ref}}""" - REQUEST_PATH = "stage2_control/s2_00_request.json" - LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" - MCP_PROTOCOL_VERSION = "2025-03-26" - INLINE_USER_HASH = "{{__user_hash__}}" - INLINE_WORKSPACE_HASH = "{{__workspace_hash__}}" - REQUEST_KEYS = frozenset({ - "schema_version", "workflow_id", "request_id", "attempt_id", - "stage1_run_root_ref", "stage1_deployment_root_ref", - }) - ID_RE = re.compile(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}\Z") + MAX_FILE_BYTES = 32 * 1024 * 1024 - class PrepareError(ValueError): - def __init__(self, code: str, message: str) -> None: - super().__init__(message) - self.code = code + MAX_RUN_BYTES = 256 * 1024 * 1024 - def canonical_json_bytes(value: Any) -> bytes: - return (json.dumps(value, ensure_ascii=False, allow_nan=False, - sort_keys=True, separators=(",", ":")) + "\n").encode("utf-8") + MAX_JSON_DEPTH = 96 - def _relative_path(value: str, *, code: str) -> str: - if not isinstance(value, str) or not value or "\x00" in value or "\\" in value: - raise PrepareError(code, "logical path is empty or malformed") - if unicodedata.normalize("NFC", value) != value: - raise PrepareError(code, "logical path must already be NFC") - path = PurePosixPath(value) - if path.is_absolute() or any(part in {"", ".", ".."} for part in path.parts): - raise PrepareError(code, "logical path must be a contained relative path") - rendered = path.as_posix() - if rendered != value: - raise PrepareError(code, "logical path is not canonical") - return rendered + MAX_JSON_ITEMS = 1_000_000 - def build_request( - request_id: str, - attempt_id: str, - stage1_run_root_ref: str, - stage1_deployment_root_ref: str, - ) -> bytes: - """Return the exact six-field canonical request; do not access localdocs.""" - if not isinstance(request_id, str) or ID_RE.fullmatch(request_id) is None: - raise PrepareError("REQUEST_ID_INVALID", "request_id contains forbidden characters") - if not isinstance(attempt_id, str) or ID_RE.fullmatch(attempt_id) is None: - raise PrepareError("ATTEMPT_ID_INVALID", "attempt_id contains forbidden characters") - request = { - "schema_version": "stage2_s2_00_execution_request.v1", - "workflow_id": "S2_00", - "request_id": request_id, - "attempt_id": attempt_id, - "stage1_run_root_ref": _relative_path( - stage1_run_root_ref, code="STAGE1_RUN_ROOT_REF_INVALID"), - "stage1_deployment_root_ref": _relative_path( - stage1_deployment_root_ref, code="STAGE1_DEPLOYMENT_ROOT_REF_INVALID"), - } - if set(request) != REQUEST_KEYS: - raise PrepareError("RUN_REQUEST_CLOSED_SHAPE", "request field set drifted") - return canonical_json_bytes(request) + SEMANTIC_SIGNAL_KINDS = frozenset({"canonical", "domain_signal"}) - class LocaldocsSession: - """Small localdocs JSON-RPC session with verified binary write/read-back.""" - - def __init__(self, user_hash: str, workspace_hash: str, *, client: Any = None) -> None: - for name, value in (("user_hash", user_hash), ("workspace_hash", workspace_hash)): - if not isinstance(value, str) or re.fullmatch(r"[a-f0-9]{64}", value) is None: - raise PrepareError("CONTEXT_HASH_INVALID", f"{name} is not a SHA-256 digest") - self.user_hash = user_hash - self.workspace_hash = workspace_hash - if client is None: - try: - import httpx - except ImportError as exc: - raise PrepareError("HTTPX_UNAVAILABLE", "httpx==0.28.1 is required") from exc - client = httpx.Client(timeout=60) - self.client = client - self.headers = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} - self.session_id: str | None = None - self.next_id = 10 - self.initialized = False - - def close(self) -> None: - self.client.close() - - def _post(self, body: Mapping[str, Any], expected_id: int | None) -> Mapping[str, Any] | None: - try: - response = self.client.post(LOCALDOCS_URL, json=dict(body), headers=dict(self.headers)) - response.raise_for_status() - except Exception as exc: - raise PrepareError("MCP_TRANSPORT_ERROR", "localdocs transport failed") from exc - session_id = response.headers.get("mcp-session-id") - if session_id: - if self.session_id is None and expected_id == 1: - self.session_id = session_id - elif session_id != self.session_id: - raise PrepareError("MCP_SESSION_ID_CHANGED", "localdocs session changed") - self.headers["mcp-session-id"] = session_id - if expected_id is None: - return None - try: - payload = response.json() - except Exception as exc: - raise PrepareError("MCP_RESPONSE_INVALID", "localdocs response is not JSON") from exc - if not isinstance(payload, dict) or payload.get("jsonrpc") != "2.0" or payload.get("id") != expected_id: - raise PrepareError("MCP_RESPONSE_INVALID", "localdocs response ID or shape mismatch") - if "error" in payload or not isinstance(payload.get("result"), dict): - raise PrepareError("MCP_TOOL_ERROR", "localdocs returned an error") - return payload["result"] - - def initialize(self) -> None: - response = self._post({"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": { - "protocolVersion": MCP_PROTOCOL_VERSION, "capabilities": {}, "clientInfo": { - "name": "liti-s2-00-prepare", "version": "1.0.0", - "user_id": self.user_hash, "workspace_id": self.workspace_hash, - }}}, 1) - if response is None or response.get("protocolVersion") != MCP_PROTOCOL_VERSION or self.session_id is None: - raise PrepareError("MCP_INITIALIZE_INVALID", "localdocs initialization failed") - self._post({"jsonrpc": "2.0", "method": "notifications/initialized"}, None) - self.initialized = True - - def call(self, name: str, arguments: Mapping[str, Any]) -> Mapping[str, Any]: - if not self.initialized: - raise PrepareError("MCP_NOT_INITIALIZED", "localdocs is not initialized") - message_id = self.next_id - self.next_id += 1 - result = self._post({"jsonrpc": "2.0", "id": message_id, "method": "tools/call", - "params": {"name": name, "arguments": dict(arguments)}}, message_id) - if result is None or result.get("isError") is True: - raise PrepareError("MCP_TOOL_ERROR", f"localdocs {name} failed") - return result - - def read_binary(self, logical_path: str) -> bytes: - path = _relative_path(logical_path, code="LOCALDOCS_READ_PATH_INVALID") - result = self.call("read_binary_doc", {"doc_name": path}) - content = result.get("content") - if not isinstance(content, list) or len(content) != 1 or not isinstance(content[0], dict) or content[0].get("type") != "text": - raise PrepareError("LOCALDOCS_READ_SHAPE", "invalid binary response") - text = content[0].get("text") - if not isinstance(text, str): - raise PrepareError("LOCALDOCS_READ_SHAPE", "missing binary envelope") - try: - envelope = json.loads(text) - if isinstance(envelope, dict) and "results" in envelope: - rows = envelope["results"] - if not isinstance(rows, list) or len(rows) != 1 or not isinstance(rows[0], dict): - raise ValueError("binary result cardinality mismatch") - inner = rows[0].get("content", rows[0].get("text")) - envelope = json.loads(inner) if isinstance(inner, str) else inner - encoded = envelope["content_base64"] - if not isinstance(encoded, str): - raise ValueError("binary content is not base64") - payload = base64.b64decode(encoded, validate=True) - size = envelope.get("byte_length", envelope.get("size")) - if size is not None and (not isinstance(size, int) or size != len(payload)): - raise ValueError("binary size mismatch") - digest = envelope.get("sha256") - if digest is not None and digest != hashlib.sha256(payload).hexdigest(): - raise ValueError("binary hash mismatch") - return payload - except (ValueError, KeyError, TypeError, base64.binascii.Error) as exc: - raise PrepareError("LOCALDOCS_READ_SHAPE", "invalid binary envelope") from exc - - def write_binary_verified(self, logical_path: str, payload: bytes) -> str: - path = _relative_path(logical_path, code="LOCALDOCS_WRITE_PATH_INVALID") - if path != REQUEST_PATH: - raise PrepareError("LOCALDOCS_WRITE_PATH_INVALID", "prepare may write only the fixed request path") - result = self.call("write_binary_file", { - "path": path, "content_base64": base64.b64encode(payload).decode("ascii"), "overwrite": True, - }) - if not isinstance(result.get("content"), list) or result.get("isError") is True: - raise PrepareError("LOCALDOCS_WRITE_FAILED", "localdocs did not acknowledge write") - observed = self.read_binary(path) - if observed != payload: - raise PrepareError("LOCALDOCS_WRITE_READBACK_MISMATCH", "request read-back differs") - return hashlib.sha256(observed).hexdigest() - - - def prepare_request( - request_id: str, - attempt_id: str, - stage1_run_root_ref: str, - stage1_deployment_root_ref: str, - *, - localdocs: Any = None, - ) -> dict[str, Any]: - """Validate, write, read back, and close an authenticated localdocs session.""" - payload = build_request(request_id, attempt_id, stage1_run_root_ref, stage1_deployment_root_ref) - session = localdocs if localdocs is not None else LocaldocsSession(INLINE_USER_HASH, INLINE_WORKSPACE_HASH) - failure: Exception | None = None - digest: str | None = None - try: - session.initialize() - digest = session.write_binary_verified(REQUEST_PATH, payload) - except Exception as exc: - failure = exc - try: - session.close() - except Exception as exc: - if failure is None: - failure = PrepareError("LOCALDOCS_CLOSE_FAILED", "localdocs session close failed") - failure.__cause__ = exc - if failure is not None: - raise failure - return {"ok": True, "workflow_id": "S2_00", "path": REQUEST_PATH, - "request_sha256": digest} - - - if __name__ == "__main__": - sys.stdout.buffer.write(canonical_json_bytes({ - "ok": False, "error": {"code": "PREPARE_ARGUMENT_BINDING_UNVERIFIED", - "message": "four caller values must be bound by the backend"}})) - raise SystemExit(2) - - task_name: Task_S2_00_deterministic_ingress - description: >- - 고정 request와 release를 hydration하고 S2_00 pure core를 실행한 뒤 - binding-derived root에 결과를 status-last 방식으로 발행한다. - mcp: code-executor - tool_name: run_code - parameters: - language: python - requirements: "httpx==0.28.1" - network: agent-network - timeout: 300 - code: |- - #!/usr/bin/env python3 - """Deterministic Stage 2 ingress for the S2_00 contract. - - The pure ingress core uses the Python standard library; the inline MCP adapter - adds only pinned ``httpx`` transport to the fixed localdocs endpoint. It does - not perform legal reasoning, model calls, external-network access, dynamic - imports, or runtime installation. The release, source, conservation, context, - cluster, bundle, hydration, and logical-publish contracts are re-checked here. - """ - - from __future__ import annotations - - import argparse - import base64 - import binascii - from collections import Counter, defaultdict - import contextlib - from dataclasses import dataclass - import hashlib - import io - import itertools - import json - import math - import os - from pathlib import Path, PurePosixPath - import re - import shutil - import stat - import sys - import tempfile - import unicodedata - from typing import Any, Callable, Iterable, Mapping, MutableMapping, Sequence - - - ALGORITHM_VERSION = "s2_00_ingress/1.2.0" - ALGORITHM_SEMANTIC_DIGEST = hashlib.sha256( - b"LITI-S2_00-INLINE-ALGORITHM\x00" + ALGORITHM_VERSION.encode("ascii") - ).hexdigest() - MAX_FILE_BYTES = 32 * 1024 * 1024 - MAX_RUN_BYTES = 256 * 1024 * 1024 - MAX_JSON_DEPTH = 96 - MAX_JSON_ITEMS = 1_000_000 - ROUTES = frozenset({"TO_S2_10", "TO_S2_10_WITH_ISSUES", "TO_S2_40_STATUS_ONLY"}) - RELEASE_MODE = { - "STRUCTURAL_FIXTURE": "DEV_FIXTURE_RELEASE", - "SUBSET_CANARY": "SUBSET_CANARY_RELEASE", - "PRODUCTION": "PRODUCTION_RELEASE", + HARD_RELATION_KINDS = frozenset( + { + "SAME_BO_ID", + "SOURCE_BO_ATTACHMENT", + "SAME_EVIDENCE_REF", + "SAME_EVENT_REF", + "EXPLICIT_CASE_RELATION", } - SEMANTIC_SIGNAL_KINDS = frozenset({"canonical", "domain_signal"}) - HARD_RELATION_KINDS = frozenset( - { - "SAME_BO_ID", - "SOURCE_BO_ATTACHMENT", - "SAME_EVIDENCE_REF", - "SAME_EVENT_REF", - "EXPLICIT_CASE_RELATION", - } - ) - CANDIDATE_RELATION_KINDS = frozenset( - {"claim_precondition", "accessory_of", "incompatible_with", "EXPLICIT_DEPENDENCY"} - ) - P1_DIGEST_KEYS = { - "evidence_indexed_sha256": "evidence_indexed", - "evidence_event_candidates_sha256": "evidence_event_candidates", - "b1_gate_sha256": "b1_evidence_indexed_gate", - "b2_gate_sha256": "b2_event_candidates_gate", - "screening_sha256": "domain_screening", - "activation_manifest_sha256": "domain_activation_manifest", - "registry_index_sha256": "stage1_domain_registry_index", - } - V2_DOMAIN_IDS = frozenset({"E-01", "E-06", "E-12", "E-16", "E-18", "E-19", "E-20", "E-21"}) - CONTEXT_SCHEMA_ID = "https://schemas.liti-agent.local/stage2/s2_00/context.schema.v2.json" - INGRESS_SCHEMA_ID = "https://schemas.liti-agent.local/stage2/s2_00/ingress.schema.v1.json" - REVIEW_SCHEMA_ID = "https://schemas.liti-agent.local/stage2/shared/review_status.schema.v1.json" - REQUIREMENT_CLASS_ENUM = { - "identity_backbone": "IDENTITY_BACKBONE", - "routing_profile_backbone": "ROUTING_PROFILE_BACKBONE", - "evidence_scope": "EVIDENCE_EVENT_SCOPE", - "event_scope": "EVIDENCE_EVENT_SCOPE", - "integrity_corroborator": "INTEGRITY_CORROBORATOR", - "optimization_context": "OPTIMIZATION_CONTEXT", - } - ADAPTER_IDS = { - "evidence_indexed": "S2A-EVIDENCE-V3-ENVELOPE-V1", - "evidence_event_candidates": "S2A-EVENTS-V1-ENVELOPE-V1", - "client_goal": "S2A-CLIENT-GOAL-V8-V1", - "domain_screening": "S2A-DOMAIN-SCREENING-V1", - "domain_activation_manifest": "S2A-DUAL-SG01-V1", - "b1_evidence_indexed_gate": "S2A-B1-GATE-V1", - "b2_event_candidates_gate": "S2A-B2-GATE-V1", - "stage1_part1_soft_gate_handoff": "S2A-P1-HANDOFF-FLAT-V1", - "bo": "S2A-BO-V8-LIST-V1", - "signal_manifest": "S2A-SIGNAL-ALL-V1", - "stage1_part2_review_handoff": "S2A-P2-HANDOFF-FLAT-V1", - "legal_effect_structures": "S2A-LES-CURRENT-V8-V1", - "stage1_part3_review_handoff": "S2A-P3-HANDOFF-WRAPPED-V1", - "fact_ledger_base": "S2A-FACT-LEDGER-CURRENT-V8-V1", - "fact_ledger_writer_report": "S2A-FACT-LEDGER-WRITER-REPORT-V1", - "stage1_part4_review_handoff": "S2A-P4-HANDOFF-WRAPPED-V1", - } - SG01_PROJECTION_FIELDS: tuple[str, ...] = ( - "schema_version", - "signal_id", - "status", - "registry_version", - "registry_index_sha256", - "screening_sha256", - "domain_entries", + ) + + + CANDIDATE_RELATION_KINDS = frozenset( + {"claim_precondition", "accessory_of", "incompatible_with", "EXPLICIT_DEPENDENCY"} + ) + + + P1_DIGEST_KEYS = { + "evidence_indexed_sha256": "evidence_indexed", + "evidence_event_candidates_sha256": "evidence_event_candidates", + "b1_gate_sha256": "b1_evidence_indexed_gate", + "b2_gate_sha256": "b2_event_candidates_gate", + "screening_sha256": "domain_screening", + "activation_manifest_sha256": "domain_activation_manifest", + "registry_index_sha256": "stage1_domain_registry_index", + } + + + REQUIREMENT_CLASS_ENUM = { + "identity_backbone": "IDENTITY_BACKBONE", + "routing_profile_backbone": "ROUTING_PROFILE_BACKBONE", + "evidence_scope": "EVIDENCE_EVENT_SCOPE", + "event_scope": "EVIDENCE_EVENT_SCOPE", + "integrity_corroborator": "INTEGRITY_CORROBORATOR", + "optimization_context": "OPTIMIZATION_CONTEXT", + } + + + ADAPTER_IDS = { + "evidence_indexed": "S2A-EVIDENCE-V3-ENVELOPE-V1", + "evidence_event_candidates": "S2A-EVENTS-V1-ENVELOPE-V1", + "client_goal": "S2A-CLIENT-GOAL-V8-V1", + "domain_screening": "S2A-DOMAIN-SCREENING-V1", + "domain_activation_manifest": "S2A-DUAL-SG01-V1", + "b1_evidence_indexed_gate": "S2A-B1-GATE-V1", + "b2_event_candidates_gate": "S2A-B2-GATE-V1", + "stage1_part1_soft_gate_handoff": "S2A-P1-HANDOFF-FLAT-V1", + "bo": "S2A-BO-V8-LIST-V1", + "signal_manifest": "S2A-SIGNAL-ALL-V1", + "stage1_part2_review_handoff": "S2A-P2-HANDOFF-FLAT-V1", + "legal_effect_structures": "S2A-LES-CURRENT-V8-V1", + "stage1_part3_review_handoff": "S2A-P3-HANDOFF-WRAPPED-V1", + "fact_ledger_base": "S2A-FACT-LEDGER-CURRENT-V8-V1", + "fact_ledger_writer_report": "S2A-FACT-LEDGER-WRITER-REPORT-V1", + "stage1_part4_review_handoff": "S2A-P4-HANDOFF-WRAPPED-V1", + } + + + SG01_PROJECTION_FIELDS: tuple[str, ...] = ( + "schema_version", + "signal_id", + "status", + "registry_version", + "registry_index_sha256", + "screening_sha256", + "domain_entries", + "active_domain_ids", + "supporting_domain_ids", + "monitor_domain_ids", + "expected_runnable_domain_ids", + "required_calculation_domains", + "unrouted_material", + "conservation_gate", + "fail_open_policy", + "review_items", + "contract_guards", + ) + + + SG01_SET_FIELDS = frozenset( + { "active_domain_ids", "supporting_domain_ids", "monitor_domain_ids", "expected_runnable_domain_ids", "required_calculation_domains", - "unrouted_material", - "conservation_gate", - "fail_open_policy", - "review_items", - "contract_guards", - ) - SG01_SET_FIELDS = frozenset( - { - "active_domain_ids", - "supporting_domain_ids", - "monitor_domain_ids", - "expected_runnable_domain_ids", - "required_calculation_domains", - } - ) - OUTPUT_SCHEMA_TARGETS: tuple[tuple[re.Pattern[str], str, str], ...] = ( - (re.compile(r"^ingress/stage1_input_manifest\.json$"), "ingress.schema.json", "#/$defs/stage1_input_manifest"), - (re.compile(r"^ingress/intake_report\.json$"), "ingress.schema.json", "#/$defs/intake_report"), - (re.compile(r"^ingress/ingress_status\.json$"), "ingress.schema.json", "#/$defs/ingress_status"), - (re.compile(r"^ingress/technical_diagnostic\.json$"), "ingress.schema.json", "#/$defs/technical_diagnostic"), - (re.compile(r"^review/issue_ledger\.base\.json$"), "review_status.schema.json", "#/$defs/issue_ledger_base"), - (re.compile(r"^context/case_context\.json$"), "context.schema.json", "#/$defs/case_context"), - (re.compile(r"^context/evidence_inventory\.json$"), "context.schema.json", "#/$defs/evidence_inventory"), - (re.compile(r"^context/object_registry\.json$"), "context.schema.json", "#/$defs/object_registry"), - (re.compile(r"^context/party_and_title_context\.json$"), "context.schema.json", "#/$defs/party_and_title_context"), - (re.compile(r"^context/slot_crosswalk\.json$"), "context.schema.json", "#/$defs/slot_crosswalk"), - (re.compile(r"^context/cluster_plan\.json$"), "context.schema.json", "#/$defs/cluster_plan"), - (re.compile(r"^context/cluster_slices/[^/]+\.json$"), "context.schema.json", "#/$defs/cluster_slice"), - (re.compile(r"^context/bundle_plan\.json$"), "context.schema.json", "#/$defs/bundle_plan"), - ) - _RAW_VALUE_UNSET = object() + } + ) - # AgentBackend substitutes these two values in the deployed Agent YAML before - # Code Executor runs the byte-identical source. They intentionally remain - # literal placeholders in the offline parity mirror and its unit tests. - INLINE_USER_HASH = "{{__user_hash__}}" - INLINE_WORKSPACE_HASH = "{{__workspace_hash__}}" - EXPECTED_STAGE2_RELEASE_SHA256 = "2363f1166ee8a3ea2b38350cde61fccfc9852a413c6b9f14a6ed6606af15a590" - INLINE_REQUEST_PATH = "stage2_control/s2_00_request.json" - INLINE_STAGE2_ASSET_ROOT = "Default_Agent/Stage_2_Clean" - INLINE_STAGE2_RELEASE_PATH = ( - f"{INLINE_STAGE2_ASSET_ROOT}/manifest/stage2_release.json" - ) - LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" - MCP_PROTOCOL_VERSION = "2025-03-26" - INLINE_CLIENT_NAME = "liti-stage2-s2-00-inline" - INLINE_CLIENT_VERSION = "1.2.0" - INLINE_SCHEMA_MODULE_IDS = frozenset( - {"SCHEMA-INGRESS", "SCHEMA-CONTEXT", "SCHEMA-REVIEW-STATUS"} - ) + _RAW_VALUE_UNSET = object() - DEFAULT_SOURCE_CONTRACTS: tuple[dict[str, Any], ...] = ( - {"logical_input_id": "evidence_indexed", "path": "evidence_indexed.json", "criticality": "evidence_scope"}, - {"logical_input_id": "evidence_event_candidates", "path": "evidence_event_candidates.json", "criticality": "event_scope"}, - {"logical_input_id": "client_goal", "path": "client_goal.json", "criticality": "optimization_context"}, - {"logical_input_id": "domain_screening", "path": "routing/domain_screening.json", "criticality": "routing_profile_backbone"}, - {"logical_input_id": "domain_activation_manifest", "path": "routing/domain_activation_manifest.json", "criticality": "routing_profile_backbone"}, - {"logical_input_id": "b1_evidence_indexed_gate", "path": "quality_gates/B1_evidence_indexed_gate.json", "criticality": "integrity_corroborator"}, - {"logical_input_id": "b2_event_candidates_gate", "path": "quality_gates/B2_event_candidates_gate.json", "criticality": "integrity_corroborator"}, - {"logical_input_id": "stage1_part1_soft_gate_handoff", "path": "quality_gates/stage1_part1_soft_gate_handoff.json", "criticality": "integrity_corroborator"}, - {"logical_input_id": "bo", "path": "BO.json", "criticality": "identity_backbone"}, - {"logical_input_id": "signal_manifest", "path": "signals/signal_manifest.json", "criticality": "routing_profile_backbone"}, - {"logical_input_id": "stage1_part2_review_handoff", "path": "quality_gates/stage1_part2_review_handoff.json", "criticality": "integrity_corroborator"}, - {"logical_input_id": "legal_effect_structures", "path": "legal_effect_structures.json", "criticality": "routing_profile_backbone"}, - {"logical_input_id": "stage1_part3_review_handoff", "path": "quality_gates/stage1_part3_review_handoff.json", "criticality": "integrity_corroborator"}, - {"logical_input_id": "fact_ledger_base", "path": "Fact_Ledger_base.json", "criticality": "identity_backbone"}, - {"logical_input_id": "fact_ledger_writer_report", "path": "stage1_tmp/fact_ledger/fact_ledger_writer_report.json", "criticality": "integrity_corroborator"}, - {"logical_input_id": "stage1_part4_review_handoff", "path": "quality_gates/stage1_part4_review_handoff.json", "criticality": "integrity_corroborator"}, - ) + DEFAULT_SOURCE_CONTRACTS: tuple[dict[str, Any], ...] = ( + {"logical_input_id": "evidence_indexed", "path": "evidence_indexed.json", "criticality": "evidence_scope"}, + {"logical_input_id": "evidence_event_candidates", "path": "evidence_event_candidates.json", "criticality": "event_scope"}, + {"logical_input_id": "client_goal", "path": "client_goal.json", "criticality": "optimization_context"}, + {"logical_input_id": "domain_screening", "path": "routing/domain_screening.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "domain_activation_manifest", "path": "routing/domain_activation_manifest.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "b1_evidence_indexed_gate", "path": "quality_gates/B1_evidence_indexed_gate.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "b2_event_candidates_gate", "path": "quality_gates/B2_event_candidates_gate.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "stage1_part1_soft_gate_handoff", "path": "quality_gates/stage1_part1_soft_gate_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "bo", "path": "BO.json", "criticality": "identity_backbone"}, + {"logical_input_id": "signal_manifest", "path": "signals/signal_manifest.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "stage1_part2_review_handoff", "path": "quality_gates/stage1_part2_review_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "legal_effect_structures", "path": "legal_effect_structures.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "stage1_part3_review_handoff", "path": "quality_gates/stage1_part3_review_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "fact_ledger_base", "path": "Fact_Ledger_base.json", "criticality": "identity_backbone"}, + {"logical_input_id": "fact_ledger_writer_report", "path": "stage1_tmp/fact_ledger/fact_ledger_writer_report.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "stage1_part4_review_handoff", "path": "quality_gates/stage1_part4_review_handoff.json", "criticality": "integrity_corroborator"}, + ) - class IngressError(RuntimeError): - """A machine-readable deterministic ingress failure.""" + class IngressError(RuntimeError): + """A machine-readable deterministic ingress failure.""" - def __init__( - self, - code: str, - message: str, - *, - logical_input_id: str | None = None, - details: Mapping[str, Any] | None = None, - ) -> None: - super().__init__(message) - self.code = code - self.logical_input_id = logical_input_id - self.details = dict(details or {}) + def __init__( + self, + code: str, + message: str, + *, + logical_input_id: str | None = None, + details: Mapping[str, Any] | None = None, + ) -> None: + super().__init__(message) + self.code = code + self.logical_input_id = logical_input_id + self.details = dict(details or {}) - def as_dict(self) -> dict[str, Any]: - result: dict[str, Any] = {"code": self.code, "message": str(self)} - if self.logical_input_id is not None: - result["logical_input_id"] = self.logical_input_id - if self.details: - result["details"] = self.details - return result - - - @dataclass(frozen=True, slots=True) - class Snapshot: - logical_input_id: str - relative_path: str - resolved_path: str - raw: bytes - raw_sha256: str - byte_length: int - device: int - inode: int - mtime_ns: int - - - def _reject_constant(value: str) -> None: - raise ValueError(f"non-finite JSON number is forbidden: {value}") - - - def _pairs_without_duplicates(pairs: Sequence[tuple[str, Any]]) -> dict[str, Any]: - result: dict[str, Any] = {} - for key, value in pairs: - if key in result: - raise ValueError(f"duplicate JSON key: {key}") - result[key] = value + def as_dict(self) -> dict[str, Any]: + result: dict[str, Any] = {"code": self.code, "message": str(self)} + if self.logical_input_id is not None: + result["logical_input_id"] = self.logical_input_id + if self.details: + result["details"] = self.details return result - def _walk_json_limits(value: Any, *, max_depth: int, max_items: int) -> int: - count = 0 - stack: list[tuple[Any, int]] = [(value, 1)] - while stack: - current, depth = stack.pop() - if depth > max_depth: - raise IngressError("JSON_DEPTH_LIMIT", "JSON nesting depth exceeded") - if isinstance(current, dict): - count += len(current) - stack.extend((item, depth + 1) for item in current.values()) - elif isinstance(current, list): - count += len(current) - stack.extend((item, depth + 1) for item in current) - if count > max_items: - raise IngressError("JSON_ITEM_LIMIT", "JSON aggregate item limit exceeded") - return count + @dataclass(frozen=True, slots=True) + class Snapshot: + logical_input_id: str + relative_path: str + resolved_path: str + raw: bytes + raw_sha256: str + byte_length: int + device: int + inode: int + mtime_ns: int - def load_json_strict( - source: Snapshot | bytes | bytearray | memoryview | str, - *, - max_depth: int = MAX_JSON_DEPTH, - max_items: int = MAX_JSON_ITEMS, - ) -> Any: - """Parse one UTF-8 JSON value, rejecting duplicate keys and non-finite numbers.""" + def _reject_constant(value: str) -> None: + raise ValueError(f"non-finite JSON number is forbidden: {value}") - if isinstance(source, Snapshot): - raw = source.raw - elif isinstance(source, str): - raw = source.encode("utf-8") + + def _pairs_without_duplicates(pairs: Sequence[tuple[str, Any]]) -> dict[str, Any]: + result: dict[str, Any] = {} + for key, value in pairs: + if key in result: + raise ValueError(f"duplicate JSON key: {key}") + result[key] = value + return result + + + def _walk_json_limits(value: Any, *, max_depth: int, max_items: int) -> int: + count = 0 + stack: list[tuple[Any, int]] = [(value, 1)] + while stack: + current, depth = stack.pop() + if depth > max_depth: + raise IngressError("JSON_DEPTH_LIMIT", "JSON nesting depth exceeded") + if isinstance(current, dict): + count += len(current) + stack.extend((item, depth + 1) for item in current.values()) + elif isinstance(current, list): + count += len(current) + stack.extend((item, depth + 1) for item in current) + if count > max_items: + raise IngressError("JSON_ITEM_LIMIT", "JSON aggregate item limit exceeded") + return count + + + def load_json_strict( + source: Snapshot | bytes | bytearray | memoryview | str, + *, + max_depth: int = MAX_JSON_DEPTH, + max_items: int = MAX_JSON_ITEMS, + ) -> Any: + """Parse one UTF-8 JSON value, rejecting duplicate keys and non-finite numbers.""" + + if isinstance(source, Snapshot): + raw = source.raw + elif isinstance(source, str): + raw = source.encode("utf-8") + else: + raw = bytes(source) + try: + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise IngressError("INVALID_UTF8", "JSON source is not strict UTF-8") from exc + try: + value = json.loads( + text, + object_pairs_hook=_pairs_without_duplicates, + parse_constant=_reject_constant, + ) + except (json.JSONDecodeError, ValueError) as exc: + message = str(exc) + code = "DUPLICATE_JSON_KEY" if "duplicate JSON key" in message else "STRICT_JSON_PARSE_FAILED" + raise IngressError(code, message) from exc + _walk_json_limits(value, max_depth=max_depth, max_items=max_items) + return value + + + def canonical_json_bytes(value: Any) -> bytes: + """Return the project canonical parsed representation without normalizing strings.""" + + def reject_nonfinite(item: Any) -> None: + if isinstance(item, float) and not math.isfinite(item): + raise IngressError("NON_FINITE_NUMBER", "NaN and Infinity are forbidden") + if isinstance(item, dict): + for nested in item.values(): + reject_nonfinite(nested) + elif isinstance(item, (list, tuple)): + for nested in item: + reject_nonfinite(nested) + + reject_nonfinite(value) + try: + rendered = json.dumps( + value, + ensure_ascii=False, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ) + except (TypeError, ValueError) as exc: + raise IngressError("CANONICAL_SERIALIZATION_FAILED", str(exc)) from exc + return (rendered + "\n").encode("utf-8") + + + def canonical_digest(value: Any) -> str: + return hashlib.sha256(canonical_json_bytes(value)).hexdigest() + + + class _SchemaViolation(ValueError): + """Internal deterministic JSON Schema validation failure.""" + + + def _json_equal(left: Any, right: Any) -> bool: + try: + return canonical_json_bytes(left) == canonical_json_bytes(right) + except IngressError: + return False + + + def _schema_pointer(document: Mapping[str, Any], fragment: str) -> Mapping[str, Any]: + if fragment in {"", "#"}: + return document + pointer = fragment[1:] if fragment.startswith("#") else fragment + if not pointer.startswith("/"): + raise _SchemaViolation(f"unsupported schema fragment: {fragment}") + current: Any = document + for token in pointer[1:].split("/"): + key = token.replace("~1", "/").replace("~0", "~") + if not isinstance(current, dict) or key not in current: + raise _SchemaViolation(f"unresolved schema pointer: {fragment}") + current = current[key] + if not isinstance(current, dict): + raise _SchemaViolation(f"schema pointer is not an object: {fragment}") + return current + + + def _schema_type_matches(value: Any, expected: str) -> bool: + return { + "object": isinstance(value, dict), + "array": isinstance(value, list), + "string": isinstance(value, str), + "integer": isinstance(value, int) and not isinstance(value, bool), + "number": isinstance(value, (int, float)) and not isinstance(value, bool), + "boolean": isinstance(value, bool), + "null": value is None, + }.get(expected, False) + + + def _validate_schema_node( + value: Any, + schema: Mapping[str, Any], + *, + root_schema: Mapping[str, Any], + schema_documents: Mapping[str, Mapping[str, Any]], + instance_path: str, + ) -> None: + reference = schema.get("$ref") + if isinstance(reference, str): + if reference.startswith("#"): + target_root = root_schema + fragment = reference else: - raw = bytes(source) - try: - text = raw.decode("utf-8", errors="strict") - except UnicodeDecodeError as exc: - raise IngressError("INVALID_UTF8", "JSON source is not strict UTF-8") from exc - try: - value = json.loads( - text, - object_pairs_hook=_pairs_without_duplicates, - parse_constant=_reject_constant, - ) - except (json.JSONDecodeError, ValueError) as exc: - message = str(exc) - code = "DUPLICATE_JSON_KEY" if "duplicate JSON key" in message else "STRICT_JSON_PARSE_FAILED" - raise IngressError(code, message) from exc - _walk_json_limits(value, max_depth=max_depth, max_items=max_items) - return value - - - def canonical_json_bytes(value: Any) -> bytes: - """Return the project canonical parsed representation without normalizing strings.""" - - def reject_nonfinite(item: Any) -> None: - if isinstance(item, float) and not math.isfinite(item): - raise IngressError("NON_FINITE_NUMBER", "NaN and Infinity are forbidden") - if isinstance(item, dict): - for nested in item.values(): - reject_nonfinite(nested) - elif isinstance(item, (list, tuple)): - for nested in item: - reject_nonfinite(nested) - - reject_nonfinite(value) - try: - rendered = json.dumps( - value, - ensure_ascii=False, - sort_keys=True, - separators=(",", ":"), - allow_nan=False, - ) - except (TypeError, ValueError) as exc: - raise IngressError("CANONICAL_SERIALIZATION_FAILED", str(exc)) from exc - return (rendered + "\n").encode("utf-8") - - - def canonical_digest(value: Any) -> str: - return hashlib.sha256(canonical_json_bytes(value)).hexdigest() - - - class _SchemaViolation(ValueError): - """Internal deterministic JSON Schema validation failure.""" - - - def _json_equal(left: Any, right: Any) -> bool: - try: - return canonical_json_bytes(left) == canonical_json_bytes(right) - except IngressError: - return False - - - def _schema_pointer(document: Mapping[str, Any], fragment: str) -> Mapping[str, Any]: - if fragment in {"", "#"}: - return document - pointer = fragment[1:] if fragment.startswith("#") else fragment - if not pointer.startswith("/"): - raise _SchemaViolation(f"unsupported schema fragment: {fragment}") - current: Any = document - for token in pointer[1:].split("/"): - key = token.replace("~1", "/").replace("~0", "~") - if not isinstance(current, dict) or key not in current: - raise _SchemaViolation(f"unresolved schema pointer: {fragment}") - current = current[key] - if not isinstance(current, dict): - raise _SchemaViolation(f"schema pointer is not an object: {fragment}") - return current - - - def _schema_type_matches(value: Any, expected: str) -> bool: - return { - "object": isinstance(value, dict), - "array": isinstance(value, list), - "string": isinstance(value, str), - "integer": isinstance(value, int) and not isinstance(value, bool), - "number": isinstance(value, (int, float)) and not isinstance(value, bool), - "boolean": isinstance(value, bool), - "null": value is None, - }.get(expected, False) - - - def _validate_schema_node( - value: Any, - schema: Mapping[str, Any], - *, - root_schema: Mapping[str, Any], - schema_documents: Mapping[str, Mapping[str, Any]], - instance_path: str, - ) -> None: - reference = schema.get("$ref") - if isinstance(reference, str): - if reference.startswith("#"): - target_root = root_schema - fragment = reference - else: - name, separator, tail = reference.partition("#") - target_root = schema_documents.get(name) - if target_root is None: - raise _SchemaViolation(f"{instance_path}: external schema ref is not release-local: {reference}") - fragment = f"#{tail}" if separator else "#" + name, separator, tail = reference.partition("#") + target_root = schema_documents.get(name) + if target_root is None: + raise _SchemaViolation(f"{instance_path}: external schema ref is not release-local: {reference}") + fragment = f"#{tail}" if separator else "#" + _validate_schema_node( + value, + _schema_pointer(target_root, fragment), + root_schema=target_root, + schema_documents=schema_documents, + instance_path=instance_path, + ) + return + if "const" in schema and not _json_equal(value, schema["const"]): + raise _SchemaViolation(f"{instance_path}: const mismatch") + if "enum" in schema and not any(_json_equal(value, candidate) for candidate in schema["enum"]): + raise _SchemaViolation(f"{instance_path}: enum mismatch") + forbidden = schema.get("not") + if isinstance(forbidden, dict) and _schema_branch_matches( + value, + forbidden, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ): + raise _SchemaViolation(f"{instance_path}: forbidden schema branch matched") + expected_type = schema.get("type") + if expected_type is not None: + alternatives = [expected_type] if isinstance(expected_type, str) else list(expected_type) + if not any(_schema_type_matches(value, item) for item in alternatives): + raise _SchemaViolation(f"{instance_path}: expected type {alternatives}") + for keyword in ("oneOf", "anyOf"): + branches = schema.get(keyword) + if isinstance(branches, list): + matches = 0 + for branch in branches: + try: + _validate_schema_node( + value, + branch, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + matches += 1 + except _SchemaViolation: + continue + required_matches = 1 if keyword == "oneOf" else None + if (required_matches is not None and matches != required_matches) or (keyword == "anyOf" and matches == 0): + raise _SchemaViolation(f"{instance_path}: {keyword} matched {matches} branches") + all_of = schema.get("allOf") + if isinstance(all_of, list): + for branch in all_of: _validate_schema_node( value, - _schema_pointer(target_root, fragment), - root_schema=target_root, + branch, + root_schema=root_schema, schema_documents=schema_documents, instance_path=instance_path, ) - return - if "const" in schema and not _json_equal(value, schema["const"]): - raise _SchemaViolation(f"{instance_path}: const mismatch") - if "enum" in schema and not any(_json_equal(value, candidate) for candidate in schema["enum"]): - raise _SchemaViolation(f"{instance_path}: enum mismatch") - forbidden = schema.get("not") - if isinstance(forbidden, dict) and _schema_branch_matches( + condition = schema.get("if") + if isinstance(condition, dict): + condition_matches = True + try: + _validate_schema_node( + value, + condition, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + except _SchemaViolation: + condition_matches = False + selected = schema.get("then" if condition_matches else "else") + if isinstance(selected, dict): + _validate_schema_node( + value, + selected, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + if isinstance(value, dict): + minimum_properties = schema.get("minProperties") + maximum_properties = schema.get("maxProperties") + if isinstance(minimum_properties, int) and len(value) < minimum_properties: + raise _SchemaViolation(f"{instance_path}: minProperties {minimum_properties}") + if isinstance(maximum_properties, int) and len(value) > maximum_properties: + raise _SchemaViolation(f"{instance_path}: maxProperties {maximum_properties}") + required = schema.get("required", []) + if isinstance(required, list): + missing = [key for key in required if key not in value] + if missing: + raise _SchemaViolation(f"{instance_path}: missing required keys {missing}") + properties = schema.get("properties", {}) + if isinstance(properties, dict): + pattern_properties = schema.get("patternProperties", {}) + matched_by_pattern: set[str] = set() + if isinstance(pattern_properties, dict): + for key, child_value in value.items(): + for pattern_text, child_schema in pattern_properties.items(): + if re.search(pattern_text, key) is not None and isinstance(child_schema, dict): + matched_by_pattern.add(key) + _validate_schema_node( + child_value, + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + extras = sorted(set(value) - set(properties) - matched_by_pattern) + additional = schema.get("additionalProperties") + if additional is False: + if extras: + raise _SchemaViolation(f"{instance_path}: additional properties {extras}") + elif isinstance(additional, dict): + for key in extras: + _validate_schema_node( + value[key], + additional, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + for key, child_schema in properties.items(): + if key in value and isinstance(child_schema, dict): + _validate_schema_node( + value[key], + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + if isinstance(value, list): + minimum = schema.get("minItems") + maximum = schema.get("maxItems") + if isinstance(minimum, int) and len(value) < minimum: + raise _SchemaViolation(f"{instance_path}: minItems {minimum}") + if isinstance(maximum, int) and len(value) > maximum: + raise _SchemaViolation(f"{instance_path}: maxItems {maximum}") + if schema.get("uniqueItems") is True: + digests = [canonical_digest(item) for item in value] + if len(digests) != len(set(digests)): + raise _SchemaViolation(f"{instance_path}: duplicate array items") + prefix_items = schema.get("prefixItems") + if isinstance(prefix_items, list): + for index, child_schema in enumerate(prefix_items): + if index < len(value) and isinstance(child_schema, dict): + _validate_schema_node( + value[index], + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + item_schema = schema.get("items") + if item_schema is False and isinstance(prefix_items, list) and len(value) > len(prefix_items): + raise _SchemaViolation(f"{instance_path}: additional array items are forbidden") + if isinstance(item_schema, dict): + start = len(prefix_items) if isinstance(prefix_items, list) else 0 + for index, item in enumerate(value[start:], start=start): + _validate_schema_node( + item, + item_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + contains = schema.get("contains") + if isinstance(contains, dict): + if not any( + _schema_branch_matches( + item, + contains, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + for index, item in enumerate(value) + ): + raise _SchemaViolation(f"{instance_path}: contains did not match") + if isinstance(value, str): + min_length = schema.get("minLength") + if isinstance(min_length, int) and len(value) < min_length: + raise _SchemaViolation(f"{instance_path}: minLength {min_length}") + max_length = schema.get("maxLength") + if isinstance(max_length, int) and len(value) > max_length: + raise _SchemaViolation(f"{instance_path}: maxLength {max_length}") + pattern = schema.get("pattern") + if isinstance(pattern, str) and re.search(pattern, value) is None: + raise _SchemaViolation(f"{instance_path}: pattern mismatch") + if isinstance(value, (int, float)) and not isinstance(value, bool): + minimum = schema.get("minimum") + if isinstance(minimum, (int, float)) and value < minimum: + raise _SchemaViolation(f"{instance_path}: minimum {minimum}") + maximum = schema.get("maximum") + if isinstance(maximum, (int, float)) and value > maximum: + raise _SchemaViolation(f"{instance_path}: maximum {maximum}") + + + def _schema_branch_matches( + value: Any, + schema: Mapping[str, Any], + *, + root_schema: Mapping[str, Any], + schema_documents: Mapping[str, Mapping[str, Any]], + instance_path: str, + ) -> bool: + try: + _validate_schema_node( value, - forbidden, + schema, root_schema=root_schema, schema_documents=schema_documents, instance_path=instance_path, - ): - raise _SchemaViolation(f"{instance_path}: forbidden schema branch matched") - expected_type = schema.get("type") - if expected_type is not None: - alternatives = [expected_type] if isinstance(expected_type, str) else list(expected_type) - if not any(_schema_type_matches(value, item) for item in alternatives): - raise _SchemaViolation(f"{instance_path}: expected type {alternatives}") - for keyword in ("oneOf", "anyOf"): - branches = schema.get(keyword) - if isinstance(branches, list): - matches = 0 - for branch in branches: - try: - _validate_schema_node( - value, - branch, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=instance_path, - ) - matches += 1 - except _SchemaViolation: - continue - required_matches = 1 if keyword == "oneOf" else None - if (required_matches is not None and matches != required_matches) or (keyword == "anyOf" and matches == 0): - raise _SchemaViolation(f"{instance_path}: {keyword} matched {matches} branches") - all_of = schema.get("allOf") - if isinstance(all_of, list): - for branch in all_of: - _validate_schema_node( - value, - branch, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=instance_path, - ) - condition = schema.get("if") - if isinstance(condition, dict): - condition_matches = True - try: - _validate_schema_node( - value, - condition, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=instance_path, - ) - except _SchemaViolation: - condition_matches = False - selected = schema.get("then" if condition_matches else "else") - if isinstance(selected, dict): - _validate_schema_node( - value, - selected, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=instance_path, - ) - if isinstance(value, dict): - minimum_properties = schema.get("minProperties") - maximum_properties = schema.get("maxProperties") - if isinstance(minimum_properties, int) and len(value) < minimum_properties: - raise _SchemaViolation(f"{instance_path}: minProperties {minimum_properties}") - if isinstance(maximum_properties, int) and len(value) > maximum_properties: - raise _SchemaViolation(f"{instance_path}: maxProperties {maximum_properties}") - required = schema.get("required", []) - if isinstance(required, list): - missing = [key for key in required if key not in value] - if missing: - raise _SchemaViolation(f"{instance_path}: missing required keys {missing}") - properties = schema.get("properties", {}) - if isinstance(properties, dict): - pattern_properties = schema.get("patternProperties", {}) - matched_by_pattern: set[str] = set() - if isinstance(pattern_properties, dict): - for key, child_value in value.items(): - for pattern_text, child_schema in pattern_properties.items(): - if re.search(pattern_text, key) is not None and isinstance(child_schema, dict): - matched_by_pattern.add(key) - _validate_schema_node( - child_value, - child_schema, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=f"{instance_path}/{key}", - ) - extras = sorted(set(value) - set(properties) - matched_by_pattern) - additional = schema.get("additionalProperties") - if additional is False: - if extras: - raise _SchemaViolation(f"{instance_path}: additional properties {extras}") - elif isinstance(additional, dict): - for key in extras: - _validate_schema_node( - value[key], - additional, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=f"{instance_path}/{key}", - ) - for key, child_schema in properties.items(): - if key in value and isinstance(child_schema, dict): - _validate_schema_node( - value[key], - child_schema, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=f"{instance_path}/{key}", - ) - if isinstance(value, list): - minimum = schema.get("minItems") - maximum = schema.get("maxItems") - if isinstance(minimum, int) and len(value) < minimum: - raise _SchemaViolation(f"{instance_path}: minItems {minimum}") - if isinstance(maximum, int) and len(value) > maximum: - raise _SchemaViolation(f"{instance_path}: maxItems {maximum}") - if schema.get("uniqueItems") is True: - digests = [canonical_digest(item) for item in value] - if len(digests) != len(set(digests)): - raise _SchemaViolation(f"{instance_path}: duplicate array items") - prefix_items = schema.get("prefixItems") - if isinstance(prefix_items, list): - for index, child_schema in enumerate(prefix_items): - if index < len(value) and isinstance(child_schema, dict): - _validate_schema_node( - value[index], - child_schema, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=f"{instance_path}/{index}", - ) - item_schema = schema.get("items") - if item_schema is False and isinstance(prefix_items, list) and len(value) > len(prefix_items): - raise _SchemaViolation(f"{instance_path}: additional array items are forbidden") - if isinstance(item_schema, dict): - start = len(prefix_items) if isinstance(prefix_items, list) else 0 - for index, item in enumerate(value[start:], start=start): - _validate_schema_node( - item, - item_schema, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=f"{instance_path}/{index}", - ) - contains = schema.get("contains") - if isinstance(contains, dict): - if not any( - _schema_branch_matches( - item, - contains, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=f"{instance_path}/{index}", - ) - for index, item in enumerate(value) - ): - raise _SchemaViolation(f"{instance_path}: contains did not match") - if isinstance(value, str): - min_length = schema.get("minLength") - if isinstance(min_length, int) and len(value) < min_length: - raise _SchemaViolation(f"{instance_path}: minLength {min_length}") - max_length = schema.get("maxLength") - if isinstance(max_length, int) and len(value) > max_length: - raise _SchemaViolation(f"{instance_path}: maxLength {max_length}") - pattern = schema.get("pattern") - if isinstance(pattern, str) and re.search(pattern, value) is None: - raise _SchemaViolation(f"{instance_path}: pattern mismatch") - if isinstance(value, (int, float)) and not isinstance(value, bool): - minimum = schema.get("minimum") - if isinstance(minimum, (int, float)) and value < minimum: - raise _SchemaViolation(f"{instance_path}: minimum {minimum}") - maximum = schema.get("maximum") - if isinstance(maximum, (int, float)) and value > maximum: - raise _SchemaViolation(f"{instance_path}: maximum {maximum}") - - - def _schema_branch_matches( - value: Any, - schema: Mapping[str, Any], - *, - root_schema: Mapping[str, Any], - schema_documents: Mapping[str, Mapping[str, Any]], - instance_path: str, - ) -> bool: - try: - _validate_schema_node( - value, - schema, - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=instance_path, - ) - return True - except _SchemaViolation: - return False - - - def _load_output_schemas(asset_root: Path) -> dict[str, Mapping[str, Any]]: - documents: dict[str, Mapping[str, Any]] = {} - for name in ("ingress.schema.json", "context.schema.json", "review_status.schema.json"): - snapshot = open_bounded_snapshot(asset_root, f"schemas/{name}", logical_input_id=f"schema:{name}") - value = load_json_strict(snapshot) - if not isinstance(value, dict): - raise IngressError("OUTPUT_SCHEMA_SHAPE", f"schema is not an object: {name}") - documents[name] = value - return documents - - - def _validate_output_artifact( - relative_path: str, - value: Any, - schema_documents: Mapping[str, Mapping[str, Any]], - ) -> None: - target = next( - ( - (schema_name, schema_pointer) - for path_pattern, schema_name, schema_pointer in OUTPUT_SCHEMA_TARGETS - if path_pattern.fullmatch(relative_path) - ), - None, ) - if target is None: - raise IngressError( - "OUTPUT_ARTIFACT_PATH_UNDECLARED", - f"no output schema target is declared for {relative_path}", - details={"path": relative_path}, - ) - schema_name, schema_pointer = target - root_schema = schema_documents[schema_name] + return True + except _SchemaViolation: + return False + + + def _safe_relative_path(relative_path: str) -> PurePosixPath: + if not isinstance(relative_path, str) or not relative_path: + raise IngressError("INVALID_SOURCE_PATH", "source path must be a non-empty string") + if "\x00" in relative_path or "\\" in relative_path: + raise IngressError("INVALID_SOURCE_PATH", "NUL and backslash are forbidden in logical paths") + logical = PurePosixPath(relative_path) + if logical.is_absolute() or any(part in {"", ".", ".."} for part in logical.parts): + raise IngressError("PATH_TRAVERSAL", f"unsafe relative path: {relative_path}") + return logical + + + def _assert_no_symlink_components(root: Path, logical: PurePosixPath) -> None: + current = root + for part in logical.parts: + current = current / part try: - _validate_schema_node( - value, - _schema_pointer(root_schema, schema_pointer), - root_schema=root_schema, - schema_documents=schema_documents, - instance_path=relative_path, - ) - except _SchemaViolation as exc: - raise IngressError( - "OUTPUT_SCHEMA_VALIDATION_FAILED", - str(exc), - details={"path": relative_path, "schema": schema_name, "schema_pointer": schema_pointer}, - ) from exc + current_stat = current.lstat() + except FileNotFoundError: + return + if stat.S_ISLNK(current_stat.st_mode): + raise IngressError("SYMLINK_ESCAPE", f"symlink component rejected: {logical}") - def _safe_relative_path(relative_path: str) -> PurePosixPath: - if not isinstance(relative_path, str) or not relative_path: - raise IngressError("INVALID_SOURCE_PATH", "source path must be a non-empty string") - if "\x00" in relative_path or "\\" in relative_path: - raise IngressError("INVALID_SOURCE_PATH", "NUL and backslash are forbidden in logical paths") - logical = PurePosixPath(relative_path) - if logical.is_absolute() or any(part in {"", ".", ".."} for part in logical.parts): - raise IngressError("PATH_TRAVERSAL", f"unsafe relative path: {relative_path}") - return logical + def open_bounded_snapshot( + approved_root: str | os.PathLike[str], + relative_path: str, + *, + logical_input_id: str = "anonymous", + max_bytes: int = MAX_FILE_BYTES, + require_single_link: bool = True, + ) -> Snapshot: + """Read one regular file once from one descriptor and verify post-read identity.""" - - def _assert_no_symlink_components(root: Path, logical: PurePosixPath) -> None: - current = root - for part in logical.parts: - current = current / part - try: - current_stat = current.lstat() - except FileNotFoundError: - return - if stat.S_ISLNK(current_stat.st_mode): - raise IngressError("SYMLINK_ESCAPE", f"symlink component rejected: {logical}") - - - def open_bounded_snapshot( - approved_root: str | os.PathLike[str], - relative_path: str, - *, - logical_input_id: str = "anonymous", - max_bytes: int = MAX_FILE_BYTES, - require_single_link: bool = True, - ) -> Snapshot: - """Read one regular file once from one descriptor and verify post-read identity.""" - - root_arg = Path(approved_root) - if root_arg.is_symlink(): - raise IngressError("SYMLINK_ROOT_REJECTED", "approved root itself may not be a symlink") - try: - root = root_arg.resolve(strict=True) - except FileNotFoundError as exc: - raise IngressError("APPROVED_ROOT_MISSING", "approved root does not exist") from exc - if not root.is_dir(): - raise IngressError("APPROVED_ROOT_NOT_DIRECTORY", "approved root must be a directory") - logical = _safe_relative_path(relative_path) - _assert_no_symlink_components(root, logical) - candidate = root.joinpath(*logical.parts) - try: - resolved = candidate.resolve(strict=True) - except FileNotFoundError as exc: - raise IngressError("SOURCE_MISSING", f"source is missing: {relative_path}", logical_input_id=logical_input_id) from exc - try: - resolved.relative_to(root) - except ValueError as exc: - raise IngressError("PATH_ESCAPE", f"resolved source escaped approved root: {relative_path}") from exc - flags = os.O_RDONLY - if hasattr(os, "O_CLOEXEC"): - flags |= os.O_CLOEXEC - if hasattr(os, "O_NOFOLLOW"): - flags |= os.O_NOFOLLOW - try: - descriptor = os.open(candidate, flags) - except OSError as exc: - raise IngressError("SOURCE_OPEN_FAILED", f"unable to open source: {relative_path}") from exc - try: - before = os.fstat(descriptor) - if not stat.S_ISREG(before.st_mode): - raise IngressError("NON_REGULAR_SOURCE", f"source is not a regular file: {relative_path}") - if require_single_link and before.st_nlink != 1: - raise IngressError("HARDLINK_POLICY_VIOLATION", f"source link count is {before.st_nlink}") - if before.st_size > max_bytes: + root_arg = Path(approved_root) + if root_arg.is_symlink(): + raise IngressError("SYMLINK_ROOT_REJECTED", "approved root itself may not be a symlink") + try: + root = root_arg.resolve(strict=True) + except FileNotFoundError as exc: + raise IngressError("APPROVED_ROOT_MISSING", "approved root does not exist") from exc + if not root.is_dir(): + raise IngressError("APPROVED_ROOT_NOT_DIRECTORY", "approved root must be a directory") + logical = _safe_relative_path(relative_path) + _assert_no_symlink_components(root, logical) + candidate = root.joinpath(*logical.parts) + try: + resolved = candidate.resolve(strict=True) + except FileNotFoundError as exc: + raise IngressError("SOURCE_MISSING", f"source is missing: {relative_path}", logical_input_id=logical_input_id) from exc + try: + resolved.relative_to(root) + except ValueError as exc: + raise IngressError("PATH_ESCAPE", f"resolved source escaped approved root: {relative_path}") from exc + flags = os.O_RDONLY + if hasattr(os, "O_CLOEXEC"): + flags |= os.O_CLOEXEC + if hasattr(os, "O_NOFOLLOW"): + flags |= os.O_NOFOLLOW + try: + descriptor = os.open(candidate, flags) + except OSError as exc: + raise IngressError("SOURCE_OPEN_FAILED", f"unable to open source: {relative_path}") from exc + try: + before = os.fstat(descriptor) + if not stat.S_ISREG(before.st_mode): + raise IngressError("NON_REGULAR_SOURCE", f"source is not a regular file: {relative_path}") + if require_single_link and before.st_nlink != 1: + raise IngressError("HARDLINK_POLICY_VIOLATION", f"source link count is {before.st_nlink}") + if before.st_size > max_bytes: + raise IngressError("SOURCE_SIZE_LIMIT", f"source exceeds {max_bytes} bytes") + chunks: list[bytes] = [] + total = 0 + while True: + chunk = os.read(descriptor, min(1024 * 1024, max_bytes + 1 - total)) + if not chunk: + break + chunks.append(chunk) + total += len(chunk) + if total > max_bytes: raise IngressError("SOURCE_SIZE_LIMIT", f"source exceeds {max_bytes} bytes") - chunks: list[bytes] = [] - total = 0 - while True: - chunk = os.read(descriptor, min(1024 * 1024, max_bytes + 1 - total)) - if not chunk: - break - chunks.append(chunk) - total += len(chunk) - if total > max_bytes: - raise IngressError("SOURCE_SIZE_LIMIT", f"source exceeds {max_bytes} bytes") - after = os.fstat(descriptor) - finally: - os.close(descriptor) - try: - path_after = candidate.stat(follow_symlinks=False) - except FileNotFoundError as exc: - raise IngressError("SOURCE_SNAPSHOT_CHANGED", "source disappeared after snapshot") from exc - identity_before = (before.st_dev, before.st_ino, before.st_size, before.st_mtime_ns) - identity_after = (after.st_dev, after.st_ino, after.st_size, after.st_mtime_ns) - path_identity = (path_after.st_dev, path_after.st_ino, path_after.st_size, path_after.st_mtime_ns) - if identity_before != identity_after or identity_after != path_identity: - raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source changed during snapshot: {relative_path}") - raw = b"".join(chunks) - return Snapshot( - logical_input_id=logical_input_id, - relative_path=logical.as_posix(), - resolved_path=str(resolved), - raw=raw, - raw_sha256=hashlib.sha256(raw).hexdigest(), - byte_length=len(raw), - device=after.st_dev, - inode=after.st_ino, - mtime_ns=after.st_mtime_ns, - ) - - - def resolve_stage1_sources( - stage1_run_root: str | os.PathLike[str], - contract_manifest: Mapping[str, Any] | None = None, - ) -> list[dict[str, Any]]: - """Resolve only approved logical kinds; a relocation manifest cannot invent kinds.""" - - root = Path(stage1_run_root).resolve(strict=True) - if not root.is_dir(): - raise IngressError("STAGE1_ROOT_NOT_DIRECTORY", "Stage 1 run root must be a directory") - contracts = [dict(row) for row in DEFAULT_SOURCE_CONTRACTS] - overrides = dict((contract_manifest or {}).get("path_overrides", {})) - approved_ids = {row["logical_input_id"] for row in contracts} - invented = sorted(set(overrides) - approved_ids) - if invented: - raise IngressError("UNAPPROVED_LOGICAL_KIND", "relocation manifest invented logical kinds", details={"ids": invented}) - seen_paths: set[str] = set() - for row in contracts: - path = overrides.get(row["logical_input_id"], row["path"]) - safe = _safe_relative_path(path).as_posix() - if safe in seen_paths: - raise IngressError("DUPLICATE_LOGICAL_MAPPING", f"duplicate physical mapping: {safe}") - seen_paths.add(safe) - row["expected_path"] = row.pop("path") - row["observed_path"] = safe - row["resolution_source"] = ( - "RELEASE_BOUND_CONTRACT_MANIFEST" - if row["logical_input_id"] in overrides - else "DEFAULT_EXACT_PATH" - ) - return contracts - - - def load_release_lock(release_path: str | os.PathLike[str]) -> dict[str, Any]: - """Load a strict release lock without interpreting a historical status label.""" - - path = Path(release_path) - snapshot = open_bounded_snapshot(path.parent, path.name, logical_input_id="stage2_release_lock") - value = load_json_strict(snapshot) - if not isinstance(value, dict): - raise IngressError("RELEASE_LOCK_SHAPE", "release lock must be an object") - release_class = value.get("release_class") - if release_class not in set(RELEASE_MODE.values()): - raise IngressError("RELEASE_CLASS_INVALID", f"unsupported release class: {release_class}") - value = dict(value) - value["_release_raw_sha256"] = snapshot.raw_sha256 - return value - - - def _issue( - code: str, - *, - impact_scope: str = "GLOBAL", - source_refs: Sequence[str] = (), - severity: str = "ERROR", - message: str | None = None, - ) -> dict[str, Any]: - return { - "issue_code": code, - "severity": severity, - "impact_scope": impact_scope, - "scope_refs": sorted(set(source_refs)), - "source_contract_row_refs": sorted(set(source_refs)), - "reason_codes": [code], - "downstream_allowed_actions": [], - "message": message or code, - } - - - def _shape_required(value: Any, keys: Sequence[str]) -> list[str]: - if not isinstance(value, dict): - return list(keys) - return [key for key in keys if key not in value] - - - def _json_pointer_value(document: Any, pointer: str | None) -> tuple[bool, Any]: - if pointer in {None, ""}: - return (pointer == "", document) - if not isinstance(pointer, str) or not pointer.startswith("/"): - return False, None - current = document - for raw_token in pointer[1:].split("/"): - token = raw_token.replace("~1", "/").replace("~0", "~") - if isinstance(current, dict) and token in current: - current = current[token] - elif isinstance(current, list) and token.isdigit() and int(token) < len(current): - current = current[int(token)] - else: - return False, None - return True, current - - - def _release_stage1_source_rows(release_lock: Mapping[str, Any]) -> list[Mapping[str, Any]]: - rows = release_lock.get("stage1_sources") - if not isinstance(rows, list): - dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) - rows = dependency.get("stage1_sources") if isinstance(dependency, dict) else None - return [row for row in rows if isinstance(row, dict)] if isinstance(rows, list) else [] - - - def _adapter_decision(release_lock: Mapping[str, Any], adapter_id: str) -> Mapping[str, Any] | None: - for row in release_lock.get("adapter_decisions", []): - if isinstance(row, dict) and row.get("adapter_id") == adapter_id and isinstance(row.get("decision"), dict): - return row["decision"] - return None - - - def _closed_adapter_shape_errors( - document: Any, - *, - logical_id: str, - adapter_id: str, - required_keys: Sequence[str], - release_lock: Mapping[str, Any], - ) -> list[str]: - errors: list[str] = [] - if required_keys: - errors.extend(f"missing root key {key}" for key in _shape_required(document, required_keys)) - decision = _adapter_decision(release_lock, adapter_id) - if decision is not None: - root_shape = decision.get("root_shape") - if root_shape == "ARRAY" and not isinstance(document, list): - errors.append("root must be an array") - elif root_shape == "OBJECT_ENVELOPE" and not isinstance(document, dict): - errors.append("root must be an object envelope") - if isinstance(document, dict): - errors.extend( - f"missing root key {key}" - for key in _shape_required(document, decision.get("required_root_fields", [])) - ) - if isinstance(document, list): - required_item_fields = decision.get("required_item_fields", decision.get("required_row_fields", [])) - if isinstance(required_item_fields, list): - for index, item in enumerate(document): - for key in _shape_required(item, required_item_fields): - errors.append(f"row {index} missing {key}") - if decision is None: - fallback_required: dict[str, tuple[str, ...]] = { - "evidence_indexed": ("schema_contract_version", "items"), - "evidence_event_candidates": ("schema_version", "items"), - "domain_activation_manifest": SG01_PROJECTION_FIELDS, - "signal_manifest": ("downstream_read_sets", "files"), - "legal_effect_structures": ("schema_version", "structure_records"), - "fact_ledger_writer_report": ( - "schema_version", - "row_count", - "gate_firings", - "domain_effect_coverage", - "calculation_readiness", - "blocked_review_items", - "conservation", - "final_sha256", - ), - } - fallback = fallback_required.get(logical_id, ()) - if fallback: - errors.extend(f"missing root key {key}" for key in _shape_required(document, fallback)) - if logical_id in {"bo", "fact_ledger_base"} and not isinstance(document, list): - errors.append("root must be an array") - return sorted(set(errors)) - - - def _schema_document_index(deployment_documents: Mapping[str, Any]) -> dict[str, Mapping[str, Any]]: - result: dict[str, Mapping[str, Any]] = {} - for path, document in deployment_documents.items(): - if not isinstance(document, dict): - continue - result[path] = document - result[PurePosixPath(path).name] = document - schema_id = document.get("$id") - if isinstance(schema_id, str): - result[schema_id] = document - return result - - - def _source_hash_index(document: Mapping[str, Any] | None) -> dict[str, str]: - result: dict[str, str] = {} - if not isinstance(document, dict): - return result - candidate_arrays: list[Any] = [] - for key in ("source_rows", "sources", "artifacts", "files", "entries"): - if isinstance(document.get(key), list): - candidate_arrays.append(document[key]) - for wrapper in ("completion_seal", "manifest", "payload", "data"): - nested = document.get(wrapper) - if isinstance(nested, dict): - for key in ("source_rows", "sources", "artifacts", "files", "entries"): - if isinstance(nested.get(key), list): - candidate_arrays.append(nested[key]) - for rows in candidate_arrays: - for row in rows: - if not isinstance(row, dict): - continue - digest = row.get("raw_sha256", row.get("sha256")) - if not isinstance(digest, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", digest) is None: - continue - for key in ("logical_input_id", "path", "observed_path", "logical_id"): - identifier = row.get(key) - if isinstance(identifier, str) and identifier: - result[identifier] = digest.lower() - return result - - - def _source_producer_index(document: Mapping[str, Any] | None) -> dict[str, str]: - """Index producer evidence carried by a bounded completion/manifest row.""" - - result: dict[str, str] = {} - if not isinstance(document, dict): - return result - candidate_arrays: list[Any] = [] - for key in ("source_rows", "sources", "artifacts", "files", "entries"): - if isinstance(document.get(key), list): - candidate_arrays.append(document[key]) - for wrapper in ("completion_seal", "manifest", "payload", "data"): - nested = document.get(wrapper) - if isinstance(nested, dict): - for key in ("source_rows", "sources", "artifacts", "files", "entries"): - if isinstance(nested.get(key), list): - candidate_arrays.append(nested[key]) - for rows in candidate_arrays: - for row in rows: - if not isinstance(row, dict): - continue - producer = next( - ( - row.get(key) - for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by") - if isinstance(row.get(key), str) and row.get(key) - ), - None, - ) - if not isinstance(producer, str): - continue - for key in ("logical_input_id", "path", "observed_path", "logical_id"): - identifier = row.get(key) - if isinstance(identifier, str) and identifier: - result[identifier] = producer - return result - - - def _producer_value(document: Any) -> str | None: - if not isinstance(document, dict): - return None - for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): - value = document.get(key) - if isinstance(value, str) and value: - return value - for wrapper in ("metadata", "meta", "handoff", "payload"): - nested = document.get(wrapper) - if isinstance(nested, dict): - for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): - value = nested.get(key) - if isinstance(value, str) and value: - return value - # P3/P4 are closed one-key wrappers in the Stage 1 v8 handoff contract. - for wrapper in ( - "stage1_part3_review_handoff", - "stage1_part4_review_handoff", - ): - nested = document.get(wrapper) - if isinstance(nested, dict): - for key in ("created_by", "finalized_by"): - value = nested.get(key) - if isinstance(value, str) and value: - return value - return None - - - def _producer_matches( - observed: str, - expected: str, - alias_id: str | None, - release_lock: Mapping[str, Any], - ) -> bool: - if observed == expected: - return True - if alias_id is None: - return False - decision = _adapter_decision(release_lock, alias_id) - if decision is None or decision.get("bidirectional_match_allowed") is not True: - return False - pair = {decision.get("schema_writer_id"), decision.get("orchestration_producer_id")} - return {observed, expected} == pair - - - def _identity_ref(document: Any, pointer: str | None, logical_id: str) -> dict[str, Any]: - if pointer is None: - return {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": logical_id} - found, value = _json_pointer_value(document, pointer) - if not found or value is None: - return {"value": None, "disposition": "MISSING", "source_ref": f"{logical_id}#{pointer}"} - return {"value": str(value), "disposition": "OBSERVED", "source_ref": f"{logical_id}#{pointer}"} - - - def validate_ingress_contracts( - snapshots: Mapping[str, Snapshot], - contracts: Sequence[Mapping[str, Any]], - release_lock: Mapping[str, Any], - *, - deployment_snapshots: Mapping[str, Snapshot] | None = None, - deployment_documents: Mapping[str, Any] | None = None, - completion_seal: Mapping[str, Any] | None = None, - contract_manifest: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - """Strictly parse sources and verify release-bound schema, producer, identity, and seal rows.""" - - documents: dict[str, Any] = {} - rows: list[dict[str, Any]] = [] - issues: list[dict[str, Any]] = [] - deployment_snapshots = deployment_snapshots or {} - deployment_documents = deployment_documents or {} - deployment_by_path = {snapshot.relative_path: snapshot for snapshot in deployment_snapshots.values()} - schema_documents = _schema_document_index(deployment_documents) - release_source_rows = _release_stage1_source_rows(release_lock) - release_ids = [str(row.get("logical_input_id")) for row in release_source_rows] - duplicate_release_ids = sorted(key for key, count in Counter(release_ids).items() if count > 1) - if duplicate_release_ids: - raise IngressError( - "RELEASE_SOURCE_CONTRACT_DUPLICATE", - "release stage1_sources contains duplicate logical_input_id rows", - details={"logical_input_ids": duplicate_release_ids}, - ) - expected_fixed = { - str(row["logical_input_id"]): str(row["path"]) - for row in DEFAULT_SOURCE_CONTRACTS - } - expected_release_ids = set(expected_fixed) | {"signal_payload_family"} - observed_release_ids = set(release_ids) - if observed_release_ids != expected_release_ids: - raise IngressError( - "RELEASE_SOURCE_CONTRACT_SET_MISMATCH", - "release stage1_sources must be the exact 16 fixed inputs plus signal_payload_family", - details={ - "missing": sorted(expected_release_ids - observed_release_ids), - "extra": sorted(observed_release_ids - expected_release_ids), - }, - ) - release_rows = {str(row.get("logical_input_id")): row for row in release_source_rows} - for logical_id, expected_path in expected_fixed.items(): - release_row = release_rows[logical_id] - if release_row.get("path") != expected_path or release_row.get("path_rule") not in {None, ""}: - raise IngressError( - "RELEASE_SOURCE_FIXED_PATH_MISMATCH", - f"fixed source path contract mismatch: {logical_id}", - ) - signal_family = release_rows["signal_payload_family"] - if ( - signal_family.get("path") is not None - or signal_family.get("path_rule") != "signals/" - or signal_family.get("adapter_id") != "S2A-SIGNAL-PAYLOAD-FAMILY-V1" - or signal_family.get("raw_hash_source") != "MANIFEST_ROW" - ): - raise IngressError( - "SIGNAL_PAYLOAD_FAMILY_CONTRACT_MISMATCH", - "signal_payload_family must use the approved manifest-expanded path contract", - ) - completion_hashes = _source_hash_index(completion_seal) - manifest_hashes = _source_hash_index(contract_manifest) - completion_producers = _source_producer_index(completion_seal) - manifest_producers = _source_producer_index(contract_manifest) - for contract in contracts: - logical_id = str(contract["logical_input_id"]) - snapshot = snapshots.get(logical_id) - release_row = release_rows.get(logical_id) - contract_missing = release_row is None - release_row = release_row or {} - alias_value = release_row.get("producer_alias", release_row.get("producer_alias_id")) - alias_id = str(alias_value) if isinstance(alias_value, str) else None - schema_ref = release_row.get("schema_ref") if isinstance(release_row.get("schema_ref"), dict) else None - row = { - "logical_input_id": logical_id, - "requirement_class": REQUIREMENT_CLASS_ENUM.get( - str(contract.get("criticality")), - "INTEGRITY_CORROBORATOR", - ), - "expected_path": contract.get("expected_path"), - "observed_path": contract.get("observed_path"), - "resolution_source": contract.get("resolution_source"), - "schema_id": schema_ref.get("$id") if schema_ref else release_row.get("schema_id"), - "schema_sha256": schema_ref.get("sha256") if schema_ref else release_row.get("schema_sha256"), - "producer_id": release_row.get("producer_id"), - "producer_alias_id": alias_id, - "adapter_id": release_row.get("adapter_id", ADAPTER_IDS.get(logical_id, "S2A-UNBOUND-V1")), - "run_identity_ref": release_row.get("run_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), - "transaction_identity_ref": release_row.get("transaction_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), - "scope_refs": [logical_id], - "source_contract_row_refs": [logical_id], - "reason_codes": [], - "downstream_allowed_actions": [], - "issue_codes": [], - } - if contract_missing: - code = "RELEASE_SOURCE_CONTRACT_MISSING" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) - declared_path = release_row.get("path") - if isinstance(declared_path, str) and declared_path != contract.get("expected_path"): - code = "RELEASE_SOURCE_PATH_MISMATCH" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) - if snapshot is None: - row.update( - { - "raw_sha256": None, - "byte_length": 0, - "parse_status": "NOT_OBSERVED", - "schema_status": "UNEVALUABLE", - "seal_status": "UNEVALUABLE", - "scope_technical_disposition": "UNAVAILABLE", - "impact_scope": "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER", - } - ) - code = "SOURCE_MISSING" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope=row["impact_scope"], source_refs=[logical_id])) - rows.append(row) - continue - row["raw_sha256"] = snapshot.raw_sha256 - row["byte_length"] = snapshot.byte_length - try: - document = load_json_strict( - snapshot, - max_depth=int(release_lock.get("limits", {}).get("max_json_depth", MAX_JSON_DEPTH)), - max_items=int(release_lock.get("limits", {}).get("max_json_items", MAX_JSON_ITEMS)), - ) - documents[logical_id] = document - row["parse_status"] = "PASS" - except IngressError as exc: - row["parse_status"] = "FAIL" - row["schema_status"] = "UNEVALUABLE" - row["seal_status"] = "UNEVALUABLE" - row["scope_technical_disposition"] = "UNAVAILABLE" - row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" - row["reason_codes"].append(exc.code) - row["issue_codes"].append(exc.code) - issues.append(_issue(exc.code, impact_scope=row["impact_scope"], source_refs=[logical_id], message=str(exc))) - rows.append(row) - continue - expected_adapter = ADAPTER_IDS.get(logical_id) - if expected_adapter is not None and release_row.get("adapter_id") not in {None, expected_adapter}: - code = "ADAPTER_ID_MISMATCH" - row["schema_status"] = "FAIL" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) - if schema_ref is not None: - schema_path = schema_ref.get("path") - schema_snapshot = deployment_by_path.get(schema_path) if isinstance(schema_path, str) else None - schema_document = deployment_documents.get(schema_path) if isinstance(schema_path, str) else None - expected_schema_hash = schema_ref.get("sha256") - expected_schema_id = schema_ref.get("$id") - if schema_snapshot is None or not isinstance(schema_document, dict): - schema_error = "SCHEMA_REF_NOT_IN_BOUNDED_DEPLOYMENT" - elif not isinstance(expected_schema_hash, str) or schema_snapshot.raw_sha256 != expected_schema_hash.lower(): - schema_error = "SCHEMA_HASH_MISMATCH" - elif expected_schema_id is not None and schema_document.get("$id") != expected_schema_id: - schema_error = "SCHEMA_ID_MISMATCH" - else: - schema_error = None - try: - _validate_schema_node( - document, - schema_document, - root_schema=schema_document, - schema_documents=schema_documents, - instance_path=logical_id, - ) - except _SchemaViolation as exc: - schema_error = "SOURCE_SCHEMA_VALIDATION_FAILED" - issues.append( - _issue( - schema_error, - impact_scope="CLUSTER", - source_refs=[logical_id], - message=str(exc), - ) - ) - if schema_error is not None: - row["schema_status"] = "FAIL" - row["reason_codes"].append(schema_error) - row["issue_codes"].append(schema_error) - if schema_error != "SOURCE_SCHEMA_VALIDATION_FAILED": - issues.append(_issue(schema_error, impact_scope="GLOBAL", source_refs=[logical_id])) - else: - row["schema_status"] = "PASS" - else: - adapter_errors = _closed_adapter_shape_errors( - document, - logical_id=logical_id, - adapter_id=str(row["adapter_id"]), - required_keys=release_row.get("required_keys", []), - release_lock=release_lock, - ) - if contract_missing: - row["schema_status"] = "UNEVALUABLE" - elif adapter_errors: - code = "ADAPTER_REQUIRED_KEY_MISSING" - row["schema_status"] = "FAIL" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append( - _issue( - code, - impact_scope="CLUSTER", - source_refs=[logical_id], - message="; ".join(adapter_errors), - ) - ) - else: - row["schema_status"] = "PASS" - expected_producer = release_row.get("producer_id") - document_producer = _producer_value(document) - sealed_producer = ( - completion_producers.get(logical_id) - or completion_producers.get(str(contract.get("observed_path"))) - or manifest_producers.get(logical_id) - or manifest_producers.get(str(contract.get("observed_path"))) - ) - if ( - document_producer is not None - and sealed_producer is not None - and document_producer != sealed_producer - ): - code = "PRODUCER_EVIDENCE_CONFLICT" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) - observed_producer = document_producer or sealed_producer - if isinstance(expected_producer, str): - if observed_producer is None: - code = "PRODUCER_ID_UNEVALUABLE" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="CLUSTER", source_refs=[logical_id])) - elif not _producer_matches(observed_producer, expected_producer, alias_id, release_lock): - code = "PRODUCER_ID_MISMATCH" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) - row["run_identity_ref"] = _identity_ref(document, release_row.get("run_identity_pointer"), logical_id) - row["transaction_identity_ref"] = _identity_ref( - document, - release_row.get("transaction_identity_pointer"), - logical_id, - ) - raw_hash_source = str(release_row.get("raw_hash_source", "NONE")) - if raw_hash_source in {"CASE_RUN_COMPLETION_SEAL", "COMPLETION_SEAL", "COMPLETION_SEAL_ROW"}: - expected_hash = completion_hashes.get(logical_id) or completion_hashes.get(str(contract.get("observed_path"))) - elif raw_hash_source in {"CONTRACT_MANIFEST", "CONTRACT_MANIFEST_ROW", "MANIFEST_ROW"}: - expected_hash = manifest_hashes.get(logical_id) or manifest_hashes.get(str(contract.get("observed_path"))) - elif raw_hash_source in {"COMPLETION_SEAL_OR_CONTRACT_MANIFEST", "SEALED_ROW"}: - expected_hash = ( - completion_hashes.get(logical_id) - or completion_hashes.get(str(contract.get("observed_path"))) - or manifest_hashes.get(logical_id) - or manifest_hashes.get(str(contract.get("observed_path"))) - ) - elif raw_hash_source in {"UNAVAILABLE_DEV", "NONE"}: - expected_hash = None - else: - expected_hash = None - code = "RAW_HASH_SOURCE_UNAPPROVED" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) - if expected_hash is not None and expected_hash != snapshot.raw_sha256: - code = "RAW_HASH_MISMATCH" - row["seal_status"] = "FAIL" - row["reason_codes"].append(code) - row["issue_codes"].append(code) - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) - else: - row["seal_status"] = "PASS" if expected_hash else "UNEVALUABLE" - row["scope_technical_disposition"] = ( - "UNAVAILABLE" - if contract_missing or any(code in row["issue_codes"] for code in {"RAW_HASH_MISMATCH", "SCHEMA_HASH_MISMATCH", "SCHEMA_ID_MISMATCH"}) - else "AVAILABLE" - if not row["issue_codes"] - else "AVAILABLE_WITH_ISSUES" - ) - row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" - rows.append(row) - for identity_kind, field in ( - ("RUN", "run_identity_ref"), - ("TRANSACTION", "transaction_identity_ref"), - ): - observed_values = { - str(row[field]["value"]) - for row in rows - if row[field].get("disposition") == "OBSERVED" and row[field].get("value") is not None - } - if len(observed_values) > 1: - code = f"{identity_kind}_IDENTITY_CONFLICT" - issues.append(_issue(code, impact_scope="GLOBAL", source_refs=sorted(observed_values))) - for row in rows: - if row[field].get("disposition") == "OBSERVED": - row["issue_codes"] = sorted(set(row["issue_codes"] + [code])) - row["reason_codes"] = sorted(set(row["reason_codes"] + [code])) - row["scope_technical_disposition"] = "UNAVAILABLE" - return {"documents": documents, "source_contract_rows": rows, "issues": issues} - - - def _records_from_signal_document(document: Any) -> list[Any]: - if isinstance(document, list): - return list(document) - if isinstance(document, dict): - for key in ("signals", "records", "items"): - value = document.get(key) - if isinstance(value, list): - return list(value) - return [document] - return [document] - - - def _record_signal_id(record: Any) -> str | None: - if not isinstance(record, dict): - return None - value = record.get("signal_id") - if isinstance(value, str) and value: - return value - for wrapper in ("domain_activation_manifest", "payload", "data"): - nested = record.get(wrapper) - if isinstance(nested, dict) and isinstance(nested.get("signal_id"), str): - return nested["signal_id"] - return None - - - def expand_stage2_signal_all( - stage1_run_root: str | os.PathLike[str], - signal_manifest: Mapping[str, Any], - *, - max_file_bytes: int = MAX_FILE_BYTES, - max_total_bytes: int = MAX_RUN_BYTES, - signal_registry: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - """Expand Stage 2 ALL while separating semantic and integrity-only universes.""" - - downstream = signal_manifest.get("downstream_read_sets", {}) - stage2 = downstream.get("stage2", []) if isinstance(downstream, dict) else [] - if stage2 != ["ALL"]: - raise IngressError("SIGNAL_ALL_CONTRACT", "downstream_read_sets.stage2 must equal ['ALL']") - files = signal_manifest.get("files") - if not isinstance(files, list): - raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest files must be an array") - transaction_id = str(signal_manifest.get("manifest_transaction_id", signal_manifest.get("transaction_id", "MISSING"))) - file_rows: list[dict[str, Any]] = [] - semantic_rows: list[dict[str, Any]] = [] - integrity_rows: list[dict[str, Any]] = [] - occurrences: list[dict[str, Any]] = [] - payload_snapshots: list[Snapshot] = [] - issues: list[dict[str, Any]] = [] - path_counter: Counter[str] = Counter() - parsed_documents: dict[str, Any] = {} - aggregate_bytes = 0 - registry_entries = { - str(row.get("file")): row - for row in (signal_registry or {}).get("entries", []) - if isinstance(row, dict) and isinstance(row.get("file"), str) - } - compatibility_files = { - str(path) - for path in (signal_registry or {}).get("compatibility_views", []) - if isinstance(path, str) - } - domain_envelope_schema = (signal_registry or {}).get("domain_envelope") - observed_registry_files: set[str] = set() - for index, entry in enumerate(files): - if not isinstance(entry, dict) or not isinstance(entry.get("path"), str): - raise IngressError("SIGNAL_FILE_ROW_SHAPE", f"invalid signal file row at index {index}") - relative_payload = _safe_relative_path(entry["path"]).as_posix() - if relative_payload.startswith("signals/"): - raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "manifest file path must not include signals/ prefix") - physical = f"signals/{relative_payload}" - snapshot = open_bounded_snapshot( - stage1_run_root, - physical, - logical_input_id=f"signal_file:{index}", - max_bytes=max_file_bytes, - ) - document = load_json_strict(snapshot) - payload_snapshots.append(snapshot) - aggregate_bytes += snapshot.byte_length - if aggregate_bytes > max_total_bytes: - raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "signal ALL payloads exceed remaining run byte budget") - parsed_documents[relative_payload] = document - kind = entry.get("kind", "canonical") - if kind not in SEMANTIC_SIGNAL_KINDS | {"compatibility_view"}: - raise IngressError("SIGNAL_KIND_UNAPPROVED", f"unapproved signal file kind: {kind}") - semantic = kind in SEMANTIC_SIGNAL_KINDS - expected_hash = entry.get( - "file_sha256", entry.get("sha256", entry.get("raw_sha256")) - ) - row = { - "manifest_index": index, - "file_path": relative_payload, - "physical_path": physical, - "kind": kind, - "raw_sha256": snapshot.raw_sha256, - "byte_length": snapshot.byte_length, - "semantic": semantic, - "manifest_declared_record_count": entry.get("record_count"), - } - if expected_hash is not None and expected_hash != snapshot.raw_sha256: - row["hash_status"] = "FAIL" - issues.append(_issue("SIGNAL_FILE_HASH_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) - else: - row["hash_status"] = "PASS" if expected_hash else "UNEVALUABLE" - records = _records_from_signal_document(document) - row["observed_record_count"] = len(records) - declared_count = entry.get("record_count") - if isinstance(declared_count, int) and declared_count != len(records): - row["record_count_status"] = "FAIL" - issues.append(_issue("SIGNAL_RECORD_COUNT_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) - else: - row["record_count_status"] = "PASS" if isinstance(declared_count, int) else "UNEVALUABLE" - registry_row = registry_entries.get(relative_payload) - if kind == "canonical": - if signal_registry is not None and registry_row is None: - issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) - elif registry_row is not None: - observed_registry_files.add(relative_payload) - declared_schema = entry.get("schema", entry.get("schema_path")) - if declared_schema is not None and declared_schema != registry_row.get("schema"): - issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) - elif kind == "compatibility_view": - if relative_payload in registry_entries: - issues.append(_issue("SIGNAL_COMPATIBILITY_SUBSTITUTION", impact_scope="SIGNAL", source_refs=[physical])) - if signal_registry is not None and relative_payload not in compatibility_files: - issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) - elif kind == "domain_signal": - declared_schema = entry.get("schema", entry.get("schema_path")) - if signal_registry is not None and declared_schema not in {None, domain_envelope_schema}: - issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) - file_rows.append(row) - path_counter[relative_payload] += 1 - if semantic: - semantic_rows.append(row) - for record_ordinal, record in enumerate(records): - signal_id = _record_signal_id(record) - occurrence_key = [transaction_id, relative_payload, record_ordinal, signal_id] - occurrences.append( - { - "occurrence_key": occurrence_key, - "occurrence_ref": f"SIGO-{canonical_digest(occurrence_key)[:24]}", - "manifest_transaction_id": transaction_id, - "file_path": relative_payload, - "record_ordinal": record_ordinal, - "signal_id": signal_id, - "disposition": "UNMAPPED" if signal_id is None else "UNUSED", - "binding_refs": [], - "raw_record_sha256": canonical_digest(record), - "record": record, - } - ) - else: - integrity_rows.append(row) - duplicates = sorted(path for path, count in path_counter.items() if count > 1) - if duplicates: - issues.append(_issue("SIGNAL_ALL_DUPLICATE_FILE_ROW", impact_scope="SIGNAL", source_refs=duplicates)) - manifest_counter = Counter((i, row["file_path"], row["kind"]) for i, row in enumerate(file_rows)) - partition_counter = Counter((row["manifest_index"], row["file_path"], row["kind"]) for row in semantic_rows + integrity_rows) - missing_registry_files = sorted(set(registry_entries) - observed_registry_files) if signal_registry is not None else [] - if missing_registry_files: - issues.append( - _issue( - "SIGNAL_REGISTRY_COVERAGE_MISMATCH", - impact_scope="SIGNAL", - source_refs=[f"signals/{path}" for path in missing_registry_files], - ) - ) - file_conservation = ( - manifest_counter == partition_counter - and not duplicates - and not missing_registry_files - and not any(row["hash_status"] == "FAIL" or row["record_count_status"] == "FAIL" for row in file_rows) - ) - record_counter = Counter(tuple(row["occurrence_key"]) for row in occurrences) - partitioned_record_counter = Counter( - tuple(row["occurrence_key"]) - for row in occurrences - if row["disposition"] in {"USED", "UNUSED", "UNMAPPED"} - ) - record_conservation = record_counter == partitioned_record_counter - return { - "manifest_transaction_id": transaction_id, - "ordered_file_rows": file_rows, - "semantic_file_rows": semantic_rows, - "integrity_only_file_rows": integrity_rows, - "record_occurrences": occurrences, - "used_record_occurrences": [], - "unused_record_occurrences": [row for row in occurrences if row["disposition"] == "UNUSED"], - "unmapped_record_occurrences": [row for row in occurrences if row["disposition"] == "UNMAPPED"], - "_parsed_documents_by_path": parsed_documents, - "_payload_snapshots": payload_snapshots, - "file_conservation_pass": file_conservation, - "record_conservation_pass": record_conservation, - "aggregate_payload_bytes": aggregate_bytes, - "issues": issues, - } - - - def _collect_values_for_keys(value: Any, keys: frozenset[str]) -> set[str]: - result: set[str] = set() - stack = [value] - while stack: - current = stack.pop() - if isinstance(current, dict): - for key, child in current.items(): - if key in keys: - if isinstance(child, list): - result.update(str(item) for item in child if item is not None) - elif child is not None: - result.add(str(child)) - stack.append(child) - elif isinstance(current, list): - stack.extend(current) - return result - - - def bind_signal_occurrences(signal_all: MutableMapping[str, Any], documents: Mapping[str, Any]) -> dict[str, Any]: - """Bind each semantic signal occurrence to explicit Stage 1 references without deduplication.""" - - explicit_signal_ids = _collect_values_for_keys( - documents, - frozenset({"signal_id", "signal_ids", "signal_refs", "emitted_signal_ids", "required_signal_ids"}), - ) - known_refs = { - "fact_id": _collect_values_for_keys(documents.get("fact_ledger_base"), frozenset({"fact_id"})), - "source_bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "source_bo_id", "source_bo_ids"})), - "bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "bo_id"})), - "structure_id": _collect_values_for_keys(documents.get("legal_effect_structures"), frozenset({"structure_id"})), - "domain_id": _collect_values_for_keys(documents, frozenset({"domain_id", "domain_ids", "active_domain_ids"})), - "evidence_id": _collect_values_for_keys(documents.get("evidence_indexed"), frozenset({"evidence_id", "id"})), - "event_id": _collect_values_for_keys(documents.get("evidence_event_candidates"), frozenset({"event_id", "id"})), - } - link_keys = { - "fact_id": ("fact_id", "fact_ids"), - "source_bo_id": ("source_bo_id", "source_bo_ids"), - "bo_id": ("bo_id", "bo_ids"), - "structure_id": ("structure_id", "structure_ids"), - "domain_id": ("domain_id", "domain_ids"), - "evidence_id": ("evidence_id", "evidence_ids"), - "event_id": ("event_id", "event_ids"), - } - for occurrence in signal_all.get("record_occurrences", []): - signal_id = occurrence.get("signal_id") - record = occurrence.get("record") - bindings: set[str] = set() - if isinstance(signal_id, str) and signal_id in explicit_signal_ids: - bindings.add(f"signal_id:{signal_id}") - for ref_kind, candidate_keys in link_keys.items(): - observed = _collect_values_for_keys(record, frozenset(candidate_keys)) - for ref in sorted(observed & known_refs[ref_kind]): - bindings.add(f"{ref_kind}:{ref}") - if not isinstance(signal_id, str) or not signal_id: - occurrence["disposition"] = "UNMAPPED" - elif bindings: - occurrence["disposition"] = "USED" - else: - occurrence["disposition"] = "UNUSED" - occurrence["binding_refs"] = sorted(bindings) - for disposition, key in ( - ("USED", "used_record_occurrences"), - ("UNUSED", "unused_record_occurrences"), - ("UNMAPPED", "unmapped_record_occurrences"), - ): - signal_all[key] = [ - row for row in signal_all.get("record_occurrences", []) if row.get("disposition") == disposition - ] - source_counter = Counter(tuple(row["occurrence_key"]) for row in signal_all.get("record_occurrences", [])) - partition_counter = Counter( - tuple(row["occurrence_key"]) - for key in ("used_record_occurrences", "unused_record_occurrences", "unmapped_record_occurrences") - for row in signal_all[key] - ) - signal_all["record_conservation_pass"] = source_counter == partition_counter - return dict(signal_all) - - - def _activation_payload(value: Mapping[str, Any]) -> Mapping[str, Any]: - for key in ("domain_activation_manifest", "activation", "payload", "data"): - nested = value.get(key) - if isinstance(nested, dict) and any(field in nested for field in SG01_PROJECTION_FIELDS): - return nested - return value - - - def verify_activation_projection( - routing_activation: Mapping[str, Any], - signal_activation: Mapping[str, Any], - *, - routing_raw_sha256: str | None = None, - signal_raw_sha256: str | None = None, - ) -> dict[str, Any]: - """Compare approved semantic SG-01 projection while retaining both raw hashes.""" - - left = _activation_payload(routing_activation) - right = _activation_payload(signal_activation) - missing_left = [field for field in SG01_PROJECTION_FIELDS if field not in left] - missing_right = [field for field in SG01_PROJECTION_FIELDS if field not in right] - if missing_left or missing_right: - raise IngressError( - "SG01_PROJECTION_SHAPE", - "both activation artifacts must expose the complete approved 17-field projection", - details={"routing_missing": missing_left, "signal_missing": missing_right}, - ) - - def project(value: Mapping[str, Any]) -> dict[str, Any]: - result: dict[str, Any] = {} - for field in SG01_PROJECTION_FIELDS: - child = value[field] - if field in SG01_SET_FIELDS: - if not isinstance(child, list): - raise IngressError("SG01_PROJECTION_SHAPE", f"{field} must be an array") - child = sorted({canonical_json_bytes(item): item for item in child}.values(), key=canonical_json_bytes) - result[field] = child - return result - - left_projection = project(left) - right_projection = project(right) - if left_projection != right_projection: - raise IngressError( - "SG01_SEMANTIC_DRIFT", - "routing activation and signal SG-01 semantic projections differ", - details={"routing_projection": left_projection, "signal_projection": right_projection}, - ) - return { - "status": "PASS", - "projection": left_projection, - "projection_sha256": canonical_digest(left_projection), - "routing_raw_sha256": routing_raw_sha256, - "signal_raw_sha256": signal_raw_sha256, - "compared_keys": list(SG01_PROJECTION_FIELDS), - } - - - def _array_rows(value: Any, preferred_keys: Sequence[str]) -> list[Any]: - if isinstance(value, list): - return list(value) - if isinstance(value, dict): - for key in preferred_keys: - candidate = value.get(key) - if isinstance(candidate, list): - return list(candidate) - return [] - - - def verify_cross_artifact_seals( - documents: Mapping[str, Any], - snapshots: Mapping[str, Snapshot], - deployment_snapshots: Mapping[str, Snapshot] | None = None, - ) -> dict[str, Any]: - """Recompute the P1 guard and current-v8 producer invariants.""" - - checks: list[dict[str, Any]] = [] - issues: list[dict[str, Any]] = [] - deployment_snapshots = deployment_snapshots or {} - p1 = documents.get("stage1_part1_soft_gate_handoff") - if isinstance(p1, dict): - digest_guard = p1.get("digest_guard") - if not isinstance(digest_guard, dict): - issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff"])) - digest_guard = {} - elif any(key not in digest_guard for key in P1_DIGEST_KEYS): - issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff#digest_guard"])) - for digest_key, logical_id in P1_DIGEST_KEYS.items(): - source = snapshots.get(logical_id) or deployment_snapshots.get(logical_id) - observed = source.raw_sha256 if source else None - expected = digest_guard.get(digest_key) - passed = expected is not None and observed is not None and expected == observed - checks.append({"check_id": f"P1:{digest_key}", "status": "PASS" if passed else "UNEVALUABLE" if source is None else "FAIL"}) - if expected is not None and observed is not None and not passed: - issues.append(_issue("P1_DIGEST_MISMATCH", source_refs=[logical_id])) - else: - issues.append(_issue("P1_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part1_soft_gate_handoff"])) - p2 = documents.get("stage1_part2_review_handoff") - if p2 is not None and not isinstance(p2, dict): - issues.append(_issue("P2_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part2_review_handoff"])) - for stage in (3, 4): - logical = f"stage1_part{stage}_review_handoff" - value = documents.get(logical) - if value is not None: - wrapper_present = isinstance(value, dict) and isinstance(value.get(logical), dict) - if not wrapper_present: - issues.append(_issue(f"P{stage}_WRAPPER_MISSING", source_refs=[logical])) - ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) - for index, row in enumerate(ledger_rows): - if not isinstance(row, dict) or "domain_effects" not in row or "calculation_requests" not in row: - issues.append(_issue("CURRENT_V8_LEDGER_EXTENSION_MISSING", impact_scope="FACT", source_refs=[f"fact_ledger_base#/{index}"])) - return {"checks": checks, "issues": issues, "passed": not any(item["severity"] == "ERROR" for item in issues)} - - - def mint_stage2_id( - namespace: str, - canonical_tuple: Any, - *, - prefix: str | None = None, - algorithm_version: str = ALGORITHM_VERSION, - collision_registry: MutableMapping[str, str] | None = None, - ) -> dict[str, str]: - """Mint a domain-separated ID and reject short-ID collisions.""" - - if not re.fullmatch(r"[A-Za-z][A-Za-z0-9_.:-]{0,127}", namespace): - raise IngressError("INVALID_ID_NAMESPACE", f"invalid namespace: {namespace}") - mint_input = { - "namespace": namespace, - "algorithm_version": algorithm_version, - "canonical_tuple": canonical_tuple, - } - full_digest = canonical_digest(mint_input) - external_prefix = prefix or namespace.upper().replace("_", "-") - external_id = f"{external_prefix}-{full_digest[:24]}" - if collision_registry is not None: - prior = collision_registry.get(external_id) - if prior is not None and prior != full_digest: - raise IngressError("MINTED_ID_COLLISION", f"collision for {external_id}") - collision_registry[external_id] = full_digest - return { - "id": external_id, - "mint_input_sha256": full_digest, - "algorithm_version": algorithm_version, - } - - - def mint_context_occurrence_id( - kind: str, - logical_artifact_id: str, - json_pointer: str, - raw_value: Any, - explicit_lineage_refs: Sequence[str], - *, - stage1_canonical_id: str | None = None, - collision_registry: MutableMapping[str, str] | None = None, - ) -> dict[str, Any]: - """Mint an occurrence/lineage context ID without fuzzy entity resolution.""" - - if kind not in {"party", "object"}: - raise IngressError("CONTEXT_KIND_INVALID", "context kind must be party or object") - raw_value_sha256 = canonical_digest(raw_value) - canonical_tuple = [ - logical_artifact_id, - json_pointer, - raw_value_sha256, - sorted(set(explicit_lineage_refs)), - ] - minted = mint_stage2_id( - f"{kind}_context_occurrence", - canonical_tuple, - prefix="PC" if kind == "party" else "OC", - collision_registry=collision_registry, - ) - source_ref = _source_ref( - logical_artifact_id, - json_pointer, - raw_value, - stage1_id=stage1_canonical_id, - ) - identity_kind = "PARTY_CONTEXT" if kind == "party" else "OBJECT_CONTEXT" - result: dict[str, Any] = { - "identity_kind": identity_kind, - f"{kind}_context_id": stage1_canonical_id or minted["id"], - f"stage1_{kind}_id": stage1_canonical_id, - "identity_disposition": ( - "PRESERVED_STAGE1_CANONICAL_ID" - if stage1_canonical_id is not None - else "IDENTITY_UNRESOLVED" - ), - "source_refs": [source_ref], - "explicit_lineage_refs": sorted(set(explicit_lineage_refs)), - "display_label_nfc": unicodedata.normalize("NFC", str(raw_value)), - "raw_display_value_sha256": raw_value_sha256, - } - if stage1_canonical_id is None: - result["derivation"] = { - "derivation_id": minted["id"], - "algorithm_version": ALGORITHM_VERSION, - "sorted_input_refs": sorted(set(explicit_lineage_refs)) or [f"{logical_artifact_id}:{json_pointer}"], - "mint_input_sha256": minted["mint_input_sha256"], - } - return result - - - def _extract_review_arrays(document: Any) -> list[Any]: - if not isinstance(document, dict): - return [] - result: list[Any] = [] - for key in ("review_items", "review_queue", "blocked_review_items", "unresolved_review_items"): - value = document.get(key) - if isinstance(value, list): - result.extend(value) - for wrapper in ( - "handoff", - "payload", - "review", - "data", - "stage1_part3_review_handoff", - "stage1_part4_review_handoff", - ): - nested = document.get(wrapper) - if isinstance(nested, dict): - result.extend(_extract_review_arrays(nested)) - return result - - - def normalize_review_items( - review_documents: Mapping[str, Any], - release_lock: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - """Preserve every raw review occurrence in exactly one normalized partition.""" - - stage_map = { - "stage1_part1_soft_gate_handoff": ("P1", "S2A-P1-HANDOFF-FLAT-V1"), - "stage1_part2_review_handoff": ("P2", "S2A-P2-HANDOFF-FLAT-V1"), - "stage1_part3_review_handoff": ("P3", "S2A-P3-HANDOFF-WRAPPED-V1"), - "stage1_part4_review_handoff": ("P4", "S2A-P4-HANDOFF-WRAPPED-V1"), - } - release_lock = release_lock or {} - mapping_decision = _adapter_decision(release_lock, "S2-REVIEW-MAP-V1") - mapping_rows = mapping_decision.get("mappings", []) if isinstance(mapping_decision, dict) else [] - closed_mappings: dict[tuple[str, str, str], Mapping[str, Any]] = {} - for mapping in mapping_rows: - if not isinstance(mapping, dict): - continue - key = ( - str(mapping.get("source_stage", "ANY")), - str(mapping.get("source_field_kind")), - str(mapping.get("source_value")), - ) - if key in closed_mappings: - raise IngressError("REVIEW_MAPPING_DUPLICATE", f"duplicate release review mapping: {key}") - closed_mappings[key] = mapping - adapter_issues: list[dict[str, Any]] = [] - if not closed_mappings: - adapter_issues.append(_issue("REVIEW_MAPPING_CONTRACT_MISSING", impact_scope="REVIEW_ITEM")) - collected: list[tuple[str, Any, str]] = [] - adapter_ids_seen: set[str] = set() - for logical_id in sorted(review_documents): - source_stage, adapter_id = stage_map.get(logical_id, (None, None)) - if source_stage is None or adapter_id is None: - continue - adapter_ids_seen.add(adapter_id) - decision = _adapter_decision(release_lock, adapter_id) - if decision is None: - adapter_issues.append(_issue("HANDOFF_ADAPTER_CONTRACT_MISSING", impact_scope="REVIEW_ITEM", source_refs=[logical_id])) - continue - document = review_documents[logical_id] - wrapper_pointer = decision.get("wrapper_json_pointer") - wrapper_found, wrapper = _json_pointer_value(document, wrapper_pointer) - if not wrapper_found or not isinstance(wrapper, dict): - code = "P2_HANDOFF_NOT_FLAT_OBJECT" if source_stage == "P2" else f"{source_stage}_WRAPPER_MISSING" - adapter_issues.append(_issue(code, impact_scope="REVIEW_ITEM", source_refs=[logical_id])) - continue - expected_version = decision.get("schema_version") - if wrapper.get("schema_version") != expected_version: - adapter_issues.append( - _issue( - f"{source_stage}_HANDOFF_SCHEMA_VERSION_MISMATCH", - impact_scope="REVIEW_ITEM", - source_refs=[logical_id], - ) - ) - items_found, raw_items = _json_pointer_value(wrapper, decision.get("review_items_json_pointer")) - if not items_found or not isinstance(raw_items, list): - adapter_issues.append( - _issue( - f"{source_stage}_REVIEW_ITEMS_SHAPE", - impact_scope="REVIEW_ITEM", - source_refs=[logical_id], - ) - ) - raw_items = [] - if decision.get("count_field_required") is True: - declared_count = wrapper.get("review_item_count") - if not isinstance(declared_count, int) or declared_count != len(raw_items): - adapter_issues.append( - _issue("REVIEW_CONSERVATION_FAILED", impact_scope="REVIEW_ITEM", source_refs=[logical_id]) - ) - for raw_item in raw_items: - raw_hash = canonical_digest(raw_item) - collected.append((logical_id, raw_item, raw_hash)) - grouped: dict[tuple[str, str], list[Any]] = defaultdict(list) - for logical_id, raw_item, raw_hash in collected: - grouped[(logical_id, raw_hash)].append(raw_item) - raw_occurrences: list[dict[str, Any]] = [] - normalized: list[dict[str, Any]] = [] - partition_counts = {key: 0 for key in ("SUPPORTED", "CONDITIONAL", "UNRESOLVED", "EXCLUDED", "UNMAPPED")} - for (logical_id, raw_hash), items in sorted(grouped.items()): - for duplicate_index, raw_item in enumerate(items): - raw_status = raw_item.get("status") if isinstance(raw_item, dict) else None - raw_severity = raw_item.get("severity") if isinstance(raw_item, dict) else None - source_stage, adapter_id = stage_map.get(logical_id, ("P1", "S2A-P1-HANDOFF-FLAT-V1")) - field_kind = "REVIEW_ITEM_STATUS" if raw_status is not None else "REVIEW_ITEM_SEVERITY" - source_value = str(raw_status if raw_status is not None else raw_severity) - mapping = closed_mappings.get((source_stage, field_kind, source_value)) or closed_mappings.get( - ("ANY", field_kind, source_value) - ) - partition = str(mapping.get("normalized_partition")) if mapping is not None else "UNMAPPED" - if partition not in partition_counts: - raise IngressError("REVIEW_MAPPING_PARTITION_INVALID", f"release mapping produced {partition}") - upstream_id = raw_item.get("review_id") if isinstance(raw_item, dict) else None - if upstream_id: - review_id = str(upstream_id) - minted_flag = False - else: - minted = mint_stage2_id( - "review_occurrence", - [logical_id, raw_hash, duplicate_index], - prefix="REV", - ) - review_id = minted["id"] - minted_flag = True - raw_occurrence = { - "source_stage": source_stage, - "logical_input_id": logical_id, - "canonical_raw_item_sha256": raw_hash, - "duplicate_occurrence_index": duplicate_index, - } - raw_occurrences.append(raw_occurrence) - reason_codes = ["UNMAPPED_REVIEW_STATUS"] if partition == "UNMAPPED" else [] - row = { - "review_key": { - "value": review_id, - "minted": minted_flag, - "source_stage": source_stage, - "logical_input_id": logical_id, - "canonical_raw_item_sha256": raw_hash, - "duplicate_occurrence_index": duplicate_index, - "upstream_review_id": str(upstream_id) if upstream_id is not None else None, - }, - "partition": partition, - "source_status_raw": str(raw_status) if raw_status is not None else None, - "source_severity_raw": str(raw_severity) if raw_severity is not None else None, - "source_item_raw_sha256": raw_hash, - "mapping_id": str(mapping.get("mapping_id")) if mapping is not None else f"S2-REVIEW-MAP-V1:{source_stage}:UNMAPPED", - "impact_scope": "REVIEW_ITEM", - "scope_refs": [review_id], - "source_contract_row_refs": [logical_id], - "reason_codes": reason_codes, - "downstream_allowed_actions": ["CARRY_FORWARD_TO_LAWYER_REVIEW"], - "_adapter_id": adapter_id, - } - normalized.append(row) - partition_counts[partition] += 1 - raw_counter = Counter((logical_id, raw_hash) for logical_id, _, raw_hash in collected) - normalized_counter = Counter( - (row["review_key"]["logical_input_id"], row["review_key"]["canonical_raw_item_sha256"]) - for row in normalized - ) - adapter_ids = sorted({row.pop("_adapter_id") for row in normalized} | adapter_ids_seen) - adapter_conservation_pass = not any( - item["issue_code"] == "REVIEW_CONSERVATION_FAILED" for item in adapter_issues - ) - return { - "schema_version": "stage2_review_normalization_receipt.v1", - "adapter_ids": adapter_ids or ["S2A-P1-HANDOFF-FLAT-V1"], - "mapping_table_version": "S2-REVIEW-MAP-V1", - "raw_occurrences": sorted( - raw_occurrences, - key=lambda row: ( - row["source_stage"], - row["logical_input_id"], - row["canonical_raw_item_sha256"], - row["duplicate_occurrence_index"], - ), - ), - "normalized_occurrences": sorted( - normalized, - key=lambda row: row["review_key"]["value"], - ), - "partition_counts": partition_counts, - "conservation_status": "PASS" if raw_counter == normalized_counter and adapter_conservation_pass else "FAIL", - "_issues": adapter_issues, - } - - - def check_conservation( - documents: Mapping[str, Any], - *, - signal_all: Mapping[str, Any] | None = None, - normalized_reviews: Mapping[str, Any] | None = None, - source_snapshots: Mapping[str, Snapshot] | None = None, - ) -> dict[str, Any]: - """Independently compute core set, cardinality, and multiset invariants.""" - - checks: list[dict[str, Any]] = [] - issues: list[dict[str, Any]] = [] - source_snapshots = source_snapshots or {} - - def add_check( - check_id: str, - passed: bool | None, - left: Sequence[Any] | Counter[Any] | None, - right: Sequence[Any] | Counter[Any] | None, - *, - issue_code: str, - impact_scope: str, - source_refs: Sequence[str], - details: Mapping[str, Any] | None = None, - ) -> None: - left_counter = left if isinstance(left, Counter) else Counter(left or []) - right_counter = right if isinstance(right, Counter) else Counter(right or []) - row: dict[str, Any] = { - "check_id": check_id, - "status": "UNEVALUABLE" if passed is None else "PASS" if passed else "FAIL", - "left_count": sum(left_counter.values()) if left is not None else None, - "right_count": sum(right_counter.values()) if right is not None else None, - "left_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in left_counter.items())) if left is not None else None, - "right_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in right_counter.items())) if right is not None else None, - } - if details: - row.update(details) - checks.append(row) - if passed is False: - issues.append(_issue(issue_code, impact_scope=impact_scope, source_refs=source_refs)) - - bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) - ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) - bo_ids = [str(row["BO_ID"]) for row in bo_rows if isinstance(row, dict) and row.get("BO_ID") is not None] - source_bo_ids = [ - str(row["source_bo_id"]) - for row in ledger_rows - if isinstance(row, dict) and row.get("source_bo_id") is not None - ] - missing_bo_id_rows = [index for index, row in enumerate(bo_rows) if not isinstance(row, dict) or row.get("BO_ID") is None] - missing_source_bo_rows = [ - index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("source_bo_id") is None - ] - bo_pass = ( - not missing_bo_id_rows - and not missing_source_bo_rows - and Counter(bo_ids) == Counter(source_bo_ids) - ) - add_check( - "BO_FACT_MULTISET", - bo_pass, - bo_ids, - source_bo_ids, - issue_code="BO_FACT_CONSERVATION_FAILED", - impact_scope="FACT", - source_refs=["bo", "fact_ledger_base"], - details={ - "missing_bo_id_rows": missing_bo_id_rows, - "missing_source_bo_id_rows": missing_source_bo_rows, - "duplicate_bo_ids": sorted(key for key, count in Counter(bo_ids).items() if count > 1), - "dangling_source_bo_ids": sorted(set(source_bo_ids) - set(bo_ids)), - }, - ) - missing_fact_id_rows = [ - index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("fact_id") is None - ] - fact_ids = [str(row["fact_id"]) for row in ledger_rows if isinstance(row, dict) and row.get("fact_id") is not None] - expected_fact_ids = [f"F-{index:03d}" for index in range(1, len(ledger_rows) + 1)] - fact_pass = not missing_fact_id_rows and fact_ids == expected_fact_ids and len(fact_ids) == len(set(fact_ids)) - add_check( - "FACT_ID_SEQUENCE", - fact_pass, - fact_ids, - expected_fact_ids, - issue_code="FACT_ID_CONSERVATION_FAILED", - impact_scope="FACT", - source_refs=["fact_ledger_base"], - details={"missing_fact_id_rows": missing_fact_id_rows, "observed": fact_ids}, - ) - extension_missing = [ - index - for index, row in enumerate(ledger_rows) - if not isinstance(row, dict) - or not isinstance(row.get("domain_effects"), dict) - or not isinstance(row.get("calculation_requests"), list) - ] - add_check( - "CURRENT_V8_LEDGER_EXTENSIONS", - not extension_missing, - list(range(len(ledger_rows))), - [index for index in range(len(ledger_rows)) if index not in extension_missing], - issue_code="CURRENT_V8_LEDGER_EXTENSION_MISSING", - impact_scope="FACT", - source_refs=["fact_ledger_base"], - details={"missing_row_indices": extension_missing}, - ) - les_rows = _array_rows( - documents.get("legal_effect_structures"), - ("structures", "structure_records", "legal_effect_structures", "rows", "items"), - ) - dangling_les: list[str] = [] - les_ids: list[str] = [] - for row in les_rows: - if not isinstance(row, dict): - continue - structure_id = row.get("structure_id", row.get("legal_effect_structure_id")) - if structure_id is not None: - les_ids.append(str(structure_id)) - refs = row.get("source_bo_ids", []) - if isinstance(refs, list): - dangling_les.extend(str(ref) for ref in refs if ref not in set(bo_ids)) - duplicate_les_ids = sorted(key for key, count in Counter(les_ids).items() if count > 1) - les_pass = not dangling_les and not duplicate_les_ids and len(les_ids) == len(les_rows) - add_check( - "LES_BO_JOIN", - les_pass, - [str(row.get("structure_id", row.get("legal_effect_structure_id"))) for row in les_rows if isinstance(row, dict)], - les_ids, - issue_code="LES_BO_JOIN_FAILED", - impact_scope="CLUSTER", - source_refs=["legal_effect_structures", "bo"], - details={"dangling_refs": sorted(dangling_les), "duplicate_structure_ids": duplicate_les_ids}, - ) - declared_les_count = None - les_document = documents.get("legal_effect_structures") - if isinstance(les_document, dict): - for key in ("declared_structure_count", "structure_count", "record_count"): - if isinstance(les_document.get(key), int): - declared_les_count = int(les_document[key]) - break - declared_les_pass = None if declared_les_count is None else declared_les_count == len(les_rows) - add_check( - "LES_DECLARED_ACTUAL_COUNT", - declared_les_pass, - [None] * declared_les_count if declared_les_count is not None else None, - [None] * len(les_rows), - issue_code="LES_DECLARED_COUNT_MISMATCH", - impact_scope="CLUSTER", - source_refs=["legal_effect_structures"], - ) - actual_domain_index: dict[str, list[str]] = defaultdict(list) - actual_bo_index: dict[str, list[str]] = defaultdict(list) - ledger_structure_refs: list[tuple[str, str, str]] = [] - ledger_type_refs: list[tuple[str, str, str]] = [] - actual_structure_refs: list[tuple[str, str, str]] = [] - actual_type_refs: list[tuple[str, str, str]] = [] - route_count_errors: list[str] = [] - for row in les_rows: - if not isinstance(row, dict): - continue - structure_id = str(row.get("structure_id", row.get("legal_effect_structure_id", "MISSING"))) - domain_id = str(row.get("domain_id", "MISSING")) - type_id = str(row.get("type_id", row.get("type", "MISSING"))) - actual_domain_index[domain_id].append(structure_id) - source_ids = row.get("source_bo_ids", []) - if isinstance(source_ids, list): - for bo_id in source_ids: - actual_bo_index[str(bo_id)].append(structure_id) - actual_structure_refs.append((str(bo_id), domain_id, structure_id)) - actual_type_refs.append((str(bo_id), domain_id, type_id)) - routes = row.get("routes", []) - if isinstance(routes, list) and row.get("route_count", len(routes)) != len(routes): - route_count_errors.append(structure_id) - for row in ledger_rows: - if not isinstance(row, dict): - continue - bo_id = str(row.get("source_bo_id", "MISSING")) - effects = row.get("domain_effects", {}) - if not isinstance(effects, dict): - continue - for domain_id, effect in effects.items(): - if not isinstance(effect, dict): - continue - for structure_id in effect.get("structure_ids", []) if isinstance(effect.get("structure_ids"), list) else []: - ledger_structure_refs.append((bo_id, str(domain_id), str(structure_id))) - for type_id in effect.get("type_ids", []) if isinstance(effect.get("type_ids"), list) else []: - ledger_type_refs.append((bo_id, str(domain_id), str(type_id))) - structure_index = les_document.get("structure_index", {}) if isinstance(les_document, dict) else {} - index_present = isinstance(structure_index, dict) and bool(structure_index) - index_ok = True - if index_present: - declared_by_domain = structure_index.get("by_domain_id", {}) - declared_by_bo = structure_index.get("by_bo_id", {}) - index_ok = ( - isinstance(declared_by_domain, dict) - and isinstance(declared_by_bo, dict) - and {str(key): Counter(map(str, value)) for key, value in declared_by_domain.items() if isinstance(value, list)} - == {key: Counter(value) for key, value in actual_domain_index.items()} - and {str(key): Counter(map(str, value)) for key, value in declared_by_bo.items() if isinstance(value, list)} - == {key: Counter(value) for key, value in actual_bo_index.items()} - ) - reverse_ok = ( - (not ledger_structure_refs or Counter(ledger_structure_refs) == Counter(actual_structure_refs)) - and (not ledger_type_refs or Counter(ledger_type_refs) == Counter(actual_type_refs)) - and not route_count_errors - and index_ok - ) - add_check( - "LES_REVERSE_INDEX", - reverse_ok, - ledger_structure_refs + ledger_type_refs, - actual_structure_refs + actual_type_refs, - issue_code="LES_REVERSE_INDEX_MISMATCH", - impact_scope="CLUSTER", - source_refs=["legal_effect_structures", "fact_ledger_base"], - details={"index_present": index_present, "route_count_errors": route_count_errors}, - ) - evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) - event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) - evidence_ids = [ - str(row.get("evidence_id", row.get("id"))) - for row in evidence_rows - if isinstance(row, dict) and (row.get("evidence_id") is not None or row.get("id") is not None) - ] - event_ids = [ - str(row.get("event_id", row.get("id"))) - for row in event_rows - if isinstance(row, dict) and (row.get("event_id") is not None or row.get("id") is not None) - ] - fact_evidence_refs: list[str] = [] - fact_event_refs: list[str] = [] - event_evidence_refs: list[str] = [] - for row in ledger_rows: - if not isinstance(row, dict): - continue - evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) - event_values = row.get("event_refs", row.get("event_ids", [])) - if isinstance(evidence_values, list): - fact_evidence_refs.extend(str(ref) for ref in evidence_values) - if isinstance(event_values, list): - fact_event_refs.extend(str(ref) for ref in event_values) - for row in event_rows: - if not isinstance(row, dict): - continue - evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) - if isinstance(evidence_values, list): - event_evidence_refs.extend(str(ref) for ref in evidence_values) - evidence_failures = sorted( - set(fact_evidence_refs + event_evidence_refs) - set(evidence_ids) - ) - duplicate_evidence_ids = sorted(key for key, count in Counter(evidence_ids).items() if count > 1) - evidence_pass = not evidence_failures and not duplicate_evidence_ids - add_check( - "EVIDENCE_REFERENCE_CONSERVATION", - evidence_pass, - fact_evidence_refs + event_evidence_refs, - evidence_ids, - issue_code="EVIDENCE_REFERENCE_CONSERVATION_FAILED", - impact_scope="EVIDENCE", - source_refs=["evidence_indexed", "fact_ledger_base", "evidence_event_candidates"], - details={"dangling_refs": evidence_failures, "duplicate_evidence_ids": duplicate_evidence_ids}, - ) - event_failures = sorted(set(fact_event_refs) - set(event_ids)) - duplicate_event_ids = sorted(key for key, count in Counter(event_ids).items() if count > 1) - event_pass = not event_failures and not duplicate_event_ids - add_check( - "EVENT_REFERENCE_CONSERVATION", - event_pass, - fact_event_refs, - event_ids, - issue_code="EVENT_REFERENCE_CONSERVATION_FAILED", - impact_scope="EVIDENCE", - source_refs=["evidence_event_candidates", "fact_ledger_base"], - details={"dangling_refs": event_failures, "duplicate_event_ids": duplicate_event_ids}, - ) - disposition_rows = [row.get("disposition") for row in event_rows if isinstance(row, dict) and "disposition" in row] - b2_gate = documents.get("b2_event_candidates_gate") - declared_dispositions = None - if isinstance(b2_gate, dict): - declared_dispositions = b2_gate.get("event_disposition_counts") - if declared_dispositions is None and isinstance(b2_gate.get("summary"), dict): - declared_dispositions = b2_gate["summary"].get("event_disposition_counts") - if isinstance(declared_dispositions, dict): - disposition_expected = Counter( - {str(key): int(value) for key, value in declared_dispositions.items() if isinstance(value, int)} - ) - disposition_actual = Counter(str(value) for value in disposition_rows) - disposition_pass: bool | None = disposition_actual == disposition_expected - elif disposition_rows: - disposition_expected = Counter(str(value) for value in disposition_rows) - disposition_actual = Counter(str(value) for value in disposition_rows) - disposition_pass = all(isinstance(value, str) and value for value in disposition_rows) - else: - disposition_expected = Counter() - disposition_actual = Counter() - disposition_pass = None - add_check( - "EVENT_DISPOSITION_CONSERVATION", - disposition_pass, - disposition_actual, - disposition_expected, - issue_code="EVENT_DISPOSITION_CONSERVATION_FAILED", - impact_scope="EVIDENCE", - source_refs=["evidence_event_candidates", "b2_event_candidates_gate"], - ) - writer_report = documents.get("fact_ledger_writer_report") - if isinstance(writer_report, dict): - observed_domain_coverage = Counter( - str(domain_id) - for row in ledger_rows - if isinstance(row, dict) and isinstance(row.get("domain_effects"), dict) - for domain_id in row["domain_effects"] - ) - declared_domain_coverage = Counter( - {str(key): int(value) for key, value in writer_report.get("domain_effect_coverage", {}).items() if isinstance(value, int)} - ) - observed_readiness = Counter( - str(request.get("operand_state")) - for row in ledger_rows - if isinstance(row, dict) and isinstance(row.get("calculation_requests"), list) - for request in row["calculation_requests"] - if isinstance(request, dict) - ) - declared_readiness = Counter( - {str(key): int(value) for key, value in writer_report.get("calculation_readiness", {}).items() if isinstance(value, int)} - ) - ledger_snapshot = source_snapshots.get("fact_ledger_base") - final_hash = writer_report.get("final_sha256") - writer_pass = ( - writer_report.get("row_count") == len(ledger_rows) - and declared_domain_coverage == observed_domain_coverage - and declared_readiness == observed_readiness - and (ledger_snapshot is None or final_hash == ledger_snapshot.raw_sha256) - ) - add_check( - "FACT_LEDGER_WRITER_REPORT_CONNECTION", - writer_pass, - [len(ledger_rows), observed_domain_coverage, observed_readiness, ledger_snapshot.raw_sha256 if ledger_snapshot else None], - [writer_report.get("row_count"), declared_domain_coverage, declared_readiness, final_hash], - issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", - impact_scope="FACT", - source_refs=["fact_ledger_base", "fact_ledger_writer_report"], - ) - else: - add_check( - "FACT_LEDGER_WRITER_REPORT_CONNECTION", - None, - None, - None, - issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", - impact_scope="FACT", - source_refs=["fact_ledger_base", "fact_ledger_writer_report"], - ) - if signal_all is not None: - file_pass = bool(signal_all.get("file_conservation_pass")) - record_pass = bool(signal_all.get("record_conservation_pass")) - checks.append({"check_id": "SIGNAL_FILE_ROW_CONSERVATION", "status": "PASS" if file_pass else "FAIL"}) - checks.append({"check_id": "SIGNAL_RECORD_OCCURRENCE_CONSERVATION", "status": "PASS" if record_pass else "FAIL"}) - issues.extend(signal_all.get("issues", [])) - if not file_pass: - issues.append(_issue("SIGNAL_FILE_CONSERVATION_FAILED", impact_scope="SIGNAL")) - if not record_pass: - issues.append(_issue("SIGNAL_RECORD_CONSERVATION_FAILED", impact_scope="SIGNAL")) - if normalized_reviews is not None: - review_pass = normalized_reviews.get("conservation_status") == "PASS" - checks.append({"check_id": "REVIEW_OCCURRENCE_CONSERVATION", "status": "PASS" if review_pass else "FAIL"}) - if not review_pass: - issues.append(_issue("REVIEW_CONSERVATION_FAILED", impact_scope="REVIEW_ITEM")) - issues.extend(normalized_reviews.get("_issues", [])) - return {"checks": checks, "issues": issues, "passed": not any(check["status"] == "FAIL" for check in checks)} - - - def _source_ref( - logical_id: str, - pointer: str, - raw_value: Any = _RAW_VALUE_UNSET, - *, - stage1_id: str | None = None, - ) -> dict[str, Any]: - """Build a truthful RFC 6901 provenance row without pointer narrowing.""" - - row: dict[str, Any] = { - "logical_artifact_id": logical_id, - "json_pointer": pointer, - "raw_value_sha256": canonical_digest( - [logical_id, pointer] - if raw_value is _RAW_VALUE_UNSET - else raw_value - ), - "source_contract_row_ref": logical_id, - } - if stage1_id is not None: - row["stage1_id"] = stage1_id - return row - - - def _dedupe_source_refs(rows: Iterable[Mapping[str, Any]]) -> list[dict[str, Any]]: - by_digest = {canonical_digest(dict(row)): dict(row) for row in rows} - return [by_digest[key] for key in sorted(by_digest)] - - - def _coerce_source_ref(value: Mapping[str, Any] | str) -> dict[str, Any]: - if isinstance(value, dict) and { - "logical_artifact_id", - "json_pointer", - "raw_value_sha256", - }.issubset(value): - return dict(value) - text = str(value) - if "#" in text: - logical_id, pointer = text.split("#", 1) - else: - logical_id, pointer = text, "" - if pointer and not pointer.startswith("/"): - pointer = "/" + pointer - return _source_ref(logical_id or "unknown", pointer, text) - - - def _artifact_header( - schema_id: str = CONTEXT_SCHEMA_ID, - *, - schema_sha256: str = "0" * 64, - run_id: str = "STRUCTURAL-FIXTURE", - input_set_digest: str = "0" * 64, - stage2_release_digest: str = "0" * 64, - algorithm_digest: str = "0" * 64, - release_class: str = "DEV_FIXTURE_RELEASE", - ) -> dict[str, Any]: - return { - "schema_id": schema_id, - "schema_sha256": schema_sha256, - "schema_version": "stage2_s2_00_context.v2", - "producer_id": "S2_00", - "run_id": run_id, - "input_set_digest": input_set_digest, - "stage2_release_digest": stage2_release_digest, - "algorithm_digest": algorithm_digest, - "release_class": release_class, - } - - - def _derivation( - namespace: str, - input_refs: Sequence[str], - payload: Any, - ) -> dict[str, Any]: - sorted_refs = sorted(set(str(item) for item in input_refs)) or ["S2_00:EMPTY_INPUT_SET"] - minted = mint_stage2_id(namespace, [sorted_refs, payload], prefix="DRV") - return { - "derivation_id": minted["id"], - "algorithm_version": ALGORITHM_VERSION, - "sorted_input_refs": sorted_refs, - "mint_input_sha256": minted["mint_input_sha256"], - } - - - def build_case_context( - documents: Mapping[str, Any], - *, - run_binding_digest: str, - issues: Sequence[Mapping[str, Any]] = (), - artifact_header: Mapping[str, Any] | None = None, - signal_all: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - routing = documents.get("domain_activation_manifest") - routing_payload = _activation_payload(routing) if isinstance(routing, dict) else {} - active = routing_payload.get("active_domain_ids", []) - expected = routing_payload.get("expected_runnable_domain_ids", []) - bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) - ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) - les_rows = _array_rows(documents.get("legal_effect_structures"), ("structures", "structure_records", "legal_effect_structures", "rows", "items")) - bo_ids = sorted( - str(row["BO_ID"]) - for row in bo_rows - if isinstance(row, dict) and row.get("BO_ID") is not None - ) - structure_ids = sorted( - str(row.get("structure_id", row.get("legal_effect_structure_id"))) - for row in les_rows - if isinstance(row, dict) and (row.get("structure_id") is not None or row.get("legal_effect_structure_id") is not None) - ) - structures_by_bo: dict[str, list[str]] = defaultdict(list) - for row in les_rows: - if not isinstance(row, dict): - continue - structure_id = row.get("structure_id", row.get("legal_effect_structure_id")) - for bo_id in row.get("source_bo_ids", []) if isinstance(row.get("source_bo_ids"), list) else []: - if structure_id is not None: - structures_by_bo[str(bo_id)].append(str(structure_id)) - fact_contexts: list[dict[str, Any]] = [] - for index, row in enumerate(ledger_rows): - if not isinstance(row, dict): - continue - if row.get("fact_id") is None or row.get("source_bo_id") is None: - continue - fact_id = str(row["fact_id"]) - source_bo_id = str(row.get("source_bo_id", "MISSING")) - evidence_refs = row.get("evidence_refs", row.get("evidence_ids", [])) - event_refs = row.get("event_refs", row.get("event_ids", [])) - calculation_requests = row.get("calculation_requests", []) - law_version_refs = row.get("law_version_refs", []) - signal_occurrence_refs = sorted( - str(occurrence.get("occurrence_ref")) - for occurrence in (signal_all or {}).get("record_occurrences", []) - if occurrence.get("disposition") == "USED" - and ( - f"fact_id:{fact_id}" in occurrence.get("binding_refs", []) - or f"source_bo_id:{source_bo_id}" in occurrence.get("binding_refs", []) - or f"bo_id:{source_bo_id}" in occurrence.get("binding_refs", []) - ) - ) - fact_contexts.append( - { - "fact_id": fact_id, - "source_bo_id": source_bo_id, - "party_context_ids": [], - "object_context_ids": [], - "structure_refs": sorted(set(structures_by_bo.get(source_bo_id, []))), - "signal_occurrence_refs": signal_occurrence_refs, - "calculation_request_refs": sorted( - canonical_digest(value) - for value in (calculation_requests if isinstance(calculation_requests, list) else []) - ), - "law_version_refs": sorted(str(value) for value in law_version_refs) if isinstance(law_version_refs, list) else [], - "evidence_refs": sorted(str(value) for value in evidence_refs) if isinstance(evidence_refs, list) else [], - "event_refs": sorted(str(value) for value in event_refs) if isinstance(event_refs, list) else [], - "review_keys": [], - "source_refs": [_source_ref("fact_ledger_base", f"/rows/{index}", row, stage1_id=fact_id)], - } - ) - source_refs = _dedupe_source_refs( - [ - _source_ref("bo", "", documents.get("bo")), - _source_ref("fact_ledger_base", "", documents.get("fact_ledger_base")), - _source_ref("legal_effect_structures", "", documents.get("legal_effect_structures")), - _source_ref("domain_activation_manifest", "", documents.get("domain_activation_manifest")), - _source_ref("client_goal", "", documents.get("client_goal")), - ] - ) - minted = mint_stage2_id( - "case_context", - [run_binding_digest, [row["fact_id"] for row in fact_contexts], bo_ids, structure_ids], - prefix="CC", - ) - unavailable_scope_refs = sorted( - { - str(scope_ref) - for item in issues - if item.get("issue_code") - for scope_ref in item.get("scope_refs", []) - } - ) - return { - "artifact_header": dict(artifact_header or _artifact_header()), - "case_context_id": minted["id"], - "facts": sorted(fact_contexts, key=lambda row: row["fact_id"]), - "bo_ids": bo_ids, - "structure_ids": structure_ids, - "active_domain_ids": sorted(set(str(value) for value in active)) if isinstance(active, list) else [], - "expected_runnable_domain_ids": sorted(set(str(value) for value in expected)) if isinstance(expected, list) else [], - "unavailable_scope_refs": unavailable_scope_refs, - "source_refs": source_refs, - "derivation": _derivation("case_context_derivation", [f"fact:{row['fact_id']}" for row in fact_contexts], minted["mint_input_sha256"]), - } - - - def build_evidence_inventory( - documents: Mapping[str, Any], - *, - run_binding_digest: str, - artifact_header: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) - event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) - ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) - event_by_evidence: dict[str, set[str]] = defaultdict(set) - for index, row in enumerate(event_rows): - if not isinstance(row, dict): - continue - event_id = str(row.get("event_id", f"EVENT-OCCURRENCE-{index}")) - refs = row.get("evidence_refs", row.get("evidence_ids", [])) - if isinstance(refs, list): - for ref in refs: - event_by_evidence[str(ref)].add(event_id) - facts_by_evidence: dict[str, set[str]] = defaultdict(set) - for row in ledger_rows: - if not isinstance(row, dict) or row.get("fact_id") is None: - continue - refs = row.get("evidence_refs", row.get("evidence_ids", [])) - if isinstance(refs, list): - for ref in refs: - facts_by_evidence[str(ref)].add(str(row["fact_id"])) - items: list[dict[str, Any]] = [] - for index, row in enumerate(evidence_rows): - if not isinstance(row, dict): - continue - evidence_id = str(row.get("evidence_id", row.get("id", f"EVIDENCE-OCCURRENCE-{index}"))) - items.append( - { - "evidence_id": evidence_id, - "event_ids": sorted(event_by_evidence.get(evidence_id, set())), - "linked_fact_ids": sorted(facts_by_evidence.get(evidence_id, set())), - "technical_disposition": "AVAILABLE", - "source_refs": [_source_ref("evidence_indexed", f"/items/{index}", row, stage1_id=evidence_id)], - } - ) - source_refs = _dedupe_source_refs( - [ - _source_ref("evidence_indexed", "", documents.get("evidence_indexed")), - _source_ref("evidence_event_candidates", "", documents.get("evidence_event_candidates")), - _source_ref("fact_ledger_base", "", documents.get("fact_ledger_base")), - ] - ) - return { - "artifact_header": dict(artifact_header or _artifact_header()), - "items": sorted(items, key=lambda row: row["evidence_id"]), - "source_refs": source_refs, - "derivation": _derivation("evidence_inventory_derivation", [f"evidence:{row['evidence_id']}" for row in items], run_binding_digest), - } - - - def _candidate_values(row: Mapping[str, Any], keys: Sequence[str]) -> list[tuple[str, Any]]: - result: list[tuple[str, Any]] = [] - for key in keys: - value = row.get(key) - if value is None: - continue - if isinstance(value, list): - result.extend((f"/{key}/{index}", item) for index, item in enumerate(value)) - else: - result.append((f"/{key}", value)) - return result - - - def build_object_registry( - documents: Mapping[str, Any], - *, - run_binding_digest: str, - artifact_header: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) - objects: list[dict[str, Any]] = [] - collision_registry: dict[str, str] = {} - for index, row in enumerate(bo_rows): - if not isinstance(row, dict): - continue - explicit_bo = row.get("BO_ID") - for pointer, value in _candidate_values(row, ("object", "objects", "asset", "subject_matter")): - stage1_object_id = value.get("object_id") if isinstance(value, dict) else None - context = mint_context_occurrence_id( - "object", - "bo", - f"/{index}{pointer}", - value, - [str(explicit_bo)] if explicit_bo is not None else [], - stage1_canonical_id=str(stage1_object_id) if stage1_object_id is not None else None, - collision_registry=collision_registry, - ) - objects.append(context) - return { - "artifact_header": dict(artifact_header or _artifact_header()), - "objects": sorted(objects, key=lambda row: row["object_context_id"]), - "source_refs": [_source_ref("bo", "", documents.get("bo"))], - "derivation": _derivation("object_registry_derivation", [row["object_context_id"] for row in objects], run_binding_digest), - } - - - def build_party_and_title_context( - documents: Mapping[str, Any], - *, - run_binding_digest: str, - artifact_header: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) - parties: list[dict[str, Any]] = [] - collision_registry: dict[str, str] = {} - for index, row in enumerate(bo_rows): - if not isinstance(row, dict): - continue - explicit_bo = row.get("BO_ID") - for pointer, value in _candidate_values(row, ("party", "parties", "creditor", "debtor", "counterparty")): - stage1_party_id = value.get("party_id") if isinstance(value, dict) else None - identity = mint_context_occurrence_id( - "party", - "bo", - f"/{index}{pointer}", - value, - [str(explicit_bo)] if explicit_bo is not None else [], - stage1_canonical_id=str(stage1_party_id) if stage1_party_id is not None else None, - collision_registry=collision_registry, - ) - source_refs = identity["source_refs"] - parties.append( - { - "party_identity": identity, - "party_title": value.get("party_title") if isinstance(value, dict) else None, - "defendant_role": value.get("defendant_role") if isinstance(value, dict) else None, - "liability_context": value.get("liability_context") if isinstance(value, dict) else None, - "client_instruction": value.get("client_instruction") if isinstance(value, dict) else None, - "recovery_information": value.get("recovery_information") if isinstance(value, dict) else None, - "source_refs": source_refs, - } - ) - return { - "artifact_header": dict(artifact_header or _artifact_header()), - "parties": sorted(parties, key=lambda row: row["party_identity"]["party_context_id"]), - "source_refs": [_source_ref("bo", "", documents.get("bo"))], - "derivation": _derivation( - "party_title_derivation", - [row["party_identity"]["party_context_id"] for row in parties], - run_binding_digest, - ), - } - - - def build_slot_crosswalk( - documents: Mapping[str, Any], - domain_configs: Mapping[str, Any] | None = None, - *, - run_binding_digest: str, - artifact_header: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - domain_configs = domain_configs or {} - rows: list[dict[str, Any]] = [] - for domain_id in sorted(domain_configs): - config = domain_configs[domain_id] - adapter_id = "S2A-DOMAIN-CONFIG-V2" if domain_id in V2_DOMAIN_IDS else "S2A-DOMAIN-CONFIG-V1" - if not isinstance(config, dict): - continue - for slot_kind, key in ( - ("ELEMENT", "element_slots"), - ("OPPOSING_FACT", "opposing_fact_slots"), - ("DEFENSE", "defense_map"), - ): - values = config.get(key, []) - if isinstance(values, dict): - values = [{"slot_id": item_key, "value": item_value} for item_key, item_value in values.items()] - if not isinstance(values, list): - continue - for index, value in enumerate(values): - rows.append( - { - "domain_id": domain_id, - "slot_ref": str(value.get("slot_id", value.get("id", f"{domain_id}:{key}:{index}"))) if isinstance(value, dict) else f"{domain_id}:{key}:{index}", - "slot_kind": slot_kind, - "fact_ids": [], - "evidence_refs": [], - "evaluation_status": "UNEVALUABLE", - "proposed_new_slot": False, - "source_refs": [_source_ref(f"domain_config:{domain_id}", f"/{key}/{index}", value)], - "_adapter_id": adapter_id, - } - ) - adapter_ids = sorted({row.pop("_adapter_id") for row in rows}) - return { - "artifact_header": dict(artifact_header or _artifact_header()), - "domain_config_adapter_ids": adapter_ids, - "rows": sorted(rows, key=lambda row: (row["domain_id"], row["slot_kind"], row["slot_ref"])), - "rebuttal_slot_synthesis_count": 0, - "source_refs": _dedupe_source_refs( - [_source_ref("fact_ledger_base", "", documents.get("fact_ledger_base"))] - + [_source_ref(f"domain_config:{key}", "", domain_configs[key]) for key in sorted(domain_configs)] - ), - "derivation": _derivation("slot_crosswalk_derivation", [row["slot_ref"] for row in rows], run_binding_digest), - } - - - def _tarjan_scc(nodes: Sequence[str], edges: Sequence[tuple[str, str]]) -> list[list[str]]: - adjacency: dict[str, list[str]] = {node: [] for node in nodes} - for source, target in edges: - adjacency.setdefault(source, []).append(target) - adjacency.setdefault(target, []) - for value in adjacency.values(): - value.sort() - index = 0 - stack: list[str] = [] - on_stack: set[str] = set() - indices: dict[str, int] = {} - lowlink: dict[str, int] = {} - components: list[list[str]] = [] - - def visit(node: str) -> None: - nonlocal index - indices[node] = index - lowlink[node] = index - index += 1 - stack.append(node) - on_stack.add(node) - for neighbor in adjacency[node]: - if neighbor not in indices: - visit(neighbor) - lowlink[node] = min(lowlink[node], lowlink[neighbor]) - elif neighbor in on_stack: - lowlink[node] = min(lowlink[node], indices[neighbor]) - if lowlink[node] == indices[node]: - component: list[str] = [] - while True: - member = stack.pop() - on_stack.remove(member) - component.append(member) - if member == node: - break - components.append(sorted(component)) - - for node in sorted(adjacency): - if node not in indices: - visit(node) - return sorted(components, key=lambda component: component[0]) - - - def compile_cluster_plan( - members: Sequence[Mapping[str, Any]], - relations: Sequence[Mapping[str, Any]], - *, - algorithm_version: str = ALGORITHM_VERSION, - artifact_header: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - """Compile claim-neutral clusters using only explicit hard relations.""" - - by_id = {str(row["member_id"]): dict(row) for row in members} - if len(by_id) != len(members): - raise IngressError("DUPLICATE_CLUSTER_MEMBER_ID", "cluster member IDs must be unique") - parent = {member_id: member_id for member_id in by_id} - - def find(value: str) -> str: - while parent[value] != value: - parent[value] = parent[parent[value]] - value = parent[value] - return value - - def union(left: str, right: str) -> None: - a, b = find(left), find(right) - if a != b: - parent[max(a, b)] = min(a, b) - - normalized_relations: list[dict[str, Any]] = [] - relation_key_map = { - "SAME_BO_ID": ("EXPLICIT_SOURCE_RELATION", "SAME_BO_ID", True), - "SOURCE_BO_ATTACHMENT": ("EXPLICIT_SOURCE_RELATION", "LES_SOURCE_BO_ID", True), - "SAME_EVIDENCE_REF": ("EXPLICIT_SOURCE_RELATION", "SAME_EVIDENCE_REF", True), - "SAME_EVENT_REF": ("EXPLICIT_SOURCE_RELATION", "SAME_EVENT_REF", True), - "EXPLICIT_CASE_RELATION": ("EXPLICIT_SOURCE_RELATION", "APPROVED_EXPLICIT_CASE_RELATION", True), - "claim_precondition": ("CANDIDATE_RELATION", "CLAIM_PRECONDITION_CANDIDATE", False), - "accessory_of": ("CANDIDATE_RELATION", "ACCESSORY_OF_CANDIDATE", False), - "incompatible_with": ("CANDIDATE_RELATION", "INCOMPATIBLE_WITH_CANDIDATE", False), - "EXPLICIT_DEPENDENCY": ("CANDIDATE_RELATION", "CLAIM_PRECONDITION_CANDIDATE", False), - } - for relation in relations: - source = str(relation.get("source_member_id")) - target = str(relation.get("target_member_id")) - kind = str(relation.get("relation_kind")) - if source not in by_id or target not in by_id: - raise IngressError("DANGLING_CLUSTER_RELATION", f"relation references unknown member: {source}->{target}") - explicit_source_refs = relation.get("source_refs") - if not isinstance(explicit_source_refs, list) or not explicit_source_refs: - raise IngressError("RELATION_SOURCE_REF_MISSING", "every cluster relation needs explicit source refs") - edge_class, relation_key, hard_join_allowed = relation_key_map.get( - kind, - ("CANDIDATE_RELATION", "CLAIM_PRECONDITION_CANDIDATE", False), - ) - relation_source_refs = _dedupe_source_refs( - _coerce_source_ref(item) for item in explicit_source_refs - ) - edge_mint = mint_stage2_id( - "cluster_edge", - [source, target, kind, relation_source_refs], - prefix="EDGE", - ) - normalized = { - "source_member_id": source, - "target_member_id": target, - "source_relation_kind": kind, - "edge_id": edge_mint["id"], - "edge_class": edge_class, - "relation_key": relation_key, - "hard_join_allowed": hard_join_allowed, - "source_refs": relation_source_refs, - } - normalized_relations.append(normalized) - if hard_join_allowed and kind in HARD_RELATION_KINDS: - union(source, target) - groups: dict[str, list[str]] = defaultdict(list) - for member_id in sorted(by_id): - groups[find(member_id)].append(member_id) - collision_registry: dict[str, str] = {} - cluster_build_rows: list[dict[str, Any]] = [] - member_to_cluster: dict[str, str] = {} - for member_ids in sorted((sorted(value) for value in groups.values()), key=lambda value: value[0]): - source_refs = _dedupe_source_refs( - _coerce_source_ref(ref) - for member_id in member_ids - for ref in by_id[member_id].get("source_refs", []) - ) - relation_keys = sorted( - canonical_digest(relation) - for relation in normalized_relations - if relation["source_member_id"] in member_ids and relation["target_member_id"] in member_ids - ) - minted = mint_stage2_id( - "cluster", - [sorted(member_ids), source_refs, relation_keys], - prefix="CL", - algorithm_version=algorithm_version, - collision_registry=collision_registry, - ) - executable = bool(source_refs) and all( - by_id[member_id].get("scope_technical_disposition", "AVAILABLE") != "UNAVAILABLE" - and not by_id[member_id].get("residual_review_only", False) - for member_id in member_ids - ) - cluster = { - "cluster_id": minted["id"], - "mint_input_sha256": minted["mint_input_sha256"], - "member_ids": member_ids, - "source_refs": source_refs, - "cluster_status": "EXECUTABLE" if executable else "RESIDUAL_REVIEW_ONLY", - } - cluster_build_rows.append(cluster) - for member_id in member_ids: - member_to_cluster[member_id] = minted["id"] - cluster_edges: list[dict[str, Any]] = [] - edge_pairs: set[tuple[str, str]] = set() - for relation in normalized_relations: - source_cluster = member_to_cluster[relation["source_member_id"]] - target_cluster = member_to_cluster[relation["target_member_id"]] - if source_cluster == target_cluster: - continue - if relation["source_relation_kind"] not in CANDIDATE_RELATION_KINDS: - continue - edge = { - "source_cluster_id": source_cluster, - "target_cluster_id": target_cluster, - "edge_id": relation["edge_id"], - "edge_class": relation["edge_class"], - "relation_key": relation["relation_key"], - "hard_join_allowed": False, - "source_refs": relation["source_refs"], - } - cluster_edges.append(edge) - edge_pairs.add((source_cluster, target_cluster)) - cluster_ids = [row["cluster_id"] for row in cluster_build_rows] - sccs = _tarjan_scc(cluster_ids, sorted(edge_pairs)) - scc_rows: list[dict[str, Any]] = [] - cluster_to_scc: dict[str, str] = {} - for members_in_scc in sccs: - minted = mint_stage2_id("cluster_scc", members_in_scc, prefix="SCC") - cycle = len(members_in_scc) > 1 or any(left == right == members_in_scc[0] for left, right in edge_pairs) - row = { - "scc_id": minted["id"], - "member_cluster_ids": members_in_scc, - "cycle_preserved": cycle, - "mint_input_sha256": minted["mint_input_sha256"], - } - scc_rows.append(row) - for cluster_id in members_in_scc: - cluster_to_scc[cluster_id] = minted["id"] - scheduling_edges = sorted( - { - (cluster_to_scc[left], cluster_to_scc[right]) - for left, right in edge_pairs - if cluster_to_scc[left] != cluster_to_scc[right] - } - ) - scc_nodes = sorted(row["scc_id"] for row in scc_rows) - indegree = {node: 0 for node in scc_nodes} - adjacency: dict[str, set[str]] = {node: set() for node in scc_nodes} - for source, target in scheduling_edges: - if target not in adjacency[source]: - adjacency[source].add(target) - indegree[target] += 1 - scheduling_waves: list[list[str]] = [] - ready = sorted(node for node, degree in indegree.items() if degree == 0) - visited: set[str] = set() - while ready: - scheduling_waves.append(ready) - next_ready: list[str] = [] - for node in ready: - visited.add(node) - for target in sorted(adjacency[node]): - indegree[target] -= 1 - if indegree[target] == 0: - next_ready.append(target) - ready = sorted(set(next_ready)) - if len(visited) != len(scc_nodes): - raise IngressError("SCC_SCHEDULING_DAG_INVALID", "SCC condensation graph must be acyclic") - executable_ids = sorted( - row["cluster_id"] for row in cluster_build_rows if row["cluster_status"] == "EXECUTABLE" - ) - residual_ids = sorted(set(cluster_ids) - set(executable_ids)) - clusters: list[dict[str, Any]] = [] - for build_row in cluster_build_rows: - cluster_id = build_row["cluster_id"] - member_ids = build_row["member_ids"] - cluster_members = [ - { - "member_ref": str(by_id[member_id].get("member_ref", member_id)), - "member_kind": str(by_id[member_id].get("member_kind", "FACT")), - "source_refs": _dedupe_source_refs( - _coerce_source_ref(ref) for ref in by_id[member_id].get("source_refs", []) - ), - } - for member_id in member_ids - ] - internal_edges: list[dict[str, Any]] = [] - for relation in normalized_relations: - if relation["source_member_id"] in member_ids and relation["target_member_id"] in member_ids: - internal_edges.append( - { - "edge_id": relation["edge_id"], - "from_ref": relation["source_member_id"], - "to_ref": relation["target_member_id"], - "edge_class": relation["edge_class"], - "relation_key": relation["relation_key"], - "hard_join_allowed": relation["hard_join_allowed"], - "source_refs": relation["source_refs"], - } - ) - for edge in cluster_edges: - if cluster_id in {edge["source_cluster_id"], edge["target_cluster_id"]}: - internal_edges.append( - { - "edge_id": edge["edge_id"], - "from_ref": edge["source_cluster_id"], - "to_ref": edge["target_cluster_id"], - "edge_class": edge["edge_class"], - "relation_key": edge["relation_key"], - "hard_join_allowed": False, - "source_refs": edge["source_refs"], - } - ) - cluster = { - "cluster_id": cluster_id, - "cluster_status": build_row["cluster_status"], - "members": sorted(cluster_members, key=lambda row: row["member_ref"]), - "edges": sorted(internal_edges, key=lambda row: row["edge_id"]), - "scc_ids": [cluster_to_scc[cluster_id]], - "slice_ref": f"cluster_slices/{cluster_id}.json" if cluster_id in executable_ids else None, - "bundle_cohort_id": None, - "source_refs": build_row["source_refs"], - "derivation": { - "derivation_id": cluster_id, - "algorithm_version": algorithm_version, - "sorted_input_refs": sorted(member_ids), - "mint_input_sha256": build_row["mint_input_sha256"], - }, - } - clusters.append(cluster) - top_source_refs = _dedupe_source_refs( - ref for row in clusters for ref in row["source_refs"] - ) if clusters else [] - return { - "artifact_header": dict(artifact_header or _artifact_header()), - "clusters": sorted(clusters, key=lambda row: row["cluster_id"]), - "executable_cluster_ids": executable_ids, - "residual_review_cluster_ids": residual_ids, - "scheduling_waves": scheduling_waves, - "source_refs": top_source_refs, - "derivation": _derivation( - "cluster_plan_derivation", - cluster_ids or ["NO_CLUSTER"], - [scheduling_edges, [[row["scc_id"], row["member_cluster_ids"]] for row in scc_rows]], - ), - } - - - def compile_cluster_slices( - cluster_plan: Mapping[str, Any], - context_artifacts: Mapping[str, Mapping[str, Any]], - ) -> dict[str, dict[str, Any]]: - """Create immutable minimal source-ref slices for every cluster.""" - - projection_sources: dict[str, dict[str, Any]] = defaultdict(dict) - case_context = context_artifacts.get("case_context", {}) - for row in case_context.get("facts", []) if isinstance(case_context, dict) else []: - if isinstance(row, dict) and row.get("fact_id") is not None: - projection_sources["FACT"][str(row["fact_id"])] = row - evidence_inventory = context_artifacts.get("evidence_inventory", {}) - for row in evidence_inventory.get("items", []) if isinstance(evidence_inventory, dict) else []: - if not isinstance(row, dict): - continue - identifier = row.get("evidence_id", row.get("id")) - if identifier is not None: - projection_sources["EVIDENCE"][str(identifier)] = row - supplied_sources = context_artifacts.get("_projection_sources", {}) - if isinstance(supplied_sources, dict): - for kind, rows in supplied_sources.items(): - if isinstance(rows, dict): - projection_sources[str(kind)].update({str(key): value for key, value in rows.items()}) - - def materialize_projection( - kind: str, - source_id: str, - ) -> dict[str, Any]: - content = projection_sources.get(kind, {}).get(source_id) - if content is None: - raise IngressError( - "SLICE_PROJECTION_SOURCE_MISSING", - f"projection source is missing: {kind}:{source_id}", - ) - if not isinstance(content, dict): - raise IngressError( - "SLICE_PROJECTION_SOURCE_MISSING", - f"projection source is not an object: {kind}:{source_id}", - ) - raw = canonical_json_bytes(content) - if len(raw) > 262144: - raise IngressError("SLICE_PROJECTION_BUDGET_EXCEEDED", f"projection exceeds 262144 bytes: {kind}:{source_id}") - source_refs = _dedupe_source_refs(content.get("source_refs", [])) - matching_refs = [ - ref - for ref in source_refs - if ref.get("stage1_id") == source_id - and isinstance(ref.get("raw_value_sha256"), str) - and re.fullmatch(r"[a-f0-9]{64}", str(ref["raw_value_sha256"])) - ] - matching_hashes = {str(ref["raw_value_sha256"]) for ref in matching_refs} - if not matching_refs or len(matching_hashes) != 1: - raise IngressError( - "SLICE_PROJECTION_PROVENANCE_MISSING", - f"projection lacks one unambiguous source_id-bound raw hash: {kind}:{source_id}", - ) - raw_source_sha256 = next(iter(matching_hashes)) - projection_mint = mint_stage2_id("content_projection", [kind, source_id, content], prefix="CP") - return { - "projection_id": projection_mint["id"], - "projection_kind": kind, - "source_id": source_id, - "content_schema_id": None, - "content": content, - "raw_source_sha256": raw_source_sha256, - "canonical_content_sha256": canonical_digest(content), - "materialized_utf8_bytes": len(raw), - "projection_policy_id": "S2-SLICE-BOUNDED-PROJECTION-V1", - "truncated": False, - "source_refs": source_refs, - } - - projection_arrays = { - "FACT": "fact_projections", - "EVIDENCE": "evidence_projections", - "EVENT": "event_projections", - "LES_STRUCTURE": "les_structure_projections", - "SIGNAL_OCCURRENCE": "signal_occurrence_projections", - "REVIEW_ITEM": "review_item_projections", - "ACTIVE_PROFILE": "profile_projections", - } - slices: dict[str, dict[str, Any]] = {} - for cluster in cluster_plan.get("clusters", []): - cluster_id = cluster["cluster_id"] - source_refs = _dedupe_source_refs(cluster.get("source_refs", [])) - member_refs_by_kind: dict[str, list[str]] = defaultdict(list) - for member in cluster.get("members", []): - member_refs_by_kind[str(member.get("member_kind"))].append(str(member.get("member_ref"))) - minted = mint_stage2_id( - "cluster_slice", - [cluster_id, source_refs, sorted(member_refs_by_kind.items())], - prefix="SL", - ) - slice_body: dict[str, Any] = { - "artifact_header": dict(cluster_plan.get("artifact_header", _artifact_header())), - "cluster_id": cluster_id, - "slice_id": minted["id"], - "slice_schema_version": "stage2_cluster_slice.v1", - "immutable": True, - "source_refs": source_refs, - "fact_refs": sorted(set(member_refs_by_kind.get("FACT", []))), - "evidence_refs": sorted(set(member_refs_by_kind.get("EVIDENCE", []))), - "event_refs": sorted(set(member_refs_by_kind.get("EVENT", []))), - "les_structure_refs": sorted(set(member_refs_by_kind.get("LES_STRUCTURE", []))), - "signal_occurrence_refs": sorted(set(member_refs_by_kind.get("SIGNAL_OCCURRENCE", []))), - "review_keys": sorted(set(member_refs_by_kind.get("REVIEW_ITEM", []))), - "profile_refs": [], - "derivation": { - "derivation_id": minted["id"], - "algorithm_version": ALGORITHM_VERSION, - "sorted_input_refs": sorted( - str(member.get("member_ref")) for member in cluster.get("members", []) - ) or [cluster_id], - "mint_input_sha256": minted["mint_input_sha256"], - }, - } - all_projections: list[dict[str, Any]] = [] - for kind, array_name in projection_arrays.items(): - source_ids = sorted(set(member_refs_by_kind.get(kind, []))) - projections = [materialize_projection(kind, source_id) for source_id in source_ids] - slice_body[array_name] = projections - all_projections.extend(projections) - observed_projection_bytes = sum(row["materialized_utf8_bytes"] for row in all_projections) - if observed_projection_bytes > 2097152: - raise IngressError("SLICE_CONTENT_BUDGET_EXCEEDED", f"slice exceeds 2097152 bytes: {cluster_id}") - slice_body["projection_budget"] = { - "policy_id": "S2-SLICE-BOUNDED-PROJECTION-V1", - "max_projection_utf8_bytes": 262144, - "max_slice_content_utf8_bytes": 2097152, - "observed_projection_count": len(all_projections), - "observed_slice_content_utf8_bytes": observed_projection_bytes, - "budget_status": "PASS", - "raw_stage1_reread_required": False, - } - slice_body["slice_digest"] = canonical_digest(slice_body) - slices[cluster_id] = slice_body - return slices - - - _PII_PATTERNS: tuple[tuple[str, re.Pattern[str]], ...] = ( - ("KOREAN_RESIDENT_ID", re.compile(r"\b\d{6}-?[1-4]\d{6}\b")), - ("EMAIL", re.compile(r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b")), - ("KOREAN_PHONE", re.compile(r"\b01[016789]-?\d{3,4}-?\d{4}\b")), + after = os.fstat(descriptor) + finally: + os.close(descriptor) + try: + path_after = candidate.stat(follow_symlinks=False) + except FileNotFoundError as exc: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", "source disappeared after snapshot") from exc + identity_before = (before.st_dev, before.st_ino, before.st_size, before.st_mtime_ns) + identity_after = (after.st_dev, after.st_ino, after.st_size, after.st_mtime_ns) + path_identity = (path_after.st_dev, path_after.st_ino, path_after.st_size, path_after.st_mtime_ns) + if identity_before != identity_after or identity_after != path_identity: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source changed during snapshot: {relative_path}") + raw = b"".join(chunks) + return Snapshot( + logical_input_id=logical_input_id, + relative_path=logical.as_posix(), + resolved_path=str(resolved), + raw=raw, + raw_sha256=hashlib.sha256(raw).hexdigest(), + byte_length=len(raw), + device=after.st_dev, + inode=after.st_ino, + mtime_ns=after.st_mtime_ns, ) - def _scan_pii(value: Any) -> list[str]: - rendered = canonical_json_bytes(value).decode("utf-8") - return [name for name, pattern in _PII_PATTERNS if pattern.search(rendered)] + def resolve_stage1_sources( + stage1_run_root: str | os.PathLike[str], + contract_manifest: Mapping[str, Any] | None = None, + ) -> list[dict[str, Any]]: + """Resolve only approved logical kinds; a relocation manifest cannot invent kinds.""" - - def compile_bundle_plan( - cluster_plan: Mapping[str, Any], - cluster_slices: Mapping[str, Mapping[str, Any]], - release_lock: Mapping[str, Any], - *, - stage2_asset_root: str | os.PathLike[str] | None = None, - ) -> dict[str, Any]: - """Compile a data-only S2_10 handoff without materializing prompts.""" - - bundle = release_lock.get("bundle") - if not isinstance(bundle, dict): - raise IngressError("BUNDLE_RELEASE_CONTRACT_MISSING", "release bundle must be an object") - mode = bundle.get("mode") - release_class = release_lock.get("release_class") - if mode not in RELEASE_MODE or RELEASE_MODE[mode] != release_class: - raise IngressError("BUNDLE_RELEASE_CLASS_MISMATCH", f"{mode} is not valid for {release_class}") - handoff_contract_version = bundle.get("handoff_contract_version") - if handoff_contract_version != "S2_00_STRUCTURED_CONTEXT_HANDOFF_V2": - raise IngressError( - "HANDOFF_CONTRACT_VERSION_MISMATCH", - "S2_00 requires S2_00_STRUCTURED_CONTEXT_HANDOFF_V2", - ) - release_ref = str( - release_lock.get( - "release_id", - release_lock.get("stage2_release_digest", "stage2_release"), - ) + root = Path(stage1_run_root).resolve(strict=True) + if not root.is_dir(): + raise IngressError("STAGE1_ROOT_NOT_DIRECTORY", "Stage 1 run root must be a directory") + contracts = [dict(row) for row in DEFAULT_SOURCE_CONTRACTS] + overrides = dict((contract_manifest or {}).get("path_overrides", {})) + approved_ids = {row["logical_input_id"] for row in contracts} + invented = sorted(set(overrides) - approved_ids) + if invented: + raise IngressError("UNAPPROVED_LOGICAL_KIND", "relocation manifest invented logical kinds", details={"ids": invented}) + seen_paths: set[str] = set() + for row in contracts: + path = overrides.get(row["logical_input_id"], row["path"]) + safe = _safe_relative_path(path).as_posix() + if safe in seen_paths: + raise IngressError("DUPLICATE_LOGICAL_MAPPING", f"duplicate physical mapping: {safe}") + seen_paths.add(safe) + row["expected_path"] = row.pop("path") + row["observed_path"] = safe + row["resolution_source"] = ( + "RELEASE_BOUND_CONTRACT_MANIFEST" + if row["logical_input_id"] in overrides + else "DEFAULT_EXACT_PATH" ) - verification_errors: list[str] = [] + return contracts - def normalize_sha256(value: Any, *, code: str) -> str: - if not isinstance(value, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", value) is None: - verification_errors.append(code) - return "0" * 64 - return value.lower() - def verify_opaque_ref(value: Any, *, label: str) -> str: - if not isinstance(value, dict): - verification_errors.append(f"{label}_REF_MISSING") - return "0" * 64 - try: - path = _safe_relative_path(value.get("path")).as_posix() - except (IngressError, TypeError): - verification_errors.append(f"{label}_REF_PATH_INVALID") - return "0" * 64 - expected = normalize_sha256( - value.get("sha256"), - code=f"{label}_HASH_UNBOUND", + def _issue( + code: str, + *, + impact_scope: str = "GLOBAL", + source_refs: Sequence[str] = (), + severity: str = "ERROR", + message: str | None = None, + ) -> dict[str, Any]: + return { + "issue_code": code, + "severity": severity, + "impact_scope": impact_scope, + "scope_refs": sorted(set(source_refs)), + "source_contract_row_refs": sorted(set(source_refs)), + "reason_codes": [code], + "downstream_allowed_actions": [], + "message": message or code, + } + + + def _shape_required(value: Any, keys: Sequence[str]) -> list[str]: + if not isinstance(value, dict): + return list(keys) + return [key for key in keys if key not in value] + + + def _json_pointer_value(document: Any, pointer: str | None) -> tuple[bool, Any]: + if pointer in {None, ""}: + return (pointer == "", document) + if not isinstance(pointer, str) or not pointer.startswith("/"): + return False, None + current = document + for raw_token in pointer[1:].split("/"): + token = raw_token.replace("~1", "/").replace("~0", "~") + if isinstance(current, dict) and token in current: + current = current[token] + elif isinstance(current, list) and token.isdigit() and int(token) < len(current): + current = current[int(token)] + else: + return False, None + return True, current + + + def _release_stage1_source_rows(release_lock: Mapping[str, Any]) -> list[Mapping[str, Any]]: + rows = release_lock.get("stage1_sources") + if not isinstance(rows, list): + dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) + rows = dependency.get("stage1_sources") if isinstance(dependency, dict) else None + return [row for row in rows if isinstance(row, dict)] if isinstance(rows, list) else [] + + + def _adapter_decision(release_lock: Mapping[str, Any], adapter_id: str) -> Mapping[str, Any] | None: + for row in release_lock.get("adapter_decisions", []): + if isinstance(row, dict) and row.get("adapter_id") == adapter_id and isinstance(row.get("decision"), dict): + return row["decision"] + return None + + + def _closed_adapter_shape_errors( + document: Any, + *, + logical_id: str, + adapter_id: str, + required_keys: Sequence[str], + release_lock: Mapping[str, Any], + ) -> list[str]: + errors: list[str] = [] + if required_keys: + errors.extend(f"missing root key {key}" for key in _shape_required(document, required_keys)) + decision = _adapter_decision(release_lock, adapter_id) + if decision is not None: + root_shape = decision.get("root_shape") + if root_shape == "ARRAY" and not isinstance(document, list): + errors.append("root must be an array") + elif root_shape == "OBJECT_ENVELOPE" and not isinstance(document, dict): + errors.append("root must be an object envelope") + if isinstance(document, dict): + errors.extend( + f"missing root key {key}" + for key in _shape_required(document, decision.get("required_root_fields", [])) ) - if stage2_asset_root is None: - verification_errors.append("STAGE2_ASSET_ROOT_MISSING") - return expected - try: - snapshot = open_bounded_snapshot( - stage2_asset_root, - path, - logical_input_id=label.lower(), - ) - except IngressError as exc: - verification_errors.append(f"{label}_ASSET_UNAVAILABLE") - return expected - if snapshot.raw_sha256 != expected: - verification_errors.append(f"{label}_HASH_MISMATCH") - return expected - - s2_10_agent_sha256 = verify_opaque_ref( - bundle.get("s2_10_agent_ref"), - label="S2_10_AGENT", - ) - s2_10_llm_binding_sha256 = verify_opaque_ref( - bundle.get("s2_10_llm_binding_ref"), - label="S2_10_LLM_BINDING", - ) - if normalize_sha256( - bundle.get("s2_10_agent_sha256"), - code="S2_10_AGENT_HASH_UNBOUND", - ) != s2_10_agent_sha256: - verification_errors.append("S2_10_AGENT_RELEASE_HASH_MISMATCH") - if normalize_sha256( - bundle.get("s2_10_llm_binding_sha256"), - code="S2_10_LLM_BINDING_HASH_UNBOUND", - ) != s2_10_llm_binding_sha256: - verification_errors.append("S2_10_LLM_BINDING_RELEASE_HASH_MISMATCH") - if bundle.get("context_cohort_policy_id") != "S2-CACHE-STRUCTURED-CONTEXT-HANDOFF-V2": - verification_errors.append("CONTEXT_COHORT_POLICY_MISMATCH") - if bundle.get("downstream_hybrid_status") != "HYBRID_RELEASE_BOUND": - verification_errors.append("S2_10_HYBRID_RELEASE_UNBOUND") - - selected_context_input = bundle.get("selected_context_refs", []) - if not isinstance(selected_context_input, list): - raise IngressError( - "SELECTED_CONTEXT_REFS_SHAPE", - "selected_context_refs must be an array", - ) - allowed_context_kinds = { - "ACTIVE_PROFILE", - "APPROVED_COMMON_AUTHORITY", - "LAW_VALUE_TOKEN", - "AUTHORITY_PROPOSITION", + if isinstance(document, list): + required_item_fields = decision.get("required_item_fields", decision.get("required_row_fields", [])) + if isinstance(required_item_fields, list): + for index, item in enumerate(document): + for key in _shape_required(item, required_item_fields): + errors.append(f"row {index} missing {key}") + if decision is None: + fallback_required: dict[str, tuple[str, ...]] = { + "evidence_indexed": ("schema_contract_version", "items"), + "evidence_event_candidates": ("schema_version", "items"), + "domain_activation_manifest": SG01_PROJECTION_FIELDS, + "signal_manifest": ("downstream_read_sets", "files"), + "legal_effect_structures": ("schema_version", "structure_records"), + "fact_ledger_writer_report": ( + "schema_version", + "row_count", + "gate_firings", + "domain_effect_coverage", + "calculation_readiness", + "blocked_review_items", + "conservation", + "final_sha256", + ), } - selected_context_refs: list[dict[str, Any]] = [] - seen_context_ids: set[str] = set() - for index, row in enumerate(selected_context_input): + fallback = fallback_required.get(logical_id, ()) + if fallback: + errors.extend(f"missing root key {key}" for key in _shape_required(document, fallback)) + if logical_id in {"bo", "fact_ledger_base"} and not isinstance(document, list): + errors.append("root must be an array") + return sorted(set(errors)) + + + def _schema_document_index(deployment_documents: Mapping[str, Any]) -> dict[str, Mapping[str, Any]]: + result: dict[str, Mapping[str, Any]] = {} + for path, document in deployment_documents.items(): + if not isinstance(document, dict): + continue + result[path] = document + result[PurePosixPath(path).name] = document + schema_id = document.get("$id") + if isinstance(schema_id, str): + result[schema_id] = document + return result + + + def _source_hash_index(document: Mapping[str, Any] | None) -> dict[str, str]: + result: dict[str, str] = {} + if not isinstance(document, dict): + return result + candidate_arrays: list[Any] = [] + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(document.get(key), list): + candidate_arrays.append(document[key]) + for wrapper in ("completion_seal", "manifest", "payload", "data"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(nested.get(key), list): + candidate_arrays.append(nested[key]) + for rows in candidate_arrays: + for row in rows: if not isinstance(row, dict): - raise IngressError( - "SELECTED_CONTEXT_REF_SHAPE", - f"selected context row {index} must be an object", - ) - context_ref_id = row.get("context_ref_id") - context_kind = row.get("context_kind") - if not isinstance(context_ref_id, str) or not context_ref_id: - raise IngressError( - "SELECTED_CONTEXT_REF_ID_INVALID", - f"selected context row {index} lacks context_ref_id", - ) - if context_ref_id in seen_context_ids: - verification_errors.append("SELECTED_CONTEXT_REF_ID_DUPLICATE") - seen_context_ids.add(context_ref_id) - if context_kind not in allowed_context_kinds: - raise IngressError( - "SELECTED_CONTEXT_KIND_INVALID", - f"unsupported context kind: {context_kind}", - ) - try: - relative_path = _safe_relative_path(row.get("path")).as_posix() - except (IngressError, TypeError) as exc: - raise IngressError( - "SELECTED_CONTEXT_PATH_INVALID", - f"invalid selected context path at row {index}", - ) from exc - raw_sha256 = normalize_sha256( - row.get("raw_sha256"), - code="SELECTED_CONTEXT_RAW_HASH_UNBOUND", + continue + digest = row.get("raw_sha256", row.get("sha256")) + if not isinstance(digest, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", digest) is None: + continue + for key in ("logical_input_id", "path", "observed_path", "logical_id"): + identifier = row.get(key) + if isinstance(identifier, str) and identifier: + result[identifier] = digest.lower() + return result + + + def _source_producer_index(document: Mapping[str, Any] | None) -> dict[str, str]: + """Index producer evidence carried by a bounded completion/manifest row.""" + + result: dict[str, str] = {} + if not isinstance(document, dict): + return result + candidate_arrays: list[Any] = [] + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(document.get(key), list): + candidate_arrays.append(document[key]) + for wrapper in ("completion_seal", "manifest", "payload", "data"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(nested.get(key), list): + candidate_arrays.append(nested[key]) + for rows in candidate_arrays: + for row in rows: + if not isinstance(row, dict): + continue + producer = next( + ( + row.get(key) + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by") + if isinstance(row.get(key), str) and row.get(key) + ), + None, ) - canonical_sha256 = normalize_sha256( - row.get("canonical_sha256"), - code="SELECTED_CONTEXT_CANONICAL_HASH_UNBOUND", + if not isinstance(producer, str): + continue + for key in ("logical_input_id", "path", "observed_path", "logical_id"): + identifier = row.get(key) + if isinstance(identifier, str) and identifier: + result[identifier] = producer + return result + + + def _producer_value(document: Any) -> str | None: + if not isinstance(document, dict): + return None + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): + value = document.get(key) + if isinstance(value, str) and value: + return value + for wrapper in ("metadata", "meta", "handoff", "payload"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): + value = nested.get(key) + if isinstance(value, str) and value: + return value + # P3/P4 are closed one-key wrappers in the Stage 1 v8 handoff contract. + for wrapper in ( + "stage1_part3_review_handoff", + "stage1_part4_review_handoff", + ): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("created_by", "finalized_by"): + value = nested.get(key) + if isinstance(value, str) and value: + return value + return None + + + def _producer_matches( + observed: str, + expected: str, + alias_id: str | None, + release_lock: Mapping[str, Any], + ) -> bool: + if observed == expected: + return True + if alias_id is None: + return False + decision = _adapter_decision(release_lock, alias_id) + if decision is None or decision.get("bidirectional_match_allowed") is not True: + return False + pair = {decision.get("schema_writer_id"), decision.get("orchestration_producer_id")} + return {observed, expected} == pair + + + def _identity_ref(document: Any, pointer: str | None, logical_id: str) -> dict[str, Any]: + if pointer is None: + return {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": logical_id} + found, value = _json_pointer_value(document, pointer) + if not found or value is None: + return {"value": None, "disposition": "MISSING", "source_ref": f"{logical_id}#{pointer}"} + return {"value": str(value), "disposition": "OBSERVED", "source_ref": f"{logical_id}#{pointer}"} + + + def validate_ingress_contracts( + snapshots: Mapping[str, Snapshot], + contracts: Sequence[Mapping[str, Any]], + release_lock: Mapping[str, Any], + *, + deployment_snapshots: Mapping[str, Snapshot] | None = None, + deployment_documents: Mapping[str, Any] | None = None, + completion_seal: Mapping[str, Any] | None = None, + contract_manifest: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Strictly parse sources and verify release-bound schema, producer, identity, and seal rows.""" + + documents: dict[str, Any] = {} + rows: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + deployment_snapshots = deployment_snapshots or {} + deployment_documents = deployment_documents or {} + deployment_by_path = {snapshot.relative_path: snapshot for snapshot in deployment_snapshots.values()} + schema_documents = _schema_document_index(deployment_documents) + release_source_rows = _release_stage1_source_rows(release_lock) + release_ids = [str(row.get("logical_input_id")) for row in release_source_rows] + duplicate_release_ids = sorted(key for key, count in Counter(release_ids).items() if count > 1) + if duplicate_release_ids: + raise IngressError( + "RELEASE_SOURCE_CONTRACT_DUPLICATE", + "release stage1_sources contains duplicate logical_input_id rows", + details={"logical_input_ids": duplicate_release_ids}, + ) + expected_fixed = { + str(row["logical_input_id"]): str(row["path"]) + for row in DEFAULT_SOURCE_CONTRACTS + } + expected_release_ids = set(expected_fixed) | {"signal_payload_family"} + observed_release_ids = set(release_ids) + if observed_release_ids != expected_release_ids: + raise IngressError( + "RELEASE_SOURCE_CONTRACT_SET_MISMATCH", + "release stage1_sources must be the exact 16 fixed inputs plus signal_payload_family", + details={ + "missing": sorted(expected_release_ids - observed_release_ids), + "extra": sorted(observed_release_ids - expected_release_ids), + }, + ) + release_rows = {str(row.get("logical_input_id")): row for row in release_source_rows} + for logical_id, expected_path in expected_fixed.items(): + release_row = release_rows[logical_id] + if release_row.get("path") != expected_path or release_row.get("path_rule") not in {None, ""}: + raise IngressError( + "RELEASE_SOURCE_FIXED_PATH_MISMATCH", + f"fixed source path contract mismatch: {logical_id}", ) - if row.get("order_index") != index: - verification_errors.append("SELECTED_CONTEXT_ORDER_INVALID") - observed_raw_sha256: str | None = None - observed_canonical_sha256: str | None = None - pii_codes: list[str] = [] - if stage2_asset_root is None: - verification_errors.append("STAGE2_ASSET_ROOT_MISSING") - else: - try: - snapshot = open_bounded_snapshot( - stage2_asset_root, - relative_path, - logical_input_id=f"selected_context:{context_ref_id}", - ) - observed_raw_sha256 = snapshot.raw_sha256 - try: - text = snapshot.raw.decode("utf-8", errors="strict") - canonical_text = unicodedata.normalize( - "NFC", - text.replace("\r\n", "\n").replace("\r", "\n"), - ) - observed_canonical_sha256 = hashlib.sha256( - canonical_text.encode("utf-8") - ).hexdigest() - pii_codes = _scan_pii(canonical_text) - except UnicodeDecodeError: - verification_errors.append("SELECTED_CONTEXT_NOT_UTF8") - except IngressError as exc: - verification_errors.append( - "SELECTED_CONTEXT_ASSET_UNAVAILABLE" - ) - if observed_raw_sha256 != raw_sha256: - verification_errors.append("SELECTED_CONTEXT_ASSET_HASH_MISMATCH") - if observed_canonical_sha256 != canonical_sha256: - verification_errors.append( - "SELECTED_CONTEXT_CANONICAL_HASH_MISMATCH" - ) - if pii_codes: - verification_errors.append("SELECTED_CONTEXT_PII") - selected_context_refs.append( + signal_family = release_rows["signal_payload_family"] + if ( + signal_family.get("path") is not None + or signal_family.get("path_rule") != "signals/" + or signal_family.get("adapter_id") != "S2A-SIGNAL-ALL-V1" + or signal_family.get("raw_hash_source") != "MANIFEST_ROW" + ): + raise IngressError( + "SIGNAL_PAYLOAD_FAMILY_CONTRACT_MISMATCH", + "signal_payload_family must use the approved manifest-expanded path contract", + ) + completion_hashes = _source_hash_index(completion_seal) + manifest_hashes = _source_hash_index(contract_manifest) + completion_producers = _source_producer_index(completion_seal) + manifest_producers = _source_producer_index(contract_manifest) + for contract in contracts: + logical_id = str(contract["logical_input_id"]) + snapshot = snapshots.get(logical_id) + release_row = release_rows.get(logical_id) + contract_missing = release_row is None + release_row = release_row or {} + alias_value = release_row.get("producer_alias", release_row.get("producer_alias_id")) + alias_id = str(alias_value) if isinstance(alias_value, str) else None + schema_ref = release_row.get("schema_ref") if isinstance(release_row.get("schema_ref"), dict) else None + row = { + "logical_input_id": logical_id, + "requirement_class": REQUIREMENT_CLASS_ENUM.get( + str(contract.get("criticality")), + "INTEGRITY_CORROBORATOR", + ), + "expected_path": contract.get("expected_path"), + "observed_path": contract.get("observed_path"), + "resolution_source": contract.get("resolution_source"), + "schema_id": schema_ref.get("$id") if schema_ref else release_row.get("schema_id"), + "schema_sha256": schema_ref.get("sha256") if schema_ref else release_row.get("schema_sha256"), + "producer_id": release_row.get("producer_id"), + "producer_alias_id": alias_id, + "adapter_id": release_row.get("adapter_id", ADAPTER_IDS.get(logical_id, "S2A-UNBOUND-V1")), + "run_identity_ref": release_row.get("run_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), + "transaction_identity_ref": release_row.get("transaction_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), + "scope_refs": [logical_id], + "source_contract_row_refs": [logical_id], + "reason_codes": [], + "downstream_allowed_actions": [], + "issue_codes": [], + } + if contract_missing: + code = "RELEASE_SOURCE_CONTRACT_MISSING" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + declared_path = release_row.get("path") + if isinstance(declared_path, str) and declared_path != contract.get("expected_path"): + code = "RELEASE_SOURCE_PATH_MISMATCH" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if snapshot is None: + row.update( { - "context_ref_id": context_ref_id, - "context_kind": context_kind, - "path": relative_path, - "raw_sha256": raw_sha256, - "canonical_sha256": canonical_sha256, - "order_index": index, - "release_ref": str(row.get("release_ref", release_ref)), + "raw_sha256": None, + "byte_length": 0, + "parse_status": "NOT_OBSERVED", + "schema_status": "UNEVALUABLE", + "seal_status": "UNEVALUABLE", + "scope_technical_disposition": "UNAVAILABLE", + "impact_scope": "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER", } ) - selected_context_ordered_refs_sha256 = canonical_digest( - selected_context_refs - ) - declared_context_digest = normalize_sha256( - bundle.get("selected_context_ordered_refs_sha256"), - code="SELECTED_CONTEXT_ORDERED_REFS_HASH_UNBOUND", - ) - if declared_context_digest != selected_context_ordered_refs_sha256: - verification_errors.append( - "SELECTED_CONTEXT_ORDERED_REFS_HASH_MISMATCH" - ) - if mode != "STRUCTURAL_FIXTURE" and not selected_context_refs: - verification_errors.append("SELECTED_CONTEXT_EMPTY") - - wave_by_scc: dict[str, int] = {} - for wave_ordinal, wave in enumerate( - cluster_plan.get("scheduling_waves", []) - ): - for scc_id in wave: - wave_by_scc[str(scc_id)] = wave_ordinal - cluster_rows = { - str(row.get("cluster_id")): row - for row in cluster_plan.get("clusters", []) - if isinstance(row, dict) and row.get("cluster_id") is not None - } - - valid_member_slices: list[dict[str, Any]] = [] - missing_member_slices: list[dict[str, Any]] = [] - for cluster_id in cluster_plan.get("executable_cluster_ids", []): - cluster_id = str(cluster_id) - cluster_row = cluster_rows.get(cluster_id, {}) - scc_ids = [ - str(value) - for value in cluster_row.get("scc_ids", []) - ] if isinstance(cluster_row, dict) else [] - dependency_wave_ordinal = min( - (wave_by_scc.get(value, 0) for value in scc_ids), - default=0, - ) - slice_body = cluster_slices.get(cluster_id) - member_slice = { - "cluster_id": cluster_id, - "path": f"context/cluster_slices/{cluster_id}.json", - "sha256": ( - canonical_digest(slice_body) - if isinstance(slice_body, Mapping) - else "0" * 64 - ), - "dependency_wave_ordinal": dependency_wave_ordinal, - } - if isinstance(slice_body, Mapping): - valid_member_slices.append(member_slice) - else: - missing_member_slices.append(member_slice) - - def build_cohort( - member_slices: Sequence[Mapping[str, Any]], - *, - extra_reason_codes: Sequence[str] = (), - ) -> dict[str, Any]: - cohort_member_cluster_ids = sorted( - str(row["cluster_id"]) for row in member_slices - ) - public_member_slices = sorted( - (dict(row) for row in member_slices), - key=lambda row: row["cluster_id"], - ) - cohort_membership_sha256 = canonical_digest( - cohort_member_cluster_ids - ) - member_slice_ref_set_sha256 = canonical_digest( - public_member_slices - ) - receipt_errors = sorted( - set(verification_errors).union(str(code) for code in extra_reason_codes) - ) - cohort_mint = mint_stage2_id( - "bundle_cohort", - [ - handoff_contract_version, - mode, - s2_10_agent_sha256, - s2_10_llm_binding_sha256, - selected_context_ordered_refs_sha256, - cohort_membership_sha256, - member_slice_ref_set_sha256, - ], - prefix="BC", - ) - materialization_input = { - "handoff_contract_version": handoff_contract_version, - "bundle_cohort_id": cohort_mint["id"], - "s2_10_agent_sha256": s2_10_agent_sha256, - "s2_10_llm_binding_sha256": s2_10_llm_binding_sha256, - "selected_context_ordered_refs_sha256": ( - selected_context_ordered_refs_sha256 - ), - "cohort_membership_sha256": cohort_membership_sha256, - "member_slice_ref_set_sha256": member_slice_ref_set_sha256, - } - receipt = { - "schema_version": ( - "stage2_s2_00_context_materialization_receipt.v2" - ), - "bundle_cohort_id": cohort_mint["id"], - "s2_10_agent_sha256": s2_10_agent_sha256, - "s2_10_llm_binding_sha256": s2_10_llm_binding_sha256, - "selected_context_ordered_refs_sha256": ( - selected_context_ordered_refs_sha256 - ), - "cohort_membership_sha256": cohort_membership_sha256, - "member_slice_ref_set_sha256": member_slice_ref_set_sha256, - "materialization_input_sha256": canonical_digest( - materialization_input - ), - "verification_status": ( - "PASS" if not receipt_errors else "NON_EXECUTABLE" - ), - "reason_codes": receipt_errors, - } - reason_codes = list(receipt_errors) - if mode == "STRUCTURAL_FIXTURE": - reason_codes.append("STRUCTURAL_FIXTURE_NOT_EXECUTABLE") - cohort_status = "STRUCTURAL_ONLY" - elif receipt_errors: - cohort_status = "NON_EXECUTABLE" - else: - cohort_status = "EXECUTABLE" - reason_codes = sorted(set(reason_codes)) - source_refs = _dedupe_source_refs( - ref - for cluster_id in cohort_member_cluster_ids - for ref in ( - cluster_rows.get(cluster_id, {}).get("source_refs", []) - if isinstance(cluster_rows.get(cluster_id), dict) - else [] - ) - ) - return { - "handoff_contract_version": handoff_contract_version, - "bundle_cohort_id": cohort_mint["id"], - "compile_mode": mode, - "release_class": release_class, - "cohort_status": cohort_status, - "selected_context_refs": list(selected_context_refs), - "selected_context_ordered_refs_sha256": ( - selected_context_ordered_refs_sha256 - ), - "member_slices": public_member_slices, - "cohort_member_cluster_ids": cohort_member_cluster_ids, - "s2_10_agent_sha256": s2_10_agent_sha256, - "s2_10_llm_binding_sha256": s2_10_llm_binding_sha256, - "context_materialization_receipt": receipt, - "forbidden_bulk_inputs_present": False, - "reason_codes": reason_codes, - "source_refs": source_refs, - "derivation": { - "derivation_id": cohort_mint["id"], - "algorithm_version": ALGORITHM_VERSION, - "sorted_input_refs": sorted( - cohort_member_cluster_ids - + [ - f"context:{selected_context_ordered_refs_sha256}", - f"agent:{s2_10_agent_sha256}", - f"binding:{s2_10_llm_binding_sha256}", - ] - ), - "mint_input_sha256": cohort_mint["mint_input_sha256"], - }, - } - - cohorts: list[dict[str, Any]] = [] - if valid_member_slices: - cohorts.append(build_cohort(valid_member_slices)) - for member_slice in missing_member_slices: - cohorts.append( - build_cohort( - [member_slice], - extra_reason_codes=["CLUSTER_SLICE_MISSING"], - ) - ) - public_cohorts = sorted( - cohorts, - key=lambda row: ( - row["cohort_member_cluster_ids"][0] - if row["cohort_member_cluster_ids"] - else row["bundle_cohort_id"] - ), - ) - executable_cohort_ids = sorted( - row["bundle_cohort_id"] - for row in public_cohorts - if row["cohort_status"] == "EXECUTABLE" - ) - non_executable_cohort_ids = sorted( - row["bundle_cohort_id"] - for row in public_cohorts - if row["cohort_status"] != "EXECUTABLE" - ) - top_refs = ( - _dedupe_source_refs( - ref for row in public_cohorts for ref in row["source_refs"] - ) - if public_cohorts - else list(cluster_plan.get("source_refs", [])) - ) - return { - "artifact_header": dict( - cluster_plan.get("artifact_header", _artifact_header()) - ), - "handoff_contract_version": handoff_contract_version, - "cohorts": public_cohorts, - "executable_bundle_cohort_ids": executable_cohort_ids, - "non_executable_bundle_cohort_ids": non_executable_cohort_ids, - "source_refs": top_refs, - "derivation": _derivation( - "bundle_plan_derivation", - [ - row["bundle_cohort_id"] for row in public_cohorts - ] or ["NO_BUNDLE_COHORT"], - [ - selected_context_ordered_refs_sha256, - s2_10_agent_sha256, - s2_10_llm_binding_sha256, - { - row["bundle_cohort_id"]: row["reason_codes"] - for row in public_cohorts - }, - ], - ), - } - - def validate_bundle_release_cohorts(bundle_plan: Mapping[str, Any]) -> dict[str, Any]: - """Recompute every cross-object invariant the JSON Schema cannot express.""" - - cohorts = bundle_plan.get("cohorts", []) - if not isinstance(cohorts, list): - raise IngressError( - "BUNDLE_COHORT_INVARIANT_FAILED", - "bundle_plan.cohorts must be an array", - ) - mode = cohorts[0].get("compile_mode") if cohorts else None - release_class = cohorts[0].get("release_class") if cohorts else None - invariant_errors: list[str] = [] - receipt_reason_codes = { - "S2_10_AGENT_REF_MISSING", - "S2_10_AGENT_REF_PATH_INVALID", - "S2_10_AGENT_HASH_UNBOUND", - "S2_10_AGENT_ASSET_UNAVAILABLE", - "S2_10_AGENT_HASH_MISMATCH", - "S2_10_AGENT_RELEASE_HASH_MISMATCH", - "S2_10_LLM_BINDING_REF_MISSING", - "S2_10_LLM_BINDING_REF_PATH_INVALID", - "S2_10_LLM_BINDING_HASH_UNBOUND", - "S2_10_LLM_BINDING_ASSET_UNAVAILABLE", - "S2_10_LLM_BINDING_HASH_MISMATCH", - "S2_10_LLM_BINDING_RELEASE_HASH_MISMATCH", - "S2_10_HYBRID_RELEASE_UNBOUND", - "STAGE2_ASSET_ROOT_MISSING", - "CONTEXT_COHORT_POLICY_MISMATCH", - "SELECTED_CONTEXT_REF_ID_DUPLICATE", - "SELECTED_CONTEXT_ORDER_INVALID", - "SELECTED_CONTEXT_RAW_HASH_UNBOUND", - "SELECTED_CONTEXT_CANONICAL_HASH_UNBOUND", - "SELECTED_CONTEXT_ASSET_UNAVAILABLE", - "SELECTED_CONTEXT_ASSET_HASH_MISMATCH", - "SELECTED_CONTEXT_CANONICAL_HASH_MISMATCH", - "SELECTED_CONTEXT_NOT_UTF8", - "SELECTED_CONTEXT_PII", - "SELECTED_CONTEXT_ORDERED_REFS_HASH_UNBOUND", - "SELECTED_CONTEXT_ORDERED_REFS_HASH_MISMATCH", - "SELECTED_CONTEXT_EMPTY", - "CLUSTER_SLICE_MISSING", - } - cohort_reason_codes = receipt_reason_codes | { - "STRUCTURAL_FIXTURE_NOT_EXECUTABLE" - } - - for index, row in enumerate(cohorts): - if not isinstance(row, dict): - invariant_errors.append(f"COHORT_{index}_SHAPE") - continue - row_mode = row.get("compile_mode") - row_release_class = row.get("release_class") - if row_mode not in RELEASE_MODE or RELEASE_MODE[row_mode] != row_release_class: - invariant_errors.append(f"COHORT_{index}_MODE_RELEASE_CLASS") - if mode is not None and row_mode != mode: - invariant_errors.append(f"COHORT_{index}_MODE_MIXED") - if release_class is not None and row_release_class != release_class: - invariant_errors.append(f"COHORT_{index}_RELEASE_CLASS_MIXED") - - selected_refs = row.get("selected_context_refs", []) - if not isinstance(selected_refs, list): - invariant_errors.append(f"COHORT_{index}_SELECTED_CONTEXT_SHAPE") - selected_refs = [] - if [item.get("order_index") for item in selected_refs if isinstance(item, dict)] != list( - range(len(selected_refs)) - ): - invariant_errors.append(f"COHORT_{index}_SELECTED_CONTEXT_ORDER") - selected_digest = canonical_digest(selected_refs) - if row.get("selected_context_ordered_refs_sha256") != selected_digest: - invariant_errors.append(f"COHORT_{index}_SELECTED_CONTEXT_DIGEST") - - member_ids = row.get("cohort_member_cluster_ids", []) - member_slices = row.get("member_slices", []) - if not isinstance(member_ids, list) or member_ids != sorted(member_ids): - invariant_errors.append(f"COHORT_{index}_MEMBER_ID_ORDER") - member_ids = list(member_ids) if isinstance(member_ids, list) else [] - if not isinstance(member_slices, list) or member_slices != sorted( - member_slices, - key=lambda item: str(item.get("cluster_id", "")) if isinstance(item, dict) else "", - ): - invariant_errors.append(f"COHORT_{index}_MEMBER_SLICE_ORDER") - member_slices = list(member_slices) if isinstance(member_slices, list) else [] - slice_ids = [ - item.get("cluster_id") - for item in member_slices - if isinstance(item, dict) - ] - if member_ids != slice_ids: - invariant_errors.append(f"COHORT_{index}_MEMBERSHIP_SET") - membership_digest = canonical_digest(member_ids) - slice_digest = canonical_digest(member_slices) - - receipt = row.get("context_materialization_receipt") - if not isinstance(receipt, dict): - invariant_errors.append(f"COHORT_{index}_RECEIPT_SHAPE") - receipt = {} - echoed_fields = { - "bundle_cohort_id": row.get("bundle_cohort_id"), - "s2_10_agent_sha256": row.get("s2_10_agent_sha256"), - "s2_10_llm_binding_sha256": row.get("s2_10_llm_binding_sha256"), - "selected_context_ordered_refs_sha256": selected_digest, - "cohort_membership_sha256": membership_digest, - "member_slice_ref_set_sha256": slice_digest, - } - for key, expected in echoed_fields.items(): - if receipt.get(key) != expected: - invariant_errors.append(f"COHORT_{index}_RECEIPT_{key.upper()}") - materialization_input = { - "handoff_contract_version": row.get("handoff_contract_version"), - **echoed_fields, - } - if receipt.get("materialization_input_sha256") != canonical_digest( - materialization_input - ): - invariant_errors.append(f"COHORT_{index}_MATERIALIZATION_DIGEST") - - expected_cohort = mint_stage2_id( - "bundle_cohort", - [ - row.get("handoff_contract_version"), - row_mode, - row.get("s2_10_agent_sha256"), - row.get("s2_10_llm_binding_sha256"), - selected_digest, - membership_digest, - slice_digest, - ], - prefix="BC", - )["id"] - if row.get("bundle_cohort_id") != expected_cohort: - invariant_errors.append(f"COHORT_{index}_ID_MINT") - - receipt_codes = receipt.get("reason_codes", []) - row_codes = row.get("reason_codes", []) - if not isinstance(receipt_codes, list) or not set(receipt_codes).issubset( - receipt_reason_codes - ): - invariant_errors.append(f"COHORT_{index}_RECEIPT_REASON_CODE") - if not isinstance(row_codes, list) or not set(row_codes).issubset( - cohort_reason_codes - ): - invariant_errors.append(f"COHORT_{index}_REASON_CODE") - - status = row.get("cohort_status") - receipt_status = receipt.get("verification_status") - if row_mode == "STRUCTURAL_FIXTURE": - if status != "STRUCTURAL_ONLY" or "STRUCTURAL_FIXTURE_NOT_EXECUTABLE" not in row_codes: - invariant_errors.append(f"COHORT_{index}_STRUCTURAL_STATUS") - elif receipt_status == "PASS": - if status != "EXECUTABLE" or row_codes: - invariant_errors.append(f"COHORT_{index}_EXECUTABLE_STATUS") - elif receipt_status == "NON_EXECUTABLE": - if status != "NON_EXECUTABLE" or not row_codes: - invariant_errors.append(f"COHORT_{index}_NON_EXECUTABLE_STATUS") - else: - invariant_errors.append(f"COHORT_{index}_RECEIPT_STATUS") - - expected_executable_ids = sorted( - str(row.get("bundle_cohort_id")) - for row in cohorts - if isinstance(row, dict) and row.get("cohort_status") == "EXECUTABLE" - ) - expected_non_executable_ids = sorted( - str(row.get("bundle_cohort_id")) - for row in cohorts - if isinstance(row, dict) and row.get("cohort_status") != "EXECUTABLE" - ) - if bundle_plan.get("executable_bundle_cohort_ids") != expected_executable_ids: - invariant_errors.append("TOP_EXECUTABLE_COHORT_PARTITION") - if bundle_plan.get("non_executable_bundle_cohort_ids") != expected_non_executable_ids: - invariant_errors.append("TOP_NON_EXECUTABLE_COHORT_PARTITION") - if invariant_errors: - raise IngressError( - "BUNDLE_COHORT_INVARIANT_FAILED", - ",".join(sorted(set(invariant_errors))), - details={"invariant_errors": sorted(set(invariant_errors))}, - ) - - mapping_pass = mode in RELEASE_MODE and RELEASE_MODE[mode] == release_class - executable = sorted( - cluster_id - for row in cohorts - if row.get("cohort_status") == "EXECUTABLE" - for cluster_id in row.get("cohort_member_cluster_ids", []) - ) - non_executable = sorted( - cluster_id - for row in cohorts - if row.get("cohort_status") == "NON_EXECUTABLE" - for cluster_id in row.get("cohort_member_cluster_ids", []) - ) - structural = sorted( - cluster_id - for row in cohorts - if row.get("cohort_status") == "STRUCTURAL_ONLY" - for cluster_id in row.get("cohort_member_cluster_ids", []) - ) - return { - "mapping_pass": mapping_pass, - "executable_cluster_ids": executable, - "non_executable_cluster_ids": non_executable, - "structural_cluster_ids": structural, - "passed": mapping_pass and (mode == "STRUCTURAL_FIXTURE" or bool(executable)), - } - - - def _write_fsynced(path: Path, payload: bytes) -> None: - path.parent.mkdir(parents=True, exist_ok=True) - flags = os.O_WRONLY | os.O_CREAT | os.O_EXCL - if hasattr(os, "O_CLOEXEC"): - flags |= os.O_CLOEXEC - descriptor = os.open(path, flags, 0o600) + code = "SOURCE_MISSING" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope=row["impact_scope"], source_refs=[logical_id])) + rows.append(row) + continue + row["raw_sha256"] = snapshot.raw_sha256 + row["byte_length"] = snapshot.byte_length try: - view = memoryview(payload) - while view: - written = os.write(descriptor, view) - view = view[written:] - os.fsync(descriptor) - finally: - os.close(descriptor) - - - def _fsync_directory(path: Path) -> None: - flags = os.O_RDONLY - if hasattr(os, "O_DIRECTORY"): - flags |= os.O_DIRECTORY - descriptor = os.open(path, flags) - try: - os.fsync(descriptor) - finally: - os.close(descriptor) - - - def _published_tree_matches( - output_dir: Path, - run_binding_digest: str, - *, - stage2_asset_root: str | os.PathLike[str] | None = None, - ) -> bool: - status_path = output_dir / "ingress" / "ingress_status.json" - if not status_path.is_file() or status_path.is_symlink(): - return False - try: - snapshot = open_bounded_snapshot(output_dir, "ingress/ingress_status.json", logical_input_id="published_status") - status_value = load_json_strict(snapshot) - except IngressError: - return False - if not isinstance(status_value, dict): - return False - binding = status_value.get("run_binding_receipt", {}) - if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: - return False - barrier = status_value.get("output_barrier", {}) - if not isinstance(barrier, dict) or barrier.get("written_last") is not True: - return False - expected_artifacts = barrier.get("artifacts", []) - if not isinstance(expected_artifacts, list): - return False - if barrier.get("artifact_set_digest") != canonical_digest(expected_artifacts): - return False - try: - schema_documents = _load_output_schemas( - Path(stage2_asset_root) - if stage2_asset_root is not None - else Path(__file__).resolve().parents[1] - ) - _validate_output_artifact( - "ingress/ingress_status.json", - status_value, - schema_documents, - ) - except IngressError: - return False - for row in expected_artifacts: - if not isinstance(row, dict): - return False - try: - artifact = open_bounded_snapshot(output_dir, row["path"], logical_input_id="published_artifact") - except IngressError: - return False - if artifact.raw_sha256 != row.get("raw_sha256"): - return False - try: - parsed_artifact = load_json_strict(artifact) - _validate_output_artifact(row["path"], parsed_artifact, schema_documents) - except IngressError: - return False - expected_paths = { - row.get("path") - for row in expected_artifacts - if isinstance(row, dict) and isinstance(row.get("path"), str) - } | {"ingress/ingress_status.json"} - observed_paths = { - path.relative_to(output_dir).as_posix() - for path in output_dir.rglob("*") - if path.is_file() and not path.is_symlink() - } - if expected_paths != observed_paths: - return False - return True - - - def _output_schema_id(relative_path: str) -> str: - if relative_path.startswith("context/"): - return CONTEXT_SCHEMA_ID - if relative_path.startswith("review/"): - return REVIEW_SCHEMA_ID - if relative_path.startswith("ingress/"): - return INGRESS_SCHEMA_ID - raise IngressError("OUTPUT_SCHEMA_FAMILY_UNKNOWN", f"no schema family for {relative_path}") - - - def publish_atomically( - output_dir: str | os.PathLike[str], - artifacts: Mapping[str, Any], - *, - run_binding_digest: str, - attempt_id: str, - stage2_asset_root: str | os.PathLike[str] | None = None, - ) -> dict[str, Any]: - """Publish a complete normal or diagnostic tree with one directory rename.""" - - output = Path(output_dir) - if not output.is_absolute(): - output = output.resolve() - if not re.fullmatch(r"[A-Za-z0-9_.-]{1,128}", attempt_id): - raise IngressError("ATTEMPT_ID_INVALID", "attempt_id contains forbidden characters") - if "ingress/ingress_status.json" not in artifacts: - raise IngressError("OUTPUT_BARRIER_MISSING", "ingress_status must be supplied") - safe_paths = {_safe_relative_path(path).as_posix(): value for path, value in artifacts.items()} - if len(safe_paths) != len(artifacts): - raise IngressError("DUPLICATE_OUTPUT_PATH", "duplicate output paths after canonicalization") - schema_root = ( - Path(stage2_asset_root) - if stage2_asset_root is not None - else Path(__file__).resolve().parents[1] - ) - schema_documents = _load_output_schemas(schema_root) - for relative_path, supplied_value in safe_paths.items(): - value = load_json_strict(supplied_value) if isinstance(supplied_value, bytes) else supplied_value - _validate_output_artifact(relative_path, value, schema_documents) - if output.exists(): - if output.is_dir() and _published_tree_matches( - output, - run_binding_digest, - stage2_asset_root=schema_root, - ): - return {"status": "IDEMPOTENT_SUCCESS", "output_dir": str(output), "run_binding_digest": run_binding_digest} - raise IngressError("RUN_TUPLE_CONFLICT", "published output exists with a different binding") - parent = output.parent - parent.mkdir(parents=True, exist_ok=True) - staging_parent = parent / ".staging" - staging_parent.mkdir(parents=True, exist_ok=True) - staging = staging_parent / f"{output.name}.{attempt_id}" - if staging.exists(): - if staging.is_symlink() or staging.parent.resolve() != staging_parent.resolve(): - raise IngressError("STAGING_PATH_UNSAFE", "staging path is not safe") - shutil.rmtree(staging) - staging.mkdir(mode=0o700) - try: - artifact_receipt: list[dict[str, Any]] = [] - for relative_path in sorted(path for path in safe_paths if path != "ingress/ingress_status.json"): - value = safe_paths[relative_path] - payload = value if isinstance(value, bytes) else canonical_json_bytes(value) - _write_fsynced(staging / relative_path, payload) - artifact_receipt.append( - { - "logical_artifact_id": relative_path, - "path": relative_path, - "schema_id": _output_schema_id(relative_path), - "raw_sha256": hashlib.sha256(payload).hexdigest(), - } - ) - status = dict(safe_paths["ingress/ingress_status.json"]) - binding = status.get("run_binding_receipt") - if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: - raise IngressError("RUN_BINDING_RECEIPT_MISMATCH", "status run binding does not match publish binding") - status["output_barrier"] = { - "barrier_id": "S2_00_INGRESS_STATUS_BARRIER", - "barrier_path": "ingress/ingress_status.json", - "publish_semantics": "STATUS_LAST_LOGICAL_COMMIT", - "canonical_output_root": f"stage2_runs/by-binding/{run_binding_digest}/", - "branch": "DIAGNOSTIC" if "ingress/technical_diagnostic.json" in safe_paths else "NORMAL", - "artifacts": artifact_receipt, - "artifact_set_digest": canonical_digest(artifact_receipt), - "non_status_artifacts_read_back_verified": True, - "written_last": True, - } - _validate_output_artifact( - "ingress/ingress_status.json", - status, - schema_documents, - ) - _write_fsynced(staging / "ingress/ingress_status.json", canonical_json_bytes(status)) - directories = sorted((path for path in staging.rglob("*") if path.is_dir()), key=lambda path: len(path.parts), reverse=True) - for directory in directories: - _fsync_directory(directory) - _fsync_directory(staging) - _fsync_directory(staging_parent) - os.rename(staging, output) - _fsync_directory(parent) - except BaseException: - if staging.exists() and staging.parent == staging_parent: - shutil.rmtree(staging) - raise - return {"status": "PUBLISHED", "output_dir": str(output), "run_binding_digest": run_binding_digest} - - - def _load_case_snapshots( - stage1_root: Path, - contracts: Sequence[Mapping[str, Any]], - release_lock: Mapping[str, Any], - ) -> tuple[dict[str, Snapshot], list[dict[str, Any]]]: - snapshots: dict[str, Snapshot] = {} - issues: list[dict[str, Any]] = [] - aggregate = 0 - max_file = int(release_lock.get("limits", {}).get("max_file_bytes", MAX_FILE_BYTES)) - max_run = int(release_lock.get("limits", {}).get("max_run_bytes", MAX_RUN_BYTES)) - for contract in contracts: - logical_id = str(contract["logical_input_id"]) - try: - snapshot = open_bounded_snapshot( - stage1_root, - str(contract["observed_path"]), - logical_input_id=logical_id, - max_bytes=max_file, - ) - except IngressError as exc: - if exc.code != "SOURCE_MISSING": - issues.append(_issue(exc.code, source_refs=[logical_id], message=str(exc))) - continue - aggregate += snapshot.byte_length - if aggregate > max_run: - raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "aggregate Stage 1 input budget exceeded") - snapshots[logical_id] = snapshot - return snapshots, issues - - - def _load_bound_completion_seal( - stage1_root: Path, - release_lock: Mapping[str, Any], - ) -> tuple[Mapping[str, Any] | None, Snapshot | None]: - dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) - ref = dependency.get("completion_seal_ref", {}) if isinstance(dependency, dict) else {} - if not isinstance(ref, dict): - return None, None - path = ref.get("path") - digest = ref.get("sha256") - if path in {None, "PENDING_SEQUENTIAL_BIND"} or digest in {None, "PENDING_SEQUENTIAL_BIND"}: - return None, None - if not isinstance(path, str) or not isinstance(digest, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", digest) is None: - raise IngressError("COMPLETION_SEAL_UNBOUND", "completion seal ref is not exactly bound") - snapshot = open_bounded_snapshot(stage1_root, path, logical_input_id="stage1_completion_seal") - if snapshot.raw_sha256 != digest.lower(): - raise IngressError("COMPLETION_SEAL_HASH_MISMATCH", "completion seal raw hash differs from release") - document = load_json_strict(snapshot) - if not isinstance(document, dict): - raise IngressError("COMPLETION_SEAL_SHAPE", "completion seal must be an object") - return document, snapshot - - - def _load_stage1_deployment_closure( - stage1_deployment_root: Path, - release_lock: Mapping[str, Any], - ) -> tuple[dict[str, Snapshot], dict[str, Any], list[dict[str, Any]]]: - """Load only the release-enumerated Stage 1 deployment closure.""" - - dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) - locked_rows = dependency.get("concrete_paths", []) if isinstance(dependency, dict) else [] - if not isinstance(locked_rows, list) or not locked_rows: - raise IngressError("STAGE1_DEPENDENCY_LOCK_MISSING", "Stage 1 deployment closure is not enumerated") - expected_count = dependency.get("expected_concrete_path_count") - if expected_count is not None and expected_count != len(locked_rows): - raise IngressError("STAGE1_DEPENDENCY_COUNT_MISMATCH", "Stage 1 dependency row count is not sealed") - snapshots: dict[str, Snapshot] = {} - documents: dict[str, Any] = {} - source_rows: list[dict[str, Any]] = [] - seen_paths: set[str] = set() - aggregate = 0 - max_file = int(release_lock.get("limits", {}).get("max_file_bytes", MAX_FILE_BYTES)) - max_run = int(release_lock.get("limits", {}).get("max_run_bytes", MAX_RUN_BYTES)) - for index, locked in enumerate(locked_rows): - if not isinstance(locked, dict) or not isinstance(locked.get("path"), str): - raise IngressError("STAGE1_DEPENDENCY_ROW_SHAPE", f"invalid dependency row {index}") - path = _safe_relative_path(locked["path"]).as_posix() - if path in seen_paths: - raise IngressError("STAGE1_DEPENDENCY_DUPLICATE_PATH", f"duplicate dependency path: {path}") - seen_paths.add(path) - expected_hash = locked.get("sha256") - if not isinstance(expected_hash, str) or not re.fullmatch(r"[A-Fa-f0-9]{64}", expected_hash): - raise IngressError("STAGE1_DEPENDENCY_UNBOUND", f"dependency hash is not bound: {path}") - lock_id = str(locked.get("lock_id", f"S1-DEPLOY-{index + 1:03d}")) - snapshot = open_bounded_snapshot( - stage1_deployment_root, - path, - logical_input_id=f"deployment:{lock_id}", - max_bytes=max_file, - ) - aggregate += snapshot.byte_length - if aggregate > max_run: - raise IngressError("AGGREGATE_DEPLOYMENT_SIZE_LIMIT", "Stage 1 deployment closure exceeds byte budget") - if snapshot.raw_sha256 != expected_hash: - raise IngressError("STAGE1_DEPENDENCY_HASH_MISMATCH", f"deployment hash mismatch: {path}") document = load_json_strict( snapshot, max_depth=int(release_lock.get("limits", {}).get("max_json_depth", MAX_JSON_DEPTH)), max_items=int(release_lock.get("limits", {}).get("max_json_items", MAX_JSON_ITEMS)), ) - snapshots[lock_id] = snapshot - documents[path] = document - schema_id = locked.get("schema_id") - source_rows.append( - { - "logical_input_id": f"deployment:{lock_id}", - "requirement_class": "UPSTREAM_DEPLOYMENT", - "expected_path": path, - "observed_path": path, - "resolution_source": "RELEASE_BOUND_CONTRACT_MANIFEST", - "schema_id": schema_id, - "schema_sha256": None, - "producer_id": None, - "producer_alias_id": None, - "adapter_id": "S2A-UPSTREAM-DEPLOYMENT-V1", - "raw_sha256": snapshot.raw_sha256, - "byte_length": snapshot.byte_length, - "run_identity_ref": {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": path}, - "transaction_identity_ref": {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": path}, - "parse_status": "PASS", - "schema_status": "UNEVALUABLE" if schema_id else "NOT_APPLICABLE", - "seal_status": "PASS", - "scope_technical_disposition": "AVAILABLE", - "impact_scope": "GLOBAL", - "scope_refs": [f"deployment:{lock_id}"], - "source_contract_row_refs": [f"deployment:{lock_id}"], - "reason_codes": [], - "downstream_allowed_actions": [], - "issue_codes": [], - } - ) - return snapshots, documents, source_rows - - - def _signal_source_contract_rows( - signal_all: Mapping[str, Any], - release_lock: Mapping[str, Any], - ) -> list[dict[str, Any]]: - rows: list[dict[str, Any]] = [] - transaction_id = str(signal_all.get("manifest_transaction_id", "MISSING")) - family_contract = next( - ( - row - for row in _release_stage1_source_rows(release_lock) - if row.get("logical_input_id") == "signal_payload_family" - ), - {}, - ) - for file_row in signal_all.get("ordered_file_rows", []): - logical_id = f"signal_file:{int(file_row['manifest_index']):03d}" - hash_status = str(file_row.get("hash_status", "UNEVALUABLE")) - issue_codes = ["SIGNAL_FILE_HASH_MISMATCH"] if hash_status == "FAIL" else [] - rows.append( - { - "logical_input_id": logical_id, - "requirement_class": "SIGNAL_PAYLOAD", - "expected_path": str(file_row["physical_path"]), - "observed_path": str(file_row["physical_path"]), - "resolution_source": "RELEASE_BOUND_CONTRACT_MANIFEST", - "schema_id": None, - "schema_sha256": None, - "producer_id": family_contract.get("producer_id"), - "producer_alias_id": family_contract.get("producer_alias_id"), - "adapter_id": family_contract.get("adapter_id", "S2A-SIGNAL-PAYLOAD-FAMILY-V1"), - "raw_sha256": str(file_row["raw_sha256"]), - "byte_length": int(file_row["byte_length"]), - "run_identity_ref": {"value": None, "disposition": "MISSING", "source_ref": logical_id}, - "transaction_identity_ref": { - "value": transaction_id, - "disposition": "OBSERVED", - "source_ref": "signal_manifest", - }, - "parse_status": "PASS", - "schema_status": "UNEVALUABLE", - "seal_status": hash_status, - "scope_technical_disposition": "AVAILABLE_WITH_ISSUES" if issue_codes else "AVAILABLE", - "impact_scope": "SIGNAL", - "scope_refs": [logical_id], - "source_contract_row_refs": [logical_id], - "reason_codes": issue_codes, - "downstream_allowed_actions": [], - "issue_codes": issue_codes, - } - ) - return rows - - - def _snapshot_state_digest(snapshots: Sequence[Snapshot]) -> str: - """Digest stable source content, not hydration-local filesystem identity.""" - - return canonical_digest( - sorted( - [ - item.logical_input_id, - item.relative_path, - item.raw_sha256, - item.byte_length, - ] - for item in snapshots - ) - ) - - - def _verify_snapshot_state(snapshots: Sequence[Snapshot]) -> str: - rows: list[list[Any]] = [] - for item in snapshots: - path = Path(item.resolved_path) - if path.is_symlink(): - raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source became a symlink: {item.relative_path}") - try: - observed = path.stat(follow_symlinks=False) - except FileNotFoundError as exc: - raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source disappeared: {item.relative_path}") from exc - identity = (observed.st_dev, observed.st_ino, observed.st_size, observed.st_mtime_ns) - expected = (item.device, item.inode, item.byte_length, item.mtime_ns) - if identity != expected: - raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source changed: {item.relative_path}") - rows.append( - [ - item.logical_input_id, - item.relative_path, - item.raw_sha256, - item.byte_length, - ] - ) - return canonical_digest(sorted(rows)) - - - def _cluster_inputs(documents: Mapping[str, Any]) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: - ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) - bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) - les_rows = _array_rows( - documents.get("legal_effect_structures"), - ("structures", "structure_records", "legal_effect_structures", "rows", "items"), - ) - evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) - event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) - members: list[dict[str, Any]] = [] - relations: list[dict[str, Any]] = [] - internal_by_kind_ref: dict[tuple[str, str], str] = {} - - def add_member( - kind: str, - member_ref: str, - logical_id: str, - pointer: str, - raw_value: Any, - ) -> str: - key = (kind, member_ref) - prior = internal_by_kind_ref.get(key) - if prior is not None: - return prior - internal_id = f"{kind}:{member_ref}" - internal_by_kind_ref[key] = internal_id - members.append( - { - "member_id": internal_id, - "member_ref": member_ref, - "member_kind": kind, - "source_refs": [_source_ref(logical_id, pointer, raw_value, stage1_id=member_ref)], - "scope_technical_disposition": "AVAILABLE", - } - ) - return internal_id - - bo_member_by_id: dict[str, str] = {} - for index, row in enumerate(bo_rows): - if not isinstance(row, dict): - continue - bo_id = str(row.get("BO_ID", f"BO-OCCURRENCE-{index}")) - bo_member_by_id[bo_id] = add_member("BO", bo_id, "bo", f"/business_objects/{index}", row) - - fact_member_by_id: dict[str, str] = {} - fact_members_by_bo: dict[str, list[str]] = defaultdict(list) - fact_members_by_evidence: dict[str, list[str]] = defaultdict(list) - fact_members_by_event: dict[str, list[str]] = defaultdict(list) - for index, row in enumerate(ledger_rows): - if not isinstance(row, dict): - continue - fact_id = str(row.get("fact_id", f"F-OCCURRENCE-{index}")) - member_id = add_member("FACT", fact_id, "fact_ledger_base", f"/facts/{index}", row) - fact_member_by_id[fact_id] = member_id - source_bo_id = row.get("source_bo_id") - if source_bo_id is not None: - fact_members_by_bo[str(source_bo_id)].append(member_id) - evidence_refs = row.get("evidence_refs", row.get("evidence_ids", [])) - if isinstance(evidence_refs, list): - for ref in evidence_refs: - fact_members_by_evidence[str(ref)].append(member_id) - event_refs = row.get("event_refs", row.get("event_ids", [])) - if isinstance(event_refs, list): - for ref in event_refs: - fact_members_by_event[str(ref)].append(member_id) - relation_arrays = [ - row.get(key) - for key in ("relations", "explicit_relations", "candidate_relations") - if isinstance(row.get(key), list) - ] - for relation_array in relation_arrays: - for relation_index, relation in enumerate(relation_array): - if not isinstance(relation, dict): - continue - target_ref = relation.get( - "target_fact_id", - relation.get("to_fact_id", relation.get("target_member_id")), + documents[logical_id] = document + row["parse_status"] = "PASS" + except IngressError as exc: + row["parse_status"] = "FAIL" + row["schema_status"] = "UNEVALUABLE" + row["seal_status"] = "UNEVALUABLE" + row["scope_technical_disposition"] = "UNAVAILABLE" + row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" + row["reason_codes"].append(exc.code) + row["issue_codes"].append(exc.code) + issues.append(_issue(exc.code, impact_scope=row["impact_scope"], source_refs=[logical_id], message=str(exc))) + rows.append(row) + continue + expected_adapter = ADAPTER_IDS.get(logical_id) + if expected_adapter is not None and release_row.get("adapter_id") not in {None, expected_adapter}: + code = "ADAPTER_ID_MISMATCH" + row["schema_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if schema_ref is not None: + schema_path = schema_ref.get("path") + schema_snapshot = deployment_by_path.get(schema_path) if isinstance(schema_path, str) else None + schema_document = deployment_documents.get(schema_path) if isinstance(schema_path, str) else None + expected_schema_hash = schema_ref.get("sha256") + expected_schema_id = schema_ref.get("$id") + if schema_snapshot is None or not isinstance(schema_document, dict): + schema_error = "SCHEMA_REF_NOT_IN_BOUNDED_DEPLOYMENT" + elif not isinstance(expected_schema_hash, str) or schema_snapshot.raw_sha256 != expected_schema_hash.lower(): + schema_error = "SCHEMA_HASH_MISMATCH" + elif expected_schema_id is not None and schema_document.get("$id") != expected_schema_id: + schema_error = "SCHEMA_ID_MISMATCH" + else: + schema_error = None + try: + _validate_schema_node( + document, + schema_document, + root_schema=schema_document, + schema_documents=schema_documents, + instance_path=logical_id, ) - relation_kind = str(relation.get("relation_kind", relation.get("kind", ""))) - if target_ref is None or relation_kind not in CANDIDATE_RELATION_KINDS | {"EXPLICIT_CASE_RELATION"}: - continue - relations.append( - { - "_deferred_source_fact_id": fact_id, - "_deferred_target_fact_id": str(target_ref), - "relation_kind": relation_kind, - "source_refs": [ - _source_ref( - "fact_ledger_base", - f"/facts/{index}/relations/{relation_index}", - relation, - ) - ], - } - ) - - for bo_id, fact_member_ids in sorted(fact_members_by_bo.items()): - bo_member = bo_member_by_id.get(bo_id) - if bo_member is None: - continue - for fact_member in sorted(set(fact_member_ids)): - relations.append( - { - "source_member_id": fact_member, - "target_member_id": bo_member, - "relation_kind": "SAME_BO_ID", - "source_refs": [_source_ref("bo", "", bo_id, stage1_id=bo_id)], - } - ) - - for index, row in enumerate(les_rows): - if not isinstance(row, dict): - continue - structure_ref = str(row.get("structure_id", row.get("legal_effect_structure_id", f"LES-OCCURRENCE-{index}"))) - les_member = add_member( - "LES_STRUCTURE", - structure_ref, - "legal_effect_structures", - f"/structures/{index}", - row, - ) - source_bo_ids = row.get("source_bo_ids", []) - if isinstance(source_bo_ids, list): - for bo_id in source_bo_ids: - bo_member = bo_member_by_id.get(str(bo_id)) - if bo_member is not None: - relations.append( - { - "source_member_id": les_member, - "target_member_id": bo_member, - "relation_kind": "SOURCE_BO_ATTACHMENT", - "source_refs": [ - _source_ref( - "legal_effect_structures", - f"/structures/{index}/source_bo_ids", - source_bo_ids, - ) - ], - } - ) - - for index, row in enumerate(evidence_rows): - if not isinstance(row, dict): - continue - evidence_id = str(row.get("evidence_id", row.get("id", f"EVIDENCE-OCCURRENCE-{index}"))) - evidence_member = add_member("EVIDENCE", evidence_id, "evidence_indexed", f"/items/{index}", row) - for fact_member in sorted(set(fact_members_by_evidence.get(evidence_id, []))): - relations.append( - { - "source_member_id": fact_member, - "target_member_id": evidence_member, - "relation_kind": "SAME_EVIDENCE_REF", - "source_refs": [_source_ref("fact_ledger_base", "", evidence_id, stage1_id=evidence_id)], - } - ) - - for index, row in enumerate(event_rows): - if not isinstance(row, dict): - continue - event_id = str(row.get("event_id", row.get("id", f"EVENT-OCCURRENCE-{index}"))) - event_member = add_member("EVENT", event_id, "evidence_event_candidates", f"/items/{index}", row) - for fact_member in sorted(set(fact_members_by_event.get(event_id, []))): - relations.append( - { - "source_member_id": fact_member, - "target_member_id": event_member, - "relation_kind": "SAME_EVENT_REF", - "source_refs": [_source_ref("fact_ledger_base", "", event_id, stage1_id=event_id)], - } - ) - - resolved_relations: list[dict[str, Any]] = [] - for relation in relations: - if "_deferred_source_fact_id" not in relation: - resolved_relations.append(relation) - continue - source_member = fact_member_by_id.get(str(relation["_deferred_source_fact_id"])) - target_member = fact_member_by_id.get(str(relation["_deferred_target_fact_id"])) - if source_member is None or target_member is None: - continue - resolved_relations.append( - { - "source_member_id": source_member, - "target_member_id": target_member, - "relation_kind": relation["relation_kind"], - "source_refs": relation["source_refs"], - } - ) - return members, resolved_relations - - - def _run_binding_digest( - snapshots: Mapping[str, Snapshot], - release_lock: Mapping[str, Any], - ) -> str: - input_set = [[key, snapshots[key].raw_sha256] for key in sorted(snapshots)] - binding = { - "input_set_digest": canonical_digest(input_set), - "stage2_release_digest": release_lock.get("_release_raw_sha256", release_lock.get("stage2_release_digest")), - "algorithm_digest": ALGORITHM_SEMANTIC_DIGEST, - "release_class": release_lock.get("release_class"), - } - return canonical_digest(binding) - - - def _input_set_digest( - snapshots: Mapping[str, Snapshot], - additional_snapshots: Sequence[Snapshot] = (), - ) -> str: - rows = [ - ["run", key, snapshots[key].relative_path, snapshots[key].raw_sha256] - for key in sorted(snapshots) - ] - rows.extend( - ["closure", item.logical_input_id, item.relative_path, item.raw_sha256] - for item in additional_snapshots - ) - return canonical_digest(sorted(rows, key=canonical_digest)) - - - def _make_run_binding_receipt( - snapshots: Mapping[str, Snapshot], - release_lock: Mapping[str, Any], - *, - request_id: str, - user_context_sha256: str, - workspace_context_sha256: str, - additional_snapshots: Sequence[Snapshot] = (), - ) -> dict[str, Any]: - input_digest = _input_set_digest(snapshots, additional_snapshots) - release_digest = str( - release_lock.get( - "_release_raw_sha256", - release_lock.get( - "release_digest", - release_lock.get("stage2_release_digest", "0" * 64), - ), - ) - ) - algorithm_digest = ALGORITHM_SEMANTIC_DIGEST - binding_digest = canonical_digest( - { - "input_set_digest": input_digest, - "stage2_release_digest": release_digest, - "algorithm_digest": algorithm_digest, - "release_class": release_lock.get("release_class"), - } - ) - return { - "schema_version": "stage2_s2_00_run_binding_receipt.v1.1", - "request_id": request_id, - "run_id": f"S2RUN-{binding_digest}", - "input_set_digest": input_digest, - "stage2_release_digest": release_digest, - "algorithm_digest": algorithm_digest, - "release_class": release_lock.get("release_class"), - "run_binding_digest": binding_digest, - "canonical_output_root": f"stage2_runs/by-binding/{binding_digest}/", - "run_identity_derivation": "RUN_ID_PREFIXED_FROM_RUN_BINDING_DIGEST", - "output_root_derivation": "stage2_runs/by-binding//", - "user_context_sha256": user_context_sha256, - "workspace_context_sha256": workspace_context_sha256, - } - - - def canonical_run_id(run_binding_digest: str) -> str: - """Derive the immutable run identifier from the complete binding digest.""" - - if re.fullmatch(r"[a-f0-9]{64}", run_binding_digest) is None: - raise IngressError( - "RUN_BINDING_DIGEST_INVALID", - "canonical run id requires one lowercase SHA-256 digest", - ) - return f"S2RUN-{run_binding_digest}" - - - def _compact_conservation_checks( - checks: Sequence[Mapping[str, Any]], - ) -> list[dict[str, Any]]: - compact: list[dict[str, Any]] = [] - for check in checks: - check_id = str(check.get("check_id", "UNNAMED_CONSERVATION_CHECK")) - status_value = str(check.get("status", "UNEVALUABLE")) - status = status_value if status_value in {"PASS", "FAIL", "UNEVALUABLE"} else "UNEVALUABLE" - observed_payload = {key: value for key, value in check.items() if key not in {"status"}} - left_digest = canonical_digest(observed_payload) if status != "UNEVALUABLE" else None - right_digest = left_digest if status == "PASS" else canonical_digest([check_id, "EXPECTED"]) if status == "FAIL" else None - left_count = next( - ( - int(check[key]) - for key in ("left_count", "bo_count", "source_bo_ref_count") - if isinstance(check.get(key), int) - ), - None, - ) - right_count = next( - ( - int(check[key]) - for key in ("right_count", "source_bo_ref_count", "bo_count") - if isinstance(check.get(key), int) - ), - None, - ) - compact.append( - { - "check_id": check_id, - "status": status, - "left_counter_digest": left_digest, - "right_counter_digest": right_digest, - "left_count": left_count, - "right_count": right_count, - "source_refs": [check_id], - "issue_codes": [f"{check_id}_FAILED"] if status == "FAIL" else [], - } - ) - return compact - - - def _compact_review_receipt( - normalized_reviews: Mapping[str, Any], - ) -> dict[str, Any]: - normalized = normalized_reviews.get("normalized_occurrences", []) - raw = normalized_reviews.get("raw_occurrences", []) - counts = dict(normalized_reviews.get("partition_counts", {})) - for key in ("SUPPORTED", "CONDITIONAL", "UNRESOLVED", "EXCLUDED", "UNMAPPED"): - counts.setdefault(key, 0) - return { - "schema_version": "stage2_review_normalization_receipt.v1", - "adapter_ids": list(normalized_reviews.get("adapter_ids", [])), - "raw_occurrence_count": len(raw), - "normalized_occurrence_count": len(normalized), - "partition_counts": counts, - "unmapped_occurrence_count": counts["UNMAPPED"], - "conservation_status": str(normalized_reviews.get("conservation_status", "FAIL")), - "ordered_occurrence_refs": [ - str(row.get("review_key", {}).get("value")) - for row in normalized - if row.get("review_key", {}).get("value") - ], - } - - - def _build_issue_ledger( - issues: Sequence[Mapping[str, Any]], - normalized_reviews: Mapping[str, Any], - *, - run_id: str, - input_set_digest: str, - ) -> dict[str, Any]: - issue_rows: list[dict[str, Any]] = [] - seen: set[str] = set() - for index, issue in enumerate(issues): - issue_code = str(issue.get("issue_code", "UNSPECIFIED_ISSUE")) - source_refs = sorted(set(str(item) for item in issue.get("scope_refs", []))) or ["S2_00"] - issue_mint = mint_stage2_id( - "issue", - [issue_code, source_refs, index], - prefix="ISS", - ) - if issue_mint["id"] in seen: - continue - seen.add(issue_mint["id"]) - raw_severity = str(issue.get("severity", "ERROR")).upper() - technical_severity = { - "INFO": "INFO", - "REVIEW": "REVIEW", - "WARNING": "HARD_WARNING", - "HARD_WARNING": "HARD_WARNING", - "ERROR": "BLOCK", - "BLOCK": "BLOCK", - }.get(raw_severity, "REVIEW") - issue_rows.append( - { - "issue_id": issue_mint["id"], - "issue_code": issue_code, - "technical_severity": technical_severity, - "review_partition": "UNRESOLVED", - "impact_scope": str(issue.get("impact_scope", "GLOBAL")), - "scope_refs": source_refs, - "source_contract_row_refs": sorted( - set(str(item) for item in issue.get("source_contract_row_refs", source_refs)) - ), - "reason_codes": sorted(set(str(item) for item in issue.get("reason_codes", [issue_code]))), - "downstream_allowed_actions": sorted( - set(str(item) for item in issue.get("downstream_allowed_actions", [])) - ), - "source_refs": source_refs, - "status": "OPEN", - "upstream_review_key": None, - "raw_item_sha256": None, - } - ) - for occurrence in normalized_reviews.get("normalized_occurrences", []): - if occurrence.get("partition") != "UNMAPPED": - continue - review_key = occurrence["review_key"]["value"] - issue_mint = mint_stage2_id( - "issue", - ["UNMAPPED_REVIEW_STATUS", review_key], - prefix="ISS", - ) - issue_rows.append( - { - "issue_id": issue_mint["id"], - "issue_code": "UNMAPPED_REVIEW_STATUS", - "technical_severity": "REVIEW", - "review_partition": "UNMAPPED", - "impact_scope": "REVIEW_ITEM", - "scope_refs": [review_key], - "source_contract_row_refs": occurrence["source_contract_row_refs"], - "reason_codes": ["UNMAPPED_REVIEW_STATUS"], - "downstream_allowed_actions": occurrence["downstream_allowed_actions"], - "source_refs": occurrence["source_contract_row_refs"], - "status": "OPEN", - "upstream_review_key": review_key, - "raw_item_sha256": occurrence["source_item_raw_sha256"], - } - ) - return { - "schema_version": "stage2_issue_ledger_base.v1", - "producer_id": "S2_00", - "run_id": run_id, - "input_set_digest": input_set_digest, - "mapping_table_version": "S2-REVIEW-MAP-V1", - "issues": sorted(issue_rows, key=lambda row: row["issue_id"]), - "raw_review_occurrence_count": len(normalized_reviews.get("raw_occurrences", [])), - "normalized_review_occurrence_count": len(normalized_reviews.get("normalized_occurrences", [])), - "review_conservation_status": str(normalized_reviews.get("conservation_status", "FAIL")), - } - - - def _route( - source_rows: Sequence[Mapping[str, Any]], - cluster_plan: Mapping[str, Any], - cohort_receipt: Mapping[str, Any], - issues: Sequence[Mapping[str, Any]], - ) -> str: - identity_unavailable = any( - row.get("requirement_class") == "IDENTITY_BACKBONE" - and row.get("scope_technical_disposition") == "UNAVAILABLE" - for row in source_rows - ) - diagnostic_issue_codes = { - "RELEASE_SOURCE_CONTRACT_MISSING", - "RELEASE_SOURCE_PATH_MISMATCH", - "RAW_HASH_MISMATCH", - "SCHEMA_HASH_MISMATCH", - "SCHEMA_ID_MISMATCH", - "RUN_IDENTITY_CONFLICT", - "TRANSACTION_IDENTITY_CONFLICT", - "BO_FACT_CONSERVATION_FAILED", - "FACT_ID_CONSERVATION_FAILED", - "CURRENT_V8_LEDGER_EXTENSION_MISSING", - } - observed_issue_codes = {str(item.get("issue_code")) for item in issues} - if identity_unavailable or observed_issue_codes & diagnostic_issue_codes or not cluster_plan.get("clusters"): - return "TO_S2_40_STATUS_ONLY" - if not cohort_receipt.get("executable_cluster_ids"): - return "TO_S2_40_STATUS_ONLY" - return "TO_S2_10_WITH_ISSUES" if issues else "TO_S2_10" - - - def execute_ingress( - stage1_run_root: str | os.PathLike[str], - release_lock: Mapping[str, Any], - *, - output_dir: str | os.PathLike[str] | None, - attempt_id: str, - run_id: str = "S2-00-REQUEST", - user_context_sha256: str = "0" * 64, - workspace_context_sha256: str = "0" * 64, - stage1_deployment_root: str | os.PathLike[str] | None = None, - stage2_asset_root: str | os.PathLike[str] | None = None, - contract_manifest: Mapping[str, Any] | None = None, - hydration_stability_receipt: Mapping[str, Any] | None = None, - ) -> dict[str, Any]: - """Execute C00→C05→C10→C15 for a canary or production case run.""" - - if release_lock.get("release_class") == "DEV_FIXTURE_RELEASE": - raise IngressError("DEV_FIXTURE_REAL_RUN_FORBIDDEN", "DEV fixture release cannot publish a case run") - if output_dir is None: - raise IngressError("OUTPUT_DIR_REQUIRED", "output_dir is required for canary/production") - if stage1_deployment_root is None: - raise IngressError("STAGE1_DEPLOYMENT_ROOT_REQUIRED", "Stage 1 deployment root is required") - if not re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}", run_id): - raise IngressError("REQUEST_ID_INVALID", "request_id contains forbidden characters") - if not isinstance(hydration_stability_receipt, dict): - raise IngressError( - "HYDRATION_STABILITY_RECEIPT_REQUIRED", - "canary and production ingress require the remote two-pass hydration receipt", - ) - for label, value in ( - ("user_context_sha256", user_context_sha256), - ("workspace_context_sha256", workspace_context_sha256), - ): - if not re.fullmatch(r"[A-Fa-f0-9]{64}", value): - raise IngressError("RUN_CONTEXT_HASH_INVALID", f"{label} must be a SHA-256 digest") - root_arg = Path(stage1_run_root) - deployment_root_arg = Path(stage1_deployment_root) - asset_root_arg = Path(stage2_asset_root) if stage2_asset_root is not None else Path(__file__).resolve().parents[1] - for candidate, code in ( - (root_arg, "SYMLINK_ROOT_REJECTED"), - (deployment_root_arg, "SYMLINK_DEPLOYMENT_ROOT_REJECTED"), - (asset_root_arg, "SYMLINK_ASSET_ROOT_REJECTED"), - ): - if candidate.is_symlink(): - raise IngressError(code, "approved root itself may not be a symlink") - root = root_arg.resolve(strict=True) - deployment_root = deployment_root_arg.resolve(strict=True) - asset_root = ( - asset_root_arg.resolve(strict=True) - if stage2_asset_root is not None - else Path(__file__).resolve().parents[1] - ) - deployment_snapshots, deployment_documents, deployment_rows = _load_stage1_deployment_closure( - deployment_root, - release_lock, - ) - contracts = resolve_stage1_sources(root, contract_manifest) - snapshots, snapshot_issues = _load_case_snapshots(root, contracts, release_lock) - completion_seal, completion_seal_snapshot = _load_bound_completion_seal(root, release_lock) - ingress = validate_ingress_contracts( - snapshots, - contracts, - release_lock, - deployment_snapshots=deployment_snapshots, - deployment_documents=deployment_documents, - completion_seal=completion_seal, - contract_manifest=contract_manifest, - ) - documents = ingress["documents"] - issues = list(snapshot_issues) + list(ingress["issues"]) - signal_all: dict[str, Any] | None = None - if isinstance(documents.get("signal_manifest"), dict): - try: - max_run = int(release_lock.get("limits", {}).get("max_run_bytes", MAX_RUN_BYTES)) - already_loaded = sum(snapshot.byte_length for snapshot in snapshots.values()) + sum( - snapshot.byte_length for snapshot in deployment_snapshots.values() - ) - remaining = max_run - already_loaded - if remaining < 0: - raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "case and deployment closure exceed aggregate run budget") - signal_all = expand_stage2_signal_all( - root, - documents["signal_manifest"], - max_file_bytes=int(release_lock.get("limits", {}).get("max_file_bytes", MAX_FILE_BYTES)), - max_total_bytes=remaining, - signal_registry=deployment_documents.get("signals/signal_registry.v2.json"), - ) - signal_all = bind_signal_occurrences(signal_all, documents) - issues.extend(signal_all.get("issues", [])) - except IngressError as exc: - issues.append(_issue(exc.code, impact_scope="SIGNAL", source_refs=["signal_manifest"], message=str(exc))) - routing_activation = documents.get("domain_activation_manifest") - if isinstance(routing_activation, dict): - try: - parsed_signal_documents = signal_all.get("_parsed_documents_by_path", {}) if signal_all else {} - signal_activation = parsed_signal_documents.get("domain_activation_manifest.json") - if signal_activation is None: - raise IngressError( - "SG01_SIGNAL_ARTIFACT_MISSING", - "signal ALL does not contain domain_activation_manifest.json", - ) - signal_activation_row = next( - ( - row - for row in signal_all.get("ordered_file_rows", []) - if row.get("file_path") == "domain_activation_manifest.json" - ), - {}, - ) - verify_activation_projection( - routing_activation, - signal_activation, - routing_raw_sha256=snapshots.get("domain_activation_manifest").raw_sha256 if snapshots.get("domain_activation_manifest") else None, - signal_raw_sha256=signal_activation_row.get("raw_sha256"), - ) - except IngressError as exc: - issues.append(_issue(exc.code, impact_scope="SIGNAL", source_refs=["domain_activation_manifest", "signal_sg01_activation"], message=str(exc))) - seal_deployment_snapshots = { - "stage1_domain_registry_index": snapshot - for lock_id, snapshot in deployment_snapshots.items() - if snapshot.relative_path == "domains/_registry_index.json" - } - seals = verify_cross_artifact_seals(documents, snapshots, seal_deployment_snapshots) - issues.extend(seals["issues"]) - review_docs = {key: value for key, value in documents.items() if "review_handoff" in key or "soft_gate_handoff" in key} - normalized_reviews = normalize_review_items(review_docs, release_lock) - issues.extend(normalized_reviews.get("_issues", [])) - conservation = check_conservation( - documents, - signal_all=signal_all, - normalized_reviews=normalized_reviews, - source_snapshots=snapshots, - ) - issues.extend(conservation["issues"]) - signal_snapshots = list(signal_all.get("_payload_snapshots", [])) if signal_all else [] - closure_snapshots = list(deployment_snapshots.values()) + signal_snapshots - all_snapshots = list(snapshots.values()) + closure_snapshots - if completion_seal_snapshot is not None: - all_snapshots.append(completion_seal_snapshot) - snapshot_start_digest = _snapshot_state_digest(all_snapshots) - input_set_digest = _input_set_digest(snapshots, closure_snapshots) - run_binding_receipt = _make_run_binding_receipt( - snapshots, - release_lock, - request_id=run_id, - user_context_sha256=user_context_sha256, - workspace_context_sha256=workspace_context_sha256, - additional_snapshots=closure_snapshots, - ) - run_binding_digest = run_binding_receipt["run_binding_digest"] - run_id = canonical_run_id(run_binding_digest) - context_schema_snapshot = open_bounded_snapshot( - asset_root, - "schemas/context.schema.json", - logical_input_id="stage2_context_schema", - ) - artifact_header = _artifact_header( - schema_sha256=context_schema_snapshot.raw_sha256, - run_id=run_id, - input_set_digest=input_set_digest, - stage2_release_digest=run_binding_receipt["stage2_release_digest"], - algorithm_digest=run_binding_receipt["algorithm_digest"], - release_class=str(release_lock["release_class"]), - ) - domain_configs = { - match.group(1): value - for path, value in deployment_documents.items() - if (match := re.fullmatch(r"domains/([^/]+)/domain_config\.json", path)) is not None - } - routing_payload = ( - _activation_payload(documents["domain_activation_manifest"]) - if isinstance(documents.get("domain_activation_manifest"), dict) - else {} - ) - active_domain_ids = routing_payload.get("active_domain_ids", []) - if isinstance(active_domain_ids, list): - for domain_id in sorted(set(str(value) for value in active_domain_ids)): - if domain_id not in domain_configs: + except _SchemaViolation as exc: + schema_error = "SOURCE_SCHEMA_VALIDATION_FAILED" issues.append( _issue( - "ACTIVE_DOMAIN_CONFIG_MISSING", + schema_error, impact_scope="CLUSTER", - source_refs=[f"domain_config:{domain_id}"], + source_refs=[logical_id], + message=str(exc), ) ) - case_context = build_case_context( - documents, - run_binding_digest=run_binding_digest, - issues=issues, - artifact_header=artifact_header, - signal_all=signal_all, - ) - evidence_inventory = build_evidence_inventory( - documents, - run_binding_digest=run_binding_digest, - artifact_header=artifact_header, - ) - object_registry = build_object_registry( - documents, - run_binding_digest=run_binding_digest, - artifact_header=artifact_header, - ) - party_context = build_party_and_title_context( - documents, - run_binding_digest=run_binding_digest, - artifact_header=artifact_header, - ) - slot_crosswalk = build_slot_crosswalk( - documents, - domain_configs, - run_binding_digest=run_binding_digest, - artifact_header=artifact_header, - ) - members, relations = _cluster_inputs(documents) - cluster_plan = compile_cluster_plan(members, relations, artifact_header=artifact_header) - contexts = { - "case_context": case_context, - "evidence_inventory": evidence_inventory, - "object_registry": object_registry, - "party_and_title_context": party_context, - "slot_crosswalk": slot_crosswalk, - "_projection_sources": { - "EVENT": { - str(row.get("event_id", row.get("id"))): row - for row in _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) - if isinstance(row, dict) and (row.get("event_id") is not None or row.get("id") is not None) - }, - "LES_STRUCTURE": { - str(row.get("structure_id", row.get("legal_effect_structure_id"))): row - for row in _array_rows(documents.get("legal_effect_structures"), ("structures", "structure_records", "rows", "items")) - if isinstance(row, dict) and (row.get("structure_id") is not None or row.get("legal_effect_structure_id") is not None) - }, - "SIGNAL_OCCURRENCE": { - str(row.get("occurrence_ref")): row - for row in (signal_all or {}).get("record_occurrences", []) - if row.get("occurrence_ref") is not None - }, - "REVIEW_ITEM": { - str(row.get("review_key", {}).get("value")): row - for row in normalized_reviews.get("normalized_occurrences", []) - if row.get("review_key", {}).get("value") is not None - }, - "ACTIVE_PROFILE": { - str(domain_id): config - for domain_id, config in domain_configs.items() - if domain_id in set(str(value) for value in active_domain_ids) - }, - }, - } - slices = compile_cluster_slices(cluster_plan, contexts) - effective_release_lock = dict(release_lock) - if not isinstance(effective_release_lock.get("bundle"), dict): - issues.append(_issue("BUNDLE_RELEASE_CONTRACT_MISSING", impact_scope="GLOBAL", source_refs=["stage2_release"])) - effective_release_lock["bundle"] = { - "mode": { - "DEV_FIXTURE_RELEASE": "STRUCTURAL_FIXTURE", - "SUBSET_CANARY_RELEASE": "SUBSET_CANARY", - "PRODUCTION_RELEASE": "PRODUCTION", - }.get(str(release_lock.get("release_class")), "STRUCTURAL_FIXTURE"), - "handoff_contract_version": "S2_00_STRUCTURED_CONTEXT_HANDOFF_V2", - "context_cohort_policy_id": "S2-CACHE-STRUCTURED-CONTEXT-HANDOFF-V2", - "s2_10_agent_ref": { - "path": "agent_scripts/Stage_2_S2_10.yml", - "sha256": "0" * 64, - }, - "s2_10_llm_binding_ref": { - "path": "deployment/stage2_s2_10_llm_binding.yml", - "sha256": "0" * 64, - }, - "s2_10_agent_sha256": "0" * 64, - "s2_10_llm_binding_sha256": "0" * 64, - "selected_context_refs": [], - "selected_context_ordered_refs_sha256": canonical_digest([]), - "downstream_hybrid_status": "PENDING_REIMPLEMENTATION_AND_RESEAL", - } - bundle_plan = compile_bundle_plan( - cluster_plan, - slices, - effective_release_lock, - stage2_asset_root=asset_root, - ) - cohort_by_cluster = { - cluster_id: row - for row in bundle_plan.get("cohorts", []) - for cluster_id in row.get("cohort_member_cluster_ids", []) - } - for cluster in cluster_plan.get("clusters", []): - cohort = cohort_by_cluster.get(cluster["cluster_id"]) - if cohort is not None: - cluster["bundle_cohort_id"] = cohort["bundle_cohort_id"] - if cohort["cohort_status"] != "EXECUTABLE" and cluster["cluster_status"] == "EXECUTABLE": - cluster["cluster_status"] = "NON_EXECUTABLE" - for code in cohort.get("reason_codes", []): - issues.append( - _issue( - str(code), - impact_scope="CLUSTER", - source_refs=[cluster["cluster_id"], cohort["bundle_cohort_id"]], - ) - ) - cohort_receipt = validate_bundle_release_cohorts(bundle_plan) - final_executable = cohort_receipt["executable_cluster_ids"] - final_residual = sorted( - set(cluster["cluster_id"] for cluster in cluster_plan.get("clusters", [])) - - set(final_executable) - ) - cluster_plan["executable_cluster_ids"] = final_executable - cluster_plan["residual_review_cluster_ids"] = final_residual - executable_scc_ids = { - scc_id - for cluster in cluster_plan.get("clusters", []) - if cluster["cluster_id"] in set(final_executable) - for scc_id in cluster.get("scc_ids", []) - } - cluster_plan["scheduling_waves"] = [ - [scc_id for scc_id in wave if scc_id in executable_scc_ids] - for wave in cluster_plan.get("scheduling_waves", []) - if any(scc_id in executable_scc_ids for scc_id in wave) - ] - route = _route(ingress["source_contract_rows"], cluster_plan, cohort_receipt, issues) - snapshot_end_digest = _verify_snapshot_state(all_snapshots) - if snapshot_end_digest != snapshot_start_digest: - raise IngressError("SOURCE_SNAPSHOT_CHANGED", "source closure changed after the one-read snapshot") - source_rows = list(ingress["source_contract_rows"]) - if signal_all: - source_rows.extend(_signal_source_contract_rows(signal_all, release_lock)) - source_rows.extend(deployment_rows) - source_rows = sorted(source_rows, key=lambda row: row["logical_input_id"]) - issue_codes = sorted(set(str(item["issue_code"]) for item in issues)) - source_counts = { - "declared": len(source_rows), - "observed": sum(row["raw_sha256"] is not None for row in source_rows), - "available": sum(row["scope_technical_disposition"] == "AVAILABLE" for row in source_rows), - "with_issues": sum(row["scope_technical_disposition"] == "AVAILABLE_WITH_ISSUES" for row in source_rows), - "unavailable": sum(row["scope_technical_disposition"] == "UNAVAILABLE" for row in source_rows), - } - intake_report = { - "schema_version": "stage2_s2_00_intake_report.v1", - "run_id": run_id, - "source_counts": source_counts, - "conservation_checks": _compact_conservation_checks(conservation["checks"]), - "review_normalization_receipt": _compact_review_receipt(normalized_reviews), - "minimum_coherent_package_possible": route != "TO_S2_40_STATUS_ONLY", - "scope_summary": [ - { - "scope_technical_disposition": row["scope_technical_disposition"], - "impact_scope": row["impact_scope"], - "scope_refs": row["scope_refs"], - "source_contract_row_refs": row["source_contract_row_refs"], - "reason_codes": row["reason_codes"], - "downstream_allowed_actions": row["downstream_allowed_actions"], - } - for row in source_rows - ], - "issue_codes": issue_codes, - } - manifest = { - "schema_version": "stage2_s2_00_stage1_input_manifest.v1.1", - "run_id": run_id, - "release_class": release_lock["release_class"], - "source_rows": source_rows, - "source_row_order": [row["logical_input_id"] for row in source_rows], - "signal_all_adapter_id": "S2A-SIGNAL-ALL-V1", - "dual_sg01_adapter_id": "S2A-DUAL-SG01-V1", - "hydration_stability_receipt": dict(hydration_stability_receipt), - "snapshot_start_digest": snapshot_start_digest, - "snapshot_end_digest": snapshot_end_digest, - "snapshot_status": "STABLE", - "input_set_digest": input_set_digest, - } - issue_ledger = _build_issue_ledger( - issues, - normalized_reviews, - run_id=run_id, - input_set_digest=input_set_digest, - ) - status = { - "schema_version": "stage2_s2_00_ingress_status.v1.1", - "run_id": run_id, - "run_binding_receipt": run_binding_receipt, - "route": route, - "executable_cluster_ids": final_executable if route != "TO_S2_40_STATUS_ONLY" else [], - "residual_review_cluster_ids": final_residual, - "issue_codes": issue_codes, - "output_barrier": { - "barrier_id": "S2_00_INGRESS_STATUS_BARRIER", - "barrier_path": "ingress/ingress_status.json", - "publish_semantics": "STATUS_LAST_LOGICAL_COMMIT", - "canonical_output_root": f"stage2_runs/by-binding/{run_binding_digest}/", - "branch": "DIAGNOSTIC" if route == "TO_S2_40_STATUS_ONLY" else "NORMAL", - "artifacts": [], - "artifact_set_digest": canonical_digest([]), - "non_status_artifacts_read_back_verified": True, - "written_last": True, - }, - } - if route == "TO_S2_40_STATUS_ONLY": - artifacts: dict[str, Any] = { - "ingress/stage1_input_manifest.json": manifest, - "ingress/intake_report.json": intake_report, - "ingress/technical_diagnostic.json": { - "schema_version": "stage2_s2_00_technical_diagnostic.v1", - "run_id": run_id, - "route": "TO_S2_40_STATUS_ONLY", - "minimum_coherent_package_possible": False, - "reason_codes": issue_codes or ["MINIMUM_COHERENT_PACKAGE_UNAVAILABLE"], - "source_contract_row_refs": sorted( - set( - ref - for item in issues - for ref in item.get("source_contract_row_refs", []) - ) - ) or ["S2_00"], - "context_published": False, - }, - "review/issue_ledger.base.json": issue_ledger, - "ingress/ingress_status.json": status, - } + if schema_error is not None: + row["schema_status"] = "FAIL" + row["reason_codes"].append(schema_error) + row["issue_codes"].append(schema_error) + if schema_error != "SOURCE_SCHEMA_VALIDATION_FAILED": + issues.append(_issue(schema_error, impact_scope="GLOBAL", source_refs=[logical_id])) + else: + row["schema_status"] = "PASS" else: - artifacts = { - "ingress/stage1_input_manifest.json": manifest, - "ingress/intake_report.json": intake_report, - "context/case_context.json": case_context, - "context/evidence_inventory.json": evidence_inventory, - "context/object_registry.json": object_registry, - "context/party_and_title_context.json": party_context, - "context/slot_crosswalk.json": slot_crosswalk, - "context/cluster_plan.json": cluster_plan, - "context/bundle_plan.json": bundle_plan, - "review/issue_ledger.base.json": issue_ledger, - "ingress/ingress_status.json": status, - } - for cluster_id, slice_body in slices.items(): - if cluster_id in set(final_executable): - artifacts[f"context/cluster_slices/{cluster_id}.json"] = slice_body - publish_receipt = publish_atomically( - output_dir, - artifacts, - run_binding_digest=run_binding_digest, - attempt_id=attempt_id, - stage2_asset_root=asset_root, - ) - return {"route": route, "run_binding_digest": run_binding_digest, "publish": publish_receipt} - - - def _execute_structural_fixture(descriptor_path: Path, release_lock: Mapping[str, Any]) -> dict[str, Any]: - snapshot = open_bounded_snapshot(descriptor_path.parent, descriptor_path.name, logical_input_id="fixture_descriptor") - descriptor = load_json_strict(snapshot) - if not isinstance(descriptor, dict) or descriptor.get("mode") != "STRUCTURAL_FIXTURE": - raise IngressError("FIXTURE_DESCRIPTOR_SHAPE", "fixture descriptor must declare STRUCTURAL_FIXTURE") - if release_lock.get("release_class") != "DEV_FIXTURE_RELEASE": - raise IngressError("FIXTURE_RELEASE_CLASS", "structural fixtures require DEV_FIXTURE_RELEASE") - required = {"fixture_id", "source_locator", "source_hashes", "source_semantics", "expected"} - missing = sorted(required - set(descriptor)) - if missing: - raise IngressError("FIXTURE_DESCRIPTOR_REQUIRED_FIELD", "fixture fields are missing", details={"missing": missing}) - if descriptor.get("source_semantics") not in {"EXPLICIT", "DERIVED"}: - raise IngressError("FIXTURE_SOURCE_SEMANTICS", "source_semantics must be EXPLICIT or DERIVED") - payloads = descriptor.get("fixture_payloads", {}) - source_hashes = descriptor.get("source_hashes", {}) - if not isinstance(payloads, dict) or not isinstance(source_hashes, dict): - raise IngressError("FIXTURE_SOURCE_HASH_SHAPE", "fixture payloads and source hashes must be objects") - for logical_id, payload in payloads.items(): - expected_hash = source_hashes.get(logical_id) - if not isinstance(expected_hash, str) or expected_hash != canonical_digest(payload): - raise IngressError( - "FIXTURE_SOURCE_HASH_MISMATCH", - f"fixture payload hash mismatch for {logical_id}", + adapter_errors = _closed_adapter_shape_errors( + document, + logical_id=logical_id, + adapter_id=str(row["adapter_id"]), + required_keys=release_row.get("required_keys", []), + release_lock=release_lock, + ) + if contract_missing: + row["schema_status"] = "UNEVALUABLE" + elif adapter_errors: + code = "ADAPTER_REQUIRED_KEY_MISSING" + row["schema_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append( + _issue( + code, + impact_scope="CLUSTER", + source_refs=[logical_id], + message="; ".join(adapter_errors), + ) ) - return { - "status": "STRUCTURAL_FIXTURE_VALIDATED", - "fixture_id": descriptor["fixture_id"], - "descriptor_sha256": snapshot.raw_sha256, - "canonical_descriptor_sha256": canonical_digest(descriptor), - "published": False, + else: + row["schema_status"] = "PASS" + expected_producer = release_row.get("producer_id") + document_producer = _producer_value(document) + sealed_producer = ( + completion_producers.get(logical_id) + or completion_producers.get(str(contract.get("observed_path"))) + or manifest_producers.get(logical_id) + or manifest_producers.get(str(contract.get("observed_path"))) + ) + if ( + document_producer is not None + and sealed_producer is not None + and document_producer != sealed_producer + ): + code = "PRODUCER_EVIDENCE_CONFLICT" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + observed_producer = document_producer or sealed_producer + if isinstance(expected_producer, str): + if observed_producer is None: + code = "PRODUCER_ID_UNEVALUABLE" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="CLUSTER", source_refs=[logical_id])) + elif not _producer_matches(observed_producer, expected_producer, alias_id, release_lock): + code = "PRODUCER_ID_MISMATCH" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + row["run_identity_ref"] = _identity_ref(document, release_row.get("run_identity_pointer"), logical_id) + row["transaction_identity_ref"] = _identity_ref( + document, + release_row.get("transaction_identity_pointer"), + logical_id, + ) + raw_hash_source = str(release_row.get("raw_hash_source", "NONE")) + if raw_hash_source in {"CASE_RUN_COMPLETION_SEAL", "COMPLETION_SEAL", "COMPLETION_SEAL_ROW"}: + expected_hash = completion_hashes.get(logical_id) or completion_hashes.get(str(contract.get("observed_path"))) + elif raw_hash_source in {"CONTRACT_MANIFEST", "CONTRACT_MANIFEST_ROW", "MANIFEST_ROW"}: + expected_hash = manifest_hashes.get(logical_id) or manifest_hashes.get(str(contract.get("observed_path"))) + elif raw_hash_source in {"COMPLETION_SEAL_OR_CONTRACT_MANIFEST", "SEALED_ROW"}: + expected_hash = ( + completion_hashes.get(logical_id) + or completion_hashes.get(str(contract.get("observed_path"))) + or manifest_hashes.get(logical_id) + or manifest_hashes.get(str(contract.get("observed_path"))) + ) + elif raw_hash_source in {"UNAVAILABLE_DEV", "NONE"}: + expected_hash = None + else: + expected_hash = None + code = "RAW_HASH_SOURCE_UNAPPROVED" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if expected_hash is not None and expected_hash != snapshot.raw_sha256: + code = "RAW_HASH_MISMATCH" + row["seal_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + else: + row["seal_status"] = "PASS" if expected_hash else "UNEVALUABLE" + row["scope_technical_disposition"] = ( + "UNAVAILABLE" + if contract_missing or any(code in row["issue_codes"] for code in {"RAW_HASH_MISMATCH", "SCHEMA_HASH_MISMATCH", "SCHEMA_ID_MISMATCH"}) + else "AVAILABLE" + if not row["issue_codes"] + else "AVAILABLE_WITH_ISSUES" + ) + row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" + rows.append(row) + for identity_kind, field in ( + ("RUN", "run_identity_ref"), + ("TRANSACTION", "transaction_identity_ref"), + ): + observed_values = { + str(row[field]["value"]) + for row in rows + if row[field].get("disposition") == "OBSERVED" and row[field].get("value") is not None } + if len(observed_values) > 1: + code = f"{identity_kind}_IDENTITY_CONFLICT" + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=sorted(observed_values))) + for row in rows: + if row[field].get("disposition") == "OBSERVED": + row["issue_codes"] = sorted(set(row["issue_codes"] + [code])) + row["reason_codes"] = sorted(set(row["reason_codes"] + [code])) + row["scope_technical_disposition"] = "UNAVAILABLE" + return {"documents": documents, "source_contract_rows": rows, "issues": issues} - def _inline_sha256(value: str, *, code: str) -> str: - if not isinstance(value, str) or re.fullmatch(r"[a-f0-9]{64}", value) is None: - raise IngressError(code, "expected one lowercase SHA-256 digest") + def _records_from_signal_document(document: Any) -> list[Any]: + if isinstance(document, list): + return list(document) + if isinstance(document, dict): + for key in ("signals", "records", "items"): + value = document.get(key) + if isinstance(value, list): + return list(value) + return [document] + return [document] + + + def _record_signal_id(record: Any) -> str | None: + if not isinstance(record, dict): + return None + value = record.get("signal_id") + if isinstance(value, str) and value: return value - - - def _inline_relative_path(value: str, *, code: str) -> str: - if not isinstance(value, str) or not value or "\x00" in value or "\\" in value: - raise IngressError(code, "logical path is empty or malformed") - if unicodedata.normalize("NFC", value) != value: - raise IngressError(code, "logical path must already be NFC") - path = PurePosixPath(value) - if path.is_absolute() or any(part in {"", ".", ".."} for part in path.parts): - raise IngressError(code, "logical path must be a contained relative path") - rendered = path.as_posix() - if rendered != value: - raise IngressError(code, "logical path is not canonical") - return rendered - - - def _inline_join(root: str, relative: str, *, code: str) -> str: - safe_root = _inline_relative_path(root, code=code) - safe_relative = _inline_relative_path(relative, code=code) - return _inline_relative_path(f"{safe_root}/{safe_relative}", code=code) - - - def _inline_parse_mcp_payload(raw: bytes, expected_id: int) -> Mapping[str, Any]: - """Parse one JSON or SSE JSON-RPC terminal response with an exact ID.""" - - candidates: list[Any] - try: - candidates = [load_json_strict(raw)] - except IngressError: - try: - text = raw.decode("utf-8", errors="strict") - except UnicodeDecodeError as exc: - raise IngressError("MCP_RESPONSE_UTF8", "MCP response is not strict UTF-8") from exc - events: list[bytes] = [] - data_lines: list[str] = [] - for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n"): - if line == "": - if data_lines: - events.append("\n".join(data_lines).encode("utf-8")) - data_lines = [] - continue - if line.startswith(":") or line.startswith("event:") or line.startswith("id:") or line.startswith("retry:"): - continue - if not line.startswith("data:"): - raise IngressError("MCP_SSE_SHAPE", "unexpected non-data SSE line") - payload = line[5:] - if payload.startswith(" "): - payload = payload[1:] - data_lines.append(payload) - if data_lines: - events.append("\n".join(data_lines).encode("utf-8")) - if not events: - raise IngressError("MCP_RESPONSE_SHAPE", "MCP response contains no JSON terminal event") - candidates = [load_json_strict(event) for event in events] - matching = [ - item - for item in candidates - if isinstance(item, dict) and item.get("id") == expected_id - ] - if len(matching) != 1: - raise IngressError( - "MCP_RESPONSE_ID_MISMATCH", - "MCP response must contain exactly one terminal result with the request ID", - ) - response = matching[0] - if response.get("jsonrpc") != "2.0": - raise IngressError("MCP_JSONRPC_VERSION", "MCP response jsonrpc must equal 2.0") - if response.get("error") is not None: - raise IngressError( - "MCP_JSONRPC_ERROR", - "MCP server returned a JSON-RPC error", - details={"rpc_error": response.get("error")}, - ) - if "result" not in response or not isinstance(response["result"], dict): - raise IngressError("MCP_RESULT_SHAPE", "MCP response result must be an object") - return response - - - def _inline_tool_text(result: Mapping[str, Any], tool_name: str) -> str: - if result.get("isError") is True: - content = result.get("content") - rendered = canonical_json_bytes(content).decode("utf-8", errors="replace") if content is not None else "" - lowered = rendered.lower() - code = ( - "LOCALDOCS_NOT_FOUND" - if any(marker in lowered for marker in ("not found", "does not exist", "no such file")) - else "MCP_TOOL_ERROR" - ) - raise IngressError(code, f"localdocs {tool_name} returned isError=true") - content = result.get("content") - if not isinstance(content, list) or len(content) != 1: - raise IngressError("MCP_CONTENT_CARDINALITY", "MCP tool result must contain exactly one content block") - block = content[0] - if not isinstance(block, dict) or block.get("type") != "text" or not isinstance(block.get("text"), str): - raise IngressError("MCP_CONTENT_SHAPE", "MCP tool result must contain one text block") - return block["text"] - - - def _inline_binary_envelope(text: str, logical_path: str) -> bytes: - value = load_json_strict(text) - if isinstance(value, dict) and "results" in value: - results = value.get("results") - if not isinstance(results, list) or len(results) != 1 or not isinstance(results[0], dict): - raise IngressError("LOCALDOCS_RESULT_CARDINALITY", "binary response must contain one result row") - inner: Any = results[0].get("content", results[0].get("text")) - value = load_json_strict(inner) if isinstance(inner, str) else inner - if not isinstance(value, dict) or not isinstance(value.get("content_base64"), str): - raise IngressError("LOCALDOCS_BINARY_ENVELOPE", "binary response lacks content_base64") - try: - payload = base64.b64decode(value["content_base64"].encode("ascii"), validate=True) - except (UnicodeEncodeError, binascii.Error, ValueError) as exc: - raise IngressError("LOCALDOCS_BASE64_INVALID", "binary response is not strict base64") from exc - declared_size = value.get("byte_length", value.get("size")) - if declared_size is not None and (not isinstance(declared_size, int) or declared_size != len(payload)): - raise IngressError("LOCALDOCS_BYTE_LENGTH_MISMATCH", f"binary length mismatch: {logical_path}") - declared_hash = value.get("sha256") - if declared_hash is not None and declared_hash != hashlib.sha256(payload).hexdigest(): - raise IngressError("LOCALDOCS_HASH_MISMATCH", f"binary hash mismatch: {logical_path}") - return payload - - - class _InlineLocaldocs: - """Minimal user/workspace-bound localdocs JSON-RPC client.""" - - def __init__( - self, - user_hash: str, - workspace_hash: str, - *, - client: Any | None = None, - timeout_seconds: int = 60, - ) -> None: - self.user_hash = _inline_sha256(user_hash, code="USER_CONTEXT_HASH_INVALID") - self.workspace_hash = _inline_sha256( - workspace_hash, - code="WORKSPACE_CONTEXT_HASH_INVALID", - ) - if client is None: - try: - import httpx # type: ignore - except ImportError as exc: - raise IngressError("HTTPX_UNAVAILABLE", "Code Executor must supply httpx==0.28.1") from exc - client = httpx.Client(timeout=timeout_seconds) - self.client = client - self.headers = { - "Content-Type": "application/json", - "Accept": "application/json, text/event-stream", - } - self._message_ids = itertools.count(10) - self._initialized = False - self._session_id: str | None = None - - def close(self) -> None: - close = getattr(self.client, "close", None) - if callable(close): - close() - - def _post(self, body: Mapping[str, Any], expected_id: int | None) -> Mapping[str, Any] | None: - try: - response = self.client.post(LOCALDOCS_URL, json=dict(body), headers=dict(self.headers)) - response.raise_for_status() - except Exception as exc: - raise IngressError("MCP_TRANSPORT_ERROR", "localdocs transport failed") from exc - session_id = response.headers.get("mcp-session-id") - if session_id: - if not isinstance(session_id, str) or not session_id.strip(): - raise IngressError("MCP_SESSION_ID_INVALID", "localdocs returned an invalid session ID") - normalized_session_id = session_id.strip() - if self._session_id is None: - if expected_id != 1: - raise IngressError( - "MCP_SESSION_ID_OUTSIDE_INITIALIZE", - "localdocs first bound a session outside initialize", - ) - self._session_id = normalized_session_id - elif normalized_session_id != self._session_id: - raise IngressError( - "MCP_SESSION_ID_CHANGED", - "localdocs changed the initialized session ID", - ) - self.headers["mcp-session-id"] = self._session_id - if expected_id is None: - return None - raw = response.content if isinstance(response.content, bytes) else bytes(response.content) - return _inline_parse_mcp_payload(raw, expected_id) - - def initialize(self) -> None: - response = self._post( - { - "jsonrpc": "2.0", - "id": 1, - "method": "initialize", - "params": { - "protocolVersion": MCP_PROTOCOL_VERSION, - "capabilities": {}, - "clientInfo": { - "name": INLINE_CLIENT_NAME, - "version": INLINE_CLIENT_VERSION, - "user_id": self.user_hash, - "workspace_id": self.workspace_hash, - }, - }, - }, - 1, - ) - if response is None: - raise IngressError("MCP_INITIALIZE_EMPTY", "localdocs initialize returned no result") - result = response.get("result") - if not isinstance(result, dict) or result.get("protocolVersion") != MCP_PROTOCOL_VERSION: - raise IngressError( - "MCP_PROTOCOL_VERSION_MISMATCH", - "localdocs did not negotiate the requested MCP protocol version", - ) - if self._session_id is None or "mcp-session-id" not in self.headers: - raise IngressError("MCP_SESSION_ID_MISSING", "localdocs initialize did not bind a session ID") - self._post( - {"jsonrpc": "2.0", "method": "notifications/initialized"}, - None, - ) - self._initialized = True - - def call(self, tool_name: str, arguments: Mapping[str, Any]) -> Mapping[str, Any]: - if not self._initialized: - raise IngressError("MCP_NOT_INITIALIZED", "localdocs session is not initialized") - message_id = next(self._message_ids) - response = self._post( - { - "jsonrpc": "2.0", - "id": message_id, - "method": "tools/call", - "params": {"name": tool_name, "arguments": dict(arguments)}, - }, - message_id, - ) - if response is None: - raise IngressError("MCP_TOOL_EMPTY", f"localdocs {tool_name} returned no result") - return response["result"] - - def read_binary(self, logical_path: str) -> bytes: - path = _inline_relative_path(logical_path, code="LOCALDOCS_READ_PATH_INVALID") - result = self.call("read_binary_doc", {"doc_name": path}) - return _inline_binary_envelope(_inline_tool_text(result, "read_binary_doc"), path) - - def read_binary_optional(self, logical_path: str) -> bytes | None: - try: - return self.read_binary(logical_path) - except IngressError as exc: - if exc.code == "LOCALDOCS_NOT_FOUND": - return None - raise - - def write_binary_verified(self, logical_path: str, payload: bytes) -> str: - path = _inline_relative_path(logical_path, code="LOCALDOCS_WRITE_PATH_INVALID") - encoded = base64.b64encode(payload).decode("ascii") - result = self.call( - "write_binary_file", - {"path": path, "content_base64": encoded, "overwrite": True}, - ) - _inline_tool_text(result, "write_binary_file") - observed = self.read_binary(path) - if observed != payload: - raise IngressError("LOCALDOCS_WRITE_READBACK_MISMATCH", f"read-back mismatch: {path}") - return hashlib.sha256(observed).hexdigest() - - - def _inline_validate_request(raw: bytes) -> dict[str, str]: - value = load_json_strict(raw) - if not isinstance(value, dict): - raise IngressError("RUN_REQUEST_SHAPE", "S2_00 request must be an object") - required = { - "schema_version", - "workflow_id", - "request_id", - "attempt_id", - "stage1_run_root_ref", - "stage1_deployment_root_ref", - } - if set(value) != required: - raise IngressError( - "RUN_REQUEST_CLOSED_SHAPE", - "S2_00 request has missing or unknown keys", - details={"missing": sorted(required - set(value)), "extra": sorted(set(value) - required)}, - ) - if value.get("schema_version") != "stage2_s2_00_execution_request.v1": - raise IngressError("RUN_REQUEST_SCHEMA_VERSION", "unsupported S2_00 request schema") - if value.get("workflow_id") != "S2_00": - raise IngressError("WORKFLOW_ID_MISMATCH", "S2_00 request workflow_id must equal S2_00") - request_id = value.get("request_id") - attempt_id = value.get("attempt_id") - if not isinstance(request_id, str) or re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}", request_id) is None: - raise IngressError("REQUEST_ID_INVALID", "request_id contains forbidden characters") - if not isinstance(attempt_id, str) or re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}", attempt_id) is None: - raise IngressError("ATTEMPT_ID_INVALID", "attempt_id contains forbidden characters") - run_root = _inline_relative_path( - value.get("stage1_run_root_ref"), - code="STAGE1_RUN_ROOT_REF_INVALID", - ) - deployment_root = _inline_relative_path( - value.get("stage1_deployment_root_ref"), - code="STAGE1_DEPLOYMENT_ROOT_REF_INVALID", - ) - return { - "schema_version": value["schema_version"], - "workflow_id": value["workflow_id"], - "request_id": request_id, - "attempt_id": attempt_id, - "stage1_run_root_ref": run_root, - "stage1_deployment_root_ref": deployment_root, - } - - - def _inline_bound_ref(row: Mapping[str, Any], *, code: str) -> tuple[str, str]: - path = _inline_relative_path(row.get("path"), code=code) - digest = _inline_sha256(str(row.get("sha256", "")), code=f"{code}_HASH") - if row.get("binding_status") not in {None, "BOUND"}: - raise IngressError(code, f"asset is not release-bound: {path}") - return path, digest - - - def _inline_release_materialization_plan( - request: Mapping[str, str], - release: Mapping[str, Any], - first_read: Callable[[str], bytes], - ) -> tuple[list[tuple[str, str, str]], dict[str, str]]: - """Return exact remote->temporary destinations and expected remote hashes.""" - - destinations: list[tuple[str, str, str]] = [] - expected_hashes: dict[str, str] = {} - seen_destinations: set[tuple[str, str]] = set() - - def add( - remote_path: str, - local_family: str, - relative_path: str, - *, - expected_sha256: str | None = None, - ) -> None: - remote = _inline_relative_path(remote_path, code="HYDRATION_REMOTE_PATH_INVALID") - relative = _inline_relative_path(relative_path, code="HYDRATION_LOCAL_PATH_INVALID") - key = (local_family, relative) - if key in seen_destinations: - raise IngressError("HYDRATION_DESTINATION_DUPLICATE", f"duplicate hydration destination: {relative}") - seen_destinations.add(key) - destinations.append((remote, local_family, relative)) - if expected_sha256 is not None: - digest = _inline_sha256(expected_sha256.lower(), code="HYDRATION_EXPECTED_HASH_INVALID") - prior = expected_hashes.get(remote) - if prior is not None and prior != digest: - raise IngressError("HYDRATION_HASH_CONFLICT", f"conflicting expected hash: {remote}") - expected_hashes[remote] = digest - - release_remote = INLINE_STAGE2_RELEASE_PATH - add( - release_remote, - "stage2_asset", - "manifest/stage2_release.json", - expected_sha256=EXPECTED_STAGE2_RELEASE_SHA256, - ) - - source_rows = _release_stage1_source_rows(release) - if not isinstance(source_rows, list): - raise IngressError("RELEASE_STAGE1_SOURCE_SHAPE", "release stage1_sources must be an array") - fixed_rows = [row for row in source_rows if row.get("logical_input_id") != "signal_payload_family"] - if len(fixed_rows) != 16: - raise IngressError("RELEASE_STAGE1_SOURCE_COUNT", "release must enumerate exactly 16 fixed inputs") - signal_manifest_relative: str | None = None - for row in fixed_rows: - if not isinstance(row, dict) or not isinstance(row.get("path"), str): - raise IngressError("RELEASE_STAGE1_SOURCE_ROW", "fixed source row lacks an exact path") - relative = _inline_relative_path(row["path"], code="STAGE1_SOURCE_PATH_INVALID") - remote = _inline_join(request["stage1_run_root_ref"], relative, code="STAGE1_SOURCE_PATH_INVALID") - add(remote, "stage1_run", relative) - first_read(remote) - if row.get("logical_input_id") == "signal_manifest": - signal_manifest_relative = relative - if signal_manifest_relative is None: - raise IngressError("SIGNAL_MANIFEST_RELEASE_ROW_MISSING", "release lacks signal_manifest") - signal_remote = _inline_join( - request["stage1_run_root_ref"], - signal_manifest_relative, - code="SIGNAL_MANIFEST_PATH_INVALID", - ) - signal_manifest = load_json_strict(first_read(signal_remote)) - if not isinstance(signal_manifest, dict) or not isinstance(signal_manifest.get("files"), list): - raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest files must be an array") - for index, row in enumerate(signal_manifest["files"]): - if not isinstance(row, dict) or not isinstance(row.get("path"), str): - raise IngressError("SIGNAL_FILE_ROW_SHAPE", f"invalid signal row: {index}") - relative_payload = _inline_relative_path(row["path"], code="SIGNAL_FILE_PATH_INVALID") - if relative_payload.startswith("signals/"): - raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "signal manifest path includes signals/") - relative = f"signals/{relative_payload}" - remote = _inline_join(request["stage1_run_root_ref"], relative, code="SIGNAL_FILE_PATH_INVALID") - expected = row.get("file_sha256", row.get("sha256", row.get("raw_sha256"))) - add(remote, "stage1_run", relative, expected_sha256=expected if isinstance(expected, str) else None) - first_read(remote) - - dependencies = release.get("dependency_locks") - stage1_dependency = dependencies.get("stage1") if isinstance(dependencies, dict) else None - closure = stage1_dependency.get("concrete_paths") if isinstance(stage1_dependency, dict) else None - if not isinstance(closure, list) or not closure: - raise IngressError("STAGE1_DEPENDENCY_LOCK_MISSING", "Stage 1 deployment closure is absent") - expected_count = stage1_dependency.get("expected_concrete_path_count") - if expected_count is not None and expected_count != len(closure): - raise IngressError("STAGE1_DEPENDENCY_COUNT_MISMATCH", "Stage 1 closure count differs from release") - for row in closure: - if not isinstance(row, dict): - raise IngressError("STAGE1_DEPENDENCY_ROW_SHAPE", "Stage 1 closure row must be an object") - relative, digest = _inline_bound_ref(row, code="STAGE1_DEPENDENCY_UNBOUND") - remote = _inline_join( - request["stage1_deployment_root_ref"], - relative, - code="STAGE1_DEPLOYMENT_PATH_INVALID", - ) - add(remote, "stage1_deployment", relative, expected_sha256=digest) - first_read(remote) - - contract_ref = stage1_dependency.get("contract_manifest_ref", {}) - if isinstance(contract_ref, dict) and isinstance(contract_ref.get("path"), str): - relative, digest = _inline_bound_ref(contract_ref, code="CONTRACT_MANIFEST_UNBOUND") - remote = _inline_join( - request["stage1_deployment_root_ref"], - relative, - code="CONTRACT_MANIFEST_PATH_INVALID", - ) - add(remote, "stage1_deployment", relative, expected_sha256=digest) - first_read(remote) - completion_ref = stage1_dependency.get("completion_seal_ref", {}) - if isinstance(completion_ref, dict) and completion_ref.get("path") not in {None, "PENDING_SEQUENTIAL_BIND"}: - relative, digest = _inline_bound_ref(completion_ref, code="COMPLETION_SEAL_UNBOUND") - remote = _inline_join(request["stage1_run_root_ref"], relative, code="COMPLETION_SEAL_PATH_INVALID") - add(remote, "stage1_run", relative, expected_sha256=digest) - first_read(remote) - - manifest_ref = release.get("module_manifest_ref") - if not isinstance(manifest_ref, dict): - raise IngressError("MODULE_MANIFEST_REF_MISSING", "release lacks module_manifest_ref") - manifest_relative, manifest_digest = _inline_bound_ref(manifest_ref, code="MODULE_MANIFEST_UNBOUND") - manifest_remote = _inline_join(INLINE_STAGE2_ASSET_ROOT, manifest_relative, code="MODULE_MANIFEST_PATH_INVALID") - add(manifest_remote, "stage2_asset", manifest_relative, expected_sha256=manifest_digest) - module_manifest = load_json_strict(first_read(manifest_remote)) - modules = module_manifest.get("modules") if isinstance(module_manifest, dict) else None - if not isinstance(modules, list): - raise IngressError("MODULE_MANIFEST_SHAPE", "module manifest lacks modules array") - found_schema_ids: set[str] = set() - for row in modules: - if not isinstance(row, dict) or row.get("module_id") not in INLINE_SCHEMA_MODULE_IDS: - continue - relative, digest = _inline_bound_ref(row, code="STAGE2_SCHEMA_UNBOUND") - remote = _inline_join(INLINE_STAGE2_ASSET_ROOT, relative, code="STAGE2_SCHEMA_PATH_INVALID") - add(remote, "stage2_asset", relative, expected_sha256=digest) - first_read(remote) - found_schema_ids.add(str(row["module_id"])) - if found_schema_ids != set(INLINE_SCHEMA_MODULE_IDS): - raise IngressError( - "STAGE2_SCHEMA_CLOSURE_INCOMPLETE", - "module manifest does not bind all S2_00 schemas", - details={"missing": sorted(set(INLINE_SCHEMA_MODULE_IDS) - found_schema_ids)}, - ) - - stage2_direct = dependencies.get("stage2_direct") if isinstance(dependencies, dict) else None - if not isinstance(stage2_direct, list): - raise IngressError("STAGE2_DIRECT_LOCK_MISSING", "release lacks Stage 2 direct closure") - for row in stage2_direct: - if not isinstance(row, dict): - raise IngressError("STAGE2_DIRECT_ROW_SHAPE", "Stage 2 direct row must be an object") - relative, digest = _inline_bound_ref(row, code="STAGE2_DIRECT_UNBOUND") - remote = _inline_join(INLINE_STAGE2_ASSET_ROOT, relative, code="STAGE2_DIRECT_PATH_INVALID") - add(remote, "stage2_asset", relative, expected_sha256=digest) - first_read(remote) - return destinations, expected_hashes - - - def _inline_hydrate( - localdocs: _InlineLocaldocs, - temp_root: Path, - ) -> tuple[dict[str, str], Mapping[str, Any], Path, Path, Path, dict[str, Any]]: - first_pass: dict[str, bytes] = {} - first_pass_bytes = 0 - - request_raw = localdocs.read_binary(INLINE_REQUEST_PATH) - if len(request_raw) > MAX_FILE_BYTES: - raise IngressError("SOURCE_SIZE_LIMIT", "S2_00 request exceeds the global pre-parse file limit") - first_pass[INLINE_REQUEST_PATH] = request_raw - first_pass_bytes += len(request_raw) - request = _inline_validate_request(request_raw) - - release_raw = localdocs.read_binary(INLINE_STAGE2_RELEASE_PATH) - if len(release_raw) > MAX_FILE_BYTES: - raise IngressError("SOURCE_SIZE_LIMIT", "Stage 2 release exceeds the global pre-parse file limit") - first_pass[INLINE_STAGE2_RELEASE_PATH] = release_raw - first_pass_bytes += len(release_raw) - if first_pass_bytes > MAX_RUN_BYTES: - raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "initial request and release exceed the global pre-parse budget") - if EXPECTED_STAGE2_RELEASE_SHA256 == "0" * 64: - raise IngressError( - "EXPECTED_STAGE2_RELEASE_SHA256_UNBOUND", - "build-time Stage 2 release pin has not been bound", - ) - expected_release = _inline_sha256( - EXPECTED_STAGE2_RELEASE_SHA256, - code="EXPECTED_STAGE2_RELEASE_SHA256_INVALID", - ) - if hashlib.sha256(release_raw).hexdigest() != expected_release: - raise IngressError("STAGE2_RELEASE_PIN_MISMATCH", "Stage 2 release bytes differ from the build-time pin") - release = load_json_strict(release_raw) - if not isinstance(release, dict): - raise IngressError("RELEASE_LOCK_SHAPE", "Stage 2 release must be an object") - limits = release.get("limits") if isinstance(release.get("limits"), dict) else {} - max_file = int(limits.get("max_file_bytes", MAX_FILE_BYTES)) - max_total = int(limits.get("max_run_bytes", MAX_RUN_BYTES)) - max_paths = int(limits.get("max_hydration_paths", 1024)) - if max_file <= 0 or max_file > MAX_FILE_BYTES: - raise IngressError("RELEASE_FILE_LIMIT_INVALID", "release max_file_bytes exceeds the build-time ceiling") - if max_total <= 0 or max_total > MAX_RUN_BYTES: - raise IngressError("RELEASE_RUN_LIMIT_INVALID", "release max_run_bytes exceeds the build-time ceiling") - if max_paths <= 0 or max_paths > 1024: - raise IngressError("RELEASE_PATH_LIMIT_INVALID", "release max_hydration_paths exceeds the build-time ceiling") - if len(request_raw) > max_file or len(release_raw) > max_file: - raise IngressError("SOURCE_SIZE_LIMIT", "initial request or release exceeds the release-bound file limit") - if first_pass_bytes > max_total: - raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "initial request and release exceed the release-bound budget") - - def first_read(path: str) -> bytes: - nonlocal first_pass_bytes - safe = _inline_relative_path(path, code="HYDRATION_REMOTE_PATH_INVALID") - if safe not in first_pass: - payload = localdocs.read_binary(safe) - if len(payload) > max_file: - raise IngressError("SOURCE_SIZE_LIMIT", f"remote source exceeds file limit: {safe}") - first_pass_bytes += len(payload) - if first_pass_bytes > max_total: - raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "first-pass hydration exceeds release budget") - first_pass[safe] = payload - return first_pass[safe] - - destinations, expected_hashes = _inline_release_materialization_plan( - request, - release, - first_read, - ) - if len(first_pass) > max_paths: - raise IngressError("HYDRATION_PATH_COUNT_LIMIT", "hydration path count exceeds release budget") - for remote, expected in expected_hashes.items(): - observed = hashlib.sha256(first_read(remote)).hexdigest() - if observed != expected: - raise IngressError("HYDRATION_BOUND_HASH_MISMATCH", f"release-bound hash mismatch: {remote}") - - second_pass_bytes = 0 - second_pass: dict[str, bytes] = {} - for remote in sorted(first_pass): - payload = localdocs.read_binary(remote) - second_pass_bytes += len(payload) - if len(payload) > max_file or second_pass_bytes > max_total: - raise IngressError("TWO_PASS_READ_BUDGET", "second-pass hydration exceeds release budget") - if payload != first_pass[remote] or hashlib.sha256(payload).digest() != hashlib.sha256(first_pass[remote]).digest(): - raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"remote source changed between bounded reads: {remote}") - second_pass[remote] = payload - - pass_1_hash_rows = [ - [path, hashlib.sha256(first_pass[path]).hexdigest()] - for path in sorted(first_pass) - ] - pass_2_hash_rows = [ - [path, hashlib.sha256(second_pass[path]).hexdigest()] - for path in sorted(second_pass) - ] - hydration_receipt = { - "schema_version": "stage2_s2_00_two_pass_hydration_receipt.v1", - "transport": "localdocs.read_binary_doc", - "read_policy": "BOUNDED_TWO_PASS_BINARY_RAW_HASH_MAP_EQUALITY", - "read_pass_count": 2, - "max_read_passes": 2, - "pass_1_raw_hash_map_digest": canonical_digest(pass_1_hash_rows), - "pass_2_raw_hash_map_digest": canonical_digest(pass_2_hash_rows), - "source_receipts": [ - { - "logical_input_id": path, - "logical_path": path, - "pass_1_raw_sha256": hashlib.sha256(first_pass[path]).hexdigest(), - "pass_2_raw_sha256": hashlib.sha256(second_pass[path]).hexdigest(), - "pass_1_byte_length": len(first_pass[path]), - "pass_2_byte_length": len(second_pass[path]), - "stability_status": "STABLE", - } - for path in sorted(first_pass) - ], - "stability_status": "STABLE", - } - - roots = { - "stage1_run": temp_root / "stage1_run", - "stage1_deployment": temp_root / "stage1_deployment", - "stage2_asset": temp_root / "stage2_asset", - } - for root in roots.values(): - root.mkdir(parents=True, exist_ok=False) - for remote, family, relative in sorted(destinations): - target = roots[family] / PurePosixPath(relative) - target.parent.mkdir(parents=True, exist_ok=True) - if target.exists(): - raise IngressError("HYDRATION_DESTINATION_EXISTS", f"duplicate materialization: {relative}") - target.write_bytes(first_pass[remote]) - return ( - request, - release, - roots["stage1_run"], - roots["stage1_deployment"], - roots["stage2_asset"], - hydration_receipt, - ) - - - def _inline_output_files(output_root: Path, run_binding_digest: str) -> dict[str, bytes]: - if not output_root.is_dir() or output_root.is_symlink(): - raise IngressError("LOCAL_OUTPUT_TREE_MISSING", "pure core did not produce an output tree") - files: dict[str, bytes] = {} - for path in sorted(output_root.rglob("*")): - if path.is_symlink(): - raise IngressError("LOCAL_OUTPUT_SYMLINK", "pure core output contains a symlink") - if not path.is_file(): - continue - relative = _safe_relative_path(path.relative_to(output_root).as_posix()).as_posix() - files[relative] = path.read_bytes() - status_raw = files.get("ingress/ingress_status.json") - if status_raw is None: - raise IngressError("OUTPUT_BARRIER_MISSING", "pure core output lacks ingress_status") - status = load_json_strict(status_raw) - binding = status.get("run_binding_receipt") if isinstance(status, dict) else None - if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: - raise IngressError("RUN_BINDING_RECEIPT_MISMATCH", "pure core status has a different binding") - barrier = status.get("output_barrier") - rows = barrier.get("artifacts") if isinstance(barrier, dict) else None - if not isinstance(rows, list) or barrier.get("written_last") is not True: - raise IngressError("OUTPUT_BARRIER_INVALID", "pure core output barrier is incomplete") - expected_paths = {"ingress/ingress_status.json"} - for row in rows: - if not isinstance(row, dict) or not isinstance(row.get("path"), str): - raise IngressError("OUTPUT_BARRIER_ROW_SHAPE", "output barrier row is malformed") - relative = _safe_relative_path(row["path"]).as_posix() - payload = files.get(relative) - if payload is None or hashlib.sha256(payload).hexdigest() != row.get("raw_sha256"): - raise IngressError("OUTPUT_BARRIER_HASH_MISMATCH", f"output barrier mismatch: {relative}") - expected_paths.add(relative) - if set(files) != expected_paths: - raise IngressError("OUTPUT_BARRIER_SET_MISMATCH", "output tree differs from its barrier set") - if barrier.get("artifact_set_digest") != canonical_digest(rows): - raise IngressError("OUTPUT_BARRIER_DIGEST_MISMATCH", "output barrier row digest is invalid") - return files - - - def _inline_verify_existing_remote( - localdocs: _InlineLocaldocs, - output_root: str, - files: Mapping[str, bytes], - run_binding_digest: str, - ) -> bool: - status_path = _inline_join(output_root, "ingress/ingress_status.json", code="OUTPUT_PATH_INVALID") - status_raw = localdocs.read_binary_optional(status_path) - if status_raw is None: - return False - status = load_json_strict(status_raw) - binding = status.get("run_binding_receipt") if isinstance(status, dict) else None - if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: - raise IngressError("RUN_ID_BINDING_CONFLICT", "existing remote barrier has a different binding") - if status_raw != files.get("ingress/ingress_status.json"): - raise IngressError("IDEMPOTENT_STATUS_MISMATCH", "existing remote status is not byte-identical") - for relative, expected in sorted(files.items()): - observed = localdocs.read_binary(_inline_join(output_root, relative, code="OUTPUT_PATH_INVALID")) - if observed != expected: - raise IngressError("IDEMPOTENT_ARTIFACT_MISMATCH", f"existing artifact differs: {relative}") - return True - - - def _inline_publish_remote( - localdocs: _InlineLocaldocs, - files: Mapping[str, bytes], - run_binding_digest: str, - ) -> dict[str, Any]: - output_root = f"stage2_runs/by-binding/{run_binding_digest}" - status_raw = files["ingress/ingress_status.json"] - status = load_json_strict(status_raw) - barrier = status.get("output_barrier") if isinstance(status, dict) else None - branch = barrier.get("branch") if isinstance(barrier, dict) else None - if branch not in {"NORMAL", "DIAGNOSTIC"}: - raise IngressError("OUTPUT_BARRIER_BRANCH_INVALID", "pure core output barrier lacks a valid branch") - artifact_rows = [ - { - "logical_artifact_id": relative, - "path": relative, - "schema_id": _output_schema_id(relative), - "raw_sha256": hashlib.sha256(payload).hexdigest(), - "byte_length": len(payload), - } - for relative, payload in sorted(files.items()) - if relative != "ingress/ingress_status.json" - ] - if _inline_verify_existing_remote(localdocs, output_root, files, run_binding_digest): - publication_status = "IDEMPOTENT_SUCCESS" - else: - for relative in sorted(path for path in files if path != "ingress/ingress_status.json"): - localdocs.write_binary_verified( - _inline_join(output_root, relative, code="OUTPUT_PATH_INVALID"), - files[relative], - ) - status_path = _inline_join(output_root, "ingress/ingress_status.json", code="OUTPUT_PATH_INVALID") - localdocs.write_binary_verified(status_path, status_raw) - if localdocs.read_binary(status_path) != status_raw: - raise IngressError("OUTPUT_BARRIER_READBACK_MISMATCH", "remote ingress_status read-back failed") - publication_status = "PUBLISHED_STATUS_LAST" - return { - "schema_version": "stage2_logical_publish_receipt.v1", - "barrier_id": "S2_00_INGRESS_STATUS_BARRIER", - "barrier_path": "ingress/ingress_status.json", - "publish_semantics": "STATUS_LAST_LOGICAL_COMMIT", - "canonical_output_root": f"{output_root}/", - "run_binding_digest": run_binding_digest, - "branch": branch, - "artifacts": artifact_rows, - "artifact_set_digest": canonical_digest(artifact_rows), - "barrier_raw_sha256": hashlib.sha256(status_raw).hexdigest(), - "non_status_artifacts_read_back_verified": True, - "barrier_written_last": True, - "downstream_consumption_allowed": True, - "publication_status": publication_status, - } - - - def _inline_context_is_bound() -> bool: - return ( - re.fullmatch(r"[a-f0-9]{64}", INLINE_USER_HASH) is not None - and re.fullmatch(r"[a-f0-9]{64}", INLINE_WORKSPACE_HASH) is not None - ) - - - def run_inline_mcp(*, client: Any | None = None) -> int: - """Execute one complete MCP Code Executor S2_00 task and emit one JSON receipt.""" - - localdocs: _InlineLocaldocs | None = None - receipt: dict[str, Any] - exit_code = 0 - try: - with contextlib.redirect_stdout(io.StringIO()): - localdocs = _InlineLocaldocs( - INLINE_USER_HASH, - INLINE_WORKSPACE_HASH, - client=client, - ) - try: - localdocs.initialize() - with tempfile.TemporaryDirectory(prefix="liti-s2-00-") as directory: - temp_root = Path(directory) - ( - request, - _release, - stage1_root, - deployment_root, - asset_root, - hydration_receipt, - ) = _inline_hydrate(localdocs, temp_root) - release_lock = load_release_lock(asset_root / "manifest" / "stage2_release.json") - contract_manifest = _load_bound_contract_manifest(deployment_root, release_lock) - local_output = temp_root / "core_output" - result = execute_ingress( - stage1_root, - release_lock, - output_dir=local_output, - attempt_id=request["attempt_id"], - run_id=request["request_id"], - user_context_sha256=INLINE_USER_HASH, - workspace_context_sha256=INLINE_WORKSPACE_HASH, - stage1_deployment_root=deployment_root, - stage2_asset_root=asset_root, - contract_manifest=contract_manifest, - hydration_stability_receipt=hydration_receipt, - ) - run_binding_digest = _inline_sha256( - str(result.get("run_binding_digest", "")), - code="CORE_RUN_BINDING_INVALID", - ) - files = _inline_output_files(local_output, run_binding_digest) - publication = _inline_publish_remote(localdocs, files, run_binding_digest) - receipt = { - "schema_version": "stage2_s2_00_inner_receipt.v1", - "workflow_id": "S2_00", - "ok": True, - "status": ( - "DIAGNOSTIC_PUBLISHED" - if result["route"] == "TO_S2_40_STATUS_ONLY" - else "SUCCEEDED" - ), - "request_id": request["request_id"], - "route": result["route"], - "run_id": canonical_run_id(run_binding_digest), - "run_binding_digest": run_binding_digest, - "expected_release_sha256": EXPECTED_STAGE2_RELEASE_SHA256, - "algorithm_digest": ALGORITHM_SEMANTIC_DIGEST, - "artifact_set_digest": publication["artifact_set_digest"], - "ingress_status_sha256": publication["barrier_raw_sha256"], - "logical_publish_receipt": publication, - } - finally: - if localdocs is not None: - localdocs.close() - except Exception as exc: - error = ( - exc.as_dict() - if isinstance(exc, IngressError) - else {"code": "INLINE_RUNTIME_ERROR", "message": str(exc)} - ) - receipt = { - "schema_version": "stage2_s2_00_inner_receipt.v1", - "workflow_id": "S2_00", - "ok": False, - "status": "FAILED_NO_BARRIER", - "expected_release_sha256": EXPECTED_STAGE2_RELEASE_SHA256, - "error": error, - } - exit_code = 2 - sys.stdout.buffer.write(canonical_json_bytes(receipt)) - return exit_code - - - class _FixedArgvParser(argparse.ArgumentParser): - def error(self, message: str) -> None: - raise IngressError("CLI_ARGUMENT_ERROR", message) - - - def _parser() -> argparse.ArgumentParser: - parser = _FixedArgvParser( - description="S2_00 deterministic Stage 1 ingress", - add_help=False, - ) - parser.add_argument("--workflow-id", required=True) - parser.add_argument("--release-ref", required=True) - parser.add_argument("--run-id", required=True) - parser.add_argument("--attempt-id", required=True) - parser.add_argument("--user-context-sha256", required=True) - parser.add_argument("--workspace-context-sha256", required=True) - parser.add_argument("--stage1-run-root-ref", required=True) - parser.add_argument("--stage1-deployment-root-ref", required=True) - return parser - - - def _project_root_from_runtime() -> Path: - for candidate in Path(__file__).resolve().parents: - if candidate.name == "Case_02_Comparison_Research": - return candidate - raise IngressError("PROJECT_ROOT_NOT_FOUND", "runtime is not located below the canonical project root") - - - def _resolve_contained_ref( - approved_root: Path, - ref: str, - *, - code: str, - ) -> Path: - if not isinstance(ref, str) or not ref or "\x00" in ref: - raise IngressError(code, "workspace reference is empty or malformed") - candidate = Path(ref) - if not candidate.is_absolute(): - candidate = approved_root / candidate - if candidate.is_symlink(): - raise IngressError(code, "workspace reference may not be a symlink") - approved_resolved = approved_root.resolve(strict=True) - try: - lexical_relative = candidate.relative_to(approved_root) - except ValueError: - lexical_relative = None - if lexical_relative is not None: - _assert_no_symlink_components(approved_resolved, PurePosixPath(lexical_relative.as_posix())) - resolved = candidate.resolve(strict=True) - try: - resolved.relative_to(approved_resolved) - except ValueError as exc: - raise IngressError(code, "workspace reference escaped its approved root") from exc - return resolved - - - def _load_bound_contract_manifest( - stage1_deployment_root: Path, - release_lock: Mapping[str, Any], - ) -> Mapping[str, Any] | None: - dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) - ref = dependency.get("contract_manifest_ref", {}) if isinstance(dependency, dict) else {} - if not isinstance(ref, dict) or not isinstance(ref.get("path"), str): - return None - expected_hash = ref.get("sha256") - if not isinstance(expected_hash, str) or not re.fullmatch(r"[A-Fa-f0-9]{64}", expected_hash): - raise IngressError("CONTRACT_MANIFEST_UNBOUND", "contract manifest hash is not bound") + for wrapper in ("domain_activation_manifest", "payload", "data"): + nested = record.get(wrapper) + if isinstance(nested, dict) and isinstance(nested.get("signal_id"), str): + return nested["signal_id"] + return None + + + def expand_stage2_signal_all( + stage1_run_root: str | os.PathLike[str], + signal_manifest: Mapping[str, Any], + *, + max_file_bytes: int = MAX_FILE_BYTES, + max_total_bytes: int = MAX_RUN_BYTES, + signal_registry: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Expand Stage 2 ALL while separating semantic and integrity-only universes.""" + + downstream = signal_manifest.get("downstream_read_sets", {}) + stage2 = downstream.get("stage2", []) if isinstance(downstream, dict) else [] + if stage2 != ["ALL"]: + raise IngressError("SIGNAL_ALL_CONTRACT", "downstream_read_sets.stage2 must equal ['ALL']") + files = signal_manifest.get("files") + if not isinstance(files, list): + raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest files must be an array") + transaction_id = str(signal_manifest.get("manifest_transaction_id", signal_manifest.get("transaction_id", "MISSING"))) + file_rows: list[dict[str, Any]] = [] + semantic_rows: list[dict[str, Any]] = [] + integrity_rows: list[dict[str, Any]] = [] + occurrences: list[dict[str, Any]] = [] + payload_snapshots: list[Snapshot] = [] + issues: list[dict[str, Any]] = [] + path_counter: Counter[str] = Counter() + parsed_documents: dict[str, Any] = {} + aggregate_bytes = 0 + registry_entries = { + str(row.get("file")): row + for row in (signal_registry or {}).get("entries", []) + if isinstance(row, dict) and isinstance(row.get("file"), str) + } + compatibility_files = { + str(path) + for path in (signal_registry or {}).get("compatibility_views", []) + if isinstance(path, str) + } + domain_envelope_schema = (signal_registry or {}).get("domain_envelope") + observed_registry_files: set[str] = set() + for index, entry in enumerate(files): + if not isinstance(entry, dict) or not isinstance(entry.get("path"), str): + raise IngressError("SIGNAL_FILE_ROW_SHAPE", f"invalid signal file row at index {index}") + relative_payload = _safe_relative_path(entry["path"]).as_posix() + if relative_payload.startswith("signals/"): + raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "manifest file path must not include signals/ prefix") + physical = f"signals/{relative_payload}" snapshot = open_bounded_snapshot( - stage1_deployment_root, - ref["path"], - logical_input_id="stage1_contract_manifest", + stage1_run_root, + physical, + logical_input_id=f"signal_file:{index}", + max_bytes=max_file_bytes, ) - if snapshot.raw_sha256 != expected_hash: - raise IngressError("CONTRACT_MANIFEST_HASH_MISMATCH", "contract manifest raw hash differs from release") - value = load_json_strict(snapshot) - if not isinstance(value, dict): - raise IngressError("CONTRACT_MANIFEST_SHAPE", "contract manifest must be an object") - return value + document = load_json_strict(snapshot) + payload_snapshots.append(snapshot) + aggregate_bytes += snapshot.byte_length + if aggregate_bytes > max_total_bytes: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "signal ALL payloads exceed remaining run byte budget") + parsed_documents[relative_payload] = document + kind = entry.get("kind", "canonical") + if kind not in SEMANTIC_SIGNAL_KINDS | {"compatibility_view"}: + raise IngressError("SIGNAL_KIND_UNAPPROVED", f"unapproved signal file kind: {kind}") + semantic = kind in SEMANTIC_SIGNAL_KINDS + expected_hash = entry.get( + "file_sha256", entry.get("sha256", entry.get("raw_sha256")) + ) + row = { + "manifest_index": index, + "file_path": relative_payload, + "physical_path": physical, + "kind": kind, + "raw_sha256": snapshot.raw_sha256, + "byte_length": snapshot.byte_length, + "semantic": semantic, + "manifest_declared_record_count": entry.get("record_count"), + } + if expected_hash is not None and expected_hash != snapshot.raw_sha256: + row["hash_status"] = "FAIL" + issues.append(_issue("SIGNAL_FILE_HASH_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + else: + row["hash_status"] = "PASS" if expected_hash else "UNEVALUABLE" + records = _records_from_signal_document(document) + row["observed_record_count"] = len(records) + declared_count = entry.get("record_count") + if isinstance(declared_count, int) and declared_count != len(records): + row["record_count_status"] = "FAIL" + issues.append(_issue("SIGNAL_RECORD_COUNT_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + else: + row["record_count_status"] = "PASS" if isinstance(declared_count, int) else "UNEVALUABLE" + registry_row = registry_entries.get(relative_payload) + if kind == "canonical": + if signal_registry is not None and registry_row is None: + issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif registry_row is not None: + observed_registry_files.add(relative_payload) + declared_schema = entry.get("schema", entry.get("schema_path")) + if declared_schema is not None and declared_schema != registry_row.get("schema"): + issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif kind == "compatibility_view": + if relative_payload in registry_entries: + issues.append(_issue("SIGNAL_COMPATIBILITY_SUBSTITUTION", impact_scope="SIGNAL", source_refs=[physical])) + if signal_registry is not None and relative_payload not in compatibility_files: + issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif kind == "domain_signal": + declared_schema = entry.get("schema", entry.get("schema_path")) + if signal_registry is not None and declared_schema not in {None, domain_envelope_schema}: + issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + file_rows.append(row) + path_counter[relative_payload] += 1 + if semantic: + semantic_rows.append(row) + for record_ordinal, record in enumerate(records): + signal_id = _record_signal_id(record) + occurrence_key = [transaction_id, relative_payload, record_ordinal, signal_id] + occurrences.append( + { + "occurrence_key": occurrence_key, + "occurrence_ref": f"SIGO-{canonical_digest(occurrence_key)[:24]}", + "manifest_transaction_id": transaction_id, + "file_path": relative_payload, + "record_ordinal": record_ordinal, + "signal_id": signal_id, + "disposition": "UNMAPPED" if signal_id is None else "UNUSED", + "binding_refs": [], + "raw_record_sha256": canonical_digest(record), + "record": record, + } + ) + else: + integrity_rows.append(row) + duplicates = sorted(path for path, count in path_counter.items() if count > 1) + if duplicates: + issues.append(_issue("SIGNAL_ALL_DUPLICATE_FILE_ROW", impact_scope="SIGNAL", source_refs=duplicates)) + manifest_counter = Counter((i, row["file_path"], row["kind"]) for i, row in enumerate(file_rows)) + partition_counter = Counter((row["manifest_index"], row["file_path"], row["kind"]) for row in semantic_rows + integrity_rows) + missing_registry_files = sorted(set(registry_entries) - observed_registry_files) if signal_registry is not None else [] + if missing_registry_files: + issues.append( + _issue( + "SIGNAL_REGISTRY_COVERAGE_MISMATCH", + impact_scope="SIGNAL", + source_refs=[f"signals/{path}" for path in missing_registry_files], + ) + ) + file_conservation = ( + manifest_counter == partition_counter + and not duplicates + and not missing_registry_files + and not any(row["hash_status"] == "FAIL" or row["record_count_status"] == "FAIL" for row in file_rows) + ) + record_counter = Counter(tuple(row["occurrence_key"]) for row in occurrences) + partitioned_record_counter = Counter( + tuple(row["occurrence_key"]) + for row in occurrences + if row["disposition"] in {"USED", "UNUSED", "UNMAPPED"} + ) + record_conservation = record_counter == partitioned_record_counter + return { + "manifest_transaction_id": transaction_id, + "ordered_file_rows": file_rows, + "semantic_file_rows": semantic_rows, + "integrity_only_file_rows": integrity_rows, + "record_occurrences": occurrences, + "used_record_occurrences": [], + "unused_record_occurrences": [row for row in occurrences if row["disposition"] == "UNUSED"], + "unmapped_record_occurrences": [row for row in occurrences if row["disposition"] == "UNMAPPED"], + "_parsed_documents_by_path": parsed_documents, + "_payload_snapshots": payload_snapshots, + "file_conservation_pass": file_conservation, + "record_conservation_pass": record_conservation, + "aggregate_payload_bytes": aggregate_bytes, + "issues": issues, + } - def main(argv: Sequence[str] | None = None) -> int: - """CLI entrypoint. Stdout contains exactly one canonical JSON object.""" + def _collect_values_for_keys(value: Any, keys: frozenset[str]) -> set[str]: + result: set[str] = set() + stack = [value] + while stack: + current = stack.pop() + if isinstance(current, dict): + for key, child in current.items(): + if key in keys: + if isinstance(child, list): + result.update(str(item) for item in child if item is not None) + elif child is not None: + result.add(str(child)) + stack.append(child) + elif isinstance(current, list): + stack.extend(current) + return result + + def bind_signal_occurrences(signal_all: MutableMapping[str, Any], documents: Mapping[str, Any]) -> dict[str, Any]: + """Bind each semantic signal occurrence to explicit Stage 1 references without deduplication.""" + + explicit_signal_ids = _collect_values_for_keys( + documents, + frozenset({"signal_id", "signal_ids", "signal_refs", "emitted_signal_ids", "required_signal_ids"}), + ) + known_refs = { + "fact_id": _collect_values_for_keys(documents.get("fact_ledger_base"), frozenset({"fact_id"})), + "source_bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "source_bo_id", "source_bo_ids"})), + "bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "bo_id"})), + "structure_id": _collect_values_for_keys(documents.get("legal_effect_structures"), frozenset({"structure_id"})), + "domain_id": _collect_values_for_keys(documents, frozenset({"domain_id", "domain_ids", "active_domain_ids"})), + "evidence_id": _collect_values_for_keys(documents.get("evidence_indexed"), frozenset({"evidence_id", "id"})), + "event_id": _collect_values_for_keys(documents.get("evidence_event_candidates"), frozenset({"event_id", "id"})), + } + link_keys = { + "fact_id": ("fact_id", "fact_ids"), + "source_bo_id": ("source_bo_id", "source_bo_ids"), + "bo_id": ("bo_id", "bo_ids"), + "structure_id": ("structure_id", "structure_ids"), + "domain_id": ("domain_id", "domain_ids"), + "evidence_id": ("evidence_id", "evidence_ids"), + "event_id": ("event_id", "event_ids"), + } + for occurrence in signal_all.get("record_occurrences", []): + signal_id = occurrence.get("signal_id") + record = occurrence.get("record") + bindings: set[str] = set() + if isinstance(signal_id, str) and signal_id in explicit_signal_ids: + bindings.add(f"signal_id:{signal_id}") + for ref_kind, candidate_keys in link_keys.items(): + observed = _collect_values_for_keys(record, frozenset(candidate_keys)) + for ref in sorted(observed & known_refs[ref_kind]): + bindings.add(f"{ref_kind}:{ref}") + if not isinstance(signal_id, str) or not signal_id: + occurrence["disposition"] = "UNMAPPED" + elif bindings: + occurrence["disposition"] = "USED" + else: + occurrence["disposition"] = "UNUSED" + occurrence["binding_refs"] = sorted(bindings) + for disposition, key in ( + ("USED", "used_record_occurrences"), + ("UNUSED", "unused_record_occurrences"), + ("UNMAPPED", "unmapped_record_occurrences"), + ): + signal_all[key] = [ + row for row in signal_all.get("record_occurrences", []) if row.get("disposition") == disposition + ] + source_counter = Counter(tuple(row["occurrence_key"]) for row in signal_all.get("record_occurrences", [])) + partition_counter = Counter( + tuple(row["occurrence_key"]) + for key in ("used_record_occurrences", "unused_record_occurrences", "unmapped_record_occurrences") + for row in signal_all[key] + ) + signal_all["record_conservation_pass"] = source_counter == partition_counter + return dict(signal_all) + + + def _activation_payload(value: Mapping[str, Any]) -> Mapping[str, Any]: + for key in ("domain_activation_manifest", "activation", "payload", "data"): + nested = value.get(key) + if isinstance(nested, dict) and any(field in nested for field in SG01_PROJECTION_FIELDS): + return nested + return value + + + def verify_activation_projection( + routing_activation: Mapping[str, Any], + signal_activation: Mapping[str, Any], + *, + routing_raw_sha256: str | None = None, + signal_raw_sha256: str | None = None, + ) -> dict[str, Any]: + """Compare approved semantic SG-01 projection while retaining both raw hashes.""" + + left = _activation_payload(routing_activation) + right = _activation_payload(signal_activation) + missing_left = [field for field in SG01_PROJECTION_FIELDS if field not in left] + missing_right = [field for field in SG01_PROJECTION_FIELDS if field not in right] + if missing_left or missing_right: + raise IngressError( + "SG01_PROJECTION_SHAPE", + "both activation artifacts must expose the complete approved 17-field projection", + details={"routing_missing": missing_left, "signal_missing": missing_right}, + ) + + def project(value: Mapping[str, Any]) -> dict[str, Any]: + result: dict[str, Any] = {} + for field in SG01_PROJECTION_FIELDS: + child = value[field] + if field in SG01_SET_FIELDS: + if not isinstance(child, list): + raise IngressError("SG01_PROJECTION_SHAPE", f"{field} must be an array") + child = sorted({canonical_json_bytes(item): item for item in child}.values(), key=canonical_json_bytes) + result[field] = child + return result + + left_projection = project(left) + right_projection = project(right) + if left_projection != right_projection: + raise IngressError( + "SG01_SEMANTIC_DRIFT", + "routing activation and signal SG-01 semantic projections differ", + details={"routing_projection": left_projection, "signal_projection": right_projection}, + ) + return { + "status": "PASS", + "projection": left_projection, + "projection_sha256": canonical_digest(left_projection), + "routing_raw_sha256": routing_raw_sha256, + "signal_raw_sha256": signal_raw_sha256, + "compared_keys": list(SG01_PROJECTION_FIELDS), + } + + + def verify_cross_artifact_seals( + documents: Mapping[str, Any], + snapshots: Mapping[str, Snapshot], + deployment_snapshots: Mapping[str, Snapshot] | None = None, + ) -> dict[str, Any]: + """Recompute the P1 guard and current-v8 producer invariants.""" + + checks: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + deployment_snapshots = deployment_snapshots or {} + p1 = documents.get("stage1_part1_soft_gate_handoff") + if isinstance(p1, dict): + digest_guard = p1.get("digest_guard") + if not isinstance(digest_guard, dict): + issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff"])) + digest_guard = {} + elif any(key not in digest_guard for key in P1_DIGEST_KEYS): + issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff#digest_guard"])) + for digest_key, logical_id in P1_DIGEST_KEYS.items(): + source = snapshots.get(logical_id) or deployment_snapshots.get(logical_id) + observed = source.raw_sha256 if source else None + expected = digest_guard.get(digest_key) + passed = expected is not None and observed is not None and expected == observed + checks.append({"check_id": f"P1:{digest_key}", "status": "PASS" if passed else "UNEVALUABLE" if source is None else "FAIL"}) + if expected is not None and observed is not None and not passed: + issues.append(_issue("P1_DIGEST_MISMATCH", source_refs=[logical_id])) + else: + issues.append(_issue("P1_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part1_soft_gate_handoff"])) + p2 = documents.get("stage1_part2_review_handoff") + if p2 is not None and not isinstance(p2, dict): + issues.append(_issue("P2_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part2_review_handoff"])) + for stage in (3, 4): + logical = f"stage1_part{stage}_review_handoff" + value = documents.get(logical) + if value is not None: + wrapper_present = isinstance(value, dict) and isinstance(value.get(logical), dict) + if not wrapper_present: + issues.append(_issue(f"P{stage}_WRAPPER_MISSING", source_refs=[logical])) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + for index, row in enumerate(ledger_rows): + if not isinstance(row, dict) or "domain_effects" not in row or "calculation_requests" not in row: + issues.append(_issue("CURRENT_V8_LEDGER_EXTENSION_MISSING", impact_scope="FACT", source_refs=[f"fact_ledger_base#/{index}"])) + return {"checks": checks, "issues": issues, "passed": not any(item["severity"] == "ERROR" for item in issues)} + + + def check_conservation( + documents: Mapping[str, Any], + *, + signal_all: Mapping[str, Any] | None = None, + normalized_reviews: Mapping[str, Any] | None = None, + source_snapshots: Mapping[str, Snapshot] | None = None, + ) -> dict[str, Any]: + """Independently compute core set, cardinality, and multiset invariants.""" + + checks: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + source_snapshots = source_snapshots or {} + + def add_check( + check_id: str, + passed: bool | None, + left: Sequence[Any] | Counter[Any] | None, + right: Sequence[Any] | Counter[Any] | None, + *, + issue_code: str, + impact_scope: str, + source_refs: Sequence[str], + details: Mapping[str, Any] | None = None, + ) -> None: + left_counter = left if isinstance(left, Counter) else Counter(canonical_digest(value) for value in (left or [])) + right_counter = right if isinstance(right, Counter) else Counter(canonical_digest(value) for value in (right or [])) + row: dict[str, Any] = { + "check_id": check_id, + "status": "UNEVALUABLE" if passed is None else "PASS" if passed else "FAIL", + "left_count": sum(left_counter.values()) if left is not None else None, + "right_count": sum(right_counter.values()) if right is not None else None, + "left_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in left_counter.items())) if left is not None else None, + "right_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in right_counter.items())) if right is not None else None, + } + if details: + row.update(details) + checks.append(row) + if passed is False: + issues.append(_issue(issue_code, impact_scope=impact_scope, source_refs=source_refs)) + + bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + bo_ids = [str(row["BO_ID"]) for row in bo_rows if isinstance(row, dict) and row.get("BO_ID") is not None] + source_bo_ids = [ + str(row["source_bo_id"]) + for row in ledger_rows + if isinstance(row, dict) and row.get("source_bo_id") is not None + ] + missing_bo_id_rows = [index for index, row in enumerate(bo_rows) if not isinstance(row, dict) or row.get("BO_ID") is None] + missing_source_bo_rows = [ + index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("source_bo_id") is None + ] + bo_pass = ( + not missing_bo_id_rows + and not missing_source_bo_rows + and Counter(bo_ids) == Counter(source_bo_ids) + ) + add_check( + "BO_FACT_MULTISET", + bo_pass, + bo_ids, + source_bo_ids, + issue_code="BO_FACT_CONSERVATION_FAILED", + impact_scope="FACT", + source_refs=["bo", "fact_ledger_base"], + details={ + "missing_bo_id_rows": missing_bo_id_rows, + "missing_source_bo_id_rows": missing_source_bo_rows, + "duplicate_bo_ids": sorted(key for key, count in Counter(bo_ids).items() if count > 1), + "dangling_source_bo_ids": sorted(set(source_bo_ids) - set(bo_ids)), + }, + ) + missing_fact_id_rows = [ + index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("fact_id") is None + ] + fact_ids = [str(row["fact_id"]) for row in ledger_rows if isinstance(row, dict) and row.get("fact_id") is not None] + expected_fact_ids = [f"F-{index:03d}" for index in range(1, len(ledger_rows) + 1)] + fact_pass = not missing_fact_id_rows and fact_ids == expected_fact_ids and len(fact_ids) == len(set(fact_ids)) + add_check( + "FACT_ID_SEQUENCE", + fact_pass, + fact_ids, + expected_fact_ids, + issue_code="FACT_ID_CONSERVATION_FAILED", + impact_scope="FACT", + source_refs=["fact_ledger_base"], + details={"missing_fact_id_rows": missing_fact_id_rows, "observed": fact_ids}, + ) + extension_missing = [ + index + for index, row in enumerate(ledger_rows) + if not isinstance(row, dict) + or not isinstance(row.get("domain_effects"), dict) + or not isinstance(row.get("calculation_requests"), list) + ] + add_check( + "CURRENT_V8_LEDGER_EXTENSIONS", + not extension_missing, + list(range(len(ledger_rows))), + [index for index in range(len(ledger_rows)) if index not in extension_missing], + issue_code="CURRENT_V8_LEDGER_EXTENSION_MISSING", + impact_scope="FACT", + source_refs=["fact_ledger_base"], + details={"missing_row_indices": extension_missing}, + ) + les_rows = _array_rows( + documents.get("legal_effect_structures"), + ("structures", "structure_records", "legal_effect_structures", "rows", "items"), + ) + dangling_les: list[str] = [] + les_ids: list[str] = [] + for row in les_rows: + if not isinstance(row, dict): + continue + structure_id = row.get("structure_id", row.get("legal_effect_structure_id")) + if structure_id is not None: + les_ids.append(str(structure_id)) + refs = row.get("source_bo_ids", []) + if isinstance(refs, list): + dangling_les.extend(str(ref) for ref in refs if ref not in set(bo_ids)) + duplicate_les_ids = sorted(key for key, count in Counter(les_ids).items() if count > 1) + les_pass = not dangling_les and not duplicate_les_ids and len(les_ids) == len(les_rows) + add_check( + "LES_BO_JOIN", + les_pass, + [str(row.get("structure_id", row.get("legal_effect_structure_id"))) for row in les_rows if isinstance(row, dict)], + les_ids, + issue_code="LES_BO_JOIN_FAILED", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures", "bo"], + details={"dangling_refs": sorted(dangling_les), "duplicate_structure_ids": duplicate_les_ids}, + ) + declared_les_count = None + les_document = documents.get("legal_effect_structures") + if isinstance(les_document, dict): + for key in ("declared_structure_count", "structure_count", "record_count"): + if isinstance(les_document.get(key), int): + declared_les_count = int(les_document[key]) + break + declared_les_pass = None if declared_les_count is None else declared_les_count == len(les_rows) + add_check( + "LES_DECLARED_ACTUAL_COUNT", + declared_les_pass, + [None] * declared_les_count if declared_les_count is not None else None, + [None] * len(les_rows), + issue_code="LES_DECLARED_COUNT_MISMATCH", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures"], + ) + actual_domain_index: dict[str, list[str]] = defaultdict(list) + actual_bo_index: dict[str, list[str]] = defaultdict(list) + ledger_structure_refs: list[tuple[str, str, str]] = [] + ledger_type_refs: list[tuple[str, str, str]] = [] + actual_structure_refs: list[tuple[str, str, str]] = [] + actual_type_refs: list[tuple[str, str, str]] = [] + route_count_errors: list[str] = [] + for row in les_rows: + if not isinstance(row, dict): + continue + structure_id = str(row.get("structure_id", row.get("legal_effect_structure_id", "MISSING"))) + domain_id = str(row.get("domain_id", "MISSING")) + type_id = str(row.get("type_id", row.get("type", "MISSING"))) + actual_domain_index[domain_id].append(structure_id) + source_ids = row.get("source_bo_ids", []) + if isinstance(source_ids, list): + for bo_id in source_ids: + actual_bo_index[str(bo_id)].append(structure_id) + actual_structure_refs.append((str(bo_id), domain_id, structure_id)) + actual_type_refs.append((str(bo_id), domain_id, type_id)) + routes = row.get("routes", []) + if isinstance(routes, list) and row.get("route_count", len(routes)) != len(routes): + route_count_errors.append(structure_id) + for row in ledger_rows: + if not isinstance(row, dict): + continue + bo_id = str(row.get("source_bo_id", "MISSING")) + effects = row.get("domain_effects", {}) + if not isinstance(effects, dict): + continue + for domain_id, effect in effects.items(): + if not isinstance(effect, dict): + continue + for structure_id in effect.get("structure_ids", []) if isinstance(effect.get("structure_ids"), list) else []: + ledger_structure_refs.append((bo_id, str(domain_id), str(structure_id))) + for type_id in effect.get("type_ids", []) if isinstance(effect.get("type_ids"), list) else []: + ledger_type_refs.append((bo_id, str(domain_id), str(type_id))) + structure_index = les_document.get("structure_index", {}) if isinstance(les_document, dict) else {} + index_present = isinstance(structure_index, dict) and bool(structure_index) + index_ok = True + if index_present: + declared_by_domain = structure_index.get("by_domain_id", {}) + declared_by_bo = structure_index.get("by_bo_id", {}) + index_ok = ( + isinstance(declared_by_domain, dict) + and isinstance(declared_by_bo, dict) + and {str(key): Counter(map(str, value)) for key, value in declared_by_domain.items() if isinstance(value, list)} + == {key: Counter(value) for key, value in actual_domain_index.items()} + and {str(key): Counter(map(str, value)) for key, value in declared_by_bo.items() if isinstance(value, list)} + == {key: Counter(value) for key, value in actual_bo_index.items()} + ) + reverse_ok = ( + (not ledger_structure_refs or Counter(ledger_structure_refs) == Counter(actual_structure_refs)) + and (not ledger_type_refs or Counter(ledger_type_refs) == Counter(actual_type_refs)) + and not route_count_errors + and index_ok + ) + add_check( + "LES_REVERSE_INDEX", + reverse_ok, + ledger_structure_refs + ledger_type_refs, + actual_structure_refs + actual_type_refs, + issue_code="LES_REVERSE_INDEX_MISMATCH", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures", "fact_ledger_base"], + details={"index_present": index_present, "route_count_errors": route_count_errors}, + ) + evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) + event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) + evidence_ids = [ + str(row.get("evidence_id", row.get("id"))) + for row in evidence_rows + if isinstance(row, dict) and (row.get("evidence_id") is not None or row.get("id") is not None) + ] + event_ids = [ + str(row.get("event_id", row.get("id"))) + for row in event_rows + if isinstance(row, dict) and (row.get("event_id") is not None or row.get("id") is not None) + ] + fact_evidence_refs: list[str] = [] + fact_event_refs: list[str] = [] + event_evidence_refs: list[str] = [] + for row in ledger_rows: + if not isinstance(row, dict): + continue + evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) + event_values = row.get("event_refs", row.get("event_ids", [])) + if isinstance(evidence_values, list): + fact_evidence_refs.extend(str(ref) for ref in evidence_values) + if isinstance(event_values, list): + fact_event_refs.extend(str(ref) for ref in event_values) + for row in event_rows: + if not isinstance(row, dict): + continue + evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) + if isinstance(evidence_values, list): + event_evidence_refs.extend(str(ref) for ref in evidence_values) + evidence_failures = sorted( + set(fact_evidence_refs + event_evidence_refs) - set(evidence_ids) + ) + duplicate_evidence_ids = sorted(key for key, count in Counter(evidence_ids).items() if count > 1) + evidence_pass = not evidence_failures and not duplicate_evidence_ids + add_check( + "EVIDENCE_REFERENCE_CONSERVATION", + evidence_pass, + fact_evidence_refs + event_evidence_refs, + evidence_ids, + issue_code="EVIDENCE_REFERENCE_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_indexed", "fact_ledger_base", "evidence_event_candidates"], + details={"dangling_refs": evidence_failures, "duplicate_evidence_ids": duplicate_evidence_ids}, + ) + event_failures = sorted(set(fact_event_refs) - set(event_ids)) + duplicate_event_ids = sorted(key for key, count in Counter(event_ids).items() if count > 1) + event_pass = not event_failures and not duplicate_event_ids + add_check( + "EVENT_REFERENCE_CONSERVATION", + event_pass, + fact_event_refs, + event_ids, + issue_code="EVENT_REFERENCE_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_event_candidates", "fact_ledger_base"], + details={"dangling_refs": event_failures, "duplicate_event_ids": duplicate_event_ids}, + ) + disposition_rows = [row.get("disposition") for row in event_rows if isinstance(row, dict) and "disposition" in row] + b2_gate = documents.get("b2_event_candidates_gate") + declared_dispositions = None + if isinstance(b2_gate, dict): + declared_dispositions = b2_gate.get("event_disposition_counts") + if declared_dispositions is None and isinstance(b2_gate.get("summary"), dict): + declared_dispositions = b2_gate["summary"].get("event_disposition_counts") + if isinstance(declared_dispositions, dict): + disposition_expected = Counter( + {str(key): int(value) for key, value in declared_dispositions.items() if isinstance(value, int)} + ) + disposition_actual = Counter(str(value) for value in disposition_rows) + disposition_pass: bool | None = disposition_actual == disposition_expected + elif disposition_rows: + disposition_expected = Counter(str(value) for value in disposition_rows) + disposition_actual = Counter(str(value) for value in disposition_rows) + disposition_pass = all(isinstance(value, str) and value for value in disposition_rows) + else: + disposition_expected = Counter() + disposition_actual = Counter() + disposition_pass = None + add_check( + "EVENT_DISPOSITION_CONSERVATION", + disposition_pass, + disposition_actual, + disposition_expected, + issue_code="EVENT_DISPOSITION_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_event_candidates", "b2_event_candidates_gate"], + ) + writer_report = documents.get("fact_ledger_writer_report") + if isinstance(writer_report, dict): + observed_domain_coverage = Counter( + str(domain_id) + for row in ledger_rows + if isinstance(row, dict) and isinstance(row.get("domain_effects"), dict) + for domain_id in row["domain_effects"] + ) + declared_domain_coverage = Counter( + {str(key): int(value) for key, value in writer_report.get("domain_effect_coverage", {}).items() if isinstance(value, int)} + ) + observed_readiness = Counter( + str(request.get("operand_state")) + for row in ledger_rows + if isinstance(row, dict) and isinstance(row.get("calculation_requests"), list) + for request in row["calculation_requests"] + if isinstance(request, dict) + ) + declared_readiness = Counter( + {str(key): int(value) for key, value in writer_report.get("calculation_readiness", {}).items() if isinstance(value, int)} + ) + ledger_snapshot = source_snapshots.get("fact_ledger_base") + final_hash = writer_report.get("final_sha256") + writer_pass = ( + writer_report.get("row_count") == len(ledger_rows) + and declared_domain_coverage == observed_domain_coverage + and declared_readiness == observed_readiness + and (ledger_snapshot is None or final_hash == ledger_snapshot.raw_sha256) + ) + add_check( + "FACT_LEDGER_WRITER_REPORT_CONNECTION", + writer_pass, + [len(ledger_rows), observed_domain_coverage, observed_readiness, ledger_snapshot.raw_sha256 if ledger_snapshot else None], + [writer_report.get("row_count"), declared_domain_coverage, declared_readiness, final_hash], + issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", + impact_scope="FACT", + source_refs=["fact_ledger_base", "fact_ledger_writer_report"], + ) + else: + add_check( + "FACT_LEDGER_WRITER_REPORT_CONNECTION", + None, + None, + None, + issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", + impact_scope="FACT", + source_refs=["fact_ledger_base", "fact_ledger_writer_report"], + ) + if signal_all is not None: + file_pass = bool(signal_all.get("file_conservation_pass")) + record_pass = bool(signal_all.get("record_conservation_pass")) + checks.append({"check_id": "SIGNAL_FILE_ROW_CONSERVATION", "status": "PASS" if file_pass else "FAIL"}) + checks.append({"check_id": "SIGNAL_RECORD_OCCURRENCE_CONSERVATION", "status": "PASS" if record_pass else "FAIL"}) + issues.extend(signal_all.get("issues", [])) + if not file_pass: + issues.append(_issue("SIGNAL_FILE_CONSERVATION_FAILED", impact_scope="SIGNAL")) + if not record_pass: + issues.append(_issue("SIGNAL_RECORD_CONSERVATION_FAILED", impact_scope="SIGNAL")) + if normalized_reviews is not None: + review_pass = normalized_reviews.get("conservation_status") == "PASS" + checks.append({"check_id": "REVIEW_OCCURRENCE_CONSERVATION", "status": "PASS" if review_pass else "FAIL"}) + if not review_pass: + issues.append(_issue("REVIEW_CONSERVATION_FAILED", impact_scope="REVIEW_ITEM")) + issues.extend(normalized_reviews.get("_issues", [])) + return {"checks": checks, "issues": issues, "passed": not any(check["status"] == "FAIL" for check in checks)} + + + def _source_ref( + logical_id: str, + pointer: str, + raw_value: Any = _RAW_VALUE_UNSET, + *, + stage1_id: str | None = None, + ) -> dict[str, Any]: + """Build a truthful RFC 6901 provenance row without pointer narrowing.""" + + row: dict[str, Any] = { + "logical_artifact_id": logical_id, + "json_pointer": pointer, + "raw_value_sha256": canonical_digest( + [logical_id, pointer] + if raw_value is _RAW_VALUE_UNSET + else raw_value + ), + "source_contract_row_ref": logical_id, + } + if stage1_id is not None: + row["stage1_id"] = stage1_id + return row + + + def _tarjan_scc(nodes: Sequence[str], edges: Sequence[tuple[str, str]]) -> list[list[str]]: + adjacency: dict[str, list[str]] = {node: [] for node in nodes} + for source, target in edges: + adjacency.setdefault(source, []).append(target) + adjacency.setdefault(target, []) + for value in adjacency.values(): + value.sort() + index = 0 + stack: list[str] = [] + on_stack: set[str] = set() + indices: dict[str, int] = {} + lowlink: dict[str, int] = {} + components: list[list[str]] = [] + + def visit(node: str) -> None: + nonlocal index + indices[node] = index + lowlink[node] = index + index += 1 + stack.append(node) + on_stack.add(node) + for neighbor in adjacency[node]: + if neighbor not in indices: + visit(neighbor) + lowlink[node] = min(lowlink[node], lowlink[neighbor]) + elif neighbor in on_stack: + lowlink[node] = min(lowlink[node], indices[neighbor]) + if lowlink[node] == indices[node]: + component: list[str] = [] + while True: + member = stack.pop() + on_stack.remove(member) + component.append(member) + if member == node: + break + components.append(sorted(component)) + + for node in sorted(adjacency): + if node not in indices: + visit(node) + return sorted(components, key=lambda component: component[0]) + + + def _inline_sha256(value: str, *, code: str) -> str: + if not isinstance(value, str) or re.fullmatch(r"[a-f0-9]{64}", value) is None: + raise IngressError(code, "expected one lowercase SHA-256 digest") + return value + + + def _inline_relative_path(value: str, *, code: str) -> str: + if not isinstance(value, str) or not value or "\x00" in value or "\\" in value: + raise IngressError(code, "logical path is empty or malformed") + if unicodedata.normalize("NFC", value) != value: + raise IngressError(code, "logical path must already be NFC") + path = PurePosixPath(value) + if path.is_absolute() or any(part in {"", ".", ".."} for part in path.parts): + raise IngressError(code, "logical path must be a contained relative path") + rendered = path.as_posix() + if rendered != value: + raise IngressError(code, "logical path is not canonical") + return rendered + + + def _inline_parse_mcp_payload(raw: bytes, expected_id: int) -> Mapping[str, Any]: + """Parse one JSON or SSE JSON-RPC terminal response with an exact ID.""" + + candidates: list[Any] + try: + candidates = [load_json_strict(raw)] + except IngressError: try: - args = _parser().parse_args(argv) - if args.workflow_id != "S2_00": - raise IngressError("WORKFLOW_ID_MISMATCH", "runtime accepts only workflow_id S2_00") - asset_root = Path(__file__).resolve().parents[1] - project_root = _project_root_from_runtime() - release_path = _resolve_contained_ref(asset_root, args.release_ref, code="RELEASE_REF_OUTSIDE_ROOT") - if not release_path.is_file(): - raise IngressError("RELEASE_REF_NOT_FILE", "release reference must identify one regular file") - stage1_run_root = _resolve_contained_ref( - project_root, - args.stage1_run_root_ref, - code="STAGE1_RUN_ROOT_OUTSIDE_WORKSPACE", - ) - stage1_deployment_root = _resolve_contained_ref( - project_root, - args.stage1_deployment_root_ref, - code="STAGE1_DEPLOYMENT_ROOT_OUTSIDE_WORKSPACE", - ) - if not stage1_run_root.is_dir() or not stage1_deployment_root.is_dir(): - raise IngressError("STAGE1_ROOT_NOT_DIRECTORY", "Stage 1 roots must be directories") - release_lock = load_release_lock(release_path) - contract_manifest = _load_bound_contract_manifest(stage1_deployment_root, release_lock) - output_dir = project_root / "stage2_runs" / args.run_id - result = execute_ingress( - stage1_run_root, - release_lock, - output_dir=output_dir, - attempt_id=args.attempt_id, - run_id=args.run_id, - user_context_sha256=args.user_context_sha256, - workspace_context_sha256=args.workspace_context_sha256, - stage1_deployment_root=stage1_deployment_root, - stage2_asset_root=asset_root, - contract_manifest=contract_manifest, - ) - sys.stdout.buffer.write(canonical_json_bytes({"ok": True, "result": result})) - return 0 - except (IngressError, FileNotFoundError, PermissionError, OSError) as exc: - error = exc.as_dict() if isinstance(exc, IngressError) else {"code": "OS_ERROR", "message": str(exc)} - sys.stdout.buffer.write(canonical_json_bytes({"ok": False, "error": error})) - return 2 + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise IngressError("MCP_RESPONSE_UTF8", "MCP response is not strict UTF-8") from exc + events: list[bytes] = [] + data_lines: list[str] = [] + for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n"): + if line == "": + if data_lines: + events.append("\n".join(data_lines).encode("utf-8")) + data_lines = [] + continue + if line.startswith(":") or line.startswith("event:") or line.startswith("id:") or line.startswith("retry:"): + continue + if not line.startswith("data:"): + raise IngressError("MCP_SSE_SHAPE", "unexpected non-data SSE line") + payload = line[5:] + if payload.startswith(" "): + payload = payload[1:] + data_lines.append(payload) + if data_lines: + events.append("\n".join(data_lines).encode("utf-8")) + if not events: + raise IngressError("MCP_RESPONSE_SHAPE", "MCP response contains no JSON terminal event") + candidates = [load_json_strict(event) for event in events] + matching = [ + item + for item in candidates + if isinstance(item, dict) and item.get("id") == expected_id + ] + if len(matching) != 1: + raise IngressError( + "MCP_RESPONSE_ID_MISMATCH", + "MCP response must contain exactly one terminal result with the JSON-RPC message ID", + ) + response = matching[0] + if response.get("jsonrpc") != "2.0": + raise IngressError("MCP_JSONRPC_VERSION", "MCP response jsonrpc must equal 2.0") + if response.get("error") is not None: + raise IngressError( + "MCP_JSONRPC_ERROR", + "MCP server returned a JSON-RPC error", + details={"rpc_error": response.get("error")}, + ) + if "result" not in response or not isinstance(response["result"], dict): + raise IngressError("MCP_RESULT_SHAPE", "MCP response result must be an object") + return response - if __name__ == "__main__": - raise SystemExit(run_inline_mcp()) - task_procedure: - IN: - nexts: - - Task_S2_00_prepare_request - wait_until: [] - Task_S2_00_prepare_request: - nexts: - - Task_S2_00_deterministic_ingress - wait_until: - - IN - Task_S2_00_deterministic_ingress: - nexts: - - OUT - wait_until: - - Task_S2_00_prepare_request - OUT: - nexts: [] - wait_until: - - Task_S2_00_deterministic_ingress + def _inline_tool_text(result: Mapping[str, Any], tool_name: str) -> str: + if result.get("isError") is True: + content = result.get("content") + rendered = canonical_json_bytes(content).decode("utf-8", errors="replace") if content is not None else "" + lowered = rendered.lower() + code = ( + "LOCALDOCS_NOT_FOUND" + if any(marker in lowered for marker in ("not found", "does not exist", "no such file")) + else "MCP_TOOL_ERROR" + ) + raise IngressError(code, f"localdocs {tool_name} returned isError=true") + content = result.get("content") + if not isinstance(content, list) or len(content) != 1: + raise IngressError("MCP_CONTENT_CARDINALITY", "MCP tool result must contain exactly one content block") + block = content[0] + if not isinstance(block, dict) or block.get("type") != "text" or not isinstance(block.get("text"), str): + raise IngressError("MCP_CONTENT_SHAPE", "MCP tool result must contain one text block") + return block["text"] + + + def _inline_binary_envelope(text: str, logical_path: str) -> bytes: + value = load_json_strict(text) + if isinstance(value, dict) and "results" in value: + results = value.get("results") + if not isinstance(results, list) or len(results) != 1 or not isinstance(results[0], dict): + raise IngressError("LOCALDOCS_RESULT_CARDINALITY", "binary response must contain one result row") + inner: Any = results[0].get("content", results[0].get("text")) + value = load_json_strict(inner) if isinstance(inner, str) else inner + if not isinstance(value, dict) or not isinstance(value.get("content_base64"), str): + raise IngressError("LOCALDOCS_BINARY_ENVELOPE", "binary response lacks content_base64") + try: + payload = base64.b64decode(value["content_base64"].encode("ascii"), validate=True) + except (UnicodeEncodeError, binascii.Error, ValueError) as exc: + raise IngressError("LOCALDOCS_BASE64_INVALID", "binary response is not strict base64") from exc + declared_size = value.get("byte_length", value.get("size")) + if declared_size is not None and (not isinstance(declared_size, int) or declared_size != len(payload)): + raise IngressError("LOCALDOCS_BYTE_LENGTH_MISMATCH", f"binary length mismatch: {logical_path}") + declared_hash = value.get("sha256") + if declared_hash is not None and declared_hash != hashlib.sha256(payload).hexdigest(): + raise IngressError("LOCALDOCS_HASH_MISMATCH", f"binary hash mismatch: {logical_path}") + return payload + + + class _InlineLocaldocs: + """Minimal user/workspace-bound localdocs JSON-RPC client.""" + + def __init__( + self, + user_hash: str, + workspace_hash: str, + *, + client: Any | None = None, + timeout_seconds: int = 60, + ) -> None: + self.user_hash = _inline_sha256(user_hash, code="USER_CONTEXT_HASH_INVALID") + self.workspace_hash = _inline_sha256( + workspace_hash, + code="WORKSPACE_CONTEXT_HASH_INVALID", + ) + if client is None: + try: + import httpx # type: ignore + except ImportError as exc: + raise IngressError("HTTPX_UNAVAILABLE", "Code Executor must supply httpx==0.28.1") from exc + client = httpx.Client(timeout=timeout_seconds) + self.client = client + self.headers = { + "Content-Type": "application/json", + "Accept": "application/json, text/event-stream", + } + self._message_ids = itertools.count(10) + self._initialized = False + self._session_id: str | None = None + + def close(self) -> None: + close = getattr(self.client, "close", None) + if callable(close): + close() + + def _post(self, body: Mapping[str, Any], expected_id: int | None) -> Mapping[str, Any] | None: + try: + response = self.client.post(LOCALDOCS_URL, json=dict(body), headers=dict(self.headers)) + response.raise_for_status() + except Exception as exc: + raise IngressError("MCP_TRANSPORT_ERROR", "localdocs transport failed") from exc + session_id = response.headers.get("mcp-session-id") + if session_id: + if not isinstance(session_id, str) or not session_id.strip(): + raise IngressError("MCP_SESSION_ID_INVALID", "localdocs returned an invalid session ID") + normalized_session_id = session_id.strip() + if self._session_id is None: + if expected_id != 1: + raise IngressError( + "MCP_SESSION_ID_OUTSIDE_INITIALIZE", + "localdocs first bound a session outside initialize", + ) + self._session_id = normalized_session_id + elif normalized_session_id != self._session_id: + raise IngressError( + "MCP_SESSION_ID_CHANGED", + "localdocs changed the initialized session ID", + ) + self.headers["mcp-session-id"] = self._session_id + if expected_id is None: + return None + raw = response.content if isinstance(response.content, bytes) else bytes(response.content) + return _inline_parse_mcp_payload(raw, expected_id) + + def initialize(self) -> None: + response = self._post( + { + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": MCP_PROTOCOL_VERSION, + "capabilities": {}, + "clientInfo": { + "name": INLINE_CLIENT_NAME, + "version": INLINE_CLIENT_VERSION, + "user_id": self.user_hash, + "workspace_id": self.workspace_hash, + }, + }, + }, + 1, + ) + if response is None: + raise IngressError("MCP_INITIALIZE_EMPTY", "localdocs initialize returned no result") + result = response.get("result") + if not isinstance(result, dict) or result.get("protocolVersion") != MCP_PROTOCOL_VERSION: + raise IngressError( + "MCP_PROTOCOL_VERSION_MISMATCH", + "localdocs did not negotiate the requested MCP protocol version", + ) + if self._session_id is None or "mcp-session-id" not in self.headers: + raise IngressError("MCP_SESSION_ID_MISSING", "localdocs initialize did not bind a session ID") + self._post( + {"jsonrpc": "2.0", "method": "notifications/initialized"}, + None, + ) + self._initialized = True + + def call(self, tool_name: str, arguments: Mapping[str, Any]) -> Mapping[str, Any]: + if not self._initialized: + raise IngressError("MCP_NOT_INITIALIZED", "localdocs session is not initialized") + message_id = next(self._message_ids) + response = self._post( + { + "jsonrpc": "2.0", + "id": message_id, + "method": "tools/call", + "params": {"name": tool_name, "arguments": dict(arguments)}, + }, + message_id, + ) + if response is None: + raise IngressError("MCP_TOOL_EMPTY", f"localdocs {tool_name} returned no result") + return response["result"] + + def read_binary(self, logical_path: str) -> bytes: + path = _inline_relative_path(logical_path, code="LOCALDOCS_READ_PATH_INVALID") + result = self.call("read_binary_doc", {"doc_name": path}) + return _inline_binary_envelope(_inline_tool_text(result, "read_binary_doc"), path) + + def read_binary_optional(self, logical_path: str) -> bytes | None: + try: + return self.read_binary(logical_path) + except IngressError as exc: + if exc.code == "LOCALDOCS_NOT_FOUND": + return None + raise + + def write_binary_verified(self, logical_path: str, payload: bytes, *, overwrite: bool = False) -> str: + path = _inline_relative_path(logical_path, code="LOCALDOCS_WRITE_PATH_INVALID") + encoded = base64.b64encode(payload).decode("ascii") + result = self.call( + "write_binary_file", + {"path": path, "content_base64": encoded, "overwrite": overwrite}, + ) + _inline_tool_text(result, "write_binary_file") + observed = self.read_binary(path) + if observed != payload: + raise IngressError("LOCALDOCS_WRITE_READBACK_MISMATCH", f"read-back mismatch: {path}") + return hashlib.sha256(observed).hexdigest() + + + SOURCE_POLICY = load_json_strict(r'''{"stage1_sources":[{"adapter_id":"S2A-EVIDENCE-V3-ENVELOPE-V1","logical_input_id":"evidence_indexed","path":"evidence_indexed.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B1_quality_gate_evidence_indexed","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_contract_version","items"],"requirement_class":"EVIDENCE_EVENT_SCOPE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-EVENTS-V1-ENVELOPE-V1","logical_input_id":"evidence_event_candidates","path":"evidence_event_candidates.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B2_quality_gate_event_candidates","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version","items"],"requirement_class":"EVIDENCE_EVENT_SCOPE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-CLIENT-GOAL-V8-V1","logical_input_id":"client_goal","path":"client_goal.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_A_client_goal","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["primary_goal","constraints","parties"],"requirement_class":"OPTIMIZATION_CONTEXT","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-DOMAIN-SCREENING-V1","logical_input_id":"domain_screening","path":"routing/domain_screening.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_A0_domain_screener_02","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["domain_screening"],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-DUAL-SG01-V1","logical_input_id":"domain_activation_manifest","path":"routing/domain_activation_manifest.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_D0_domain_activation_gate","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["domain_activation_manifest"],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/s5/domain_activation_manifest.schema.json","path":"signals/schemas/domain_activation_manifest.schema.json","sha256":"013a6ebd230ebe46dda665af9f6c4448b267444b44e7b8f701f2fae80a2ee92a"},"transaction_identity_pointer":null},{"adapter_id":"S2A-B1-GATE-V1","logical_input_id":"b1_evidence_indexed_gate","path":"quality_gates/B1_evidence_indexed_gate.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B12_gate_audit_finalizer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_contract_version","gate_id","overall_severity","hard_gate_findings","review_findings","stage2_auto_progression_allowed"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-B2-GATE-V1","logical_input_id":"b2_event_candidates_gate","path":"quality_gates/B2_event_candidates_gate.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B12_gate_audit_finalizer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_contract_version","gate_id","overall_severity","hard_gate_findings","review_findings","stage2_auto_progression_allowed"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-P1-HANDOFF-FLAT-V1","logical_input_id":"stage1_part1_soft_gate_handoff","path":"quality_gates/stage1_part1_soft_gate_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B2_SHA256_soft_gate_handoff_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version","handoff_status","review_items","stage2_auto_progression_allowed","hard_gate_summary","review_item_conservation","digest_guard"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-BO-V8-LIST-V1","logical_input_id":"bo","path":"BO.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_BO_F0_final_bo_compiler_gate_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":[],"requirement_class":"IDENTITY_BACKBONE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-SIGNAL-ALL-V1","logical_input_id":"signal_manifest","path":"signals/signal_manifest.json","path_rule":null,"producer_alias_id":"PA-SG-COMPILER-001","producer_id":"Task_C_BO_S0_signal_bundle_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["files","downstream_read_sets"],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/s5/signal_manifest.schema.json","path":"signals/schemas/signal_manifest.schema.json","sha256":"5e72084780b82b29582c9ffcf48f3e4894d7c0b152e5ce8df394583c07dde681"},"transaction_identity_pointer":"/transaction_id"},{"adapter_id":"S2A-P2-HANDOFF-FLAT-V1","logical_input_id":"stage1_part2_review_handoff","path":"quality_gates/stage1_part2_review_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_BO_F0_final_bo_compiler_gate_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version","status","review_items"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-LES-CURRENT-V8-V1","logical_input_id":"legal_effect_structures","path":"legal_effect_structures.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_LE_L2_final_structure_index_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":[],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/part3/legal_effect_structures.schema.json","path":"platform/schemas/legal_effect_structures.schema.json","sha256":"fc962e8ae39f9bede64ba017297eded6413689204a065e00c3b3bdca8f1854df"},"transaction_identity_pointer":"/signal_manifest_transaction_id"},{"adapter_id":"S2A-P3-HANDOFF-WRAPPED-V1","logical_input_id":"stage1_part3_review_handoff","path":"quality_gates/stage1_part3_review_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_LE_L2_final_structure_index_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["stage1_part3_review_handoff"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-FACT-LEDGER-CURRENT-V8-V1","logical_input_id":"fact_ledger_base","path":"Fact_Ledger_base.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_FL_F2_final_fact_ledger_gate_and_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":[],"requirement_class":"IDENTITY_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/part4/fact_ledger_base.schema.json","path":"platform/schemas/fact_ledger_base.schema.json","sha256":"b3f0e79ecb4c2f720f3e07e89154aadbd2327e4129cc703569fb5635240d2fe8"},"transaction_identity_pointer":null},{"adapter_id":"S2A-FACT-LEDGER-WRITER-REPORT-V1","logical_input_id":"fact_ledger_writer_report","path":"stage1_tmp/fact_ledger/fact_ledger_writer_report.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_FL_F2_final_fact_ledger_gate_and_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-P4-HANDOFF-WRAPPED-V1","logical_input_id":"stage1_part4_review_handoff","path":"quality_gates/stage1_part4_review_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_FL_F2_final_fact_ledger_gate_and_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["stage1_part4_review_handoff"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-SIGNAL-ALL-V1","logical_input_id":"signal_payload_family","path":null,"path_rule":"signals/","producer_alias_id":"PA-SG-COMPILER-001","producer_id":"Task_C_BO_S0_signal_bundle_writer","raw_hash_source":"MANIFEST_ROW","required_keys":[],"requirement_class":"SIGNAL_PAYLOAD","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null}],"dependency_locks":{"stage1":{"closure_scope":"REFERENCED_55_ONLY_NOT_FULL_STAGE1_RUNTIME_RELEASE","closure_snapshot_date":"2026-08-29","concrete_paths":[{"binding_status":"BOUND","lock_id":"S1-DEPLOY-001","path":"runtime_manifest.json","schema_id":"stage1_runtime_manifest.v1","sha256":"8964593a64a9b1bc90122054bb09eb3911827a06ed62dab0d6b7c745e7e18f54","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-002","path":"domains/_registry_index.json","schema_id":null,"sha256":"9f177ebf8860e20e05483967a2037f3baa09c2ac92c69ddeb260c04ca31ebf39","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-003","path":"signals/signal_registry.v2.json","schema_id":"signal_registry.v2","sha256":"4392b40da458102f8dd11b40b40ae3f694b7b5911849b050e2e4118c569e5ab0","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-004","path":"domains/E-00/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"5919f7ea1d7be02666b0c48aa6a66445e6d454fc2fbb21d7fe5154b0a1e68f6f","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-005","path":"domains/E-01/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"b557e92cd1b093bf31792dbcf5b62cab8ad064a65c4421e79f141e06b4cc2192","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-006","path":"domains/E-02/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"be407c980c28226a15406f85b5861b04a4e19a13870513ac6626349fc05ac434","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-007","path":"domains/E-03/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"7d3814f9b50cd5b33ef65a4eb778693552b3685bd369e765e9ac032734ebe23e","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-008","path":"domains/E-04/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"95a8c600cdce5a687f766788af0f763ee1b6a895e6ed80934afd28fe9a107e25","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-009","path":"domains/E-05/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"e50541018f47aa27de2f8b13ec3fa52210cf8456a6feed8356af78c1f1da144a","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-010","path":"domains/E-06/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"1e2bee36cb3c24dd37fc3beb3cf70236d531126c4f62ee97b5b42e55f4b0745c","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-011","path":"domains/E-07/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"22ac562084b1ce231b7257d099c18b6a4619defa2fc42504d590e0bcc494c5f8","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-012","path":"domains/E-08/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"4d545306d42120e8552dd953d4336ef6de828ea827779944ba73acfda3d3a8bb","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-013","path":"domains/E-09/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"7ef7340750094efeb372c397eb3134e21d62dda6988b1f6fa0a197b9963868e0","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-014","path":"domains/E-10/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"e8d4f45fa76ea9e09333256dd4ea36cd3dd963bf60c04814a2cb8dc90d152f0a","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-015","path":"domains/E-11/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"eb78d0188a0a2400307b1c34c8f8703c54cd86dd06b942c1709c44a8630a68e1","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-016","path":"domains/E-12/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"f336accdc6de10cdcc28c1190328054bca402fb77a2a9859d59fbaf5e84dd170","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-017","path":"domains/E-13/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"e27e2e3855b5868a3ec12c7093b872434702f2e465b73c0bfc948b516aa0fc35","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-018","path":"domains/E-14/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"a273148cc17f07d90cda500fa5cb7df30c9cd253f7b86048cf4f63495d36a156","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-019","path":"domains/E-15/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"7f37edddc101a08ed8a0e25f3a2c638e72571edc212ac91d96c1e33c51202a69","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-020","path":"domains/E-16/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"ae9af46ee31b6ef0dafedd35ccd7959a941d67d1a3dcc70e0b13647896873323","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-021","path":"domains/E-17/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"5c61f4486bdc968e4b30734b3c045404f0a711f47ea3abbe6c5c64652fb7f68c","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-022","path":"domains/E-18/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"74ff76148d175929bdeeeced77e9a9922d29b51ad3d00ae6711c43c55717692c","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-023","path":"domains/E-19/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"a8578f54a3fead3bbd35c62d7199b0f8aafb77f5d409a87236a55d2550fbfd37","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-024","path":"domains/E-20/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"29ee14cfe7789f004e6b6978d5360cf1bebe33bf88715df0cb47257993a11d00","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-025","path":"domains/E-21/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"4e1684a843d9e0c5af82f45332ad85abe94eda0c3ff0d578235d892aee39b908","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-026","path":"domains/EC-00/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"fe74de112b73289485dcead7e0fc7d270c794b3cf8a29ee00fab1eb64ba13861","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-027","path":"domains/X1/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"ad2fee7d206018f9a1f66e5fdf40dd67b686f6938099bad1ffc5d538db14ac57","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-028","path":"domains/X2/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"eba7d4671546bd66f1350d144ae0884f8beffb0dffdc147b45b9d5292676d46f","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-029","path":"domains/X3/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"8b67a638ae4a86aca3a2216974242b11ec39790162c9f366edfa91b02c3d270a","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-030","path":"platform/schemas/client_goal_domain_profiles.schema.json","schema_id":null,"sha256":"ae2bfe0d754a09cbae16b2c15bf1518fc23f9e1bda8fa1f5f949606c8e42c010","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-031","path":"platform/schemas/domain_fanout_plan.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s1/domain_fanout_plan.schema.json","sha256":"3b0948613a5996028b9c030a99f0b51d682f6035e019557756b1a15d43971113","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-032","path":"platform/schemas/domain_seed_output.schema.v3.json","schema_id":"https://schemas.liti-agent.local/stage1/s1/domain_seed_output.schema.v3.json","sha256":"992acf05dbccb34c65ead4e8c592f424e3b91672dc109cbd1bfa76a0a71a13c9","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-033","path":"platform/schemas/domain_slice.schema.v2.json","schema_id":"https://schemas.liti-agent.local/stage1/s1/domain_slice.schema.v2.json","sha256":"212a405088e7cf7ba2c65528a1c716938c946df7fe3bae3256b613051ed31aa3","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-034","path":"platform/schemas/fact_exception_pack.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part4/fact_exception_pack.schema.json","sha256":"4eba7e51ed46a99e3704bc2333169749f4a16935de26a8c8027c1cac98ea58cf","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-035","path":"platform/schemas/fact_ledger_base.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part4/fact_ledger_base.schema.json","sha256":"b3f0e79ecb4c2f720f3e07e89154aadbd2327e4129cc703569fb5635240d2fe8","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-036","path":"platform/schemas/fact_ledger_candidate_bundle.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part4/fact_ledger_candidate_bundle.schema.json","sha256":"4e481504fb795b2be510680a8fa88124f5763a8124462f7870be4125ed9a7730","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-037","path":"platform/schemas/legal_effect_structures.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part3/legal_effect_structures.schema.json","sha256":"fc962e8ae39f9bede64ba017297eded6413689204a065e00c3b3bdca8f1854df","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-038","path":"platform/schemas/structure_seed_bundle.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part3/structure_seed_bundle.schema.json","sha256":"b7af9e422b6ac3876cffea39ec4f617eea76a631a57dfdfcd57d3785a83c667a","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-039","path":"signals/_common/evidence_slot_status.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/evidence_slot_status.schema.json","sha256":"292b03960b187cef668b8635a8d7539fde7c31f0c20d01af52c4f6ff8519d7b1","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-040","path":"signals/_common/signal_item.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/signal_item.schema.json","sha256":"de8695f98041c06cf50c0d8d2ebc31e7b3c518ca9d39a27da940438704c58bb1","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-041","path":"signals/schemas/domain_activation_manifest.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/domain_activation_manifest.schema.json","sha256":"013a6ebd230ebe46dda665af9f6c4448b267444b44e7b8f701f2fae80a2ee92a","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-042","path":"signals/schemas/procedural_posture_relief_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/procedural_posture_relief_signals.schema.json","sha256":"fefb4317ad63088919b61777c71fe75ee6aa507b9f599dcf455d2af63dfc5e0d","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-043","path":"signals/schemas/party_capacity_standing_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/party_capacity_standing_signals.schema.json","sha256":"66de89ac53964166f6caabd50cbc03eb82dede0acf702d5e6d825c1d82ef81d0","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-044","path":"signals/schemas/governing_law_version_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/governing_law_version_signals.schema.json","sha256":"13a3f62f03356090d2cb24de2da0ba217928dfe8eb3c111d0f5e87c7df3119ee","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-045","path":"signals/schemas/legal_relation_lifecycle_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/legal_relation_lifecycle_signals.schema.json","sha256":"420613a5900c4360487b89b978efedde58f5ddc61644130e4b9e63ef8ab33d8b","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-046","path":"signals/schemas/timeline_notice_condition_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/timeline_notice_condition_signals.schema.json","sha256":"99c66208524155cea6bbd5e24fd26998cc9b653c89b24b569c793e36f1623d35","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-047","path":"signals/schemas/asset_right_state_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/asset_right_state_signals.schema.json","sha256":"fc34fbb3d33a284c3d57f3c278cbda8b3555ef26ee2f06b803fd2410ebce38b6","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-048","path":"signals/schemas/liability_causation_damage_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/liability_causation_damage_signals.schema.json","sha256":"34102cb8eeda80773eb62a5ee61e3d714bf90424ed5350dcac4b7bf873a72c5a","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-049","path":"signals/schemas/defense_exception_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/defense_exception_signals.schema.json","sha256":"010148c15e60e4d112b142f80b1723c06e34ba22b3edefae9e4371f2353b053e","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-050","path":"signals/schemas/evidence_proof_conflict_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/evidence_proof_conflict_signals.schema.json","sha256":"c419f568e28c06c629bc715aff7b0737b77e9c4871c91d4fae8f6ecf04196390","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-051","path":"signals/schemas/calculation_requirements.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/calculation_requirements.schema.json","sha256":"7fdb5ef0f50d7af22ac417abc4022cd238f5ab0dc942866420616729a9e3571f","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-052","path":"signals/schemas/remedy_enforcement_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/remedy_enforcement_signals.schema.json","sha256":"999e1969b983748f209e9b5239f7edd0ec43bc642d9ea8fd7edbf34f9ce653f3","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-053","path":"signals/schemas/legal_effect_routes.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/legal_effect_routes.schema.json","sha256":"c24cb740c370aa2477199a0225be8291164ef5c787601fd962a370c642cc3cc0","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-054","path":"signals/schemas/domain_signal_envelope.schema.v2.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/domain_signal_envelope.schema.v2.json","sha256":"1483d6c5f98083f59172feff9b7c15b44d3ed789db5b6172d0de05f06e9d3fbc","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-055","path":"signals/schemas/signal_manifest.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/signal_manifest.schema.json","sha256":"5e72084780b82b29582c9ffcf48f3e4894d7c0b152e5ce8df394583c07dde681","source_manifest":"signals/signal_registry.v2.json"}],"contract_manifest_ref":{"mode":"CONDITIONAL_RELOCATION_ONLY","path":null,"sha256":null,"status":"NOT_REQUIRED_DEFAULT_PATHS"},"expected_concrete_path_count":55,"full_stage1_runtime_release_status":"STAGE1_NOT_RELEASE_READY"}},"adapter_decisions":[{"adapter_id":"S2A-SIGNAL-ALL-V1","decision":{"file_conservation_equation":"semantic_file_rows + integrity_only_file_rows = Counter(signal_manifest.files[])","global_signal_id_uniqueness_assumed":false,"integrity_only_kinds":["compatibility_view"],"manifest_selector":"/downstream_read_sets/stage2","physical_path_rule":"U/signals/","record_conservation_equation":"used_record_occurrences + unused_record_occurrences + unmapped_record_occurrences = records_from_semantic_files","record_occurrence_key":["manifest_transaction_id","file_path","record_ordinal","signal_id"],"row_order":"PRESERVE_MANIFEST_ORDER","row_source":"/files","semantic_kinds":["canonical","domain_signal"],"sentinel":["ALL"]}},{"adapter_id":"S2A-DUAL-SG01-V1","decision":{"comparison":"PARSED_CANONICAL_PROJECTION_EQUAL","payload_root":"/domain_activation_manifest","projection_json_pointers":["/schema_version","/signal_id","/status","/registry_version","/registry_index_sha256","/screening_sha256","/domain_entries","/active_domain_ids","/supporting_domain_ids","/monitor_domain_ids","/expected_runnable_domain_ids","/required_calculation_domains","/unrouted_material","/conservation_gate","/fail_open_policy","/review_items","/contract_guards"],"raw_hash_policy":"PRESERVE_AND_VERIFY_SEPARATELY","routing_path":"routing/domain_activation_manifest.json","set_semantics_json_pointers":["/active_domain_ids","/supporting_domain_ids","/monitor_domain_ids","/expected_runnable_domain_ids","/required_calculation_domains"],"signal_path":"signals/domain_activation_manifest.json"}},{"adapter_id":"S2A-P1-HANDOFF-FLAT-V1","decision":{"count_field_required":false,"logical_input_id":"P1_REVIEW_HANDOFF","p1_digest_keys":["evidence_indexed_sha256","evidence_event_candidates_sha256","b1_gate_sha256","b2_gate_sha256","screening_sha256","activation_manifest_sha256","registry_index_sha256"],"review_items_json_pointer":"/review_items","schema_version":"stage1_part1_soft_gate_handoff.v1","seal_sources":["routing/domain_screening.json","routing/domain_activation_manifest.json","domains/_registry_index.json"],"source_stage":"P1","status_json_pointer":"/handoff_status","wrapper_json_pointer":""}},{"adapter_id":"S2A-P2-HANDOFF-FLAT-V1","decision":{"count_field_required":false,"logical_input_id":"P2_REVIEW_HANDOFF","review_items_json_pointer":"/review_items","schema_version":"stage1_part2_review_handoff.v1","seal_sources":["BO.json","signals/signal_manifest.json"],"source_stage":"P2","status_json_pointer":"/status","wrapper_json_pointer":""}},{"adapter_id":"S2A-P3-HANDOFF-WRAPPED-V1","decision":{"count_field_required":true,"logical_input_id":"P3_REVIEW_HANDOFF","review_items_json_pointer":"/review_items","schema_version":"stage1_part3_review_handoff.v1","seal_sources":["legal_effect_structures.json","validation_assets/routing/part3_receipt.json"],"source_stage":"P3","status_json_pointer":"/status","wrapper_json_pointer":"/stage1_part3_review_handoff"}},{"adapter_id":"S2A-P4-HANDOFF-WRAPPED-V1","decision":{"count_field_required":true,"logical_input_id":"P4_REVIEW_HANDOFF","review_items_json_pointer":"/review_items","schema_version":"stage1_part4_review_handoff.v1","seal_sources":["Fact_Ledger_base.json","validation_assets/routing/part4_receipt.json","stage1_tmp/fact_ledger/fact_ledger_writer_report.json"],"source_stage":"P4","status_json_pointer":"/status","wrapper_json_pointer":"/stage1_part4_review_handoff"}},{"adapter_id":"S2-REVIEW-MAP-V1","decision":{"aggregate_handoff_status_never_resolves_item":true,"handoff_status_mappings":[{"source_stage":"P1","source_value":"READY_NO_REVIEW","technical_disposition":"AVAILABLE"},{"source_stage":"P1","source_value":"READY_WITH_REVIEW","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P1","source_value":"BLOCKED","technical_disposition":"UNAVAILABLE"},{"source_stage":"P2","source_value":"PENDING_FINALIZE","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P2","source_value":"FINALIZED","technical_disposition":"AVAILABLE"},{"source_stage":"P3","source_value":"OPEN","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P3","source_value":"FINALIZED","technical_disposition":"AVAILABLE"},{"source_stage":"P4","source_value":"OPEN","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P4","source_value":"FINALIZED","technical_disposition":"AVAILABLE"}],"mappings":[{"mapping_id":"S2RM-001","normalized_partition":"SUPPORTED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"SUPPORTED","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-002","normalized_partition":"CONDITIONAL","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"CONDITIONAL","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-003","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"UNRESOLVED","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-004","normalized_partition":"EXCLUDED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"EXCLUDED","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-005","normalized_partition":"SUPPORTED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"observed","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-006","normalized_partition":"CONDITIONAL","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"inferred","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-007","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"contested","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-008","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"missing_required","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-009","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"review","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-010","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"NO_SUPPORT","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-011","normalized_partition":"CONDITIONAL","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"info","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-012","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"review","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-013","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"SOFT_WARNING","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-014","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"hard_warning","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-015","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"HARD_WARNING","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-016","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"block","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-017","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"BLOCK","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"}],"normalized_partitions":["SUPPORTED","CONDITIONAL","UNRESOLVED","EXCLUDED","UNMAPPED"],"resolution_inference_allowed":false,"unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"}},{"adapter_id":"S2A-BO-V8-LIST-V1","decision":{"logical_input_id":"BO","open_source_fields_policy":"PRESERVE_UNMODELED_FIELDS_WITH_RAW_HASH","producer_id":"Task_C_BO_F0_final_bo_compiler_gate_writer","required_item_fields":["BO_ID","id","BOType","ActionType","JuristicAct","Action","Reason","PriorAct","ReasonRefs","Legal_Keywords","core_field_base","amount","EvidenceTitles","Evidence","source_evidence_indexes","provenance","downstream_seed_refs","extensions"],"required_root_fields":[],"root_shape":"ARRAY","schema_contract_version":null}},{"adapter_id":"S2A-EVIDENCE-V3-ENVELOPE-V1","decision":{"logical_input_id":"EVIDENCE_INDEXED","open_source_fields_policy":"PRESERVE_UNMODELED_FIELDS_WITH_RAW_HASH","producer_id":"Task_B1_quality_gate_evidence_indexed","required_root_fields":["schema_contract_version","items"],"root_shape":"OBJECT_ENVELOPE","schema_contract_version":"evidence_indexed.v3"}},{"adapter_id":"S2A-EVENTS-V1-ENVELOPE-V1","decision":{"logical_input_id":"EVIDENCE_EVENT_CANDIDATES","open_source_fields_policy":"PRESERVE_UNMODELED_FIELDS_WITH_RAW_HASH","producer_id":"Task_B2_quality_gate_event_candidates","required_root_fields":["schema_version","items"],"root_shape":"OBJECT_ENVELOPE","schema_contract_version":"evidence_event_candidates.v1"}},{"adapter_id":"S2A-DOMAIN-CONFIG-V1","decision":{"accepted_schema_version":"stage1_domain_config.v1","depends_on_legal_dependency_allowed":false,"rebuttal_slot_synthesis_allowed":false,"required_slot_fields":["element_slots","opposing_fact_slots","defense_map","calculation_bindings","emits_signals"],"undeclared_slot_policy":"PRESERVE_AS_PROPOSED_NEW_SLOT_ISSUE"}},{"adapter_id":"S2A-DOMAIN-CONFIG-V2","decision":{"accepted_schema_version":"stage1_domain_config.v2","depends_on_legal_dependency_allowed":false,"rebuttal_slot_synthesis_allowed":false,"required_slot_fields":["element_slots","opposing_fact_slots","defense_map","calculation_bindings","emits_signals"],"undeclared_slot_policy":"PRESERVE_AS_PROPOSED_NEW_SLOT_ISSUE"}},{"adapter_id":"S2A-FACT-LEDGER-CURRENT-V8-V1","decision":{"bo_source_bo_id_multiset_equality_required":true,"fact_id_pattern":"^F-[0-9]{3,}$","legacy_adapter_status":"DISABLED_NO_APPROVED_ADAPTER","producer_generation":"CURRENT_V8","required_row_fields":["fact_id","source_bo_id","domain_effects","calculation_requests"],"root_shape":"ARRAY"}},{"adapter_id":"PA-SG-COMPILER-001","decision":{"bidirectional_match_allowed":true,"global_alias_allowed":false,"orchestration_producer_id":"Task_C_BO_S0_signal_bundle_writer","schema_writer_id":"Task_C_BO_S0_canonical_signal_compiler","scope":"STAGE1_PART2_SIGNAL_TRANSACTION_ONLY"}}],"release_class":"DEV_FIXTURE_RELEASE","limits":{"max_file_bytes":33554432,"max_run_bytes":268435456,"max_json_depth":96,"max_json_items":1000000}}''') + + + RAW_STAGE1_RESULTS = { + 'evidence_indexed': r"""{{prev.evidence_indexed.json}}""", + 'evidence_event_candidates': r"""{{prev.evidence_event_candidates.json}}""", + 'client_goal': r"""{{prev.client_goal.json}}""", + 'domain_screening': r"""{{prev.routing/domain_screening.json}}""", + 'domain_activation_manifest': r"""{{prev.routing/domain_activation_manifest.json}}""", + 'b1_evidence_indexed_gate': r"""{{prev.quality_gates/B1_evidence_indexed_gate.json}}""", + 'b2_event_candidates_gate': r"""{{prev.quality_gates/B2_event_candidates_gate.json}}""", + 'stage1_part1_soft_gate_handoff': r"""{{prev.quality_gates/stage1_part1_soft_gate_handoff.json}}""", + 'bo': r"""{{prev.BO.json}}""", + 'signal_manifest': r"""{{prev.signals/signal_manifest.json}}""", + 'stage1_part2_review_handoff': r"""{{prev.quality_gates/stage1_part2_review_handoff.json}}""", + 'legal_effect_structures': r"""{{prev.legal_effect_structures.json}}""", + 'stage1_part3_review_handoff': r"""{{prev.quality_gates/stage1_part3_review_handoff.json}}""", + 'fact_ledger_base': r"""{{prev.Fact_Ledger_base.json}}""", + 'fact_ledger_writer_report': r"""{{prev.stage1_tmp/fact_ledger/fact_ledger_writer_report.json}}""", + 'stage1_part4_review_handoff': r"""{{prev.quality_gates/stage1_part4_review_handoff.json}}""", + } + + + STATUS_PATH = "ingress/ingress_status.json" + NORMAL_PATHS = frozenset({"ingress/stage1_input_manifest.json", "ingress/intake_report.json", "review/issue_ledger.base.json", "context/case_context.json", STATUS_PATH}) + BLOCKED_PATHS = frozenset({"ingress/stage1_input_manifest.json", "ingress/intake_report.json", "review/issue_ledger.base.json", "ingress/technical_diagnostic.json", STATUS_PATH}) + ROW_KEYS = { + "bo": ("business_objects", "BO", "rows", "items"), + "fact_ledger_base": ("facts", "fact_ledger", "rows", "items"), + "legal_effect_structures": ("structures", "structure_records", "legal_effect_structures", "rows", "items"), + "evidence_indexed": ("evidence", "evidence_items", "rows", "items"), + "evidence_event_candidates": ("events", "event_candidates", "rows", "items"), + } + WRAPPER_KEYS = ("payload", "data", "fact_ledger_base", "Fact_Ledger_base", "legal_effect_structures") + REVIEW_ARRAY_KEYS = frozenset({"review_items", "review_queue", "blocked_review_items", "unresolved_review_items", "review_findings", "hard_gate_findings"}) + + + def _pointer_token(value: str) -> str: + return value.replace("~", "~0").replace("/", "~1") + + + def _row_locations(document: Any, keys: Sequence[str], pointer: str = "") -> list[tuple[str, Any]]: + if isinstance(document, list): + return [(f"{pointer}/{i}", row) for i, row in enumerate(document)] + if not isinstance(document, dict): + raise IngressError("SOURCE_ROWS_SHAPE", "record source must be an array or approved envelope") + arrays = [(key, document[key]) for key in keys if isinstance(document.get(key), list)] + if len(arrays) > 1: + raise IngressError("SOURCE_ROWS_AMBIGUOUS", "multiple record arrays in one source envelope") + if arrays: + key, rows = arrays[0] + return [(f"{pointer}/{_pointer_token(key)}/{i}", row) for i, row in enumerate(rows)] + nested = [key for key in WRAPPER_KEYS if isinstance(document.get(key), dict)] + if len(nested) != 1: + raise IngressError("SOURCE_ROWS_SHAPE", "approved record array is missing or ambiguous") + key = nested[0] + return _row_locations(document[key], keys, f"{pointer}/{_pointer_token(key)}") + + + def _array_rows(document: Any, keys: Sequence[str]) -> list[Any]: + if document is None: + return [] + return [row for _, row in _row_locations(document, keys)] + + + def _json_value(raw: Any) -> Any: + if not isinstance(raw, (str, bytes)): + return raw + if isinstance(raw, str) and re.fullmatch(r"\s*\{\{[^{}]+\}\}\s*", raw): + raise IngressError("PREV_REFERENCE_UNRESOLVED", "required backend result reference was not resolved") + try: + return load_json_strict(raw) + except IngressError: + if isinstance(raw, str) and raw.strip() and not raw.lstrip().startswith(("{", "[", '"')): + return raw.strip() + raise + + + def validate_direct_roots(run_root: Any, deployment_root: Any) -> dict[str, str]: + result = { + "stage1_run_root_ref": _inline_relative_path(_json_value(run_root), code="STAGE1_RUN_ROOT_INVALID"), + "stage1_deployment_root_ref": _inline_relative_path(_json_value(deployment_root), code="STAGE1_DEPLOYMENT_ROOT_INVALID"), + } + output = f"stage2_runs/from-stage1/{result['stage1_run_root_ref']}/s2_00" + out = PurePosixPath(output) + for value in result.values(): + original = PurePosixPath(value) + if out == original or original in out.parents or out in original.parents: + raise IngressError("OUTPUT_SOURCE_OVERLAP", "output and source roots must be disjoint") + result["output_root"] = output + return result + + + def _previous_source(value: Any, expected_path: str) -> tuple[bytes | None, Any | None]: + """Return exact bytes when supplied; otherwise retain parsed value for comparison.""" + if isinstance(value, bytes): + load_json_strict(value) + return value, None + parsed = _json_value(value) + if isinstance(parsed, dict) and "content_base64" in parsed: + return _inline_binary_envelope(canonical_json_bytes(parsed).decode(), expected_path), None + if isinstance(parsed, dict) and set(parsed).issubset({"path", "content", "text", "name", "doc_name", "sha256", "byte_length"}): + supplied_path = parsed.get("path", parsed.get("doc_name", parsed.get("name"))) + if supplied_path is not None and supplied_path != expected_path: + raise IngressError("PREV_SOURCE_PATH_MISMATCH", "backend result names a different source file") + content = parsed.get("content", parsed.get("text")) + if isinstance(content, str): + raw = content.encode("utf-8") + load_json_strict(raw) + return raw, None + if content is not None: + return None, content + if supplied_path is not None: + return None, None + if isinstance(parsed, str): + if parsed != expected_path: + raise IngressError("PREV_SOURCE_PATH_MISMATCH", "backend result names a different source file") + return None, None + if isinstance(parsed, (dict, list)): + return None, parsed + raise IngressError("PREV_SOURCE_SHAPE", "backend result must provide source JSON, raw bytes, or its exact path") + + + def _copy_to_temp(root: Path, path: str, raw: bytes) -> None: + safe = _safe_relative_path(path) + target = root.joinpath(*safe.parts) + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + + + def _walk_values(value: Any, pointer: str = "") -> Iterable[tuple[str, Any]]: + yield pointer, value + if isinstance(value, dict): + for key, item in value.items(): + yield from _walk_values(item, f"{pointer}/{_pointer_token(key)}") + elif isinstance(value, list): + for index, item in enumerate(value): + yield from _walk_values(item, f"{pointer}/{index}") + + + def _schema_dependencies(document: Mapping[str, Any], current_path: str, locks: Mapping[str, Any]) -> set[str]: + dependencies = set() + for _, item in _walk_values(document): + if not isinstance(item, dict) or not isinstance(item.get("$ref"), str): + continue + ref = item["$ref"].split("#", 1)[0] + if not ref: + continue + candidates = [path for path, row in locks.items() if row.get("schema_id") == ref] + if not candidates and "://" not in ref: + relative = posixpath.normpath(posixpath.join(posixpath.dirname(current_path), ref)) + if relative in locks: + candidates = [relative] + elif ref in locks: + candidates = [ref] + if not candidates: + candidates = [path for path in locks if PurePosixPath(path).name == PurePosixPath(ref).name] + if len(candidates) != 1: + raise IngressError("SCHEMA_DEPENDENCY_UNBOUND", "schema reference is not uniquely bound to Stage 1 deployment") + dependencies.add(candidates[0]) + return dependencies + + + def hydrate_stage1(localdocs: _InlineLocaldocs, temp_root: Path, roots: Mapping[str, str], stage1_results: Mapping[str, Any], policy: Mapping[str, Any]) -> dict[str, Any]: + """Reuse prev results; fetch only missing raw bytes and needed upstream dependencies.""" + stage1_root = temp_root / "stage1" + deployment_root = temp_root / "deployment" + stage1_root.mkdir(); deployment_root.mkdir() + observed: dict[str, bytes] = {} + documents: dict[str, Any] = {} + source_snapshots: dict[str, Snapshot] = {} + issues = [] + total = 0 + def remember(path: str, raw: bytes) -> None: + nonlocal total + if path in observed: + if observed[path] != raw: + raise IngressError("SOURCE_PATH_CONTENT_CONFLICT", "one source path has conflicting results") + return + if len(raw) > MAX_FILE_BYTES: + raise IngressError("SOURCE_SIZE_LIMIT", "input exceeds per-file byte limit") + total += len(raw) + if total > MAX_RUN_BYTES: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "input set exceeds byte limit") + observed[path] = raw + for contract in DEFAULT_SOURCE_CONTRACTS: + logical = contract["logical_input_id"] + relative = contract["path"] + logical_path = f"{roots['stage1_run_root_ref']}/{relative}" + if logical not in stage1_results: + raise IngressError("PREV_SOURCE_MISSING", "required Stage 1 result reference is missing", logical_input_id=logical) + exact, parsed = _previous_source(stage1_results[logical], logical_path) + raw = exact if exact is not None else localdocs.read_binary_optional(logical_path) + if raw is None: + issues.append(_issue("SOURCE_MISSING", source_refs=[logical])) + continue + value = load_json_strict(raw) + if parsed is not None and not _json_equal(value, parsed): + raise IngressError("PREV_SOURCE_CONTENT_MISMATCH", "backend result differs from the original file", logical_input_id=logical) + remember(logical_path, raw) + _copy_to_temp(stage1_root, relative, raw) + documents[logical] = value + source_snapshots[logical] = open_bounded_snapshot(stage1_root, relative, logical_input_id=logical) + manifest = documents.get("signal_manifest") + if isinstance(manifest, dict): + files = manifest.get("files") + if not isinstance(files, list): + raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest must contain its actual files array") + for index, row in enumerate(files): + if not isinstance(row, dict) or not isinstance(row.get("path"), str): + raise IngressError("SIGNAL_FILE_ROW_SHAPE", "signal manifest row is malformed") + relative = _safe_relative_path(row["path"]).as_posix() + if relative.startswith("signals/"): + raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "signal row path must not repeat signals/") + relative = f"signals/{relative}" + path = f"{roots['stage1_run_root_ref']}/{relative}" + # A dynamic file is an existing Stage 1 path selected by its manifest. + # Provided raw results can be reused; no new result-list request is created. + provided = stage1_results.get(relative) + if provided is not None: + exact, parsed = _previous_source(provided, path) + else: + exact, parsed = None, None + raw = exact if exact is not None else localdocs.read_binary(path) + if parsed is not None and not _json_equal(load_json_strict(raw), parsed): + raise IngressError("PREV_SOURCE_CONTENT_MISMATCH", "dynamic result differs from its source") + remember(path, raw) + _copy_to_temp(stage1_root, relative, raw) + locks = {row["path"]: row for row in policy["dependency_locks"]["stage1"]["concrete_paths"]} + if len(locks) != len(policy["dependency_locks"]["stage1"]["concrete_paths"]): + raise IngressError("STAGE1_DEPENDENCY_DUPLICATE_PATH", "upstream dependency table contains duplicate paths") + deployment_snapshots: dict[str, Snapshot] = {} + deployment_documents: dict[str, Any] = {} + needed = {"domains/_registry_index.json", "signals/signal_registry.v2.json"} + needed.update(row["schema_ref"]["path"] for row in policy["stage1_sources"] if isinstance(row.get("schema_ref"), dict)) + activation = documents.get("domain_activation_manifest") + payload = _activation_payload(activation) if isinstance(activation, dict) else {} + for domain in payload.get("active_domain_ids", []): + needed.add(f"domains/{_safe_relative_path(str(domain)).as_posix()}/domain_config.json") + while needed: + relative = min(needed); needed.remove(relative) + if relative in deployment_documents: + continue + row = locks.get(relative) + if row is None: + raise IngressError("STAGE1_DEPENDENCY_UNBOUND", "required upstream dependency is not pinned") + expected = _inline_sha256(row.get("sha256"), code="STAGE1_DEPENDENCY_UNBOUND") + path = f"{roots['stage1_deployment_root_ref']}/{relative}" + raw = localdocs.read_binary(path) + if hashlib.sha256(raw).hexdigest() != expected: + raise IngressError("STAGE1_DEPENDENCY_HASH_MISMATCH", "upstream deployment file differs from its pin") + remember(path, raw) + value = load_json_strict(raw) + _copy_to_temp(deployment_root, relative, raw) + deployment_snapshots[relative] = open_bounded_snapshot(deployment_root, relative, logical_input_id=f"deployment:{relative}") + deployment_documents[relative] = value + if isinstance(value, dict): + needed.update(_schema_dependencies(value, relative, locks) - deployment_documents.keys()) + if relative == "signals/signal_registry.v2.json" and isinstance(value, dict): + for entry in value.get("entries", []): + if isinstance(entry, dict) and isinstance(entry.get("schema"), str): + schema = entry["schema"] + needed.add(schema if schema.startswith("signals/") else f"signals/{schema}") + envelope = value.get("domain_envelope") + if isinstance(envelope, str): + needed.add(envelope if envelope.startswith("signals/") else f"signals/{envelope}") + return {"stage1_root": stage1_root, "deployment_root": deployment_root, "snapshots": source_snapshots, "documents": documents, "deployment_snapshots": deployment_snapshots, "deployment_documents": deployment_documents, "observed": observed, "issues": issues} + + + def verify_remote_stability(localdocs: _InlineLocaldocs, observed: Mapping[str, bytes]) -> None: + for path, expected in sorted(observed.items()): + if localdocs.read_binary(path) != expected: + raise IngressError("HYDRATION_SOURCE_CHANGED", "source differs from the first read/result reference") + + + def _provenance(logical: str, pointer: str, documents: Mapping[str, Any]) -> dict[str, Any]: + found, value = _json_pointer_value(documents[logical], pointer) + if not found: + raise IngressError("SOURCE_POINTER_INVALID", "projection pointer does not address the original") + return _source_ref(logical, pointer, value) + + + def normalize_review_items(review_documents: Mapping[str, Any], release_lock: Mapping[str, Any] | None = None) -> dict[str, Any]: + """Preserve every review/gate occurrence, its content, exact pointer, and blocking state.""" + mapping = _adapter_decision(release_lock or {}, "S2-REVIEW-MAP-V1") or {} + table = {(row.get("source_stage", "ANY"), row.get("source_field_kind"), str(row.get("source_value"))): row.get("normalized_partition") for row in mapping.get("mappings", [])} + rows = [] + partitions = Counter() + adapter_issues = [] + for stage_number in range(1, 5): + logical = 'stage1_part1_soft_gate_handoff' if stage_number == 1 else f'stage1_part{stage_number}_review_handoff' + if logical not in review_documents: + continue + adapter = f'S2A-P{stage_number}-HANDOFF-' + ('FLAT-V1' if stage_number < 3 else 'WRAPPED-V1') + decision = _adapter_decision(release_lock or {}, adapter) + if not isinstance(decision, dict): + adapter_issues.append(_issue('HANDOFF_ADAPTER_CONTRACT_MISSING', source_refs=[logical])) + continue + found, wrapper = _json_pointer_value(review_documents[logical], decision.get('wrapper_json_pointer')) + if not found or not isinstance(wrapper, dict): + adapter_issues.append(_issue(f'P{stage_number}_WRAPPER_MISSING', source_refs=[logical])) + continue + if wrapper.get('schema_version') != decision.get('schema_version'): + adapter_issues.append(_issue(f'P{stage_number}_HANDOFF_SCHEMA_VERSION_MISMATCH', source_refs=[logical])) + found, handoff_items = _json_pointer_value(wrapper, decision.get('review_items_json_pointer')) + if not found or not isinstance(handoff_items, list): + adapter_issues.append(_issue(f'P{stage_number}_REVIEW_ITEMS_SHAPE', source_refs=[logical])) + elif decision.get('count_field_required') is True and wrapper.get('review_item_count') != len(handoff_items): + adapter_issues.append(_issue('REVIEW_CONSERVATION_FAILED', source_refs=[logical])) + for logical, document in sorted(review_documents.items()): + stage_match = re.search(r"part([1-4])", logical) + stage = f"P{stage_match.group(1)}" if stage_match else "ANY" + for pointer, value in _walk_values(document): + if not isinstance(value, dict): + continue + for key in sorted(REVIEW_ARRAY_KEYS): + items = value.get(key) + if not isinstance(items, list): + continue + for index, item in enumerate(items): + item_pointer = f"{pointer}/{_pointer_token(key)}/{index}" + raw_status = item.get("status") if isinstance(item, dict) else None + raw_severity = item.get("severity") if isinstance(item, dict) else None + kind = "REVIEW_ITEM_STATUS" if raw_status is not None else "REVIEW_ITEM_SEVERITY" + raw_value = str(raw_status if raw_status is not None else raw_severity) + partition = table.get((stage, kind, raw_value), table.get(("ANY", kind, raw_value), "UNMAPPED")) + explicit_block = key == "blocked_review_items" or isinstance(item, dict) and (item.get("blocking") is True or item.get("blocked") is True or str(item.get("status", "")).upper() == "BLOCKED" or str(item.get("severity", "")).upper() in {"BLOCKING", "CRITICAL", "FATAL"}) + row = {"review_ref": f"{logical}#{item_pointer}", "source_ref": _provenance(logical, item_pointer, review_documents), "source_status_raw": raw_status, "source_severity_raw": raw_severity, "partition": partition, "blocking": bool(explicit_block), "content": item} + rows.append(row); partitions[partition] += 1 + # Count source occurrences independently; duplicates remain distinct by pointer. + expected = sum(len(v[k]) for doc in review_documents.values() for _, v in _walk_values(doc) if isinstance(v, dict) for k in REVIEW_ARRAY_KEYS if isinstance(v.get(k), list)) + return {"normalized_occurrences": rows, "partition_counts": dict(partitions), "conservation_status": "PASS" if expected == len(rows) and len({r['review_ref'] for r in rows}) == expected and not any(x["issue_code"] == "REVIEW_CONSERVATION_FAILED" for x in adapter_issues) else "FAIL", "_issues": adapter_issues} + + + def _project_content(value: Any) -> Any: + if not isinstance(value, dict): + return value + # Envelope/protocol metadata remains reachable through provenance instead of copying files. + return {key: item for key, item in value.items() if key not in {"schema_version", "schema_contract_version", "producer_id", "created_by", "finalized_by", "metadata", "meta"}} + + + def compile_case_context(documents: Mapping[str, Any], signal_all: Mapping[str, Any], reviews: Mapping[str, Any], deployment_documents: Mapping[str, Any]) -> dict[str, Any]: + """Normalize original records once and group only explicit source relationships.""" + members = [] + lookup = {} + identities = {"bo": ("BO", ("BO_ID",)), "fact_ledger_base": ("FACT", ("fact_id",)), "legal_effect_structures": ("LES", ("structure_id", "legal_effect_structure_id")), "evidence_indexed": ("EVIDENCE", ("evidence_id", "id")), "evidence_event_candidates": ("EVENT", ("event_id", "id"))} + raw_rows = {} + for logical, keys in ROW_KEYS.items(): + for pointer, value in _row_locations(documents[logical], keys): + if not isinstance(value, dict): + raise IngressError("SOURCE_RECORD_SHAPE", "original record must be an object") + kind, id_keys = identities[logical] + identifier = next((str(value[k]) for k in id_keys if value.get(k) is not None), None) + ref = f"{logical}#{pointer}" + if identifier is not None: + if (kind, identifier) in lookup: + raise IngressError("SOURCE_RECORD_ID_DUPLICATE", "original record ID occurs more than once") + lookup[(kind, identifier)] = ref + member = {"member_ref": ref, "kind": kind, "stage1_id": identifier, "source_ref": _provenance(logical, pointer, documents), "field_refs": {key: _provenance(logical, f"{pointer}/{_pointer_token(key)}", documents) for key in value}, "projection": _project_content(value)} + members.append(member); raw_rows[ref] = (logical, pointer, value) + relationships = []; candidates = []; unresolved = [] + parent = {m['member_ref']: m['member_ref'] for m in members} + def find(ref): + while parent[ref] != ref: + parent[ref] = parent[parent[ref]]; ref = parent[ref] + return ref + def join(a,b): + a,b=find(a),find(b) + if a!=b:parent[max(a,b)]=min(a,b) + def edge(source, kind, identifier, relation, pointer, *, hard=True): + logical, _, _ = raw_rows[source] + target = lookup.get((kind, str(identifier))) + row = {"from_ref": source, "to_ref": target, "target_stage1_id": str(identifier), "relation_kind": relation, "source_ref": _provenance(logical, pointer, documents), "hard_join_allowed": hard, "disposition": "OBSERVED" if target else "UNEVALUABLE"} + if target is None: + unresolved.append(row) + elif hard: + relationships.append(row); join(source,target) + else: + candidates.append(row) + for member in members: + ref=member['member_ref']; logical,pointer,row=raw_rows[ref] + if member['kind']=='FACT': + if row.get('source_bo_id') is not None:edge(ref,'BO',row['source_bo_id'],'SAME_BO_ID',f"{pointer}/source_bo_id") + for keys,kind,relation in [(('evidence_refs','evidence_ids'),'EVIDENCE','SAME_EVIDENCE_REF'),(('event_refs','event_ids'),'EVENT','SAME_EVENT_REF')]: + key=next((k for k in keys if isinstance(row.get(k),list)),None) + if key: + for index,identifier in enumerate(row[key]):edge(ref,kind,identifier,relation,f"{pointer}/{key}/{index}") + for key in ('relations','explicit_relations','candidate_relations'): + for index,item in enumerate(row.get(key,[]) if isinstance(row.get(key),list) else []): + if not isinstance(item,dict):continue + target=item.get('target_fact_id',item.get('to_fact_id')) + relation=str(item.get('relation_kind',item.get('kind','UNCLASSIFIED'))) + if target is not None:edge(ref,'FACT',target,relation,f"{pointer}/{key}/{index}",hard=relation=='EXPLICIT_CASE_RELATION') + elif member['kind']=='LES': + for index,identifier in enumerate(row.get('source_bo_ids',[]) if isinstance(row.get('source_bo_ids'),list) else []):edge(ref,'BO',identifier,'SOURCE_BO_ATTACHMENT',f"{pointer}/source_bo_ids/{index}") + elif member['kind']=='EVENT': + key=next((k for k in ('evidence_refs','evidence_ids') if isinstance(row.get(k),list)),None) + if key: + for index,identifier in enumerate(row[key]):edge(ref,'EVIDENCE',identifier,'SAME_EVIDENCE_REF',f"{pointer}/{key}/{index}") + member_by_ref = {row['member_ref']: row for row in members} + grouped=defaultdict(list) + for ref in sorted(parent):grouped[find(ref)].append(ref) + clusters=[]; membership={} + for index,refs in enumerate(sorted(grouped.values(),key=lambda v:v[0]),1): + cluster_ref=f"CL-{index:03d}" + clusters.append({'cluster_ref':cluster_ref,'member_refs':refs,'source_refs':[member_by_ref[ref]['source_ref'] for ref in refs]}) + for ref in refs:membership[ref]=cluster_ref + cluster_edges=sorted({(membership[r['from_ref']],membership[r['to_ref']]) for r in candidates if r['relation_kind'] in CANDIDATE_RELATION_KINDS and membership[r['from_ref']]!=membership[r['to_ref']]}) + sccs=_tarjan_scc([c['cluster_ref'] for c in clusters],cluster_edges) + component={ref:index for index,group in enumerate(sccs) for ref in group} + indegree={i:0 for i in range(len(sccs))}; adjacency=defaultdict(set) + for left,right in cluster_edges: + a,b=component[left],component[right] + if a!=b and b not in adjacency[a]:adjacency[a].add(b); indegree[b]+=1 + ready=sorted(i for i in indegree if indegree[i]==0); waves=[] + while ready: + waves.append([sccs[i] for i in ready]); upcoming=[] + for i in ready: + for j in sorted(adjacency[i]): + indegree[j]-=1 + if indegree[j]==0:upcoming.append(j) + ready=sorted(set(upcoming)) + signal_refs=[] + for occurrence in signal_all.get('record_occurrences',[]): + logical=f"signal:{occurrence['file_path']}" + document=documents[logical] + locations=_record_locations_for_signal(document) + ordinal=occurrence['record_ordinal'] + pointer,value=locations[ordinal] + signal_refs.append({'source_ref':_provenance(logical,pointer,documents),'signal_id':occurrence['signal_id'],'disposition':occurrence['disposition'],'binding_refs':occurrence.get('binding_refs',[]),'projection':_project_content(value)}) + for cluster in clusters: + member_set=set(cluster['member_refs']) + cluster_members = [member_by_ref[ref] for ref in cluster['member_refs']] + bound_ids={f"{m['kind']}:{m['stage1_id']}" for m in cluster_members if m['stage1_id'] is not None} + selected=[] + for index,row in enumerate(signal_refs): + tokens={t.replace('fact_id:','FACT:').replace('source_bo_id:','BO:').replace('bo_id:','BO:').replace('evidence_id:','EVIDENCE:').replace('event_id:','EVENT:') for t in row['binding_refs']} + if tokens & bound_ids:selected.append(index) + cluster['signal_indexes']=selected + cluster['review_refs']=[r['review_ref'] for r in reviews['normalized_occurrences'] if any(str(m['stage1_id']) in _collect_values_for_keys(r['content'], {'fact_id','fact_ids','BO_ID','bo_id','bo_ids','source_bo_id','source_bo_ids','evidence_id','evidence_ids','event_id','event_ids'}) for m in cluster_members if m['stage1_id'] is not None)] + cluster['bundle']={'member_refs':cluster['member_refs'],'signal_indexes':selected,'review_refs':cluster['review_refs']} + slot_links=[]; party_object_refs=[] + for logical,document in documents.items(): + if logical.startswith('deployment:'):continue + for pointer,value in _walk_values(document): + if not isinstance(value,dict):continue + if any(k in value for k in ('slot_id','slot_ref','evidence_slot_id')): + slot_links.append({'source_ref':_provenance(logical,pointer,documents),'projection':_project_content(value),'disposition':'OBSERVED'}) + for key in ('parties','party_refs','object_refs','objects','title_refs'): + if isinstance(value.get(key),(list,dict)): + party_object_refs.append({'kind':key,'source_ref':_provenance(logical,f"{pointer}/{key}",documents)}) + return {'source_documents':[_provenance(logical,'',documents) for logical in sorted(documents) if not logical.startswith('deployment:')], 'members':members,'relationships':relationships,'candidate_dependencies':candidates,'unresolved_relationships':unresolved,'clusters':clusters,'scheduling_waves':waves,'client_goal':{'source_ref':_provenance('client_goal','',documents),'projection':_project_content(documents['client_goal'])},'routing':{'source_ref':_provenance('domain_activation_manifest','',documents),'projection':_activation_payload(documents['domain_activation_manifest'])},'signals':signal_refs,'global_review_refs':[r['review_ref'] for r in reviews['normalized_occurrences']],'object_and_party_refs':party_object_refs,'slot_links':slot_links,'slot_link_status':'OBSERVED' if slot_links else 'UNEVALUABLE','active_profiles':[{'path':path,'sha256':canonical_digest(value),'profile':value} for path,value in sorted(deployment_documents.items()) if re.fullmatch(r'domains/[^/]+/domain_config\.json',path)]} + + + def _record_locations_for_signal(document: Any) -> list[tuple[str, Any]]: + rows=_records_from_signal_document(document) + if isinstance(document,list):return [(f'/{i}',v) for i,v in enumerate(document)] + if not isinstance(document,dict):return [] + if rows == [document]:return [('',document)] + candidates=[(p,v) for p,v in _walk_values(document) if isinstance(v,list) and v==rows] + if len(candidates)!=1: + raise IngressError('SIGNAL_RECORD_POINTER_AMBIGUOUS','signal record array cannot be located uniquely') + p,v=candidates[0] + return [(f'{p}/{i}',item) for i,item in enumerate(v)] + + + def _validate_provenance(value: Any, documents: Mapping[str, Any]) -> None: + for _,row in _walk_values(value): + if not isinstance(row,dict) or not {'logical_artifact_id','json_pointer','raw_value_sha256'}.issubset(row):continue + logical=row['logical_artifact_id'] + if logical not in documents:raise IngressError('SOURCE_REF_UNKNOWN','output refers to an unknown source') + found,raw=_json_pointer_value(documents[logical],row['json_pointer']) + if not found or canonical_digest(raw)!=row['raw_value_sha256']: + raise IngressError('SOURCE_REF_HASH_MISMATCH','output provenance does not match original content') + + + def _clean_issues(issues: Sequence[Mapping[str, Any]]) -> list[dict[str, Any]]: + rows=[]; seen=set() + for row in issues: + cleaned={k:row[k] for k in ('issue_code','severity','impact_scope','scope_refs','source_refs','message') if k in row} + key=canonical_digest(cleaned) + if key not in seen:seen.add(key); rows.append(cleaned) + return sorted(rows,key=canonical_digest) + + + def execute_ingress(hydrated: Mapping[str, Any], roots: Mapping[str, str], *, policy: Mapping[str, Any] = SOURCE_POLICY, fixture: bool = False) -> dict[str, Any]: + """Pure C00-C15 core. Fixture evaluation never enables remote publication.""" + if not fixture and policy.get('release_class')=='DEV_FIXTURE_RELEASE': + raise IngressError('DEV_FIXTURE_REAL_RUN_FORBIDDEN','DEV fixture admission cannot publish a real case') + snapshots=hydrated['snapshots']; deployment=hydrated['deployment_documents']; dep_snapshots=hydrated['deployment_snapshots'] + contracts=resolve_stage1_sources(hydrated['stage1_root']) + ingress=validate_ingress_contracts(snapshots,contracts,policy,deployment_snapshots=dep_snapshots,deployment_documents=deployment) + documents=ingress['documents']; issues=list(hydrated['issues'])+ingress['issues']; checks=[] + signal_all={}; reviews={'normalized_occurrences':[],'partition_counts':{},'conservation_status':'PASS','_issues':[]} + try: + if set(documents)!={r['logical_input_id'] for r in DEFAULT_SOURCE_CONTRACTS}: + raise IngressError('SOURCE_SET_INCOMPLETE','required Stage 1 sources are unavailable') + signal_all=expand_stage2_signal_all(hydrated['stage1_root'],documents['signal_manifest'],signal_registry=deployment.get('signals/signal_registry.v2.json')) + signal_all=bind_signal_occurrences(signal_all,documents) + issues.extend(signal_all['issues']) + for row in signal_all['ordered_file_rows']: + logical=f"signal:{row['file_path']}" + document=signal_all['_parsed_documents_by_path'][row['file_path']] + documents[logical]=document + manifest_row=documents['signal_manifest']['files'][row['manifest_index']] + schema_path=manifest_row.get('schema',manifest_row.get('schema_path')) + if isinstance(schema_path,str): + if not schema_path.startswith('signals/'):schema_path=f'signals/{schema_path}' + schema=deployment.get(schema_path) + if not isinstance(schema,dict):raise IngressError('SIGNAL_SCHEMA_UNBOUND','signal schema is not in the selected upstream closure') + try:_validate_schema_node(document,schema,root_schema=schema,schema_documents=_schema_document_index(deployment),instance_path=logical) + except _SchemaViolation as exc:raise IngressError('SIGNAL_SCHEMA_VALIDATION_FAILED',str(exc)) from exc + activation=signal_all['_parsed_documents_by_path'].get('domain_activation_manifest.json') + if activation is None:raise IngressError('SG01_SIGNAL_ARTIFACT_MISSING','signal ALL lacks domain activation') + verify_activation_projection(documents['domain_activation_manifest'],activation) + seals=verify_cross_artifact_seals(documents,snapshots,{'stage1_domain_registry_index':dep_snapshots['domains/_registry_index.json']} if 'domains/_registry_index.json' in dep_snapshots else {}) + checks.extend(seals['checks']); issues.extend(seals['issues']) + reviews=normalize_review_items(documents,policy) + conserved=check_conservation(documents,signal_all=signal_all,normalized_reviews=reviews,source_snapshots=snapshots) + checks.extend(conserved['checks']); issues.extend(conserved['issues']) + for logical,doc in documents.items(): + if logical.startswith('signal:'):continue + for pointer,value in _walk_values(doc): + if not isinstance(value,dict):continue + if value.get('stage2_auto_progression_allowed') is False or value.get('blocking') is True or value.get('blocked') is True or str(value.get('status',value.get('handoff_status',''))).upper()=='BLOCKED': + issues.append(_issue('UPSTREAM_BLOCKING_GATE',source_refs=[f'{logical}#{pointer}'])) + if any(r['blocking'] for r in reviews['normalized_occurrences']):issues.append(_issue('UPSTREAM_BLOCKING_REVIEW')) + except (IngressError,_SchemaViolation) as exc: + code=exc.code if isinstance(exc,IngressError) else 'SOURCE_SCHEMA_VALIDATION_FAILED' + issues.append(_issue(code,message=str(exc))) + issues=_clean_issues(issues) + serious=any(row.get('severity')=='ERROR' and row.get('issue_code') not in {'PRODUCER_ID_UNEVALUABLE','UNMAPPED_REVIEW_STATUS'} for row in issues) + if any(row.get('parse_status')!='PASS' or row.get('schema_status')=='FAIL' or row.get('seal_status')=='FAIL' for row in ingress['source_contract_rows']):serious=True + if any(c.get('status')=='FAIL' for c in checks):serious=True + status='BLOCKED' if serious else 'READY_WITH_ISSUES' if issues or any(r['partition'] in {'UNRESOLVED','CONDITIONAL','UNMAPPED'} for r in reviews['normalized_occurrences']) or any(r.get('seal_status')=='UNEVALUABLE' for r in ingress['source_contract_rows']) else 'READY' + context=None + if status!='BLOCKED': + try: + context=compile_case_context(documents,signal_all,reviews,deployment) + if not context['clusters']: + raise IngressError('NO_COHERENT_CLUSTER', 'no source records form a usable case context') + if context['unresolved_relationships']: + issues=_clean_issues(issues+[_issue('RELATION_TARGET_UNEVALUABLE',severity='WARNING')]); status='READY_WITH_ISSUES' + _validate_provenance(context,documents) + except IngressError as exc: + issues=_clean_issues(issues+[_issue(exc.code,message=str(exc))]); status='BLOCKED'; context=None + _validate_provenance(reviews['normalized_occurrences'],documents) + header={'schema_version':'stage2_s2_00_direct.v3','algorithm_version':ALGORITHM_VERSION,'stage1_run_root_ref':roots['stage1_run_root_ref'],'stage1_deployment_root_ref':roots['stage1_deployment_root_ref']} + manifest_rows=[{'logical_input_id':row['logical_input_id'],'path':snapshots[row['logical_input_id']].relative_path if row['logical_input_id'] in snapshots else row.get('expected_path'),'raw_sha256':row.get('raw_sha256'),'byte_length':row.get('byte_length'),'parse_status':row.get('parse_status'),'schema_status':row.get('schema_status'),'seal_status':row.get('seal_status'),'run_identity_ref':row.get('run_identity_ref'),'transaction_identity_ref':row.get('transaction_identity_ref')} for row in ingress['source_contract_rows']] + for row in signal_all.get('ordered_file_rows',[]):manifest_rows.append({'logical_input_id':f"signal:{row['file_path']}",'path':row['physical_path'],'raw_sha256':row['raw_sha256'],'byte_length':row['byte_length'],'hash_status':row['hash_status'],'record_count_status':row['record_count_status']}) + deployment_rows=[{'path':path,'raw_sha256':snap.raw_sha256,'byte_length':snap.byte_length} for path,snap in sorted(dep_snapshots.items())] + source_hashes={path:hashlib.sha256(raw).hexdigest() for path,raw in sorted(hydrated['observed'].items())} + files={ + 'ingress/stage1_input_manifest.json':{**header,'sources':manifest_rows,'deployment_sources':deployment_rows}, + 'ingress/intake_report.json':{**header,'checks':checks,'issues':issues,'source_contract_rows':[{k:v for k,v in row.items() if k!='downstream_allowed_actions'} for row in ingress['source_contract_rows']]}, + 'review/issue_ledger.base.json':{**header,'review_items':reviews['normalized_occurrences'],'partition_counts':reviews['partition_counts'],'conservation_status':reviews['conservation_status'],'issues':issues}, + } + if status=='BLOCKED':files['ingress/technical_diagnostic.json']={**header,'status':status,'issues':issues,'checks':checks} + else:files['context/case_context.json']={**header,**context} + serialized={path:canonical_json_bytes(value)+b'\n' for path,value in files.items()} + artifact_rows=[{'path':path,'raw_sha256':hashlib.sha256(raw).hexdigest(),'byte_length':len(raw)} for path,raw in sorted(serialized.items())] + files[STATUS_PATH]={**header,'status':status,'output_root':roots['output_root'],'source_hashes':source_hashes,'artifacts':artifact_rows,'written_last':True,'publication_semantics':'STATUS_LAST_LOGICAL_COMMIT'} + serialized[STATUS_PATH]=canonical_json_bytes(files[STATUS_PATH])+b'\n' + validate_output_files(serialized,roots) + return {'status':status,'files':serialized,'documents':documents} + + + def validate_output_files(files: Mapping[str, bytes], roots: Mapping[str, str]) -> dict[str, Any]: + status=load_json_strict(files.get(STATUS_PATH,b'')) + allowed=NORMAL_PATHS if status.get('status') in {'READY','READY_WITH_ISSUES'} else BLOCKED_PATHS if status.get('status')=='BLOCKED' else frozenset() + if set(files)!=allowed:raise IngressError('OUTPUT_ARTIFACT_SET_INVALID','output set differs from its processing state') + common={'schema_version','algorithm_version','stage1_run_root_ref','stage1_deployment_root_ref'} + fields={ + 'ingress/stage1_input_manifest.json':{'sources','deployment_sources'}, + 'ingress/intake_report.json':{'checks','issues','source_contract_rows'}, + 'review/issue_ledger.base.json':{'review_items','partition_counts','conservation_status','issues'}, + 'context/case_context.json':{'source_documents','members','relationships','candidate_dependencies','unresolved_relationships','clusters','scheduling_waves','client_goal','routing','signals','global_review_refs','object_and_party_refs','slot_links','slot_link_status','active_profiles'}, + 'ingress/technical_diagnostic.json':{'status','issues','checks'}, + STATUS_PATH:{'status','output_root','source_hashes','artifacts','written_last','publication_semantics'}, + } + for path,raw in files.items(): + value=load_json_strict(raw) + if not isinstance(value,dict) or set(value)!=common|fields[path]:raise IngressError('OUTPUT_CLOSED_SCHEMA_INVALID','output fields do not match the inline contract') + if value['algorithm_version']!=ALGORITHM_VERSION or value['schema_version']!='stage2_s2_00_direct.v3':raise IngressError('OUTPUT_VERSION_INVALID','output algorithm/schema version differs') + if any(value[key]!=roots[key] for key in ('stage1_run_root_ref','stage1_deployment_root_ref')):raise IngressError('OUTPUT_SOURCE_BINDING_INVALID','output roots differ from inputs') + if status['output_root']!=roots['output_root'] or status['written_last'] is not True or status['publication_semantics']!='STATUS_LAST_LOGICAL_COMMIT':raise IngressError('OUTPUT_STATUS_INVALID','status does not identify the logical completion boundary') + rows=status['artifacts'] + if not isinstance(rows,list) or len(rows)!=len(files)-1 or {r.get('path') for r in rows}!=set(files)-{STATUS_PATH}:raise IngressError('OUTPUT_STATUS_SET_INVALID','status inventory differs from actual outputs') + for row in rows: + raw=files[row['path']] + if set(row)!={'path','raw_sha256','byte_length'} or row['raw_sha256']!=hashlib.sha256(raw).hexdigest() or row['byte_length']!=len(raw):raise IngressError('OUTPUT_STATUS_HASH_INVALID','status inventory does not match output bytes') + return status + + + def publish_result(localdocs: _InlineLocaldocs, roots: Mapping[str,str], files: Mapping[str,bytes]) -> dict[str,Any]: + """No overwrite, exact completed-result reuse, and status-last publication.""" + status=validate_output_files(files,roots); output=roots['output_root'] + existing=localdocs.read_binary_optional(f'{output}/{STATUS_PATH}') + if existing is not None: + if existing!=files[STATUS_PATH]:raise IngressError('EXISTING_OUTPUT_CONFLICT','existing completed output differs in source, version, status, or inventory') + for relative,raw in sorted(files.items()): + if localdocs.read_binary(f'{output}/{relative}')!=raw:raise IngressError('EXISTING_OUTPUT_CORRUPT','existing artifact differs from completed status') + publication='REUSED_COMPLETED_OUTPUT' + else: + for relative in sorted(NORMAL_PATHS|BLOCKED_PATHS): + if relative!=STATUS_PATH and localdocs.read_binary_optional(f'{output}/{relative}') is not None:raise IngressError('PARTIAL_OUTPUT_CONFLICT','unfinished output requires explicit recovery; no overwrite') + for relative in sorted(set(files)-{STATUS_PATH}):localdocs.write_binary_verified(f'{output}/{relative}',files[relative],overwrite=False) + localdocs.write_binary_verified(f'{output}/{STATUS_PATH}',files[STATUS_PATH],overwrite=False) + publication='PUBLISHED_STATUS_LAST' + return {'ok':status['status']!='BLOCKED','status':status['status'],'output_root':output,'publication':publication,'ingress_status_sha256':hashlib.sha256(files[STATUS_PATH]).hexdigest()} + + + def run_inline_mcp(run_root: Any = RAW_RUN_ROOT, deployment_root: Any = RAW_DEPLOYMENT_ROOT, *, stage1_results: Mapping[str,Any] | None = None, client: Any | None = None) -> int: + localdocs=None + try: + roots=validate_direct_roots(run_root,deployment_root) + supplied=RAW_STAGE1_RESULTS if stage1_results is None else stage1_results + # Validate all required references before making a remote call. + for row in DEFAULT_SOURCE_CONTRACTS: + logical=row['logical_input_id'] + if logical not in supplied:raise IngressError('PREV_SOURCE_MISSING','required backend result is absent',logical_input_id=logical) + _previous_source(supplied[logical],f"{roots['stage1_run_root_ref']}/{row['path']}") + if SOURCE_POLICY['release_class']=='DEV_FIXTURE_RELEASE':raise IngressError('DEV_FIXTURE_REAL_RUN_FORBIDDEN','DEV fixture admission cannot publish a real case') + localdocs=_InlineLocaldocs(INLINE_USER_HASH,INLINE_WORKSPACE_HASH,client=client) + localdocs.initialize() + with tempfile.TemporaryDirectory(prefix='liti-s2-00-') as directory: + hydrated=hydrate_stage1(localdocs,Path(directory),roots,supplied,SOURCE_POLICY) + result=execute_ingress(hydrated,roots,policy=SOURCE_POLICY) + verify_remote_stability(localdocs,hydrated['observed']) + receipt=publish_result(localdocs,roots,result['files']) + print(json.dumps(receipt,ensure_ascii=False,separators=(',',':'))) + return 0 if receipt['ok'] else 2 + except Exception as exc: + error=exc.as_dict() if isinstance(exc,IngressError) else {'code':'S2_00_RUNTIME_ERROR','message':str(exc)} + # Failure does not assert that a status already written remotely is absent. + print(json.dumps({'ok':False,'status':'FAILED','error':error},ensure_ascii=False,separators=(',',':'))) + return 2 + finally: + if localdocs is not None:localdocs.close() + + + if __name__ == '__main__': + raise SystemExit(run_inline_mcp()) + task_procedure: + IN: + nexts: + - Task_S2_00_deterministic_ingress + wait_until: [] + Task_S2_00_deterministic_ingress: + nexts: + - OUT + wait_until: + - IN + OUT: + nexts: [] + wait_until: + - Task_S2_00_deterministic_ingress diff --git a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00_10_02.yml b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00_10_02.yml new file mode 100644 index 00000000..df39cc88 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00_10_02.yml @@ -0,0 +1,6502 @@ +Agent: + name: Stage_2_S2_00_v2 + version: "1.2.0" + description: >- + Stage 1 Part 1-4의 고정 입력을 검증·보존·정규화하고 claim-neutral + cluster slice와 S2_10 structured-context handoff plan을 생성하는 + S2_00 deterministic ingress Agent. + metadata: + workflow_id: S2_00 + execution_class: NON-LLM-DETERMINISTIC + execution_authority: MCP_CODE_EXECUTOR_INLINE + release_ref: Default_Agent/Stage_2_Clean/manifest/stage2_release.json + workflow_contract_ref: Default_Agent/Stage_2_Clean/workflows/S2_00_stage1_ingress_normalize_and_bundle_compile.yml + deployment_binding_ref: Default_Agent/Stage_2_Clean/deployment/stage2_code_executor_binding.yml + implementation_status: S2_00_PREPARE_OFFLINE_VERIFIED_BACKEND_ARGUMENT_GATE_AND_SERIALIZATION_UNVERIFIED_LIVE_ADMISSION_PENDING + Stages: + - name: S2_00 + description: >- + 첫 Code Executor task에서 고정 request를 준비하고, 다음 ingress + task의 한 호출 안에서 C00, C05, C10, C15를 순차 실행한다. + localdocs binary IO와 status-last 논리 배리어로 결과를 발행한다. + prevs: [] + nexts: [] + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + tasks: + - task_name: Task_S2_00_prepare_request + description: >- + 네 외부 문자열 인자 request_id, attempt_id, stage1_run_root_ref, + stage1_deployment_root_ref를 검증하고 고정 localdocs request 경로에 + 저장·read-back한다. Backend 인자 결속 및 성공 게이트는 미검증이며 + 인자가 결속되지 않은 직접 실행은 실패로 종료한다. + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx==0.28.1" + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + """Prepare the fixed S2_00 control request from four explicit caller values. + + The backend argument-delivery contract is not bound. Direct execution fails + closed; callers must invoke ``prepare_request`` with the four named values. + """ + + from __future__ import annotations + + import base64 + import hashlib + import json + from pathlib import PurePosixPath + import re + import sys + import unicodedata + from typing import Any, Mapping + + + REQUEST_PATH = "stage2_control/s2_00_request.json" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_PROTOCOL_VERSION = "2025-03-26" + INLINE_USER_HASH = "{{__user_hash__}}" + INLINE_WORKSPACE_HASH = "{{__workspace_hash__}}" + REQUEST_KEYS = frozenset({ + "schema_version", "workflow_id", "request_id", "attempt_id", + "stage1_run_root_ref", "stage1_deployment_root_ref", + }) + ID_RE = re.compile(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}\Z") + + + class PrepareError(ValueError): + def __init__(self, code: str, message: str) -> None: + super().__init__(message) + self.code = code + + + def canonical_json_bytes(value: Any) -> bytes: + return (json.dumps(value, ensure_ascii=False, allow_nan=False, + sort_keys=True, separators=(",", ":")) + "\n").encode("utf-8") + + + def _relative_path(value: str, *, code: str) -> str: + if not isinstance(value, str) or not value or "\x00" in value or "\\" in value: + raise PrepareError(code, "logical path is empty or malformed") + if unicodedata.normalize("NFC", value) != value: + raise PrepareError(code, "logical path must already be NFC") + path = PurePosixPath(value) + if path.is_absolute() or any(part in {"", ".", ".."} for part in path.parts): + raise PrepareError(code, "logical path must be a contained relative path") + rendered = path.as_posix() + if rendered != value: + raise PrepareError(code, "logical path is not canonical") + return rendered + + + def build_request( + request_id: str, + attempt_id: str, + stage1_run_root_ref: str, + stage1_deployment_root_ref: str, + ) -> bytes: + """Return the exact six-field canonical request; do not access localdocs.""" + if not isinstance(request_id, str) or ID_RE.fullmatch(request_id) is None: + raise PrepareError("REQUEST_ID_INVALID", "request_id contains forbidden characters") + if not isinstance(attempt_id, str) or ID_RE.fullmatch(attempt_id) is None: + raise PrepareError("ATTEMPT_ID_INVALID", "attempt_id contains forbidden characters") + request = { + "schema_version": "stage2_s2_00_execution_request.v1", + "workflow_id": "S2_00", + "request_id": request_id, + "attempt_id": attempt_id, + "stage1_run_root_ref": _relative_path( + stage1_run_root_ref, code="STAGE1_RUN_ROOT_REF_INVALID"), + "stage1_deployment_root_ref": _relative_path( + stage1_deployment_root_ref, code="STAGE1_DEPLOYMENT_ROOT_REF_INVALID"), + } + if set(request) != REQUEST_KEYS: + raise PrepareError("RUN_REQUEST_CLOSED_SHAPE", "request field set drifted") + return canonical_json_bytes(request) + + + class LocaldocsSession: + """Small localdocs JSON-RPC session with verified binary write/read-back.""" + + def __init__(self, user_hash: str, workspace_hash: str, *, client: Any = None) -> None: + for name, value in (("user_hash", user_hash), ("workspace_hash", workspace_hash)): + if not isinstance(value, str) or re.fullmatch(r"[a-f0-9]{64}", value) is None: + raise PrepareError("CONTEXT_HASH_INVALID", f"{name} is not a SHA-256 digest") + self.user_hash = user_hash + self.workspace_hash = workspace_hash + if client is None: + try: + import httpx + except ImportError as exc: + raise PrepareError("HTTPX_UNAVAILABLE", "httpx==0.28.1 is required") from exc + client = httpx.Client(timeout=60) + self.client = client + self.headers = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"} + self.session_id: str | None = None + self.next_id = 10 + self.initialized = False + + def close(self) -> None: + self.client.close() + + def _post(self, body: Mapping[str, Any], expected_id: int | None) -> Mapping[str, Any] | None: + try: + response = self.client.post(LOCALDOCS_URL, json=dict(body), headers=dict(self.headers)) + response.raise_for_status() + except Exception as exc: + raise PrepareError("MCP_TRANSPORT_ERROR", "localdocs transport failed") from exc + session_id = response.headers.get("mcp-session-id") + if session_id: + if self.session_id is None and expected_id == 1: + self.session_id = session_id + elif session_id != self.session_id: + raise PrepareError("MCP_SESSION_ID_CHANGED", "localdocs session changed") + self.headers["mcp-session-id"] = session_id + if expected_id is None: + return None + try: + payload = response.json() + except Exception as exc: + raise PrepareError("MCP_RESPONSE_INVALID", "localdocs response is not JSON") from exc + if not isinstance(payload, dict) or payload.get("jsonrpc") != "2.0" or payload.get("id") != expected_id: + raise PrepareError("MCP_RESPONSE_INVALID", "localdocs response ID or shape mismatch") + if "error" in payload or not isinstance(payload.get("result"), dict): + raise PrepareError("MCP_TOOL_ERROR", "localdocs returned an error") + return payload["result"] + + def initialize(self) -> None: + response = self._post({"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": { + "protocolVersion": MCP_PROTOCOL_VERSION, "capabilities": {}, "clientInfo": { + "name": "liti-s2-00-prepare", "version": "1.0.0", + "user_id": self.user_hash, "workspace_id": self.workspace_hash, + }}}, 1) + if response is None or response.get("protocolVersion") != MCP_PROTOCOL_VERSION or self.session_id is None: + raise PrepareError("MCP_INITIALIZE_INVALID", "localdocs initialization failed") + self._post({"jsonrpc": "2.0", "method": "notifications/initialized"}, None) + self.initialized = True + + def call(self, name: str, arguments: Mapping[str, Any]) -> Mapping[str, Any]: + if not self.initialized: + raise PrepareError("MCP_NOT_INITIALIZED", "localdocs is not initialized") + message_id = self.next_id + self.next_id += 1 + result = self._post({"jsonrpc": "2.0", "id": message_id, "method": "tools/call", + "params": {"name": name, "arguments": dict(arguments)}}, message_id) + if result is None or result.get("isError") is True: + raise PrepareError("MCP_TOOL_ERROR", f"localdocs {name} failed") + return result + + def read_binary(self, logical_path: str) -> bytes: + path = _relative_path(logical_path, code="LOCALDOCS_READ_PATH_INVALID") + result = self.call("read_binary_doc", {"doc_name": path}) + content = result.get("content") + if not isinstance(content, list) or len(content) != 1 or not isinstance(content[0], dict) or content[0].get("type") != "text": + raise PrepareError("LOCALDOCS_READ_SHAPE", "invalid binary response") + text = content[0].get("text") + if not isinstance(text, str): + raise PrepareError("LOCALDOCS_READ_SHAPE", "missing binary envelope") + try: + envelope = json.loads(text) + if isinstance(envelope, dict) and "results" in envelope: + rows = envelope["results"] + if not isinstance(rows, list) or len(rows) != 1 or not isinstance(rows[0], dict): + raise ValueError("binary result cardinality mismatch") + inner = rows[0].get("content", rows[0].get("text")) + envelope = json.loads(inner) if isinstance(inner, str) else inner + encoded = envelope["content_base64"] + if not isinstance(encoded, str): + raise ValueError("binary content is not base64") + payload = base64.b64decode(encoded, validate=True) + size = envelope.get("byte_length", envelope.get("size")) + if size is not None and (not isinstance(size, int) or size != len(payload)): + raise ValueError("binary size mismatch") + digest = envelope.get("sha256") + if digest is not None and digest != hashlib.sha256(payload).hexdigest(): + raise ValueError("binary hash mismatch") + return payload + except (ValueError, KeyError, TypeError, base64.binascii.Error) as exc: + raise PrepareError("LOCALDOCS_READ_SHAPE", "invalid binary envelope") from exc + + def write_binary_verified(self, logical_path: str, payload: bytes) -> str: + path = _relative_path(logical_path, code="LOCALDOCS_WRITE_PATH_INVALID") + if path != REQUEST_PATH: + raise PrepareError("LOCALDOCS_WRITE_PATH_INVALID", "prepare may write only the fixed request path") + result = self.call("write_binary_file", { + "path": path, "content_base64": base64.b64encode(payload).decode("ascii"), "overwrite": True, + }) + if not isinstance(result.get("content"), list) or result.get("isError") is True: + raise PrepareError("LOCALDOCS_WRITE_FAILED", "localdocs did not acknowledge write") + observed = self.read_binary(path) + if observed != payload: + raise PrepareError("LOCALDOCS_WRITE_READBACK_MISMATCH", "request read-back differs") + return hashlib.sha256(observed).hexdigest() + + + def prepare_request( + request_id: str, + attempt_id: str, + stage1_run_root_ref: str, + stage1_deployment_root_ref: str, + *, + localdocs: Any = None, + ) -> dict[str, Any]: + """Validate, write, read back, and close an authenticated localdocs session.""" + payload = build_request(request_id, attempt_id, stage1_run_root_ref, stage1_deployment_root_ref) + session = localdocs if localdocs is not None else LocaldocsSession(INLINE_USER_HASH, INLINE_WORKSPACE_HASH) + failure: Exception | None = None + digest: str | None = None + try: + session.initialize() + digest = session.write_binary_verified(REQUEST_PATH, payload) + except Exception as exc: + failure = exc + try: + session.close() + except Exception as exc: + if failure is None: + failure = PrepareError("LOCALDOCS_CLOSE_FAILED", "localdocs session close failed") + failure.__cause__ = exc + if failure is not None: + raise failure + return {"ok": True, "workflow_id": "S2_00", "path": REQUEST_PATH, + "request_sha256": digest} + + + if __name__ == "__main__": + sys.stdout.buffer.write(canonical_json_bytes({ + "ok": False, "error": {"code": "PREPARE_ARGUMENT_BINDING_UNVERIFIED", + "message": "four caller values must be bound by the backend"}})) + raise SystemExit(2) + - task_name: Task_S2_00_deterministic_ingress + description: >- + 고정 request와 release를 hydration하고 S2_00 pure core를 실행한 뒤 + binding-derived root에 결과를 status-last 방식으로 발행한다. + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: "httpx==0.28.1" + network: agent-network + timeout: 300 + code: |- + #!/usr/bin/env python3 + """Deterministic Stage 2 ingress for the S2_00 contract. + + The pure ingress core uses the Python standard library; the inline MCP adapter + adds only pinned ``httpx`` transport to the fixed localdocs endpoint. It does + not perform legal reasoning, model calls, external-network access, dynamic + imports, or runtime installation. The release, source, conservation, context, + cluster, bundle, hydration, and logical-publish contracts are re-checked here. + """ + + from __future__ import annotations + + import argparse + import base64 + import binascii + from collections import Counter, defaultdict + import contextlib + from dataclasses import dataclass + import hashlib + import io + import itertools + import json + import math + import os + from pathlib import Path, PurePosixPath + import re + import shutil + import stat + import sys + import tempfile + import unicodedata + from typing import Any, Callable, Iterable, Mapping, MutableMapping, Sequence + + + ALGORITHM_VERSION = "s2_00_ingress/1.2.0" + ALGORITHM_SEMANTIC_DIGEST = hashlib.sha256( + b"LITI-S2_00-INLINE-ALGORITHM\x00" + ALGORITHM_VERSION.encode("ascii") + ).hexdigest() + MAX_FILE_BYTES = 32 * 1024 * 1024 + MAX_RUN_BYTES = 256 * 1024 * 1024 + MAX_JSON_DEPTH = 96 + MAX_JSON_ITEMS = 1_000_000 + ROUTES = frozenset({"TO_S2_10", "TO_S2_10_WITH_ISSUES", "TO_S2_40_STATUS_ONLY"}) + RELEASE_MODE = { + "STRUCTURAL_FIXTURE": "DEV_FIXTURE_RELEASE", + "SUBSET_CANARY": "SUBSET_CANARY_RELEASE", + "PRODUCTION": "PRODUCTION_RELEASE", + } + SEMANTIC_SIGNAL_KINDS = frozenset({"canonical", "domain_signal"}) + HARD_RELATION_KINDS = frozenset( + { + "SAME_BO_ID", + "SOURCE_BO_ATTACHMENT", + "SAME_EVIDENCE_REF", + "SAME_EVENT_REF", + "EXPLICIT_CASE_RELATION", + } + ) + CANDIDATE_RELATION_KINDS = frozenset( + {"claim_precondition", "accessory_of", "incompatible_with", "EXPLICIT_DEPENDENCY"} + ) + P1_DIGEST_KEYS = { + "evidence_indexed_sha256": "evidence_indexed", + "evidence_event_candidates_sha256": "evidence_event_candidates", + "b1_gate_sha256": "b1_evidence_indexed_gate", + "b2_gate_sha256": "b2_event_candidates_gate", + "screening_sha256": "domain_screening", + "activation_manifest_sha256": "domain_activation_manifest", + "registry_index_sha256": "stage1_domain_registry_index", + } + V2_DOMAIN_IDS = frozenset({"E-01", "E-06", "E-12", "E-16", "E-18", "E-19", "E-20", "E-21"}) + CONTEXT_SCHEMA_ID = "https://schemas.liti-agent.local/stage2/s2_00/context.schema.v2.json" + INGRESS_SCHEMA_ID = "https://schemas.liti-agent.local/stage2/s2_00/ingress.schema.v1.json" + REVIEW_SCHEMA_ID = "https://schemas.liti-agent.local/stage2/shared/review_status.schema.v1.json" + REQUIREMENT_CLASS_ENUM = { + "identity_backbone": "IDENTITY_BACKBONE", + "routing_profile_backbone": "ROUTING_PROFILE_BACKBONE", + "evidence_scope": "EVIDENCE_EVENT_SCOPE", + "event_scope": "EVIDENCE_EVENT_SCOPE", + "integrity_corroborator": "INTEGRITY_CORROBORATOR", + "optimization_context": "OPTIMIZATION_CONTEXT", + } + ADAPTER_IDS = { + "evidence_indexed": "S2A-EVIDENCE-V3-ENVELOPE-V1", + "evidence_event_candidates": "S2A-EVENTS-V1-ENVELOPE-V1", + "client_goal": "S2A-CLIENT-GOAL-V8-V1", + "domain_screening": "S2A-DOMAIN-SCREENING-V1", + "domain_activation_manifest": "S2A-DUAL-SG01-V1", + "b1_evidence_indexed_gate": "S2A-B1-GATE-V1", + "b2_event_candidates_gate": "S2A-B2-GATE-V1", + "stage1_part1_soft_gate_handoff": "S2A-P1-HANDOFF-FLAT-V1", + "bo": "S2A-BO-V8-LIST-V1", + "signal_manifest": "S2A-SIGNAL-ALL-V1", + "stage1_part2_review_handoff": "S2A-P2-HANDOFF-FLAT-V1", + "legal_effect_structures": "S2A-LES-CURRENT-V8-V1", + "stage1_part3_review_handoff": "S2A-P3-HANDOFF-WRAPPED-V1", + "fact_ledger_base": "S2A-FACT-LEDGER-CURRENT-V8-V1", + "fact_ledger_writer_report": "S2A-FACT-LEDGER-WRITER-REPORT-V1", + "stage1_part4_review_handoff": "S2A-P4-HANDOFF-WRAPPED-V1", + } + SG01_PROJECTION_FIELDS: tuple[str, ...] = ( + "schema_version", + "signal_id", + "status", + "registry_version", + "registry_index_sha256", + "screening_sha256", + "domain_entries", + "active_domain_ids", + "supporting_domain_ids", + "monitor_domain_ids", + "expected_runnable_domain_ids", + "required_calculation_domains", + "unrouted_material", + "conservation_gate", + "fail_open_policy", + "review_items", + "contract_guards", + ) + SG01_SET_FIELDS = frozenset( + { + "active_domain_ids", + "supporting_domain_ids", + "monitor_domain_ids", + "expected_runnable_domain_ids", + "required_calculation_domains", + } + ) + OUTPUT_SCHEMA_TARGETS: tuple[tuple[re.Pattern[str], str, str], ...] = ( + (re.compile(r"^ingress/stage1_input_manifest\.json$"), "ingress.schema.json", "#/$defs/stage1_input_manifest"), + (re.compile(r"^ingress/intake_report\.json$"), "ingress.schema.json", "#/$defs/intake_report"), + (re.compile(r"^ingress/ingress_status\.json$"), "ingress.schema.json", "#/$defs/ingress_status"), + (re.compile(r"^ingress/technical_diagnostic\.json$"), "ingress.schema.json", "#/$defs/technical_diagnostic"), + (re.compile(r"^review/issue_ledger\.base\.json$"), "review_status.schema.json", "#/$defs/issue_ledger_base"), + (re.compile(r"^context/case_context\.json$"), "context.schema.json", "#/$defs/case_context"), + (re.compile(r"^context/evidence_inventory\.json$"), "context.schema.json", "#/$defs/evidence_inventory"), + (re.compile(r"^context/object_registry\.json$"), "context.schema.json", "#/$defs/object_registry"), + (re.compile(r"^context/party_and_title_context\.json$"), "context.schema.json", "#/$defs/party_and_title_context"), + (re.compile(r"^context/slot_crosswalk\.json$"), "context.schema.json", "#/$defs/slot_crosswalk"), + (re.compile(r"^context/cluster_plan\.json$"), "context.schema.json", "#/$defs/cluster_plan"), + (re.compile(r"^context/cluster_slices/[^/]+\.json$"), "context.schema.json", "#/$defs/cluster_slice"), + (re.compile(r"^context/bundle_plan\.json$"), "context.schema.json", "#/$defs/bundle_plan"), + ) + _RAW_VALUE_UNSET = object() + + + # AgentBackend substitutes these two values in the deployed Agent YAML before + # Code Executor runs the byte-identical source. They intentionally remain + # literal placeholders in the offline parity mirror and its unit tests. + INLINE_USER_HASH = "{{__user_hash__}}" + INLINE_WORKSPACE_HASH = "{{__workspace_hash__}}" + EXPECTED_STAGE2_RELEASE_SHA256 = "2363f1166ee8a3ea2b38350cde61fccfc9852a413c6b9f14a6ed6606af15a590" + INLINE_REQUEST_PATH = "stage2_control/s2_00_request.json" + INLINE_STAGE2_ASSET_ROOT = "Default_Agent/Stage_2_Clean" + INLINE_STAGE2_RELEASE_PATH = ( + f"{INLINE_STAGE2_ASSET_ROOT}/manifest/stage2_release.json" + ) + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_PROTOCOL_VERSION = "2025-03-26" + INLINE_CLIENT_NAME = "liti-stage2-s2-00-inline" + INLINE_CLIENT_VERSION = "1.2.0" + INLINE_SCHEMA_MODULE_IDS = frozenset( + {"SCHEMA-INGRESS", "SCHEMA-CONTEXT", "SCHEMA-REVIEW-STATUS"} + ) + + + DEFAULT_SOURCE_CONTRACTS: tuple[dict[str, Any], ...] = ( + {"logical_input_id": "evidence_indexed", "path": "evidence_indexed.json", "criticality": "evidence_scope"}, + {"logical_input_id": "evidence_event_candidates", "path": "evidence_event_candidates.json", "criticality": "event_scope"}, + {"logical_input_id": "client_goal", "path": "client_goal.json", "criticality": "optimization_context"}, + {"logical_input_id": "domain_screening", "path": "routing/domain_screening.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "domain_activation_manifest", "path": "routing/domain_activation_manifest.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "b1_evidence_indexed_gate", "path": "quality_gates/B1_evidence_indexed_gate.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "b2_event_candidates_gate", "path": "quality_gates/B2_event_candidates_gate.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "stage1_part1_soft_gate_handoff", "path": "quality_gates/stage1_part1_soft_gate_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "bo", "path": "BO.json", "criticality": "identity_backbone"}, + {"logical_input_id": "signal_manifest", "path": "signals/signal_manifest.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "stage1_part2_review_handoff", "path": "quality_gates/stage1_part2_review_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "legal_effect_structures", "path": "legal_effect_structures.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "stage1_part3_review_handoff", "path": "quality_gates/stage1_part3_review_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "fact_ledger_base", "path": "Fact_Ledger_base.json", "criticality": "identity_backbone"}, + {"logical_input_id": "fact_ledger_writer_report", "path": "stage1_tmp/fact_ledger/fact_ledger_writer_report.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "stage1_part4_review_handoff", "path": "quality_gates/stage1_part4_review_handoff.json", "criticality": "integrity_corroborator"}, + ) + + + class IngressError(RuntimeError): + """A machine-readable deterministic ingress failure.""" + + def __init__( + self, + code: str, + message: str, + *, + logical_input_id: str | None = None, + details: Mapping[str, Any] | None = None, + ) -> None: + super().__init__(message) + self.code = code + self.logical_input_id = logical_input_id + self.details = dict(details or {}) + + def as_dict(self) -> dict[str, Any]: + result: dict[str, Any] = {"code": self.code, "message": str(self)} + if self.logical_input_id is not None: + result["logical_input_id"] = self.logical_input_id + if self.details: + result["details"] = self.details + return result + + + @dataclass(frozen=True, slots=True) + class Snapshot: + logical_input_id: str + relative_path: str + resolved_path: str + raw: bytes + raw_sha256: str + byte_length: int + device: int + inode: int + mtime_ns: int + + + def _reject_constant(value: str) -> None: + raise ValueError(f"non-finite JSON number is forbidden: {value}") + + + def _pairs_without_duplicates(pairs: Sequence[tuple[str, Any]]) -> dict[str, Any]: + result: dict[str, Any] = {} + for key, value in pairs: + if key in result: + raise ValueError(f"duplicate JSON key: {key}") + result[key] = value + return result + + + def _walk_json_limits(value: Any, *, max_depth: int, max_items: int) -> int: + count = 0 + stack: list[tuple[Any, int]] = [(value, 1)] + while stack: + current, depth = stack.pop() + if depth > max_depth: + raise IngressError("JSON_DEPTH_LIMIT", "JSON nesting depth exceeded") + if isinstance(current, dict): + count += len(current) + stack.extend((item, depth + 1) for item in current.values()) + elif isinstance(current, list): + count += len(current) + stack.extend((item, depth + 1) for item in current) + if count > max_items: + raise IngressError("JSON_ITEM_LIMIT", "JSON aggregate item limit exceeded") + return count + + + def load_json_strict( + source: Snapshot | bytes | bytearray | memoryview | str, + *, + max_depth: int = MAX_JSON_DEPTH, + max_items: int = MAX_JSON_ITEMS, + ) -> Any: + """Parse one UTF-8 JSON value, rejecting duplicate keys and non-finite numbers.""" + + if isinstance(source, Snapshot): + raw = source.raw + elif isinstance(source, str): + raw = source.encode("utf-8") + else: + raw = bytes(source) + try: + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise IngressError("INVALID_UTF8", "JSON source is not strict UTF-8") from exc + try: + value = json.loads( + text, + object_pairs_hook=_pairs_without_duplicates, + parse_constant=_reject_constant, + ) + except (json.JSONDecodeError, ValueError) as exc: + message = str(exc) + code = "DUPLICATE_JSON_KEY" if "duplicate JSON key" in message else "STRICT_JSON_PARSE_FAILED" + raise IngressError(code, message) from exc + _walk_json_limits(value, max_depth=max_depth, max_items=max_items) + return value + + + def canonical_json_bytes(value: Any) -> bytes: + """Return the project canonical parsed representation without normalizing strings.""" + + def reject_nonfinite(item: Any) -> None: + if isinstance(item, float) and not math.isfinite(item): + raise IngressError("NON_FINITE_NUMBER", "NaN and Infinity are forbidden") + if isinstance(item, dict): + for nested in item.values(): + reject_nonfinite(nested) + elif isinstance(item, (list, tuple)): + for nested in item: + reject_nonfinite(nested) + + reject_nonfinite(value) + try: + rendered = json.dumps( + value, + ensure_ascii=False, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ) + except (TypeError, ValueError) as exc: + raise IngressError("CANONICAL_SERIALIZATION_FAILED", str(exc)) from exc + return (rendered + "\n").encode("utf-8") + + + def canonical_digest(value: Any) -> str: + return hashlib.sha256(canonical_json_bytes(value)).hexdigest() + + + class _SchemaViolation(ValueError): + """Internal deterministic JSON Schema validation failure.""" + + + def _json_equal(left: Any, right: Any) -> bool: + try: + return canonical_json_bytes(left) == canonical_json_bytes(right) + except IngressError: + return False + + + def _schema_pointer(document: Mapping[str, Any], fragment: str) -> Mapping[str, Any]: + if fragment in {"", "#"}: + return document + pointer = fragment[1:] if fragment.startswith("#") else fragment + if not pointer.startswith("/"): + raise _SchemaViolation(f"unsupported schema fragment: {fragment}") + current: Any = document + for token in pointer[1:].split("/"): + key = token.replace("~1", "/").replace("~0", "~") + if not isinstance(current, dict) or key not in current: + raise _SchemaViolation(f"unresolved schema pointer: {fragment}") + current = current[key] + if not isinstance(current, dict): + raise _SchemaViolation(f"schema pointer is not an object: {fragment}") + return current + + + def _schema_type_matches(value: Any, expected: str) -> bool: + return { + "object": isinstance(value, dict), + "array": isinstance(value, list), + "string": isinstance(value, str), + "integer": isinstance(value, int) and not isinstance(value, bool), + "number": isinstance(value, (int, float)) and not isinstance(value, bool), + "boolean": isinstance(value, bool), + "null": value is None, + }.get(expected, False) + + + def _validate_schema_node( + value: Any, + schema: Mapping[str, Any], + *, + root_schema: Mapping[str, Any], + schema_documents: Mapping[str, Mapping[str, Any]], + instance_path: str, + ) -> None: + reference = schema.get("$ref") + if isinstance(reference, str): + if reference.startswith("#"): + target_root = root_schema + fragment = reference + else: + name, separator, tail = reference.partition("#") + target_root = schema_documents.get(name) + if target_root is None: + raise _SchemaViolation(f"{instance_path}: external schema ref is not release-local: {reference}") + fragment = f"#{tail}" if separator else "#" + _validate_schema_node( + value, + _schema_pointer(target_root, fragment), + root_schema=target_root, + schema_documents=schema_documents, + instance_path=instance_path, + ) + return + if "const" in schema and not _json_equal(value, schema["const"]): + raise _SchemaViolation(f"{instance_path}: const mismatch") + if "enum" in schema and not any(_json_equal(value, candidate) for candidate in schema["enum"]): + raise _SchemaViolation(f"{instance_path}: enum mismatch") + forbidden = schema.get("not") + if isinstance(forbidden, dict) and _schema_branch_matches( + value, + forbidden, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ): + raise _SchemaViolation(f"{instance_path}: forbidden schema branch matched") + expected_type = schema.get("type") + if expected_type is not None: + alternatives = [expected_type] if isinstance(expected_type, str) else list(expected_type) + if not any(_schema_type_matches(value, item) for item in alternatives): + raise _SchemaViolation(f"{instance_path}: expected type {alternatives}") + for keyword in ("oneOf", "anyOf"): + branches = schema.get(keyword) + if isinstance(branches, list): + matches = 0 + for branch in branches: + try: + _validate_schema_node( + value, + branch, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + matches += 1 + except _SchemaViolation: + continue + required_matches = 1 if keyword == "oneOf" else None + if (required_matches is not None and matches != required_matches) or (keyword == "anyOf" and matches == 0): + raise _SchemaViolation(f"{instance_path}: {keyword} matched {matches} branches") + all_of = schema.get("allOf") + if isinstance(all_of, list): + for branch in all_of: + _validate_schema_node( + value, + branch, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + condition = schema.get("if") + if isinstance(condition, dict): + condition_matches = True + try: + _validate_schema_node( + value, + condition, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + except _SchemaViolation: + condition_matches = False + selected = schema.get("then" if condition_matches else "else") + if isinstance(selected, dict): + _validate_schema_node( + value, + selected, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + if isinstance(value, dict): + minimum_properties = schema.get("minProperties") + maximum_properties = schema.get("maxProperties") + if isinstance(minimum_properties, int) and len(value) < minimum_properties: + raise _SchemaViolation(f"{instance_path}: minProperties {minimum_properties}") + if isinstance(maximum_properties, int) and len(value) > maximum_properties: + raise _SchemaViolation(f"{instance_path}: maxProperties {maximum_properties}") + required = schema.get("required", []) + if isinstance(required, list): + missing = [key for key in required if key not in value] + if missing: + raise _SchemaViolation(f"{instance_path}: missing required keys {missing}") + properties = schema.get("properties", {}) + if isinstance(properties, dict): + pattern_properties = schema.get("patternProperties", {}) + matched_by_pattern: set[str] = set() + if isinstance(pattern_properties, dict): + for key, child_value in value.items(): + for pattern_text, child_schema in pattern_properties.items(): + if re.search(pattern_text, key) is not None and isinstance(child_schema, dict): + matched_by_pattern.add(key) + _validate_schema_node( + child_value, + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + extras = sorted(set(value) - set(properties) - matched_by_pattern) + additional = schema.get("additionalProperties") + if additional is False: + if extras: + raise _SchemaViolation(f"{instance_path}: additional properties {extras}") + elif isinstance(additional, dict): + for key in extras: + _validate_schema_node( + value[key], + additional, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + for key, child_schema in properties.items(): + if key in value and isinstance(child_schema, dict): + _validate_schema_node( + value[key], + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + if isinstance(value, list): + minimum = schema.get("minItems") + maximum = schema.get("maxItems") + if isinstance(minimum, int) and len(value) < minimum: + raise _SchemaViolation(f"{instance_path}: minItems {minimum}") + if isinstance(maximum, int) and len(value) > maximum: + raise _SchemaViolation(f"{instance_path}: maxItems {maximum}") + if schema.get("uniqueItems") is True: + digests = [canonical_digest(item) for item in value] + if len(digests) != len(set(digests)): + raise _SchemaViolation(f"{instance_path}: duplicate array items") + prefix_items = schema.get("prefixItems") + if isinstance(prefix_items, list): + for index, child_schema in enumerate(prefix_items): + if index < len(value) and isinstance(child_schema, dict): + _validate_schema_node( + value[index], + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + item_schema = schema.get("items") + if item_schema is False and isinstance(prefix_items, list) and len(value) > len(prefix_items): + raise _SchemaViolation(f"{instance_path}: additional array items are forbidden") + if isinstance(item_schema, dict): + start = len(prefix_items) if isinstance(prefix_items, list) else 0 + for index, item in enumerate(value[start:], start=start): + _validate_schema_node( + item, + item_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + contains = schema.get("contains") + if isinstance(contains, dict): + if not any( + _schema_branch_matches( + item, + contains, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + for index, item in enumerate(value) + ): + raise _SchemaViolation(f"{instance_path}: contains did not match") + if isinstance(value, str): + min_length = schema.get("minLength") + if isinstance(min_length, int) and len(value) < min_length: + raise _SchemaViolation(f"{instance_path}: minLength {min_length}") + max_length = schema.get("maxLength") + if isinstance(max_length, int) and len(value) > max_length: + raise _SchemaViolation(f"{instance_path}: maxLength {max_length}") + pattern = schema.get("pattern") + if isinstance(pattern, str) and re.search(pattern, value) is None: + raise _SchemaViolation(f"{instance_path}: pattern mismatch") + if isinstance(value, (int, float)) and not isinstance(value, bool): + minimum = schema.get("minimum") + if isinstance(minimum, (int, float)) and value < minimum: + raise _SchemaViolation(f"{instance_path}: minimum {minimum}") + maximum = schema.get("maximum") + if isinstance(maximum, (int, float)) and value > maximum: + raise _SchemaViolation(f"{instance_path}: maximum {maximum}") + + + def _schema_branch_matches( + value: Any, + schema: Mapping[str, Any], + *, + root_schema: Mapping[str, Any], + schema_documents: Mapping[str, Mapping[str, Any]], + instance_path: str, + ) -> bool: + try: + _validate_schema_node( + value, + schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + return True + except _SchemaViolation: + return False + + + def _load_output_schemas(asset_root: Path) -> dict[str, Mapping[str, Any]]: + documents: dict[str, Mapping[str, Any]] = {} + for name in ("ingress.schema.json", "context.schema.json", "review_status.schema.json"): + snapshot = open_bounded_snapshot(asset_root, f"schemas/{name}", logical_input_id=f"schema:{name}") + value = load_json_strict(snapshot) + if not isinstance(value, dict): + raise IngressError("OUTPUT_SCHEMA_SHAPE", f"schema is not an object: {name}") + documents[name] = value + return documents + + + def _validate_output_artifact( + relative_path: str, + value: Any, + schema_documents: Mapping[str, Mapping[str, Any]], + ) -> None: + target = next( + ( + (schema_name, schema_pointer) + for path_pattern, schema_name, schema_pointer in OUTPUT_SCHEMA_TARGETS + if path_pattern.fullmatch(relative_path) + ), + None, + ) + if target is None: + raise IngressError( + "OUTPUT_ARTIFACT_PATH_UNDECLARED", + f"no output schema target is declared for {relative_path}", + details={"path": relative_path}, + ) + schema_name, schema_pointer = target + root_schema = schema_documents[schema_name] + try: + _validate_schema_node( + value, + _schema_pointer(root_schema, schema_pointer), + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=relative_path, + ) + except _SchemaViolation as exc: + raise IngressError( + "OUTPUT_SCHEMA_VALIDATION_FAILED", + str(exc), + details={"path": relative_path, "schema": schema_name, "schema_pointer": schema_pointer}, + ) from exc + + + def _safe_relative_path(relative_path: str) -> PurePosixPath: + if not isinstance(relative_path, str) or not relative_path: + raise IngressError("INVALID_SOURCE_PATH", "source path must be a non-empty string") + if "\x00" in relative_path or "\\" in relative_path: + raise IngressError("INVALID_SOURCE_PATH", "NUL and backslash are forbidden in logical paths") + logical = PurePosixPath(relative_path) + if logical.is_absolute() or any(part in {"", ".", ".."} for part in logical.parts): + raise IngressError("PATH_TRAVERSAL", f"unsafe relative path: {relative_path}") + return logical + + + def _assert_no_symlink_components(root: Path, logical: PurePosixPath) -> None: + current = root + for part in logical.parts: + current = current / part + try: + current_stat = current.lstat() + except FileNotFoundError: + return + if stat.S_ISLNK(current_stat.st_mode): + raise IngressError("SYMLINK_ESCAPE", f"symlink component rejected: {logical}") + + + def open_bounded_snapshot( + approved_root: str | os.PathLike[str], + relative_path: str, + *, + logical_input_id: str = "anonymous", + max_bytes: int = MAX_FILE_BYTES, + require_single_link: bool = True, + ) -> Snapshot: + """Read one regular file once from one descriptor and verify post-read identity.""" + + root_arg = Path(approved_root) + if root_arg.is_symlink(): + raise IngressError("SYMLINK_ROOT_REJECTED", "approved root itself may not be a symlink") + try: + root = root_arg.resolve(strict=True) + except FileNotFoundError as exc: + raise IngressError("APPROVED_ROOT_MISSING", "approved root does not exist") from exc + if not root.is_dir(): + raise IngressError("APPROVED_ROOT_NOT_DIRECTORY", "approved root must be a directory") + logical = _safe_relative_path(relative_path) + _assert_no_symlink_components(root, logical) + candidate = root.joinpath(*logical.parts) + try: + resolved = candidate.resolve(strict=True) + except FileNotFoundError as exc: + raise IngressError("SOURCE_MISSING", f"source is missing: {relative_path}", logical_input_id=logical_input_id) from exc + try: + resolved.relative_to(root) + except ValueError as exc: + raise IngressError("PATH_ESCAPE", f"resolved source escaped approved root: {relative_path}") from exc + flags = os.O_RDONLY + if hasattr(os, "O_CLOEXEC"): + flags |= os.O_CLOEXEC + if hasattr(os, "O_NOFOLLOW"): + flags |= os.O_NOFOLLOW + try: + descriptor = os.open(candidate, flags) + except OSError as exc: + raise IngressError("SOURCE_OPEN_FAILED", f"unable to open source: {relative_path}") from exc + try: + before = os.fstat(descriptor) + if not stat.S_ISREG(before.st_mode): + raise IngressError("NON_REGULAR_SOURCE", f"source is not a regular file: {relative_path}") + if require_single_link and before.st_nlink != 1: + raise IngressError("HARDLINK_POLICY_VIOLATION", f"source link count is {before.st_nlink}") + if before.st_size > max_bytes: + raise IngressError("SOURCE_SIZE_LIMIT", f"source exceeds {max_bytes} bytes") + chunks: list[bytes] = [] + total = 0 + while True: + chunk = os.read(descriptor, min(1024 * 1024, max_bytes + 1 - total)) + if not chunk: + break + chunks.append(chunk) + total += len(chunk) + if total > max_bytes: + raise IngressError("SOURCE_SIZE_LIMIT", f"source exceeds {max_bytes} bytes") + after = os.fstat(descriptor) + finally: + os.close(descriptor) + try: + path_after = candidate.stat(follow_symlinks=False) + except FileNotFoundError as exc: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", "source disappeared after snapshot") from exc + identity_before = (before.st_dev, before.st_ino, before.st_size, before.st_mtime_ns) + identity_after = (after.st_dev, after.st_ino, after.st_size, after.st_mtime_ns) + path_identity = (path_after.st_dev, path_after.st_ino, path_after.st_size, path_after.st_mtime_ns) + if identity_before != identity_after or identity_after != path_identity: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source changed during snapshot: {relative_path}") + raw = b"".join(chunks) + return Snapshot( + logical_input_id=logical_input_id, + relative_path=logical.as_posix(), + resolved_path=str(resolved), + raw=raw, + raw_sha256=hashlib.sha256(raw).hexdigest(), + byte_length=len(raw), + device=after.st_dev, + inode=after.st_ino, + mtime_ns=after.st_mtime_ns, + ) + + + def resolve_stage1_sources( + stage1_run_root: str | os.PathLike[str], + contract_manifest: Mapping[str, Any] | None = None, + ) -> list[dict[str, Any]]: + """Resolve only approved logical kinds; a relocation manifest cannot invent kinds.""" + + root = Path(stage1_run_root).resolve(strict=True) + if not root.is_dir(): + raise IngressError("STAGE1_ROOT_NOT_DIRECTORY", "Stage 1 run root must be a directory") + contracts = [dict(row) for row in DEFAULT_SOURCE_CONTRACTS] + overrides = dict((contract_manifest or {}).get("path_overrides", {})) + approved_ids = {row["logical_input_id"] for row in contracts} + invented = sorted(set(overrides) - approved_ids) + if invented: + raise IngressError("UNAPPROVED_LOGICAL_KIND", "relocation manifest invented logical kinds", details={"ids": invented}) + seen_paths: set[str] = set() + for row in contracts: + path = overrides.get(row["logical_input_id"], row["path"]) + safe = _safe_relative_path(path).as_posix() + if safe in seen_paths: + raise IngressError("DUPLICATE_LOGICAL_MAPPING", f"duplicate physical mapping: {safe}") + seen_paths.add(safe) + row["expected_path"] = row.pop("path") + row["observed_path"] = safe + row["resolution_source"] = ( + "RELEASE_BOUND_CONTRACT_MANIFEST" + if row["logical_input_id"] in overrides + else "DEFAULT_EXACT_PATH" + ) + return contracts + + + def load_release_lock(release_path: str | os.PathLike[str]) -> dict[str, Any]: + """Load a strict release lock without interpreting a historical status label.""" + + path = Path(release_path) + snapshot = open_bounded_snapshot(path.parent, path.name, logical_input_id="stage2_release_lock") + value = load_json_strict(snapshot) + if not isinstance(value, dict): + raise IngressError("RELEASE_LOCK_SHAPE", "release lock must be an object") + release_class = value.get("release_class") + if release_class not in set(RELEASE_MODE.values()): + raise IngressError("RELEASE_CLASS_INVALID", f"unsupported release class: {release_class}") + value = dict(value) + value["_release_raw_sha256"] = snapshot.raw_sha256 + return value + + + def _issue( + code: str, + *, + impact_scope: str = "GLOBAL", + source_refs: Sequence[str] = (), + severity: str = "ERROR", + message: str | None = None, + ) -> dict[str, Any]: + return { + "issue_code": code, + "severity": severity, + "impact_scope": impact_scope, + "scope_refs": sorted(set(source_refs)), + "source_contract_row_refs": sorted(set(source_refs)), + "reason_codes": [code], + "downstream_allowed_actions": [], + "message": message or code, + } + + + def _shape_required(value: Any, keys: Sequence[str]) -> list[str]: + if not isinstance(value, dict): + return list(keys) + return [key for key in keys if key not in value] + + + def _json_pointer_value(document: Any, pointer: str | None) -> tuple[bool, Any]: + if pointer in {None, ""}: + return (pointer == "", document) + if not isinstance(pointer, str) or not pointer.startswith("/"): + return False, None + current = document + for raw_token in pointer[1:].split("/"): + token = raw_token.replace("~1", "/").replace("~0", "~") + if isinstance(current, dict) and token in current: + current = current[token] + elif isinstance(current, list) and token.isdigit() and int(token) < len(current): + current = current[int(token)] + else: + return False, None + return True, current + + + def _release_stage1_source_rows(release_lock: Mapping[str, Any]) -> list[Mapping[str, Any]]: + rows = release_lock.get("stage1_sources") + if not isinstance(rows, list): + dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) + rows = dependency.get("stage1_sources") if isinstance(dependency, dict) else None + return [row for row in rows if isinstance(row, dict)] if isinstance(rows, list) else [] + + + def _adapter_decision(release_lock: Mapping[str, Any], adapter_id: str) -> Mapping[str, Any] | None: + for row in release_lock.get("adapter_decisions", []): + if isinstance(row, dict) and row.get("adapter_id") == adapter_id and isinstance(row.get("decision"), dict): + return row["decision"] + return None + + + def _closed_adapter_shape_errors( + document: Any, + *, + logical_id: str, + adapter_id: str, + required_keys: Sequence[str], + release_lock: Mapping[str, Any], + ) -> list[str]: + errors: list[str] = [] + if required_keys: + errors.extend(f"missing root key {key}" for key in _shape_required(document, required_keys)) + decision = _adapter_decision(release_lock, adapter_id) + if decision is not None: + root_shape = decision.get("root_shape") + if root_shape == "ARRAY" and not isinstance(document, list): + errors.append("root must be an array") + elif root_shape == "OBJECT_ENVELOPE" and not isinstance(document, dict): + errors.append("root must be an object envelope") + if isinstance(document, dict): + errors.extend( + f"missing root key {key}" + for key in _shape_required(document, decision.get("required_root_fields", [])) + ) + if isinstance(document, list): + required_item_fields = decision.get("required_item_fields", decision.get("required_row_fields", [])) + if isinstance(required_item_fields, list): + for index, item in enumerate(document): + for key in _shape_required(item, required_item_fields): + errors.append(f"row {index} missing {key}") + if decision is None: + fallback_required: dict[str, tuple[str, ...]] = { + "evidence_indexed": ("schema_contract_version", "items"), + "evidence_event_candidates": ("schema_version", "items"), + "domain_activation_manifest": SG01_PROJECTION_FIELDS, + "signal_manifest": ("downstream_read_sets", "files"), + "legal_effect_structures": ("schema_version", "structure_records"), + "fact_ledger_writer_report": ( + "schema_version", + "row_count", + "gate_firings", + "domain_effect_coverage", + "calculation_readiness", + "blocked_review_items", + "conservation", + "final_sha256", + ), + } + fallback = fallback_required.get(logical_id, ()) + if fallback: + errors.extend(f"missing root key {key}" for key in _shape_required(document, fallback)) + if logical_id in {"bo", "fact_ledger_base"} and not isinstance(document, list): + errors.append("root must be an array") + return sorted(set(errors)) + + + def _schema_document_index(deployment_documents: Mapping[str, Any]) -> dict[str, Mapping[str, Any]]: + result: dict[str, Mapping[str, Any]] = {} + for path, document in deployment_documents.items(): + if not isinstance(document, dict): + continue + result[path] = document + result[PurePosixPath(path).name] = document + schema_id = document.get("$id") + if isinstance(schema_id, str): + result[schema_id] = document + return result + + + def _source_hash_index(document: Mapping[str, Any] | None) -> dict[str, str]: + result: dict[str, str] = {} + if not isinstance(document, dict): + return result + candidate_arrays: list[Any] = [] + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(document.get(key), list): + candidate_arrays.append(document[key]) + for wrapper in ("completion_seal", "manifest", "payload", "data"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(nested.get(key), list): + candidate_arrays.append(nested[key]) + for rows in candidate_arrays: + for row in rows: + if not isinstance(row, dict): + continue + digest = row.get("raw_sha256", row.get("sha256")) + if not isinstance(digest, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", digest) is None: + continue + for key in ("logical_input_id", "path", "observed_path", "logical_id"): + identifier = row.get(key) + if isinstance(identifier, str) and identifier: + result[identifier] = digest.lower() + return result + + + def _source_producer_index(document: Mapping[str, Any] | None) -> dict[str, str]: + """Index producer evidence carried by a bounded completion/manifest row.""" + + result: dict[str, str] = {} + if not isinstance(document, dict): + return result + candidate_arrays: list[Any] = [] + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(document.get(key), list): + candidate_arrays.append(document[key]) + for wrapper in ("completion_seal", "manifest", "payload", "data"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(nested.get(key), list): + candidate_arrays.append(nested[key]) + for rows in candidate_arrays: + for row in rows: + if not isinstance(row, dict): + continue + producer = next( + ( + row.get(key) + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by") + if isinstance(row.get(key), str) and row.get(key) + ), + None, + ) + if not isinstance(producer, str): + continue + for key in ("logical_input_id", "path", "observed_path", "logical_id"): + identifier = row.get(key) + if isinstance(identifier, str) and identifier: + result[identifier] = producer + return result + + + def _producer_value(document: Any) -> str | None: + if not isinstance(document, dict): + return None + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): + value = document.get(key) + if isinstance(value, str) and value: + return value + for wrapper in ("metadata", "meta", "handoff", "payload"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): + value = nested.get(key) + if isinstance(value, str) and value: + return value + # P3/P4 are closed one-key wrappers in the Stage 1 v8 handoff contract. + for wrapper in ( + "stage1_part3_review_handoff", + "stage1_part4_review_handoff", + ): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("created_by", "finalized_by"): + value = nested.get(key) + if isinstance(value, str) and value: + return value + return None + + + def _producer_matches( + observed: str, + expected: str, + alias_id: str | None, + release_lock: Mapping[str, Any], + ) -> bool: + if observed == expected: + return True + if alias_id is None: + return False + decision = _adapter_decision(release_lock, alias_id) + if decision is None or decision.get("bidirectional_match_allowed") is not True: + return False + pair = {decision.get("schema_writer_id"), decision.get("orchestration_producer_id")} + return {observed, expected} == pair + + + def _identity_ref(document: Any, pointer: str | None, logical_id: str) -> dict[str, Any]: + if pointer is None: + return {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": logical_id} + found, value = _json_pointer_value(document, pointer) + if not found or value is None: + return {"value": None, "disposition": "MISSING", "source_ref": f"{logical_id}#{pointer}"} + return {"value": str(value), "disposition": "OBSERVED", "source_ref": f"{logical_id}#{pointer}"} + + + def validate_ingress_contracts( + snapshots: Mapping[str, Snapshot], + contracts: Sequence[Mapping[str, Any]], + release_lock: Mapping[str, Any], + *, + deployment_snapshots: Mapping[str, Snapshot] | None = None, + deployment_documents: Mapping[str, Any] | None = None, + completion_seal: Mapping[str, Any] | None = None, + contract_manifest: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Strictly parse sources and verify release-bound schema, producer, identity, and seal rows.""" + + documents: dict[str, Any] = {} + rows: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + deployment_snapshots = deployment_snapshots or {} + deployment_documents = deployment_documents or {} + deployment_by_path = {snapshot.relative_path: snapshot for snapshot in deployment_snapshots.values()} + schema_documents = _schema_document_index(deployment_documents) + release_source_rows = _release_stage1_source_rows(release_lock) + release_ids = [str(row.get("logical_input_id")) for row in release_source_rows] + duplicate_release_ids = sorted(key for key, count in Counter(release_ids).items() if count > 1) + if duplicate_release_ids: + raise IngressError( + "RELEASE_SOURCE_CONTRACT_DUPLICATE", + "release stage1_sources contains duplicate logical_input_id rows", + details={"logical_input_ids": duplicate_release_ids}, + ) + expected_fixed = { + str(row["logical_input_id"]): str(row["path"]) + for row in DEFAULT_SOURCE_CONTRACTS + } + expected_release_ids = set(expected_fixed) | {"signal_payload_family"} + observed_release_ids = set(release_ids) + if observed_release_ids != expected_release_ids: + raise IngressError( + "RELEASE_SOURCE_CONTRACT_SET_MISMATCH", + "release stage1_sources must be the exact 16 fixed inputs plus signal_payload_family", + details={ + "missing": sorted(expected_release_ids - observed_release_ids), + "extra": sorted(observed_release_ids - expected_release_ids), + }, + ) + release_rows = {str(row.get("logical_input_id")): row for row in release_source_rows} + for logical_id, expected_path in expected_fixed.items(): + release_row = release_rows[logical_id] + if release_row.get("path") != expected_path or release_row.get("path_rule") not in {None, ""}: + raise IngressError( + "RELEASE_SOURCE_FIXED_PATH_MISMATCH", + f"fixed source path contract mismatch: {logical_id}", + ) + signal_family = release_rows["signal_payload_family"] + if ( + signal_family.get("path") is not None + or signal_family.get("path_rule") != "signals/" + or signal_family.get("adapter_id") != "S2A-SIGNAL-PAYLOAD-FAMILY-V1" + or signal_family.get("raw_hash_source") != "MANIFEST_ROW" + ): + raise IngressError( + "SIGNAL_PAYLOAD_FAMILY_CONTRACT_MISMATCH", + "signal_payload_family must use the approved manifest-expanded path contract", + ) + completion_hashes = _source_hash_index(completion_seal) + manifest_hashes = _source_hash_index(contract_manifest) + completion_producers = _source_producer_index(completion_seal) + manifest_producers = _source_producer_index(contract_manifest) + for contract in contracts: + logical_id = str(contract["logical_input_id"]) + snapshot = snapshots.get(logical_id) + release_row = release_rows.get(logical_id) + contract_missing = release_row is None + release_row = release_row or {} + alias_value = release_row.get("producer_alias", release_row.get("producer_alias_id")) + alias_id = str(alias_value) if isinstance(alias_value, str) else None + schema_ref = release_row.get("schema_ref") if isinstance(release_row.get("schema_ref"), dict) else None + row = { + "logical_input_id": logical_id, + "requirement_class": REQUIREMENT_CLASS_ENUM.get( + str(contract.get("criticality")), + "INTEGRITY_CORROBORATOR", + ), + "expected_path": contract.get("expected_path"), + "observed_path": contract.get("observed_path"), + "resolution_source": contract.get("resolution_source"), + "schema_id": schema_ref.get("$id") if schema_ref else release_row.get("schema_id"), + "schema_sha256": schema_ref.get("sha256") if schema_ref else release_row.get("schema_sha256"), + "producer_id": release_row.get("producer_id"), + "producer_alias_id": alias_id, + "adapter_id": release_row.get("adapter_id", ADAPTER_IDS.get(logical_id, "S2A-UNBOUND-V1")), + "run_identity_ref": release_row.get("run_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), + "transaction_identity_ref": release_row.get("transaction_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), + "scope_refs": [logical_id], + "source_contract_row_refs": [logical_id], + "reason_codes": [], + "downstream_allowed_actions": [], + "issue_codes": [], + } + if contract_missing: + code = "RELEASE_SOURCE_CONTRACT_MISSING" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + declared_path = release_row.get("path") + if isinstance(declared_path, str) and declared_path != contract.get("expected_path"): + code = "RELEASE_SOURCE_PATH_MISMATCH" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if snapshot is None: + row.update( + { + "raw_sha256": None, + "byte_length": 0, + "parse_status": "NOT_OBSERVED", + "schema_status": "UNEVALUABLE", + "seal_status": "UNEVALUABLE", + "scope_technical_disposition": "UNAVAILABLE", + "impact_scope": "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER", + } + ) + code = "SOURCE_MISSING" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope=row["impact_scope"], source_refs=[logical_id])) + rows.append(row) + continue + row["raw_sha256"] = snapshot.raw_sha256 + row["byte_length"] = snapshot.byte_length + try: + document = load_json_strict( + snapshot, + max_depth=int(release_lock.get("limits", {}).get("max_json_depth", MAX_JSON_DEPTH)), + max_items=int(release_lock.get("limits", {}).get("max_json_items", MAX_JSON_ITEMS)), + ) + documents[logical_id] = document + row["parse_status"] = "PASS" + except IngressError as exc: + row["parse_status"] = "FAIL" + row["schema_status"] = "UNEVALUABLE" + row["seal_status"] = "UNEVALUABLE" + row["scope_technical_disposition"] = "UNAVAILABLE" + row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" + row["reason_codes"].append(exc.code) + row["issue_codes"].append(exc.code) + issues.append(_issue(exc.code, impact_scope=row["impact_scope"], source_refs=[logical_id], message=str(exc))) + rows.append(row) + continue + expected_adapter = ADAPTER_IDS.get(logical_id) + if expected_adapter is not None and release_row.get("adapter_id") not in {None, expected_adapter}: + code = "ADAPTER_ID_MISMATCH" + row["schema_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if schema_ref is not None: + schema_path = schema_ref.get("path") + schema_snapshot = deployment_by_path.get(schema_path) if isinstance(schema_path, str) else None + schema_document = deployment_documents.get(schema_path) if isinstance(schema_path, str) else None + expected_schema_hash = schema_ref.get("sha256") + expected_schema_id = schema_ref.get("$id") + if schema_snapshot is None or not isinstance(schema_document, dict): + schema_error = "SCHEMA_REF_NOT_IN_BOUNDED_DEPLOYMENT" + elif not isinstance(expected_schema_hash, str) or schema_snapshot.raw_sha256 != expected_schema_hash.lower(): + schema_error = "SCHEMA_HASH_MISMATCH" + elif expected_schema_id is not None and schema_document.get("$id") != expected_schema_id: + schema_error = "SCHEMA_ID_MISMATCH" + else: + schema_error = None + try: + _validate_schema_node( + document, + schema_document, + root_schema=schema_document, + schema_documents=schema_documents, + instance_path=logical_id, + ) + except _SchemaViolation as exc: + schema_error = "SOURCE_SCHEMA_VALIDATION_FAILED" + issues.append( + _issue( + schema_error, + impact_scope="CLUSTER", + source_refs=[logical_id], + message=str(exc), + ) + ) + if schema_error is not None: + row["schema_status"] = "FAIL" + row["reason_codes"].append(schema_error) + row["issue_codes"].append(schema_error) + if schema_error != "SOURCE_SCHEMA_VALIDATION_FAILED": + issues.append(_issue(schema_error, impact_scope="GLOBAL", source_refs=[logical_id])) + else: + row["schema_status"] = "PASS" + else: + adapter_errors = _closed_adapter_shape_errors( + document, + logical_id=logical_id, + adapter_id=str(row["adapter_id"]), + required_keys=release_row.get("required_keys", []), + release_lock=release_lock, + ) + if contract_missing: + row["schema_status"] = "UNEVALUABLE" + elif adapter_errors: + code = "ADAPTER_REQUIRED_KEY_MISSING" + row["schema_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append( + _issue( + code, + impact_scope="CLUSTER", + source_refs=[logical_id], + message="; ".join(adapter_errors), + ) + ) + else: + row["schema_status"] = "PASS" + expected_producer = release_row.get("producer_id") + document_producer = _producer_value(document) + sealed_producer = ( + completion_producers.get(logical_id) + or completion_producers.get(str(contract.get("observed_path"))) + or manifest_producers.get(logical_id) + or manifest_producers.get(str(contract.get("observed_path"))) + ) + if ( + document_producer is not None + and sealed_producer is not None + and document_producer != sealed_producer + ): + code = "PRODUCER_EVIDENCE_CONFLICT" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + observed_producer = document_producer or sealed_producer + if isinstance(expected_producer, str): + if observed_producer is None: + code = "PRODUCER_ID_UNEVALUABLE" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="CLUSTER", source_refs=[logical_id])) + elif not _producer_matches(observed_producer, expected_producer, alias_id, release_lock): + code = "PRODUCER_ID_MISMATCH" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + row["run_identity_ref"] = _identity_ref(document, release_row.get("run_identity_pointer"), logical_id) + row["transaction_identity_ref"] = _identity_ref( + document, + release_row.get("transaction_identity_pointer"), + logical_id, + ) + raw_hash_source = str(release_row.get("raw_hash_source", "NONE")) + if raw_hash_source in {"CASE_RUN_COMPLETION_SEAL", "COMPLETION_SEAL", "COMPLETION_SEAL_ROW"}: + expected_hash = completion_hashes.get(logical_id) or completion_hashes.get(str(contract.get("observed_path"))) + elif raw_hash_source in {"CONTRACT_MANIFEST", "CONTRACT_MANIFEST_ROW", "MANIFEST_ROW"}: + expected_hash = manifest_hashes.get(logical_id) or manifest_hashes.get(str(contract.get("observed_path"))) + elif raw_hash_source in {"COMPLETION_SEAL_OR_CONTRACT_MANIFEST", "SEALED_ROW"}: + expected_hash = ( + completion_hashes.get(logical_id) + or completion_hashes.get(str(contract.get("observed_path"))) + or manifest_hashes.get(logical_id) + or manifest_hashes.get(str(contract.get("observed_path"))) + ) + elif raw_hash_source in {"UNAVAILABLE_DEV", "NONE"}: + expected_hash = None + else: + expected_hash = None + code = "RAW_HASH_SOURCE_UNAPPROVED" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if expected_hash is not None and expected_hash != snapshot.raw_sha256: + code = "RAW_HASH_MISMATCH" + row["seal_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + else: + row["seal_status"] = "PASS" if expected_hash else "UNEVALUABLE" + row["scope_technical_disposition"] = ( + "UNAVAILABLE" + if contract_missing or any(code in row["issue_codes"] for code in {"RAW_HASH_MISMATCH", "SCHEMA_HASH_MISMATCH", "SCHEMA_ID_MISMATCH"}) + else "AVAILABLE" + if not row["issue_codes"] + else "AVAILABLE_WITH_ISSUES" + ) + row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" + rows.append(row) + for identity_kind, field in ( + ("RUN", "run_identity_ref"), + ("TRANSACTION", "transaction_identity_ref"), + ): + observed_values = { + str(row[field]["value"]) + for row in rows + if row[field].get("disposition") == "OBSERVED" and row[field].get("value") is not None + } + if len(observed_values) > 1: + code = f"{identity_kind}_IDENTITY_CONFLICT" + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=sorted(observed_values))) + for row in rows: + if row[field].get("disposition") == "OBSERVED": + row["issue_codes"] = sorted(set(row["issue_codes"] + [code])) + row["reason_codes"] = sorted(set(row["reason_codes"] + [code])) + row["scope_technical_disposition"] = "UNAVAILABLE" + return {"documents": documents, "source_contract_rows": rows, "issues": issues} + + + def _records_from_signal_document(document: Any) -> list[Any]: + if isinstance(document, list): + return list(document) + if isinstance(document, dict): + for key in ("signals", "records", "items"): + value = document.get(key) + if isinstance(value, list): + return list(value) + return [document] + return [document] + + + def _record_signal_id(record: Any) -> str | None: + if not isinstance(record, dict): + return None + value = record.get("signal_id") + if isinstance(value, str) and value: + return value + for wrapper in ("domain_activation_manifest", "payload", "data"): + nested = record.get(wrapper) + if isinstance(nested, dict) and isinstance(nested.get("signal_id"), str): + return nested["signal_id"] + return None + + + def expand_stage2_signal_all( + stage1_run_root: str | os.PathLike[str], + signal_manifest: Mapping[str, Any], + *, + max_file_bytes: int = MAX_FILE_BYTES, + max_total_bytes: int = MAX_RUN_BYTES, + signal_registry: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Expand Stage 2 ALL while separating semantic and integrity-only universes.""" + + downstream = signal_manifest.get("downstream_read_sets", {}) + stage2 = downstream.get("stage2", []) if isinstance(downstream, dict) else [] + if stage2 != ["ALL"]: + raise IngressError("SIGNAL_ALL_CONTRACT", "downstream_read_sets.stage2 must equal ['ALL']") + files = signal_manifest.get("files") + if not isinstance(files, list): + raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest files must be an array") + transaction_id = str(signal_manifest.get("manifest_transaction_id", signal_manifest.get("transaction_id", "MISSING"))) + file_rows: list[dict[str, Any]] = [] + semantic_rows: list[dict[str, Any]] = [] + integrity_rows: list[dict[str, Any]] = [] + occurrences: list[dict[str, Any]] = [] + payload_snapshots: list[Snapshot] = [] + issues: list[dict[str, Any]] = [] + path_counter: Counter[str] = Counter() + parsed_documents: dict[str, Any] = {} + aggregate_bytes = 0 + registry_entries = { + str(row.get("file")): row + for row in (signal_registry or {}).get("entries", []) + if isinstance(row, dict) and isinstance(row.get("file"), str) + } + compatibility_files = { + str(path) + for path in (signal_registry or {}).get("compatibility_views", []) + if isinstance(path, str) + } + domain_envelope_schema = (signal_registry or {}).get("domain_envelope") + observed_registry_files: set[str] = set() + for index, entry in enumerate(files): + if not isinstance(entry, dict) or not isinstance(entry.get("path"), str): + raise IngressError("SIGNAL_FILE_ROW_SHAPE", f"invalid signal file row at index {index}") + relative_payload = _safe_relative_path(entry["path"]).as_posix() + if relative_payload.startswith("signals/"): + raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "manifest file path must not include signals/ prefix") + physical = f"signals/{relative_payload}" + snapshot = open_bounded_snapshot( + stage1_run_root, + physical, + logical_input_id=f"signal_file:{index}", + max_bytes=max_file_bytes, + ) + document = load_json_strict(snapshot) + payload_snapshots.append(snapshot) + aggregate_bytes += snapshot.byte_length + if aggregate_bytes > max_total_bytes: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "signal ALL payloads exceed remaining run byte budget") + parsed_documents[relative_payload] = document + kind = entry.get("kind", "canonical") + if kind not in SEMANTIC_SIGNAL_KINDS | {"compatibility_view"}: + raise IngressError("SIGNAL_KIND_UNAPPROVED", f"unapproved signal file kind: {kind}") + semantic = kind in SEMANTIC_SIGNAL_KINDS + expected_hash = entry.get( + "file_sha256", entry.get("sha256", entry.get("raw_sha256")) + ) + row = { + "manifest_index": index, + "file_path": relative_payload, + "physical_path": physical, + "kind": kind, + "raw_sha256": snapshot.raw_sha256, + "byte_length": snapshot.byte_length, + "semantic": semantic, + "manifest_declared_record_count": entry.get("record_count"), + } + if expected_hash is not None and expected_hash != snapshot.raw_sha256: + row["hash_status"] = "FAIL" + issues.append(_issue("SIGNAL_FILE_HASH_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + else: + row["hash_status"] = "PASS" if expected_hash else "UNEVALUABLE" + records = _records_from_signal_document(document) + row["observed_record_count"] = len(records) + declared_count = entry.get("record_count") + if isinstance(declared_count, int) and declared_count != len(records): + row["record_count_status"] = "FAIL" + issues.append(_issue("SIGNAL_RECORD_COUNT_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + else: + row["record_count_status"] = "PASS" if isinstance(declared_count, int) else "UNEVALUABLE" + registry_row = registry_entries.get(relative_payload) + if kind == "canonical": + if signal_registry is not None and registry_row is None: + issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif registry_row is not None: + observed_registry_files.add(relative_payload) + declared_schema = entry.get("schema", entry.get("schema_path")) + if declared_schema is not None and declared_schema != registry_row.get("schema"): + issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif kind == "compatibility_view": + if relative_payload in registry_entries: + issues.append(_issue("SIGNAL_COMPATIBILITY_SUBSTITUTION", impact_scope="SIGNAL", source_refs=[physical])) + if signal_registry is not None and relative_payload not in compatibility_files: + issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif kind == "domain_signal": + declared_schema = entry.get("schema", entry.get("schema_path")) + if signal_registry is not None and declared_schema not in {None, domain_envelope_schema}: + issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + file_rows.append(row) + path_counter[relative_payload] += 1 + if semantic: + semantic_rows.append(row) + for record_ordinal, record in enumerate(records): + signal_id = _record_signal_id(record) + occurrence_key = [transaction_id, relative_payload, record_ordinal, signal_id] + occurrences.append( + { + "occurrence_key": occurrence_key, + "occurrence_ref": f"SIGO-{canonical_digest(occurrence_key)[:24]}", + "manifest_transaction_id": transaction_id, + "file_path": relative_payload, + "record_ordinal": record_ordinal, + "signal_id": signal_id, + "disposition": "UNMAPPED" if signal_id is None else "UNUSED", + "binding_refs": [], + "raw_record_sha256": canonical_digest(record), + "record": record, + } + ) + else: + integrity_rows.append(row) + duplicates = sorted(path for path, count in path_counter.items() if count > 1) + if duplicates: + issues.append(_issue("SIGNAL_ALL_DUPLICATE_FILE_ROW", impact_scope="SIGNAL", source_refs=duplicates)) + manifest_counter = Counter((i, row["file_path"], row["kind"]) for i, row in enumerate(file_rows)) + partition_counter = Counter((row["manifest_index"], row["file_path"], row["kind"]) for row in semantic_rows + integrity_rows) + missing_registry_files = sorted(set(registry_entries) - observed_registry_files) if signal_registry is not None else [] + if missing_registry_files: + issues.append( + _issue( + "SIGNAL_REGISTRY_COVERAGE_MISMATCH", + impact_scope="SIGNAL", + source_refs=[f"signals/{path}" for path in missing_registry_files], + ) + ) + file_conservation = ( + manifest_counter == partition_counter + and not duplicates + and not missing_registry_files + and not any(row["hash_status"] == "FAIL" or row["record_count_status"] == "FAIL" for row in file_rows) + ) + record_counter = Counter(tuple(row["occurrence_key"]) for row in occurrences) + partitioned_record_counter = Counter( + tuple(row["occurrence_key"]) + for row in occurrences + if row["disposition"] in {"USED", "UNUSED", "UNMAPPED"} + ) + record_conservation = record_counter == partitioned_record_counter + return { + "manifest_transaction_id": transaction_id, + "ordered_file_rows": file_rows, + "semantic_file_rows": semantic_rows, + "integrity_only_file_rows": integrity_rows, + "record_occurrences": occurrences, + "used_record_occurrences": [], + "unused_record_occurrences": [row for row in occurrences if row["disposition"] == "UNUSED"], + "unmapped_record_occurrences": [row for row in occurrences if row["disposition"] == "UNMAPPED"], + "_parsed_documents_by_path": parsed_documents, + "_payload_snapshots": payload_snapshots, + "file_conservation_pass": file_conservation, + "record_conservation_pass": record_conservation, + "aggregate_payload_bytes": aggregate_bytes, + "issues": issues, + } + + + def _collect_values_for_keys(value: Any, keys: frozenset[str]) -> set[str]: + result: set[str] = set() + stack = [value] + while stack: + current = stack.pop() + if isinstance(current, dict): + for key, child in current.items(): + if key in keys: + if isinstance(child, list): + result.update(str(item) for item in child if item is not None) + elif child is not None: + result.add(str(child)) + stack.append(child) + elif isinstance(current, list): + stack.extend(current) + return result + + + def bind_signal_occurrences(signal_all: MutableMapping[str, Any], documents: Mapping[str, Any]) -> dict[str, Any]: + """Bind each semantic signal occurrence to explicit Stage 1 references without deduplication.""" + + explicit_signal_ids = _collect_values_for_keys( + documents, + frozenset({"signal_id", "signal_ids", "signal_refs", "emitted_signal_ids", "required_signal_ids"}), + ) + known_refs = { + "fact_id": _collect_values_for_keys(documents.get("fact_ledger_base"), frozenset({"fact_id"})), + "source_bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "source_bo_id", "source_bo_ids"})), + "bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "bo_id"})), + "structure_id": _collect_values_for_keys(documents.get("legal_effect_structures"), frozenset({"structure_id"})), + "domain_id": _collect_values_for_keys(documents, frozenset({"domain_id", "domain_ids", "active_domain_ids"})), + "evidence_id": _collect_values_for_keys(documents.get("evidence_indexed"), frozenset({"evidence_id", "id"})), + "event_id": _collect_values_for_keys(documents.get("evidence_event_candidates"), frozenset({"event_id", "id"})), + } + link_keys = { + "fact_id": ("fact_id", "fact_ids"), + "source_bo_id": ("source_bo_id", "source_bo_ids"), + "bo_id": ("bo_id", "bo_ids"), + "structure_id": ("structure_id", "structure_ids"), + "domain_id": ("domain_id", "domain_ids"), + "evidence_id": ("evidence_id", "evidence_ids"), + "event_id": ("event_id", "event_ids"), + } + for occurrence in signal_all.get("record_occurrences", []): + signal_id = occurrence.get("signal_id") + record = occurrence.get("record") + bindings: set[str] = set() + if isinstance(signal_id, str) and signal_id in explicit_signal_ids: + bindings.add(f"signal_id:{signal_id}") + for ref_kind, candidate_keys in link_keys.items(): + observed = _collect_values_for_keys(record, frozenset(candidate_keys)) + for ref in sorted(observed & known_refs[ref_kind]): + bindings.add(f"{ref_kind}:{ref}") + if not isinstance(signal_id, str) or not signal_id: + occurrence["disposition"] = "UNMAPPED" + elif bindings: + occurrence["disposition"] = "USED" + else: + occurrence["disposition"] = "UNUSED" + occurrence["binding_refs"] = sorted(bindings) + for disposition, key in ( + ("USED", "used_record_occurrences"), + ("UNUSED", "unused_record_occurrences"), + ("UNMAPPED", "unmapped_record_occurrences"), + ): + signal_all[key] = [ + row for row in signal_all.get("record_occurrences", []) if row.get("disposition") == disposition + ] + source_counter = Counter(tuple(row["occurrence_key"]) for row in signal_all.get("record_occurrences", [])) + partition_counter = Counter( + tuple(row["occurrence_key"]) + for key in ("used_record_occurrences", "unused_record_occurrences", "unmapped_record_occurrences") + for row in signal_all[key] + ) + signal_all["record_conservation_pass"] = source_counter == partition_counter + return dict(signal_all) + + + def _activation_payload(value: Mapping[str, Any]) -> Mapping[str, Any]: + for key in ("domain_activation_manifest", "activation", "payload", "data"): + nested = value.get(key) + if isinstance(nested, dict) and any(field in nested for field in SG01_PROJECTION_FIELDS): + return nested + return value + + + def verify_activation_projection( + routing_activation: Mapping[str, Any], + signal_activation: Mapping[str, Any], + *, + routing_raw_sha256: str | None = None, + signal_raw_sha256: str | None = None, + ) -> dict[str, Any]: + """Compare approved semantic SG-01 projection while retaining both raw hashes.""" + + left = _activation_payload(routing_activation) + right = _activation_payload(signal_activation) + missing_left = [field for field in SG01_PROJECTION_FIELDS if field not in left] + missing_right = [field for field in SG01_PROJECTION_FIELDS if field not in right] + if missing_left or missing_right: + raise IngressError( + "SG01_PROJECTION_SHAPE", + "both activation artifacts must expose the complete approved 17-field projection", + details={"routing_missing": missing_left, "signal_missing": missing_right}, + ) + + def project(value: Mapping[str, Any]) -> dict[str, Any]: + result: dict[str, Any] = {} + for field in SG01_PROJECTION_FIELDS: + child = value[field] + if field in SG01_SET_FIELDS: + if not isinstance(child, list): + raise IngressError("SG01_PROJECTION_SHAPE", f"{field} must be an array") + child = sorted({canonical_json_bytes(item): item for item in child}.values(), key=canonical_json_bytes) + result[field] = child + return result + + left_projection = project(left) + right_projection = project(right) + if left_projection != right_projection: + raise IngressError( + "SG01_SEMANTIC_DRIFT", + "routing activation and signal SG-01 semantic projections differ", + details={"routing_projection": left_projection, "signal_projection": right_projection}, + ) + return { + "status": "PASS", + "projection": left_projection, + "projection_sha256": canonical_digest(left_projection), + "routing_raw_sha256": routing_raw_sha256, + "signal_raw_sha256": signal_raw_sha256, + "compared_keys": list(SG01_PROJECTION_FIELDS), + } + + + def _array_rows(value: Any, preferred_keys: Sequence[str]) -> list[Any]: + if isinstance(value, list): + return list(value) + if isinstance(value, dict): + for key in preferred_keys: + candidate = value.get(key) + if isinstance(candidate, list): + return list(candidate) + return [] + + + def verify_cross_artifact_seals( + documents: Mapping[str, Any], + snapshots: Mapping[str, Snapshot], + deployment_snapshots: Mapping[str, Snapshot] | None = None, + ) -> dict[str, Any]: + """Recompute the P1 guard and current-v8 producer invariants.""" + + checks: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + deployment_snapshots = deployment_snapshots or {} + p1 = documents.get("stage1_part1_soft_gate_handoff") + if isinstance(p1, dict): + digest_guard = p1.get("digest_guard") + if not isinstance(digest_guard, dict): + issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff"])) + digest_guard = {} + elif any(key not in digest_guard for key in P1_DIGEST_KEYS): + issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff#digest_guard"])) + for digest_key, logical_id in P1_DIGEST_KEYS.items(): + source = snapshots.get(logical_id) or deployment_snapshots.get(logical_id) + observed = source.raw_sha256 if source else None + expected = digest_guard.get(digest_key) + passed = expected is not None and observed is not None and expected == observed + checks.append({"check_id": f"P1:{digest_key}", "status": "PASS" if passed else "UNEVALUABLE" if source is None else "FAIL"}) + if expected is not None and observed is not None and not passed: + issues.append(_issue("P1_DIGEST_MISMATCH", source_refs=[logical_id])) + else: + issues.append(_issue("P1_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part1_soft_gate_handoff"])) + p2 = documents.get("stage1_part2_review_handoff") + if p2 is not None and not isinstance(p2, dict): + issues.append(_issue("P2_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part2_review_handoff"])) + for stage in (3, 4): + logical = f"stage1_part{stage}_review_handoff" + value = documents.get(logical) + if value is not None: + wrapper_present = isinstance(value, dict) and isinstance(value.get(logical), dict) + if not wrapper_present: + issues.append(_issue(f"P{stage}_WRAPPER_MISSING", source_refs=[logical])) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + for index, row in enumerate(ledger_rows): + if not isinstance(row, dict) or "domain_effects" not in row or "calculation_requests" not in row: + issues.append(_issue("CURRENT_V8_LEDGER_EXTENSION_MISSING", impact_scope="FACT", source_refs=[f"fact_ledger_base#/{index}"])) + return {"checks": checks, "issues": issues, "passed": not any(item["severity"] == "ERROR" for item in issues)} + + + def mint_stage2_id( + namespace: str, + canonical_tuple: Any, + *, + prefix: str | None = None, + algorithm_version: str = ALGORITHM_VERSION, + collision_registry: MutableMapping[str, str] | None = None, + ) -> dict[str, str]: + """Mint a domain-separated ID and reject short-ID collisions.""" + + if not re.fullmatch(r"[A-Za-z][A-Za-z0-9_.:-]{0,127}", namespace): + raise IngressError("INVALID_ID_NAMESPACE", f"invalid namespace: {namespace}") + mint_input = { + "namespace": namespace, + "algorithm_version": algorithm_version, + "canonical_tuple": canonical_tuple, + } + full_digest = canonical_digest(mint_input) + external_prefix = prefix or namespace.upper().replace("_", "-") + external_id = f"{external_prefix}-{full_digest[:24]}" + if collision_registry is not None: + prior = collision_registry.get(external_id) + if prior is not None and prior != full_digest: + raise IngressError("MINTED_ID_COLLISION", f"collision for {external_id}") + collision_registry[external_id] = full_digest + return { + "id": external_id, + "mint_input_sha256": full_digest, + "algorithm_version": algorithm_version, + } + + + def mint_context_occurrence_id( + kind: str, + logical_artifact_id: str, + json_pointer: str, + raw_value: Any, + explicit_lineage_refs: Sequence[str], + *, + stage1_canonical_id: str | None = None, + collision_registry: MutableMapping[str, str] | None = None, + ) -> dict[str, Any]: + """Mint an occurrence/lineage context ID without fuzzy entity resolution.""" + + if kind not in {"party", "object"}: + raise IngressError("CONTEXT_KIND_INVALID", "context kind must be party or object") + raw_value_sha256 = canonical_digest(raw_value) + canonical_tuple = [ + logical_artifact_id, + json_pointer, + raw_value_sha256, + sorted(set(explicit_lineage_refs)), + ] + minted = mint_stage2_id( + f"{kind}_context_occurrence", + canonical_tuple, + prefix="PC" if kind == "party" else "OC", + collision_registry=collision_registry, + ) + source_ref = _source_ref( + logical_artifact_id, + json_pointer, + raw_value, + stage1_id=stage1_canonical_id, + ) + identity_kind = "PARTY_CONTEXT" if kind == "party" else "OBJECT_CONTEXT" + result: dict[str, Any] = { + "identity_kind": identity_kind, + f"{kind}_context_id": stage1_canonical_id or minted["id"], + f"stage1_{kind}_id": stage1_canonical_id, + "identity_disposition": ( + "PRESERVED_STAGE1_CANONICAL_ID" + if stage1_canonical_id is not None + else "IDENTITY_UNRESOLVED" + ), + "source_refs": [source_ref], + "explicit_lineage_refs": sorted(set(explicit_lineage_refs)), + "display_label_nfc": unicodedata.normalize("NFC", str(raw_value)), + "raw_display_value_sha256": raw_value_sha256, + } + if stage1_canonical_id is None: + result["derivation"] = { + "derivation_id": minted["id"], + "algorithm_version": ALGORITHM_VERSION, + "sorted_input_refs": sorted(set(explicit_lineage_refs)) or [f"{logical_artifact_id}:{json_pointer}"], + "mint_input_sha256": minted["mint_input_sha256"], + } + return result + + + def _extract_review_arrays(document: Any) -> list[Any]: + if not isinstance(document, dict): + return [] + result: list[Any] = [] + for key in ("review_items", "review_queue", "blocked_review_items", "unresolved_review_items"): + value = document.get(key) + if isinstance(value, list): + result.extend(value) + for wrapper in ( + "handoff", + "payload", + "review", + "data", + "stage1_part3_review_handoff", + "stage1_part4_review_handoff", + ): + nested = document.get(wrapper) + if isinstance(nested, dict): + result.extend(_extract_review_arrays(nested)) + return result + + + def normalize_review_items( + review_documents: Mapping[str, Any], + release_lock: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Preserve every raw review occurrence in exactly one normalized partition.""" + + stage_map = { + "stage1_part1_soft_gate_handoff": ("P1", "S2A-P1-HANDOFF-FLAT-V1"), + "stage1_part2_review_handoff": ("P2", "S2A-P2-HANDOFF-FLAT-V1"), + "stage1_part3_review_handoff": ("P3", "S2A-P3-HANDOFF-WRAPPED-V1"), + "stage1_part4_review_handoff": ("P4", "S2A-P4-HANDOFF-WRAPPED-V1"), + } + release_lock = release_lock or {} + mapping_decision = _adapter_decision(release_lock, "S2-REVIEW-MAP-V1") + mapping_rows = mapping_decision.get("mappings", []) if isinstance(mapping_decision, dict) else [] + closed_mappings: dict[tuple[str, str, str], Mapping[str, Any]] = {} + for mapping in mapping_rows: + if not isinstance(mapping, dict): + continue + key = ( + str(mapping.get("source_stage", "ANY")), + str(mapping.get("source_field_kind")), + str(mapping.get("source_value")), + ) + if key in closed_mappings: + raise IngressError("REVIEW_MAPPING_DUPLICATE", f"duplicate release review mapping: {key}") + closed_mappings[key] = mapping + adapter_issues: list[dict[str, Any]] = [] + if not closed_mappings: + adapter_issues.append(_issue("REVIEW_MAPPING_CONTRACT_MISSING", impact_scope="REVIEW_ITEM")) + collected: list[tuple[str, Any, str]] = [] + adapter_ids_seen: set[str] = set() + for logical_id in sorted(review_documents): + source_stage, adapter_id = stage_map.get(logical_id, (None, None)) + if source_stage is None or adapter_id is None: + continue + adapter_ids_seen.add(adapter_id) + decision = _adapter_decision(release_lock, adapter_id) + if decision is None: + adapter_issues.append(_issue("HANDOFF_ADAPTER_CONTRACT_MISSING", impact_scope="REVIEW_ITEM", source_refs=[logical_id])) + continue + document = review_documents[logical_id] + wrapper_pointer = decision.get("wrapper_json_pointer") + wrapper_found, wrapper = _json_pointer_value(document, wrapper_pointer) + if not wrapper_found or not isinstance(wrapper, dict): + code = "P2_HANDOFF_NOT_FLAT_OBJECT" if source_stage == "P2" else f"{source_stage}_WRAPPER_MISSING" + adapter_issues.append(_issue(code, impact_scope="REVIEW_ITEM", source_refs=[logical_id])) + continue + expected_version = decision.get("schema_version") + if wrapper.get("schema_version") != expected_version: + adapter_issues.append( + _issue( + f"{source_stage}_HANDOFF_SCHEMA_VERSION_MISMATCH", + impact_scope="REVIEW_ITEM", + source_refs=[logical_id], + ) + ) + items_found, raw_items = _json_pointer_value(wrapper, decision.get("review_items_json_pointer")) + if not items_found or not isinstance(raw_items, list): + adapter_issues.append( + _issue( + f"{source_stage}_REVIEW_ITEMS_SHAPE", + impact_scope="REVIEW_ITEM", + source_refs=[logical_id], + ) + ) + raw_items = [] + if decision.get("count_field_required") is True: + declared_count = wrapper.get("review_item_count") + if not isinstance(declared_count, int) or declared_count != len(raw_items): + adapter_issues.append( + _issue("REVIEW_CONSERVATION_FAILED", impact_scope="REVIEW_ITEM", source_refs=[logical_id]) + ) + for raw_item in raw_items: + raw_hash = canonical_digest(raw_item) + collected.append((logical_id, raw_item, raw_hash)) + grouped: dict[tuple[str, str], list[Any]] = defaultdict(list) + for logical_id, raw_item, raw_hash in collected: + grouped[(logical_id, raw_hash)].append(raw_item) + raw_occurrences: list[dict[str, Any]] = [] + normalized: list[dict[str, Any]] = [] + partition_counts = {key: 0 for key in ("SUPPORTED", "CONDITIONAL", "UNRESOLVED", "EXCLUDED", "UNMAPPED")} + for (logical_id, raw_hash), items in sorted(grouped.items()): + for duplicate_index, raw_item in enumerate(items): + raw_status = raw_item.get("status") if isinstance(raw_item, dict) else None + raw_severity = raw_item.get("severity") if isinstance(raw_item, dict) else None + source_stage, adapter_id = stage_map.get(logical_id, ("P1", "S2A-P1-HANDOFF-FLAT-V1")) + field_kind = "REVIEW_ITEM_STATUS" if raw_status is not None else "REVIEW_ITEM_SEVERITY" + source_value = str(raw_status if raw_status is not None else raw_severity) + mapping = closed_mappings.get((source_stage, field_kind, source_value)) or closed_mappings.get( + ("ANY", field_kind, source_value) + ) + partition = str(mapping.get("normalized_partition")) if mapping is not None else "UNMAPPED" + if partition not in partition_counts: + raise IngressError("REVIEW_MAPPING_PARTITION_INVALID", f"release mapping produced {partition}") + upstream_id = raw_item.get("review_id") if isinstance(raw_item, dict) else None + if upstream_id: + review_id = str(upstream_id) + minted_flag = False + else: + minted = mint_stage2_id( + "review_occurrence", + [logical_id, raw_hash, duplicate_index], + prefix="REV", + ) + review_id = minted["id"] + minted_flag = True + raw_occurrence = { + "source_stage": source_stage, + "logical_input_id": logical_id, + "canonical_raw_item_sha256": raw_hash, + "duplicate_occurrence_index": duplicate_index, + } + raw_occurrences.append(raw_occurrence) + reason_codes = ["UNMAPPED_REVIEW_STATUS"] if partition == "UNMAPPED" else [] + row = { + "review_key": { + "value": review_id, + "minted": minted_flag, + "source_stage": source_stage, + "logical_input_id": logical_id, + "canonical_raw_item_sha256": raw_hash, + "duplicate_occurrence_index": duplicate_index, + "upstream_review_id": str(upstream_id) if upstream_id is not None else None, + }, + "partition": partition, + "source_status_raw": str(raw_status) if raw_status is not None else None, + "source_severity_raw": str(raw_severity) if raw_severity is not None else None, + "source_item_raw_sha256": raw_hash, + "mapping_id": str(mapping.get("mapping_id")) if mapping is not None else f"S2-REVIEW-MAP-V1:{source_stage}:UNMAPPED", + "impact_scope": "REVIEW_ITEM", + "scope_refs": [review_id], + "source_contract_row_refs": [logical_id], + "reason_codes": reason_codes, + "downstream_allowed_actions": ["CARRY_FORWARD_TO_LAWYER_REVIEW"], + "_adapter_id": adapter_id, + } + normalized.append(row) + partition_counts[partition] += 1 + raw_counter = Counter((logical_id, raw_hash) for logical_id, _, raw_hash in collected) + normalized_counter = Counter( + (row["review_key"]["logical_input_id"], row["review_key"]["canonical_raw_item_sha256"]) + for row in normalized + ) + adapter_ids = sorted({row.pop("_adapter_id") for row in normalized} | adapter_ids_seen) + adapter_conservation_pass = not any( + item["issue_code"] == "REVIEW_CONSERVATION_FAILED" for item in adapter_issues + ) + return { + "schema_version": "stage2_review_normalization_receipt.v1", + "adapter_ids": adapter_ids or ["S2A-P1-HANDOFF-FLAT-V1"], + "mapping_table_version": "S2-REVIEW-MAP-V1", + "raw_occurrences": sorted( + raw_occurrences, + key=lambda row: ( + row["source_stage"], + row["logical_input_id"], + row["canonical_raw_item_sha256"], + row["duplicate_occurrence_index"], + ), + ), + "normalized_occurrences": sorted( + normalized, + key=lambda row: row["review_key"]["value"], + ), + "partition_counts": partition_counts, + "conservation_status": "PASS" if raw_counter == normalized_counter and adapter_conservation_pass else "FAIL", + "_issues": adapter_issues, + } + + + def check_conservation( + documents: Mapping[str, Any], + *, + signal_all: Mapping[str, Any] | None = None, + normalized_reviews: Mapping[str, Any] | None = None, + source_snapshots: Mapping[str, Snapshot] | None = None, + ) -> dict[str, Any]: + """Independently compute core set, cardinality, and multiset invariants.""" + + checks: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + source_snapshots = source_snapshots or {} + + def add_check( + check_id: str, + passed: bool | None, + left: Sequence[Any] | Counter[Any] | None, + right: Sequence[Any] | Counter[Any] | None, + *, + issue_code: str, + impact_scope: str, + source_refs: Sequence[str], + details: Mapping[str, Any] | None = None, + ) -> None: + left_counter = left if isinstance(left, Counter) else Counter(left or []) + right_counter = right if isinstance(right, Counter) else Counter(right or []) + row: dict[str, Any] = { + "check_id": check_id, + "status": "UNEVALUABLE" if passed is None else "PASS" if passed else "FAIL", + "left_count": sum(left_counter.values()) if left is not None else None, + "right_count": sum(right_counter.values()) if right is not None else None, + "left_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in left_counter.items())) if left is not None else None, + "right_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in right_counter.items())) if right is not None else None, + } + if details: + row.update(details) + checks.append(row) + if passed is False: + issues.append(_issue(issue_code, impact_scope=impact_scope, source_refs=source_refs)) + + bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + bo_ids = [str(row["BO_ID"]) for row in bo_rows if isinstance(row, dict) and row.get("BO_ID") is not None] + source_bo_ids = [ + str(row["source_bo_id"]) + for row in ledger_rows + if isinstance(row, dict) and row.get("source_bo_id") is not None + ] + missing_bo_id_rows = [index for index, row in enumerate(bo_rows) if not isinstance(row, dict) or row.get("BO_ID") is None] + missing_source_bo_rows = [ + index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("source_bo_id") is None + ] + bo_pass = ( + not missing_bo_id_rows + and not missing_source_bo_rows + and Counter(bo_ids) == Counter(source_bo_ids) + ) + add_check( + "BO_FACT_MULTISET", + bo_pass, + bo_ids, + source_bo_ids, + issue_code="BO_FACT_CONSERVATION_FAILED", + impact_scope="FACT", + source_refs=["bo", "fact_ledger_base"], + details={ + "missing_bo_id_rows": missing_bo_id_rows, + "missing_source_bo_id_rows": missing_source_bo_rows, + "duplicate_bo_ids": sorted(key for key, count in Counter(bo_ids).items() if count > 1), + "dangling_source_bo_ids": sorted(set(source_bo_ids) - set(bo_ids)), + }, + ) + missing_fact_id_rows = [ + index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("fact_id") is None + ] + fact_ids = [str(row["fact_id"]) for row in ledger_rows if isinstance(row, dict) and row.get("fact_id") is not None] + expected_fact_ids = [f"F-{index:03d}" for index in range(1, len(ledger_rows) + 1)] + fact_pass = not missing_fact_id_rows and fact_ids == expected_fact_ids and len(fact_ids) == len(set(fact_ids)) + add_check( + "FACT_ID_SEQUENCE", + fact_pass, + fact_ids, + expected_fact_ids, + issue_code="FACT_ID_CONSERVATION_FAILED", + impact_scope="FACT", + source_refs=["fact_ledger_base"], + details={"missing_fact_id_rows": missing_fact_id_rows, "observed": fact_ids}, + ) + extension_missing = [ + index + for index, row in enumerate(ledger_rows) + if not isinstance(row, dict) + or not isinstance(row.get("domain_effects"), dict) + or not isinstance(row.get("calculation_requests"), list) + ] + add_check( + "CURRENT_V8_LEDGER_EXTENSIONS", + not extension_missing, + list(range(len(ledger_rows))), + [index for index in range(len(ledger_rows)) if index not in extension_missing], + issue_code="CURRENT_V8_LEDGER_EXTENSION_MISSING", + impact_scope="FACT", + source_refs=["fact_ledger_base"], + details={"missing_row_indices": extension_missing}, + ) + les_rows = _array_rows( + documents.get("legal_effect_structures"), + ("structures", "structure_records", "legal_effect_structures", "rows", "items"), + ) + dangling_les: list[str] = [] + les_ids: list[str] = [] + for row in les_rows: + if not isinstance(row, dict): + continue + structure_id = row.get("structure_id", row.get("legal_effect_structure_id")) + if structure_id is not None: + les_ids.append(str(structure_id)) + refs = row.get("source_bo_ids", []) + if isinstance(refs, list): + dangling_les.extend(str(ref) for ref in refs if ref not in set(bo_ids)) + duplicate_les_ids = sorted(key for key, count in Counter(les_ids).items() if count > 1) + les_pass = not dangling_les and not duplicate_les_ids and len(les_ids) == len(les_rows) + add_check( + "LES_BO_JOIN", + les_pass, + [str(row.get("structure_id", row.get("legal_effect_structure_id"))) for row in les_rows if isinstance(row, dict)], + les_ids, + issue_code="LES_BO_JOIN_FAILED", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures", "bo"], + details={"dangling_refs": sorted(dangling_les), "duplicate_structure_ids": duplicate_les_ids}, + ) + declared_les_count = None + les_document = documents.get("legal_effect_structures") + if isinstance(les_document, dict): + for key in ("declared_structure_count", "structure_count", "record_count"): + if isinstance(les_document.get(key), int): + declared_les_count = int(les_document[key]) + break + declared_les_pass = None if declared_les_count is None else declared_les_count == len(les_rows) + add_check( + "LES_DECLARED_ACTUAL_COUNT", + declared_les_pass, + [None] * declared_les_count if declared_les_count is not None else None, + [None] * len(les_rows), + issue_code="LES_DECLARED_COUNT_MISMATCH", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures"], + ) + actual_domain_index: dict[str, list[str]] = defaultdict(list) + actual_bo_index: dict[str, list[str]] = defaultdict(list) + ledger_structure_refs: list[tuple[str, str, str]] = [] + ledger_type_refs: list[tuple[str, str, str]] = [] + actual_structure_refs: list[tuple[str, str, str]] = [] + actual_type_refs: list[tuple[str, str, str]] = [] + route_count_errors: list[str] = [] + for row in les_rows: + if not isinstance(row, dict): + continue + structure_id = str(row.get("structure_id", row.get("legal_effect_structure_id", "MISSING"))) + domain_id = str(row.get("domain_id", "MISSING")) + type_id = str(row.get("type_id", row.get("type", "MISSING"))) + actual_domain_index[domain_id].append(structure_id) + source_ids = row.get("source_bo_ids", []) + if isinstance(source_ids, list): + for bo_id in source_ids: + actual_bo_index[str(bo_id)].append(structure_id) + actual_structure_refs.append((str(bo_id), domain_id, structure_id)) + actual_type_refs.append((str(bo_id), domain_id, type_id)) + routes = row.get("routes", []) + if isinstance(routes, list) and row.get("route_count", len(routes)) != len(routes): + route_count_errors.append(structure_id) + for row in ledger_rows: + if not isinstance(row, dict): + continue + bo_id = str(row.get("source_bo_id", "MISSING")) + effects = row.get("domain_effects", {}) + if not isinstance(effects, dict): + continue + for domain_id, effect in effects.items(): + if not isinstance(effect, dict): + continue + for structure_id in effect.get("structure_ids", []) if isinstance(effect.get("structure_ids"), list) else []: + ledger_structure_refs.append((bo_id, str(domain_id), str(structure_id))) + for type_id in effect.get("type_ids", []) if isinstance(effect.get("type_ids"), list) else []: + ledger_type_refs.append((bo_id, str(domain_id), str(type_id))) + structure_index = les_document.get("structure_index", {}) if isinstance(les_document, dict) else {} + index_present = isinstance(structure_index, dict) and bool(structure_index) + index_ok = True + if index_present: + declared_by_domain = structure_index.get("by_domain_id", {}) + declared_by_bo = structure_index.get("by_bo_id", {}) + index_ok = ( + isinstance(declared_by_domain, dict) + and isinstance(declared_by_bo, dict) + and {str(key): Counter(map(str, value)) for key, value in declared_by_domain.items() if isinstance(value, list)} + == {key: Counter(value) for key, value in actual_domain_index.items()} + and {str(key): Counter(map(str, value)) for key, value in declared_by_bo.items() if isinstance(value, list)} + == {key: Counter(value) for key, value in actual_bo_index.items()} + ) + reverse_ok = ( + (not ledger_structure_refs or Counter(ledger_structure_refs) == Counter(actual_structure_refs)) + and (not ledger_type_refs or Counter(ledger_type_refs) == Counter(actual_type_refs)) + and not route_count_errors + and index_ok + ) + add_check( + "LES_REVERSE_INDEX", + reverse_ok, + ledger_structure_refs + ledger_type_refs, + actual_structure_refs + actual_type_refs, + issue_code="LES_REVERSE_INDEX_MISMATCH", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures", "fact_ledger_base"], + details={"index_present": index_present, "route_count_errors": route_count_errors}, + ) + evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) + event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) + evidence_ids = [ + str(row.get("evidence_id", row.get("id"))) + for row in evidence_rows + if isinstance(row, dict) and (row.get("evidence_id") is not None or row.get("id") is not None) + ] + event_ids = [ + str(row.get("event_id", row.get("id"))) + for row in event_rows + if isinstance(row, dict) and (row.get("event_id") is not None or row.get("id") is not None) + ] + fact_evidence_refs: list[str] = [] + fact_event_refs: list[str] = [] + event_evidence_refs: list[str] = [] + for row in ledger_rows: + if not isinstance(row, dict): + continue + evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) + event_values = row.get("event_refs", row.get("event_ids", [])) + if isinstance(evidence_values, list): + fact_evidence_refs.extend(str(ref) for ref in evidence_values) + if isinstance(event_values, list): + fact_event_refs.extend(str(ref) for ref in event_values) + for row in event_rows: + if not isinstance(row, dict): + continue + evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) + if isinstance(evidence_values, list): + event_evidence_refs.extend(str(ref) for ref in evidence_values) + evidence_failures = sorted( + set(fact_evidence_refs + event_evidence_refs) - set(evidence_ids) + ) + duplicate_evidence_ids = sorted(key for key, count in Counter(evidence_ids).items() if count > 1) + evidence_pass = not evidence_failures and not duplicate_evidence_ids + add_check( + "EVIDENCE_REFERENCE_CONSERVATION", + evidence_pass, + fact_evidence_refs + event_evidence_refs, + evidence_ids, + issue_code="EVIDENCE_REFERENCE_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_indexed", "fact_ledger_base", "evidence_event_candidates"], + details={"dangling_refs": evidence_failures, "duplicate_evidence_ids": duplicate_evidence_ids}, + ) + event_failures = sorted(set(fact_event_refs) - set(event_ids)) + duplicate_event_ids = sorted(key for key, count in Counter(event_ids).items() if count > 1) + event_pass = not event_failures and not duplicate_event_ids + add_check( + "EVENT_REFERENCE_CONSERVATION", + event_pass, + fact_event_refs, + event_ids, + issue_code="EVENT_REFERENCE_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_event_candidates", "fact_ledger_base"], + details={"dangling_refs": event_failures, "duplicate_event_ids": duplicate_event_ids}, + ) + disposition_rows = [row.get("disposition") for row in event_rows if isinstance(row, dict) and "disposition" in row] + b2_gate = documents.get("b2_event_candidates_gate") + declared_dispositions = None + if isinstance(b2_gate, dict): + declared_dispositions = b2_gate.get("event_disposition_counts") + if declared_dispositions is None and isinstance(b2_gate.get("summary"), dict): + declared_dispositions = b2_gate["summary"].get("event_disposition_counts") + if isinstance(declared_dispositions, dict): + disposition_expected = Counter( + {str(key): int(value) for key, value in declared_dispositions.items() if isinstance(value, int)} + ) + disposition_actual = Counter(str(value) for value in disposition_rows) + disposition_pass: bool | None = disposition_actual == disposition_expected + elif disposition_rows: + disposition_expected = Counter(str(value) for value in disposition_rows) + disposition_actual = Counter(str(value) for value in disposition_rows) + disposition_pass = all(isinstance(value, str) and value for value in disposition_rows) + else: + disposition_expected = Counter() + disposition_actual = Counter() + disposition_pass = None + add_check( + "EVENT_DISPOSITION_CONSERVATION", + disposition_pass, + disposition_actual, + disposition_expected, + issue_code="EVENT_DISPOSITION_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_event_candidates", "b2_event_candidates_gate"], + ) + writer_report = documents.get("fact_ledger_writer_report") + if isinstance(writer_report, dict): + observed_domain_coverage = Counter( + str(domain_id) + for row in ledger_rows + if isinstance(row, dict) and isinstance(row.get("domain_effects"), dict) + for domain_id in row["domain_effects"] + ) + declared_domain_coverage = Counter( + {str(key): int(value) for key, value in writer_report.get("domain_effect_coverage", {}).items() if isinstance(value, int)} + ) + observed_readiness = Counter( + str(request.get("operand_state")) + for row in ledger_rows + if isinstance(row, dict) and isinstance(row.get("calculation_requests"), list) + for request in row["calculation_requests"] + if isinstance(request, dict) + ) + declared_readiness = Counter( + {str(key): int(value) for key, value in writer_report.get("calculation_readiness", {}).items() if isinstance(value, int)} + ) + ledger_snapshot = source_snapshots.get("fact_ledger_base") + final_hash = writer_report.get("final_sha256") + writer_pass = ( + writer_report.get("row_count") == len(ledger_rows) + and declared_domain_coverage == observed_domain_coverage + and declared_readiness == observed_readiness + and (ledger_snapshot is None or final_hash == ledger_snapshot.raw_sha256) + ) + add_check( + "FACT_LEDGER_WRITER_REPORT_CONNECTION", + writer_pass, + [len(ledger_rows), observed_domain_coverage, observed_readiness, ledger_snapshot.raw_sha256 if ledger_snapshot else None], + [writer_report.get("row_count"), declared_domain_coverage, declared_readiness, final_hash], + issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", + impact_scope="FACT", + source_refs=["fact_ledger_base", "fact_ledger_writer_report"], + ) + else: + add_check( + "FACT_LEDGER_WRITER_REPORT_CONNECTION", + None, + None, + None, + issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", + impact_scope="FACT", + source_refs=["fact_ledger_base", "fact_ledger_writer_report"], + ) + if signal_all is not None: + file_pass = bool(signal_all.get("file_conservation_pass")) + record_pass = bool(signal_all.get("record_conservation_pass")) + checks.append({"check_id": "SIGNAL_FILE_ROW_CONSERVATION", "status": "PASS" if file_pass else "FAIL"}) + checks.append({"check_id": "SIGNAL_RECORD_OCCURRENCE_CONSERVATION", "status": "PASS" if record_pass else "FAIL"}) + issues.extend(signal_all.get("issues", [])) + if not file_pass: + issues.append(_issue("SIGNAL_FILE_CONSERVATION_FAILED", impact_scope="SIGNAL")) + if not record_pass: + issues.append(_issue("SIGNAL_RECORD_CONSERVATION_FAILED", impact_scope="SIGNAL")) + if normalized_reviews is not None: + review_pass = normalized_reviews.get("conservation_status") == "PASS" + checks.append({"check_id": "REVIEW_OCCURRENCE_CONSERVATION", "status": "PASS" if review_pass else "FAIL"}) + if not review_pass: + issues.append(_issue("REVIEW_CONSERVATION_FAILED", impact_scope="REVIEW_ITEM")) + issues.extend(normalized_reviews.get("_issues", [])) + return {"checks": checks, "issues": issues, "passed": not any(check["status"] == "FAIL" for check in checks)} + + + def _source_ref( + logical_id: str, + pointer: str, + raw_value: Any = _RAW_VALUE_UNSET, + *, + stage1_id: str | None = None, + ) -> dict[str, Any]: + """Build a truthful RFC 6901 provenance row without pointer narrowing.""" + + row: dict[str, Any] = { + "logical_artifact_id": logical_id, + "json_pointer": pointer, + "raw_value_sha256": canonical_digest( + [logical_id, pointer] + if raw_value is _RAW_VALUE_UNSET + else raw_value + ), + "source_contract_row_ref": logical_id, + } + if stage1_id is not None: + row["stage1_id"] = stage1_id + return row + + + def _dedupe_source_refs(rows: Iterable[Mapping[str, Any]]) -> list[dict[str, Any]]: + by_digest = {canonical_digest(dict(row)): dict(row) for row in rows} + return [by_digest[key] for key in sorted(by_digest)] + + + def _coerce_source_ref(value: Mapping[str, Any] | str) -> dict[str, Any]: + if isinstance(value, dict) and { + "logical_artifact_id", + "json_pointer", + "raw_value_sha256", + }.issubset(value): + return dict(value) + text = str(value) + if "#" in text: + logical_id, pointer = text.split("#", 1) + else: + logical_id, pointer = text, "" + if pointer and not pointer.startswith("/"): + pointer = "/" + pointer + return _source_ref(logical_id or "unknown", pointer, text) + + + def _artifact_header( + schema_id: str = CONTEXT_SCHEMA_ID, + *, + schema_sha256: str = "0" * 64, + run_id: str = "STRUCTURAL-FIXTURE", + input_set_digest: str = "0" * 64, + stage2_release_digest: str = "0" * 64, + algorithm_digest: str = "0" * 64, + release_class: str = "DEV_FIXTURE_RELEASE", + ) -> dict[str, Any]: + return { + "schema_id": schema_id, + "schema_sha256": schema_sha256, + "schema_version": "stage2_s2_00_context.v2", + "producer_id": "S2_00", + "run_id": run_id, + "input_set_digest": input_set_digest, + "stage2_release_digest": stage2_release_digest, + "algorithm_digest": algorithm_digest, + "release_class": release_class, + } + + + def _derivation( + namespace: str, + input_refs: Sequence[str], + payload: Any, + ) -> dict[str, Any]: + sorted_refs = sorted(set(str(item) for item in input_refs)) or ["S2_00:EMPTY_INPUT_SET"] + minted = mint_stage2_id(namespace, [sorted_refs, payload], prefix="DRV") + return { + "derivation_id": minted["id"], + "algorithm_version": ALGORITHM_VERSION, + "sorted_input_refs": sorted_refs, + "mint_input_sha256": minted["mint_input_sha256"], + } + + + def build_case_context( + documents: Mapping[str, Any], + *, + run_binding_digest: str, + issues: Sequence[Mapping[str, Any]] = (), + artifact_header: Mapping[str, Any] | None = None, + signal_all: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + routing = documents.get("domain_activation_manifest") + routing_payload = _activation_payload(routing) if isinstance(routing, dict) else {} + active = routing_payload.get("active_domain_ids", []) + expected = routing_payload.get("expected_runnable_domain_ids", []) + bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + les_rows = _array_rows(documents.get("legal_effect_structures"), ("structures", "structure_records", "legal_effect_structures", "rows", "items")) + bo_ids = sorted( + str(row["BO_ID"]) + for row in bo_rows + if isinstance(row, dict) and row.get("BO_ID") is not None + ) + structure_ids = sorted( + str(row.get("structure_id", row.get("legal_effect_structure_id"))) + for row in les_rows + if isinstance(row, dict) and (row.get("structure_id") is not None or row.get("legal_effect_structure_id") is not None) + ) + structures_by_bo: dict[str, list[str]] = defaultdict(list) + for row in les_rows: + if not isinstance(row, dict): + continue + structure_id = row.get("structure_id", row.get("legal_effect_structure_id")) + for bo_id in row.get("source_bo_ids", []) if isinstance(row.get("source_bo_ids"), list) else []: + if structure_id is not None: + structures_by_bo[str(bo_id)].append(str(structure_id)) + fact_contexts: list[dict[str, Any]] = [] + for index, row in enumerate(ledger_rows): + if not isinstance(row, dict): + continue + if row.get("fact_id") is None or row.get("source_bo_id") is None: + continue + fact_id = str(row["fact_id"]) + source_bo_id = str(row.get("source_bo_id", "MISSING")) + evidence_refs = row.get("evidence_refs", row.get("evidence_ids", [])) + event_refs = row.get("event_refs", row.get("event_ids", [])) + calculation_requests = row.get("calculation_requests", []) + law_version_refs = row.get("law_version_refs", []) + signal_occurrence_refs = sorted( + str(occurrence.get("occurrence_ref")) + for occurrence in (signal_all or {}).get("record_occurrences", []) + if occurrence.get("disposition") == "USED" + and ( + f"fact_id:{fact_id}" in occurrence.get("binding_refs", []) + or f"source_bo_id:{source_bo_id}" in occurrence.get("binding_refs", []) + or f"bo_id:{source_bo_id}" in occurrence.get("binding_refs", []) + ) + ) + fact_contexts.append( + { + "fact_id": fact_id, + "source_bo_id": source_bo_id, + "party_context_ids": [], + "object_context_ids": [], + "structure_refs": sorted(set(structures_by_bo.get(source_bo_id, []))), + "signal_occurrence_refs": signal_occurrence_refs, + "calculation_request_refs": sorted( + canonical_digest(value) + for value in (calculation_requests if isinstance(calculation_requests, list) else []) + ), + "law_version_refs": sorted(str(value) for value in law_version_refs) if isinstance(law_version_refs, list) else [], + "evidence_refs": sorted(str(value) for value in evidence_refs) if isinstance(evidence_refs, list) else [], + "event_refs": sorted(str(value) for value in event_refs) if isinstance(event_refs, list) else [], + "review_keys": [], + "source_refs": [_source_ref("fact_ledger_base", f"/rows/{index}", row, stage1_id=fact_id)], + } + ) + source_refs = _dedupe_source_refs( + [ + _source_ref("bo", "", documents.get("bo")), + _source_ref("fact_ledger_base", "", documents.get("fact_ledger_base")), + _source_ref("legal_effect_structures", "", documents.get("legal_effect_structures")), + _source_ref("domain_activation_manifest", "", documents.get("domain_activation_manifest")), + _source_ref("client_goal", "", documents.get("client_goal")), + ] + ) + minted = mint_stage2_id( + "case_context", + [run_binding_digest, [row["fact_id"] for row in fact_contexts], bo_ids, structure_ids], + prefix="CC", + ) + unavailable_scope_refs = sorted( + { + str(scope_ref) + for item in issues + if item.get("issue_code") + for scope_ref in item.get("scope_refs", []) + } + ) + return { + "artifact_header": dict(artifact_header or _artifact_header()), + "case_context_id": minted["id"], + "facts": sorted(fact_contexts, key=lambda row: row["fact_id"]), + "bo_ids": bo_ids, + "structure_ids": structure_ids, + "active_domain_ids": sorted(set(str(value) for value in active)) if isinstance(active, list) else [], + "expected_runnable_domain_ids": sorted(set(str(value) for value in expected)) if isinstance(expected, list) else [], + "unavailable_scope_refs": unavailable_scope_refs, + "source_refs": source_refs, + "derivation": _derivation("case_context_derivation", [f"fact:{row['fact_id']}" for row in fact_contexts], minted["mint_input_sha256"]), + } + + + def build_evidence_inventory( + documents: Mapping[str, Any], + *, + run_binding_digest: str, + artifact_header: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) + event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + event_by_evidence: dict[str, set[str]] = defaultdict(set) + for index, row in enumerate(event_rows): + if not isinstance(row, dict): + continue + event_id = str(row.get("event_id", f"EVENT-OCCURRENCE-{index}")) + refs = row.get("evidence_refs", row.get("evidence_ids", [])) + if isinstance(refs, list): + for ref in refs: + event_by_evidence[str(ref)].add(event_id) + facts_by_evidence: dict[str, set[str]] = defaultdict(set) + for row in ledger_rows: + if not isinstance(row, dict) or row.get("fact_id") is None: + continue + refs = row.get("evidence_refs", row.get("evidence_ids", [])) + if isinstance(refs, list): + for ref in refs: + facts_by_evidence[str(ref)].add(str(row["fact_id"])) + items: list[dict[str, Any]] = [] + for index, row in enumerate(evidence_rows): + if not isinstance(row, dict): + continue + evidence_id = str(row.get("evidence_id", row.get("id", f"EVIDENCE-OCCURRENCE-{index}"))) + items.append( + { + "evidence_id": evidence_id, + "event_ids": sorted(event_by_evidence.get(evidence_id, set())), + "linked_fact_ids": sorted(facts_by_evidence.get(evidence_id, set())), + "technical_disposition": "AVAILABLE", + "source_refs": [_source_ref("evidence_indexed", f"/items/{index}", row, stage1_id=evidence_id)], + } + ) + source_refs = _dedupe_source_refs( + [ + _source_ref("evidence_indexed", "", documents.get("evidence_indexed")), + _source_ref("evidence_event_candidates", "", documents.get("evidence_event_candidates")), + _source_ref("fact_ledger_base", "", documents.get("fact_ledger_base")), + ] + ) + return { + "artifact_header": dict(artifact_header or _artifact_header()), + "items": sorted(items, key=lambda row: row["evidence_id"]), + "source_refs": source_refs, + "derivation": _derivation("evidence_inventory_derivation", [f"evidence:{row['evidence_id']}" for row in items], run_binding_digest), + } + + + def _candidate_values(row: Mapping[str, Any], keys: Sequence[str]) -> list[tuple[str, Any]]: + result: list[tuple[str, Any]] = [] + for key in keys: + value = row.get(key) + if value is None: + continue + if isinstance(value, list): + result.extend((f"/{key}/{index}", item) for index, item in enumerate(value)) + else: + result.append((f"/{key}", value)) + return result + + + def build_object_registry( + documents: Mapping[str, Any], + *, + run_binding_digest: str, + artifact_header: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) + objects: list[dict[str, Any]] = [] + collision_registry: dict[str, str] = {} + for index, row in enumerate(bo_rows): + if not isinstance(row, dict): + continue + explicit_bo = row.get("BO_ID") + for pointer, value in _candidate_values(row, ("object", "objects", "asset", "subject_matter")): + stage1_object_id = value.get("object_id") if isinstance(value, dict) else None + context = mint_context_occurrence_id( + "object", + "bo", + f"/{index}{pointer}", + value, + [str(explicit_bo)] if explicit_bo is not None else [], + stage1_canonical_id=str(stage1_object_id) if stage1_object_id is not None else None, + collision_registry=collision_registry, + ) + objects.append(context) + return { + "artifact_header": dict(artifact_header or _artifact_header()), + "objects": sorted(objects, key=lambda row: row["object_context_id"]), + "source_refs": [_source_ref("bo", "", documents.get("bo"))], + "derivation": _derivation("object_registry_derivation", [row["object_context_id"] for row in objects], run_binding_digest), + } + + + def build_party_and_title_context( + documents: Mapping[str, Any], + *, + run_binding_digest: str, + artifact_header: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) + parties: list[dict[str, Any]] = [] + collision_registry: dict[str, str] = {} + for index, row in enumerate(bo_rows): + if not isinstance(row, dict): + continue + explicit_bo = row.get("BO_ID") + for pointer, value in _candidate_values(row, ("party", "parties", "creditor", "debtor", "counterparty")): + stage1_party_id = value.get("party_id") if isinstance(value, dict) else None + identity = mint_context_occurrence_id( + "party", + "bo", + f"/{index}{pointer}", + value, + [str(explicit_bo)] if explicit_bo is not None else [], + stage1_canonical_id=str(stage1_party_id) if stage1_party_id is not None else None, + collision_registry=collision_registry, + ) + source_refs = identity["source_refs"] + parties.append( + { + "party_identity": identity, + "party_title": value.get("party_title") if isinstance(value, dict) else None, + "defendant_role": value.get("defendant_role") if isinstance(value, dict) else None, + "liability_context": value.get("liability_context") if isinstance(value, dict) else None, + "client_instruction": value.get("client_instruction") if isinstance(value, dict) else None, + "recovery_information": value.get("recovery_information") if isinstance(value, dict) else None, + "source_refs": source_refs, + } + ) + return { + "artifact_header": dict(artifact_header or _artifact_header()), + "parties": sorted(parties, key=lambda row: row["party_identity"]["party_context_id"]), + "source_refs": [_source_ref("bo", "", documents.get("bo"))], + "derivation": _derivation( + "party_title_derivation", + [row["party_identity"]["party_context_id"] for row in parties], + run_binding_digest, + ), + } + + + def build_slot_crosswalk( + documents: Mapping[str, Any], + domain_configs: Mapping[str, Any] | None = None, + *, + run_binding_digest: str, + artifact_header: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + domain_configs = domain_configs or {} + rows: list[dict[str, Any]] = [] + for domain_id in sorted(domain_configs): + config = domain_configs[domain_id] + adapter_id = "S2A-DOMAIN-CONFIG-V2" if domain_id in V2_DOMAIN_IDS else "S2A-DOMAIN-CONFIG-V1" + if not isinstance(config, dict): + continue + for slot_kind, key in ( + ("ELEMENT", "element_slots"), + ("OPPOSING_FACT", "opposing_fact_slots"), + ("DEFENSE", "defense_map"), + ): + values = config.get(key, []) + if isinstance(values, dict): + values = [{"slot_id": item_key, "value": item_value} for item_key, item_value in values.items()] + if not isinstance(values, list): + continue + for index, value in enumerate(values): + rows.append( + { + "domain_id": domain_id, + "slot_ref": str(value.get("slot_id", value.get("id", f"{domain_id}:{key}:{index}"))) if isinstance(value, dict) else f"{domain_id}:{key}:{index}", + "slot_kind": slot_kind, + "fact_ids": [], + "evidence_refs": [], + "evaluation_status": "UNEVALUABLE", + "proposed_new_slot": False, + "source_refs": [_source_ref(f"domain_config:{domain_id}", f"/{key}/{index}", value)], + "_adapter_id": adapter_id, + } + ) + adapter_ids = sorted({row.pop("_adapter_id") for row in rows}) + return { + "artifact_header": dict(artifact_header or _artifact_header()), + "domain_config_adapter_ids": adapter_ids, + "rows": sorted(rows, key=lambda row: (row["domain_id"], row["slot_kind"], row["slot_ref"])), + "rebuttal_slot_synthesis_count": 0, + "source_refs": _dedupe_source_refs( + [_source_ref("fact_ledger_base", "", documents.get("fact_ledger_base"))] + + [_source_ref(f"domain_config:{key}", "", domain_configs[key]) for key in sorted(domain_configs)] + ), + "derivation": _derivation("slot_crosswalk_derivation", [row["slot_ref"] for row in rows], run_binding_digest), + } + + + def _tarjan_scc(nodes: Sequence[str], edges: Sequence[tuple[str, str]]) -> list[list[str]]: + adjacency: dict[str, list[str]] = {node: [] for node in nodes} + for source, target in edges: + adjacency.setdefault(source, []).append(target) + adjacency.setdefault(target, []) + for value in adjacency.values(): + value.sort() + index = 0 + stack: list[str] = [] + on_stack: set[str] = set() + indices: dict[str, int] = {} + lowlink: dict[str, int] = {} + components: list[list[str]] = [] + + def visit(node: str) -> None: + nonlocal index + indices[node] = index + lowlink[node] = index + index += 1 + stack.append(node) + on_stack.add(node) + for neighbor in adjacency[node]: + if neighbor not in indices: + visit(neighbor) + lowlink[node] = min(lowlink[node], lowlink[neighbor]) + elif neighbor in on_stack: + lowlink[node] = min(lowlink[node], indices[neighbor]) + if lowlink[node] == indices[node]: + component: list[str] = [] + while True: + member = stack.pop() + on_stack.remove(member) + component.append(member) + if member == node: + break + components.append(sorted(component)) + + for node in sorted(adjacency): + if node not in indices: + visit(node) + return sorted(components, key=lambda component: component[0]) + + + def compile_cluster_plan( + members: Sequence[Mapping[str, Any]], + relations: Sequence[Mapping[str, Any]], + *, + algorithm_version: str = ALGORITHM_VERSION, + artifact_header: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Compile claim-neutral clusters using only explicit hard relations.""" + + by_id = {str(row["member_id"]): dict(row) for row in members} + if len(by_id) != len(members): + raise IngressError("DUPLICATE_CLUSTER_MEMBER_ID", "cluster member IDs must be unique") + parent = {member_id: member_id for member_id in by_id} + + def find(value: str) -> str: + while parent[value] != value: + parent[value] = parent[parent[value]] + value = parent[value] + return value + + def union(left: str, right: str) -> None: + a, b = find(left), find(right) + if a != b: + parent[max(a, b)] = min(a, b) + + normalized_relations: list[dict[str, Any]] = [] + relation_key_map = { + "SAME_BO_ID": ("EXPLICIT_SOURCE_RELATION", "SAME_BO_ID", True), + "SOURCE_BO_ATTACHMENT": ("EXPLICIT_SOURCE_RELATION", "LES_SOURCE_BO_ID", True), + "SAME_EVIDENCE_REF": ("EXPLICIT_SOURCE_RELATION", "SAME_EVIDENCE_REF", True), + "SAME_EVENT_REF": ("EXPLICIT_SOURCE_RELATION", "SAME_EVENT_REF", True), + "EXPLICIT_CASE_RELATION": ("EXPLICIT_SOURCE_RELATION", "APPROVED_EXPLICIT_CASE_RELATION", True), + "claim_precondition": ("CANDIDATE_RELATION", "CLAIM_PRECONDITION_CANDIDATE", False), + "accessory_of": ("CANDIDATE_RELATION", "ACCESSORY_OF_CANDIDATE", False), + "incompatible_with": ("CANDIDATE_RELATION", "INCOMPATIBLE_WITH_CANDIDATE", False), + "EXPLICIT_DEPENDENCY": ("CANDIDATE_RELATION", "CLAIM_PRECONDITION_CANDIDATE", False), + } + for relation in relations: + source = str(relation.get("source_member_id")) + target = str(relation.get("target_member_id")) + kind = str(relation.get("relation_kind")) + if source not in by_id or target not in by_id: + raise IngressError("DANGLING_CLUSTER_RELATION", f"relation references unknown member: {source}->{target}") + explicit_source_refs = relation.get("source_refs") + if not isinstance(explicit_source_refs, list) or not explicit_source_refs: + raise IngressError("RELATION_SOURCE_REF_MISSING", "every cluster relation needs explicit source refs") + edge_class, relation_key, hard_join_allowed = relation_key_map.get( + kind, + ("CANDIDATE_RELATION", "CLAIM_PRECONDITION_CANDIDATE", False), + ) + relation_source_refs = _dedupe_source_refs( + _coerce_source_ref(item) for item in explicit_source_refs + ) + edge_mint = mint_stage2_id( + "cluster_edge", + [source, target, kind, relation_source_refs], + prefix="EDGE", + ) + normalized = { + "source_member_id": source, + "target_member_id": target, + "source_relation_kind": kind, + "edge_id": edge_mint["id"], + "edge_class": edge_class, + "relation_key": relation_key, + "hard_join_allowed": hard_join_allowed, + "source_refs": relation_source_refs, + } + normalized_relations.append(normalized) + if hard_join_allowed and kind in HARD_RELATION_KINDS: + union(source, target) + groups: dict[str, list[str]] = defaultdict(list) + for member_id in sorted(by_id): + groups[find(member_id)].append(member_id) + collision_registry: dict[str, str] = {} + cluster_build_rows: list[dict[str, Any]] = [] + member_to_cluster: dict[str, str] = {} + for member_ids in sorted((sorted(value) for value in groups.values()), key=lambda value: value[0]): + source_refs = _dedupe_source_refs( + _coerce_source_ref(ref) + for member_id in member_ids + for ref in by_id[member_id].get("source_refs", []) + ) + relation_keys = sorted( + canonical_digest(relation) + for relation in normalized_relations + if relation["source_member_id"] in member_ids and relation["target_member_id"] in member_ids + ) + minted = mint_stage2_id( + "cluster", + [sorted(member_ids), source_refs, relation_keys], + prefix="CL", + algorithm_version=algorithm_version, + collision_registry=collision_registry, + ) + executable = bool(source_refs) and all( + by_id[member_id].get("scope_technical_disposition", "AVAILABLE") != "UNAVAILABLE" + and not by_id[member_id].get("residual_review_only", False) + for member_id in member_ids + ) + cluster = { + "cluster_id": minted["id"], + "mint_input_sha256": minted["mint_input_sha256"], + "member_ids": member_ids, + "source_refs": source_refs, + "cluster_status": "EXECUTABLE" if executable else "RESIDUAL_REVIEW_ONLY", + } + cluster_build_rows.append(cluster) + for member_id in member_ids: + member_to_cluster[member_id] = minted["id"] + cluster_edges: list[dict[str, Any]] = [] + edge_pairs: set[tuple[str, str]] = set() + for relation in normalized_relations: + source_cluster = member_to_cluster[relation["source_member_id"]] + target_cluster = member_to_cluster[relation["target_member_id"]] + if source_cluster == target_cluster: + continue + if relation["source_relation_kind"] not in CANDIDATE_RELATION_KINDS: + continue + edge = { + "source_cluster_id": source_cluster, + "target_cluster_id": target_cluster, + "edge_id": relation["edge_id"], + "edge_class": relation["edge_class"], + "relation_key": relation["relation_key"], + "hard_join_allowed": False, + "source_refs": relation["source_refs"], + } + cluster_edges.append(edge) + edge_pairs.add((source_cluster, target_cluster)) + cluster_ids = [row["cluster_id"] for row in cluster_build_rows] + sccs = _tarjan_scc(cluster_ids, sorted(edge_pairs)) + scc_rows: list[dict[str, Any]] = [] + cluster_to_scc: dict[str, str] = {} + for members_in_scc in sccs: + minted = mint_stage2_id("cluster_scc", members_in_scc, prefix="SCC") + cycle = len(members_in_scc) > 1 or any(left == right == members_in_scc[0] for left, right in edge_pairs) + row = { + "scc_id": minted["id"], + "member_cluster_ids": members_in_scc, + "cycle_preserved": cycle, + "mint_input_sha256": minted["mint_input_sha256"], + } + scc_rows.append(row) + for cluster_id in members_in_scc: + cluster_to_scc[cluster_id] = minted["id"] + scheduling_edges = sorted( + { + (cluster_to_scc[left], cluster_to_scc[right]) + for left, right in edge_pairs + if cluster_to_scc[left] != cluster_to_scc[right] + } + ) + scc_nodes = sorted(row["scc_id"] for row in scc_rows) + indegree = {node: 0 for node in scc_nodes} + adjacency: dict[str, set[str]] = {node: set() for node in scc_nodes} + for source, target in scheduling_edges: + if target not in adjacency[source]: + adjacency[source].add(target) + indegree[target] += 1 + scheduling_waves: list[list[str]] = [] + ready = sorted(node for node, degree in indegree.items() if degree == 0) + visited: set[str] = set() + while ready: + scheduling_waves.append(ready) + next_ready: list[str] = [] + for node in ready: + visited.add(node) + for target in sorted(adjacency[node]): + indegree[target] -= 1 + if indegree[target] == 0: + next_ready.append(target) + ready = sorted(set(next_ready)) + if len(visited) != len(scc_nodes): + raise IngressError("SCC_SCHEDULING_DAG_INVALID", "SCC condensation graph must be acyclic") + executable_ids = sorted( + row["cluster_id"] for row in cluster_build_rows if row["cluster_status"] == "EXECUTABLE" + ) + residual_ids = sorted(set(cluster_ids) - set(executable_ids)) + clusters: list[dict[str, Any]] = [] + for build_row in cluster_build_rows: + cluster_id = build_row["cluster_id"] + member_ids = build_row["member_ids"] + cluster_members = [ + { + "member_ref": str(by_id[member_id].get("member_ref", member_id)), + "member_kind": str(by_id[member_id].get("member_kind", "FACT")), + "source_refs": _dedupe_source_refs( + _coerce_source_ref(ref) for ref in by_id[member_id].get("source_refs", []) + ), + } + for member_id in member_ids + ] + internal_edges: list[dict[str, Any]] = [] + for relation in normalized_relations: + if relation["source_member_id"] in member_ids and relation["target_member_id"] in member_ids: + internal_edges.append( + { + "edge_id": relation["edge_id"], + "from_ref": relation["source_member_id"], + "to_ref": relation["target_member_id"], + "edge_class": relation["edge_class"], + "relation_key": relation["relation_key"], + "hard_join_allowed": relation["hard_join_allowed"], + "source_refs": relation["source_refs"], + } + ) + for edge in cluster_edges: + if cluster_id in {edge["source_cluster_id"], edge["target_cluster_id"]}: + internal_edges.append( + { + "edge_id": edge["edge_id"], + "from_ref": edge["source_cluster_id"], + "to_ref": edge["target_cluster_id"], + "edge_class": edge["edge_class"], + "relation_key": edge["relation_key"], + "hard_join_allowed": False, + "source_refs": edge["source_refs"], + } + ) + cluster = { + "cluster_id": cluster_id, + "cluster_status": build_row["cluster_status"], + "members": sorted(cluster_members, key=lambda row: row["member_ref"]), + "edges": sorted(internal_edges, key=lambda row: row["edge_id"]), + "scc_ids": [cluster_to_scc[cluster_id]], + "slice_ref": f"cluster_slices/{cluster_id}.json" if cluster_id in executable_ids else None, + "bundle_cohort_id": None, + "source_refs": build_row["source_refs"], + "derivation": { + "derivation_id": cluster_id, + "algorithm_version": algorithm_version, + "sorted_input_refs": sorted(member_ids), + "mint_input_sha256": build_row["mint_input_sha256"], + }, + } + clusters.append(cluster) + top_source_refs = _dedupe_source_refs( + ref for row in clusters for ref in row["source_refs"] + ) if clusters else [] + return { + "artifact_header": dict(artifact_header or _artifact_header()), + "clusters": sorted(clusters, key=lambda row: row["cluster_id"]), + "executable_cluster_ids": executable_ids, + "residual_review_cluster_ids": residual_ids, + "scheduling_waves": scheduling_waves, + "source_refs": top_source_refs, + "derivation": _derivation( + "cluster_plan_derivation", + cluster_ids or ["NO_CLUSTER"], + [scheduling_edges, [[row["scc_id"], row["member_cluster_ids"]] for row in scc_rows]], + ), + } + + + def compile_cluster_slices( + cluster_plan: Mapping[str, Any], + context_artifacts: Mapping[str, Mapping[str, Any]], + ) -> dict[str, dict[str, Any]]: + """Create immutable minimal source-ref slices for every cluster.""" + + projection_sources: dict[str, dict[str, Any]] = defaultdict(dict) + case_context = context_artifacts.get("case_context", {}) + for row in case_context.get("facts", []) if isinstance(case_context, dict) else []: + if isinstance(row, dict) and row.get("fact_id") is not None: + projection_sources["FACT"][str(row["fact_id"])] = row + evidence_inventory = context_artifacts.get("evidence_inventory", {}) + for row in evidence_inventory.get("items", []) if isinstance(evidence_inventory, dict) else []: + if not isinstance(row, dict): + continue + identifier = row.get("evidence_id", row.get("id")) + if identifier is not None: + projection_sources["EVIDENCE"][str(identifier)] = row + supplied_sources = context_artifacts.get("_projection_sources", {}) + if isinstance(supplied_sources, dict): + for kind, rows in supplied_sources.items(): + if isinstance(rows, dict): + projection_sources[str(kind)].update({str(key): value for key, value in rows.items()}) + + def materialize_projection( + kind: str, + source_id: str, + ) -> dict[str, Any]: + content = projection_sources.get(kind, {}).get(source_id) + if content is None: + raise IngressError( + "SLICE_PROJECTION_SOURCE_MISSING", + f"projection source is missing: {kind}:{source_id}", + ) + if not isinstance(content, dict): + raise IngressError( + "SLICE_PROJECTION_SOURCE_MISSING", + f"projection source is not an object: {kind}:{source_id}", + ) + raw = canonical_json_bytes(content) + if len(raw) > 262144: + raise IngressError("SLICE_PROJECTION_BUDGET_EXCEEDED", f"projection exceeds 262144 bytes: {kind}:{source_id}") + source_refs = _dedupe_source_refs(content.get("source_refs", [])) + matching_refs = [ + ref + for ref in source_refs + if ref.get("stage1_id") == source_id + and isinstance(ref.get("raw_value_sha256"), str) + and re.fullmatch(r"[a-f0-9]{64}", str(ref["raw_value_sha256"])) + ] + matching_hashes = {str(ref["raw_value_sha256"]) for ref in matching_refs} + if not matching_refs or len(matching_hashes) != 1: + raise IngressError( + "SLICE_PROJECTION_PROVENANCE_MISSING", + f"projection lacks one unambiguous source_id-bound raw hash: {kind}:{source_id}", + ) + raw_source_sha256 = next(iter(matching_hashes)) + projection_mint = mint_stage2_id("content_projection", [kind, source_id, content], prefix="CP") + return { + "projection_id": projection_mint["id"], + "projection_kind": kind, + "source_id": source_id, + "content_schema_id": None, + "content": content, + "raw_source_sha256": raw_source_sha256, + "canonical_content_sha256": canonical_digest(content), + "materialized_utf8_bytes": len(raw), + "projection_policy_id": "S2-SLICE-BOUNDED-PROJECTION-V1", + "truncated": False, + "source_refs": source_refs, + } + + projection_arrays = { + "FACT": "fact_projections", + "EVIDENCE": "evidence_projections", + "EVENT": "event_projections", + "LES_STRUCTURE": "les_structure_projections", + "SIGNAL_OCCURRENCE": "signal_occurrence_projections", + "REVIEW_ITEM": "review_item_projections", + "ACTIVE_PROFILE": "profile_projections", + } + slices: dict[str, dict[str, Any]] = {} + for cluster in cluster_plan.get("clusters", []): + cluster_id = cluster["cluster_id"] + source_refs = _dedupe_source_refs(cluster.get("source_refs", [])) + member_refs_by_kind: dict[str, list[str]] = defaultdict(list) + for member in cluster.get("members", []): + member_refs_by_kind[str(member.get("member_kind"))].append(str(member.get("member_ref"))) + minted = mint_stage2_id( + "cluster_slice", + [cluster_id, source_refs, sorted(member_refs_by_kind.items())], + prefix="SL", + ) + slice_body: dict[str, Any] = { + "artifact_header": dict(cluster_plan.get("artifact_header", _artifact_header())), + "cluster_id": cluster_id, + "slice_id": minted["id"], + "slice_schema_version": "stage2_cluster_slice.v1", + "immutable": True, + "source_refs": source_refs, + "fact_refs": sorted(set(member_refs_by_kind.get("FACT", []))), + "evidence_refs": sorted(set(member_refs_by_kind.get("EVIDENCE", []))), + "event_refs": sorted(set(member_refs_by_kind.get("EVENT", []))), + "les_structure_refs": sorted(set(member_refs_by_kind.get("LES_STRUCTURE", []))), + "signal_occurrence_refs": sorted(set(member_refs_by_kind.get("SIGNAL_OCCURRENCE", []))), + "review_keys": sorted(set(member_refs_by_kind.get("REVIEW_ITEM", []))), + "profile_refs": [], + "derivation": { + "derivation_id": minted["id"], + "algorithm_version": ALGORITHM_VERSION, + "sorted_input_refs": sorted( + str(member.get("member_ref")) for member in cluster.get("members", []) + ) or [cluster_id], + "mint_input_sha256": minted["mint_input_sha256"], + }, + } + all_projections: list[dict[str, Any]] = [] + for kind, array_name in projection_arrays.items(): + source_ids = sorted(set(member_refs_by_kind.get(kind, []))) + projections = [materialize_projection(kind, source_id) for source_id in source_ids] + slice_body[array_name] = projections + all_projections.extend(projections) + observed_projection_bytes = sum(row["materialized_utf8_bytes"] for row in all_projections) + if observed_projection_bytes > 2097152: + raise IngressError("SLICE_CONTENT_BUDGET_EXCEEDED", f"slice exceeds 2097152 bytes: {cluster_id}") + slice_body["projection_budget"] = { + "policy_id": "S2-SLICE-BOUNDED-PROJECTION-V1", + "max_projection_utf8_bytes": 262144, + "max_slice_content_utf8_bytes": 2097152, + "observed_projection_count": len(all_projections), + "observed_slice_content_utf8_bytes": observed_projection_bytes, + "budget_status": "PASS", + "raw_stage1_reread_required": False, + } + slice_body["slice_digest"] = canonical_digest(slice_body) + slices[cluster_id] = slice_body + return slices + + + _PII_PATTERNS: tuple[tuple[str, re.Pattern[str]], ...] = ( + ("KOREAN_RESIDENT_ID", re.compile(r"\b\d{6}-?[1-4]\d{6}\b")), + ("EMAIL", re.compile(r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b")), + ("KOREAN_PHONE", re.compile(r"\b01[016789]-?\d{3,4}-?\d{4}\b")), + ) + + + def _scan_pii(value: Any) -> list[str]: + rendered = canonical_json_bytes(value).decode("utf-8") + return [name for name, pattern in _PII_PATTERNS if pattern.search(rendered)] + + + def compile_bundle_plan( + cluster_plan: Mapping[str, Any], + cluster_slices: Mapping[str, Mapping[str, Any]], + release_lock: Mapping[str, Any], + *, + stage2_asset_root: str | os.PathLike[str] | None = None, + ) -> dict[str, Any]: + """Compile a data-only S2_10 handoff without materializing prompts.""" + + bundle = release_lock.get("bundle") + if not isinstance(bundle, dict): + raise IngressError("BUNDLE_RELEASE_CONTRACT_MISSING", "release bundle must be an object") + mode = bundle.get("mode") + release_class = release_lock.get("release_class") + if mode not in RELEASE_MODE or RELEASE_MODE[mode] != release_class: + raise IngressError("BUNDLE_RELEASE_CLASS_MISMATCH", f"{mode} is not valid for {release_class}") + handoff_contract_version = bundle.get("handoff_contract_version") + if handoff_contract_version != "S2_00_STRUCTURED_CONTEXT_HANDOFF_V2": + raise IngressError( + "HANDOFF_CONTRACT_VERSION_MISMATCH", + "S2_00 requires S2_00_STRUCTURED_CONTEXT_HANDOFF_V2", + ) + release_ref = str( + release_lock.get( + "release_id", + release_lock.get("stage2_release_digest", "stage2_release"), + ) + ) + verification_errors: list[str] = [] + + def normalize_sha256(value: Any, *, code: str) -> str: + if not isinstance(value, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", value) is None: + verification_errors.append(code) + return "0" * 64 + return value.lower() + + def verify_opaque_ref(value: Any, *, label: str) -> str: + if not isinstance(value, dict): + verification_errors.append(f"{label}_REF_MISSING") + return "0" * 64 + try: + path = _safe_relative_path(value.get("path")).as_posix() + except (IngressError, TypeError): + verification_errors.append(f"{label}_REF_PATH_INVALID") + return "0" * 64 + expected = normalize_sha256( + value.get("sha256"), + code=f"{label}_HASH_UNBOUND", + ) + if stage2_asset_root is None: + verification_errors.append("STAGE2_ASSET_ROOT_MISSING") + return expected + try: + snapshot = open_bounded_snapshot( + stage2_asset_root, + path, + logical_input_id=label.lower(), + ) + except IngressError as exc: + verification_errors.append(f"{label}_ASSET_UNAVAILABLE") + return expected + if snapshot.raw_sha256 != expected: + verification_errors.append(f"{label}_HASH_MISMATCH") + return expected + + s2_10_agent_sha256 = verify_opaque_ref( + bundle.get("s2_10_agent_ref"), + label="S2_10_AGENT", + ) + s2_10_llm_binding_sha256 = verify_opaque_ref( + bundle.get("s2_10_llm_binding_ref"), + label="S2_10_LLM_BINDING", + ) + if normalize_sha256( + bundle.get("s2_10_agent_sha256"), + code="S2_10_AGENT_HASH_UNBOUND", + ) != s2_10_agent_sha256: + verification_errors.append("S2_10_AGENT_RELEASE_HASH_MISMATCH") + if normalize_sha256( + bundle.get("s2_10_llm_binding_sha256"), + code="S2_10_LLM_BINDING_HASH_UNBOUND", + ) != s2_10_llm_binding_sha256: + verification_errors.append("S2_10_LLM_BINDING_RELEASE_HASH_MISMATCH") + if bundle.get("context_cohort_policy_id") != "S2-CACHE-STRUCTURED-CONTEXT-HANDOFF-V2": + verification_errors.append("CONTEXT_COHORT_POLICY_MISMATCH") + if bundle.get("downstream_hybrid_status") != "HYBRID_RELEASE_BOUND": + verification_errors.append("S2_10_HYBRID_RELEASE_UNBOUND") + + selected_context_input = bundle.get("selected_context_refs", []) + if not isinstance(selected_context_input, list): + raise IngressError( + "SELECTED_CONTEXT_REFS_SHAPE", + "selected_context_refs must be an array", + ) + allowed_context_kinds = { + "ACTIVE_PROFILE", + "APPROVED_COMMON_AUTHORITY", + "LAW_VALUE_TOKEN", + "AUTHORITY_PROPOSITION", + } + selected_context_refs: list[dict[str, Any]] = [] + seen_context_ids: set[str] = set() + for index, row in enumerate(selected_context_input): + if not isinstance(row, dict): + raise IngressError( + "SELECTED_CONTEXT_REF_SHAPE", + f"selected context row {index} must be an object", + ) + context_ref_id = row.get("context_ref_id") + context_kind = row.get("context_kind") + if not isinstance(context_ref_id, str) or not context_ref_id: + raise IngressError( + "SELECTED_CONTEXT_REF_ID_INVALID", + f"selected context row {index} lacks context_ref_id", + ) + if context_ref_id in seen_context_ids: + verification_errors.append("SELECTED_CONTEXT_REF_ID_DUPLICATE") + seen_context_ids.add(context_ref_id) + if context_kind not in allowed_context_kinds: + raise IngressError( + "SELECTED_CONTEXT_KIND_INVALID", + f"unsupported context kind: {context_kind}", + ) + try: + relative_path = _safe_relative_path(row.get("path")).as_posix() + except (IngressError, TypeError) as exc: + raise IngressError( + "SELECTED_CONTEXT_PATH_INVALID", + f"invalid selected context path at row {index}", + ) from exc + raw_sha256 = normalize_sha256( + row.get("raw_sha256"), + code="SELECTED_CONTEXT_RAW_HASH_UNBOUND", + ) + canonical_sha256 = normalize_sha256( + row.get("canonical_sha256"), + code="SELECTED_CONTEXT_CANONICAL_HASH_UNBOUND", + ) + if row.get("order_index") != index: + verification_errors.append("SELECTED_CONTEXT_ORDER_INVALID") + observed_raw_sha256: str | None = None + observed_canonical_sha256: str | None = None + pii_codes: list[str] = [] + if stage2_asset_root is None: + verification_errors.append("STAGE2_ASSET_ROOT_MISSING") + else: + try: + snapshot = open_bounded_snapshot( + stage2_asset_root, + relative_path, + logical_input_id=f"selected_context:{context_ref_id}", + ) + observed_raw_sha256 = snapshot.raw_sha256 + try: + text = snapshot.raw.decode("utf-8", errors="strict") + canonical_text = unicodedata.normalize( + "NFC", + text.replace("\r\n", "\n").replace("\r", "\n"), + ) + observed_canonical_sha256 = hashlib.sha256( + canonical_text.encode("utf-8") + ).hexdigest() + pii_codes = _scan_pii(canonical_text) + except UnicodeDecodeError: + verification_errors.append("SELECTED_CONTEXT_NOT_UTF8") + except IngressError as exc: + verification_errors.append( + "SELECTED_CONTEXT_ASSET_UNAVAILABLE" + ) + if observed_raw_sha256 != raw_sha256: + verification_errors.append("SELECTED_CONTEXT_ASSET_HASH_MISMATCH") + if observed_canonical_sha256 != canonical_sha256: + verification_errors.append( + "SELECTED_CONTEXT_CANONICAL_HASH_MISMATCH" + ) + if pii_codes: + verification_errors.append("SELECTED_CONTEXT_PII") + selected_context_refs.append( + { + "context_ref_id": context_ref_id, + "context_kind": context_kind, + "path": relative_path, + "raw_sha256": raw_sha256, + "canonical_sha256": canonical_sha256, + "order_index": index, + "release_ref": str(row.get("release_ref", release_ref)), + } + ) + selected_context_ordered_refs_sha256 = canonical_digest( + selected_context_refs + ) + declared_context_digest = normalize_sha256( + bundle.get("selected_context_ordered_refs_sha256"), + code="SELECTED_CONTEXT_ORDERED_REFS_HASH_UNBOUND", + ) + if declared_context_digest != selected_context_ordered_refs_sha256: + verification_errors.append( + "SELECTED_CONTEXT_ORDERED_REFS_HASH_MISMATCH" + ) + if mode != "STRUCTURAL_FIXTURE" and not selected_context_refs: + verification_errors.append("SELECTED_CONTEXT_EMPTY") + + wave_by_scc: dict[str, int] = {} + for wave_ordinal, wave in enumerate( + cluster_plan.get("scheduling_waves", []) + ): + for scc_id in wave: + wave_by_scc[str(scc_id)] = wave_ordinal + cluster_rows = { + str(row.get("cluster_id")): row + for row in cluster_plan.get("clusters", []) + if isinstance(row, dict) and row.get("cluster_id") is not None + } + + valid_member_slices: list[dict[str, Any]] = [] + missing_member_slices: list[dict[str, Any]] = [] + for cluster_id in cluster_plan.get("executable_cluster_ids", []): + cluster_id = str(cluster_id) + cluster_row = cluster_rows.get(cluster_id, {}) + scc_ids = [ + str(value) + for value in cluster_row.get("scc_ids", []) + ] if isinstance(cluster_row, dict) else [] + dependency_wave_ordinal = min( + (wave_by_scc.get(value, 0) for value in scc_ids), + default=0, + ) + slice_body = cluster_slices.get(cluster_id) + member_slice = { + "cluster_id": cluster_id, + "path": f"context/cluster_slices/{cluster_id}.json", + "sha256": ( + canonical_digest(slice_body) + if isinstance(slice_body, Mapping) + else "0" * 64 + ), + "dependency_wave_ordinal": dependency_wave_ordinal, + } + if isinstance(slice_body, Mapping): + valid_member_slices.append(member_slice) + else: + missing_member_slices.append(member_slice) + + def build_cohort( + member_slices: Sequence[Mapping[str, Any]], + *, + extra_reason_codes: Sequence[str] = (), + ) -> dict[str, Any]: + cohort_member_cluster_ids = sorted( + str(row["cluster_id"]) for row in member_slices + ) + public_member_slices = sorted( + (dict(row) for row in member_slices), + key=lambda row: row["cluster_id"], + ) + cohort_membership_sha256 = canonical_digest( + cohort_member_cluster_ids + ) + member_slice_ref_set_sha256 = canonical_digest( + public_member_slices + ) + receipt_errors = sorted( + set(verification_errors).union(str(code) for code in extra_reason_codes) + ) + cohort_mint = mint_stage2_id( + "bundle_cohort", + [ + handoff_contract_version, + mode, + s2_10_agent_sha256, + s2_10_llm_binding_sha256, + selected_context_ordered_refs_sha256, + cohort_membership_sha256, + member_slice_ref_set_sha256, + ], + prefix="BC", + ) + materialization_input = { + "handoff_contract_version": handoff_contract_version, + "bundle_cohort_id": cohort_mint["id"], + "s2_10_agent_sha256": s2_10_agent_sha256, + "s2_10_llm_binding_sha256": s2_10_llm_binding_sha256, + "selected_context_ordered_refs_sha256": ( + selected_context_ordered_refs_sha256 + ), + "cohort_membership_sha256": cohort_membership_sha256, + "member_slice_ref_set_sha256": member_slice_ref_set_sha256, + } + receipt = { + "schema_version": ( + "stage2_s2_00_context_materialization_receipt.v2" + ), + "bundle_cohort_id": cohort_mint["id"], + "s2_10_agent_sha256": s2_10_agent_sha256, + "s2_10_llm_binding_sha256": s2_10_llm_binding_sha256, + "selected_context_ordered_refs_sha256": ( + selected_context_ordered_refs_sha256 + ), + "cohort_membership_sha256": cohort_membership_sha256, + "member_slice_ref_set_sha256": member_slice_ref_set_sha256, + "materialization_input_sha256": canonical_digest( + materialization_input + ), + "verification_status": ( + "PASS" if not receipt_errors else "NON_EXECUTABLE" + ), + "reason_codes": receipt_errors, + } + reason_codes = list(receipt_errors) + if mode == "STRUCTURAL_FIXTURE": + reason_codes.append("STRUCTURAL_FIXTURE_NOT_EXECUTABLE") + cohort_status = "STRUCTURAL_ONLY" + elif receipt_errors: + cohort_status = "NON_EXECUTABLE" + else: + cohort_status = "EXECUTABLE" + reason_codes = sorted(set(reason_codes)) + source_refs = _dedupe_source_refs( + ref + for cluster_id in cohort_member_cluster_ids + for ref in ( + cluster_rows.get(cluster_id, {}).get("source_refs", []) + if isinstance(cluster_rows.get(cluster_id), dict) + else [] + ) + ) + return { + "handoff_contract_version": handoff_contract_version, + "bundle_cohort_id": cohort_mint["id"], + "compile_mode": mode, + "release_class": release_class, + "cohort_status": cohort_status, + "selected_context_refs": list(selected_context_refs), + "selected_context_ordered_refs_sha256": ( + selected_context_ordered_refs_sha256 + ), + "member_slices": public_member_slices, + "cohort_member_cluster_ids": cohort_member_cluster_ids, + "s2_10_agent_sha256": s2_10_agent_sha256, + "s2_10_llm_binding_sha256": s2_10_llm_binding_sha256, + "context_materialization_receipt": receipt, + "forbidden_bulk_inputs_present": False, + "reason_codes": reason_codes, + "source_refs": source_refs, + "derivation": { + "derivation_id": cohort_mint["id"], + "algorithm_version": ALGORITHM_VERSION, + "sorted_input_refs": sorted( + cohort_member_cluster_ids + + [ + f"context:{selected_context_ordered_refs_sha256}", + f"agent:{s2_10_agent_sha256}", + f"binding:{s2_10_llm_binding_sha256}", + ] + ), + "mint_input_sha256": cohort_mint["mint_input_sha256"], + }, + } + + cohorts: list[dict[str, Any]] = [] + if valid_member_slices: + cohorts.append(build_cohort(valid_member_slices)) + for member_slice in missing_member_slices: + cohorts.append( + build_cohort( + [member_slice], + extra_reason_codes=["CLUSTER_SLICE_MISSING"], + ) + ) + public_cohorts = sorted( + cohorts, + key=lambda row: ( + row["cohort_member_cluster_ids"][0] + if row["cohort_member_cluster_ids"] + else row["bundle_cohort_id"] + ), + ) + executable_cohort_ids = sorted( + row["bundle_cohort_id"] + for row in public_cohorts + if row["cohort_status"] == "EXECUTABLE" + ) + non_executable_cohort_ids = sorted( + row["bundle_cohort_id"] + for row in public_cohorts + if row["cohort_status"] != "EXECUTABLE" + ) + top_refs = ( + _dedupe_source_refs( + ref for row in public_cohorts for ref in row["source_refs"] + ) + if public_cohorts + else list(cluster_plan.get("source_refs", [])) + ) + return { + "artifact_header": dict( + cluster_plan.get("artifact_header", _artifact_header()) + ), + "handoff_contract_version": handoff_contract_version, + "cohorts": public_cohorts, + "executable_bundle_cohort_ids": executable_cohort_ids, + "non_executable_bundle_cohort_ids": non_executable_cohort_ids, + "source_refs": top_refs, + "derivation": _derivation( + "bundle_plan_derivation", + [ + row["bundle_cohort_id"] for row in public_cohorts + ] or ["NO_BUNDLE_COHORT"], + [ + selected_context_ordered_refs_sha256, + s2_10_agent_sha256, + s2_10_llm_binding_sha256, + { + row["bundle_cohort_id"]: row["reason_codes"] + for row in public_cohorts + }, + ], + ), + } + + def validate_bundle_release_cohorts(bundle_plan: Mapping[str, Any]) -> dict[str, Any]: + """Recompute every cross-object invariant the JSON Schema cannot express.""" + + cohorts = bundle_plan.get("cohorts", []) + if not isinstance(cohorts, list): + raise IngressError( + "BUNDLE_COHORT_INVARIANT_FAILED", + "bundle_plan.cohorts must be an array", + ) + mode = cohorts[0].get("compile_mode") if cohorts else None + release_class = cohorts[0].get("release_class") if cohorts else None + invariant_errors: list[str] = [] + receipt_reason_codes = { + "S2_10_AGENT_REF_MISSING", + "S2_10_AGENT_REF_PATH_INVALID", + "S2_10_AGENT_HASH_UNBOUND", + "S2_10_AGENT_ASSET_UNAVAILABLE", + "S2_10_AGENT_HASH_MISMATCH", + "S2_10_AGENT_RELEASE_HASH_MISMATCH", + "S2_10_LLM_BINDING_REF_MISSING", + "S2_10_LLM_BINDING_REF_PATH_INVALID", + "S2_10_LLM_BINDING_HASH_UNBOUND", + "S2_10_LLM_BINDING_ASSET_UNAVAILABLE", + "S2_10_LLM_BINDING_HASH_MISMATCH", + "S2_10_LLM_BINDING_RELEASE_HASH_MISMATCH", + "S2_10_HYBRID_RELEASE_UNBOUND", + "STAGE2_ASSET_ROOT_MISSING", + "CONTEXT_COHORT_POLICY_MISMATCH", + "SELECTED_CONTEXT_REF_ID_DUPLICATE", + "SELECTED_CONTEXT_ORDER_INVALID", + "SELECTED_CONTEXT_RAW_HASH_UNBOUND", + "SELECTED_CONTEXT_CANONICAL_HASH_UNBOUND", + "SELECTED_CONTEXT_ASSET_UNAVAILABLE", + "SELECTED_CONTEXT_ASSET_HASH_MISMATCH", + "SELECTED_CONTEXT_CANONICAL_HASH_MISMATCH", + "SELECTED_CONTEXT_NOT_UTF8", + "SELECTED_CONTEXT_PII", + "SELECTED_CONTEXT_ORDERED_REFS_HASH_UNBOUND", + "SELECTED_CONTEXT_ORDERED_REFS_HASH_MISMATCH", + "SELECTED_CONTEXT_EMPTY", + "CLUSTER_SLICE_MISSING", + } + cohort_reason_codes = receipt_reason_codes | { + "STRUCTURAL_FIXTURE_NOT_EXECUTABLE" + } + + for index, row in enumerate(cohorts): + if not isinstance(row, dict): + invariant_errors.append(f"COHORT_{index}_SHAPE") + continue + row_mode = row.get("compile_mode") + row_release_class = row.get("release_class") + if row_mode not in RELEASE_MODE or RELEASE_MODE[row_mode] != row_release_class: + invariant_errors.append(f"COHORT_{index}_MODE_RELEASE_CLASS") + if mode is not None and row_mode != mode: + invariant_errors.append(f"COHORT_{index}_MODE_MIXED") + if release_class is not None and row_release_class != release_class: + invariant_errors.append(f"COHORT_{index}_RELEASE_CLASS_MIXED") + + selected_refs = row.get("selected_context_refs", []) + if not isinstance(selected_refs, list): + invariant_errors.append(f"COHORT_{index}_SELECTED_CONTEXT_SHAPE") + selected_refs = [] + if [item.get("order_index") for item in selected_refs if isinstance(item, dict)] != list( + range(len(selected_refs)) + ): + invariant_errors.append(f"COHORT_{index}_SELECTED_CONTEXT_ORDER") + selected_digest = canonical_digest(selected_refs) + if row.get("selected_context_ordered_refs_sha256") != selected_digest: + invariant_errors.append(f"COHORT_{index}_SELECTED_CONTEXT_DIGEST") + + member_ids = row.get("cohort_member_cluster_ids", []) + member_slices = row.get("member_slices", []) + if not isinstance(member_ids, list) or member_ids != sorted(member_ids): + invariant_errors.append(f"COHORT_{index}_MEMBER_ID_ORDER") + member_ids = list(member_ids) if isinstance(member_ids, list) else [] + if not isinstance(member_slices, list) or member_slices != sorted( + member_slices, + key=lambda item: str(item.get("cluster_id", "")) if isinstance(item, dict) else "", + ): + invariant_errors.append(f"COHORT_{index}_MEMBER_SLICE_ORDER") + member_slices = list(member_slices) if isinstance(member_slices, list) else [] + slice_ids = [ + item.get("cluster_id") + for item in member_slices + if isinstance(item, dict) + ] + if member_ids != slice_ids: + invariant_errors.append(f"COHORT_{index}_MEMBERSHIP_SET") + membership_digest = canonical_digest(member_ids) + slice_digest = canonical_digest(member_slices) + + receipt = row.get("context_materialization_receipt") + if not isinstance(receipt, dict): + invariant_errors.append(f"COHORT_{index}_RECEIPT_SHAPE") + receipt = {} + echoed_fields = { + "bundle_cohort_id": row.get("bundle_cohort_id"), + "s2_10_agent_sha256": row.get("s2_10_agent_sha256"), + "s2_10_llm_binding_sha256": row.get("s2_10_llm_binding_sha256"), + "selected_context_ordered_refs_sha256": selected_digest, + "cohort_membership_sha256": membership_digest, + "member_slice_ref_set_sha256": slice_digest, + } + for key, expected in echoed_fields.items(): + if receipt.get(key) != expected: + invariant_errors.append(f"COHORT_{index}_RECEIPT_{key.upper()}") + materialization_input = { + "handoff_contract_version": row.get("handoff_contract_version"), + **echoed_fields, + } + if receipt.get("materialization_input_sha256") != canonical_digest( + materialization_input + ): + invariant_errors.append(f"COHORT_{index}_MATERIALIZATION_DIGEST") + + expected_cohort = mint_stage2_id( + "bundle_cohort", + [ + row.get("handoff_contract_version"), + row_mode, + row.get("s2_10_agent_sha256"), + row.get("s2_10_llm_binding_sha256"), + selected_digest, + membership_digest, + slice_digest, + ], + prefix="BC", + )["id"] + if row.get("bundle_cohort_id") != expected_cohort: + invariant_errors.append(f"COHORT_{index}_ID_MINT") + + receipt_codes = receipt.get("reason_codes", []) + row_codes = row.get("reason_codes", []) + if not isinstance(receipt_codes, list) or not set(receipt_codes).issubset( + receipt_reason_codes + ): + invariant_errors.append(f"COHORT_{index}_RECEIPT_REASON_CODE") + if not isinstance(row_codes, list) or not set(row_codes).issubset( + cohort_reason_codes + ): + invariant_errors.append(f"COHORT_{index}_REASON_CODE") + + status = row.get("cohort_status") + receipt_status = receipt.get("verification_status") + if row_mode == "STRUCTURAL_FIXTURE": + if status != "STRUCTURAL_ONLY" or "STRUCTURAL_FIXTURE_NOT_EXECUTABLE" not in row_codes: + invariant_errors.append(f"COHORT_{index}_STRUCTURAL_STATUS") + elif receipt_status == "PASS": + if status != "EXECUTABLE" or row_codes: + invariant_errors.append(f"COHORT_{index}_EXECUTABLE_STATUS") + elif receipt_status == "NON_EXECUTABLE": + if status != "NON_EXECUTABLE" or not row_codes: + invariant_errors.append(f"COHORT_{index}_NON_EXECUTABLE_STATUS") + else: + invariant_errors.append(f"COHORT_{index}_RECEIPT_STATUS") + + expected_executable_ids = sorted( + str(row.get("bundle_cohort_id")) + for row in cohorts + if isinstance(row, dict) and row.get("cohort_status") == "EXECUTABLE" + ) + expected_non_executable_ids = sorted( + str(row.get("bundle_cohort_id")) + for row in cohorts + if isinstance(row, dict) and row.get("cohort_status") != "EXECUTABLE" + ) + if bundle_plan.get("executable_bundle_cohort_ids") != expected_executable_ids: + invariant_errors.append("TOP_EXECUTABLE_COHORT_PARTITION") + if bundle_plan.get("non_executable_bundle_cohort_ids") != expected_non_executable_ids: + invariant_errors.append("TOP_NON_EXECUTABLE_COHORT_PARTITION") + if invariant_errors: + raise IngressError( + "BUNDLE_COHORT_INVARIANT_FAILED", + ",".join(sorted(set(invariant_errors))), + details={"invariant_errors": sorted(set(invariant_errors))}, + ) + + mapping_pass = mode in RELEASE_MODE and RELEASE_MODE[mode] == release_class + executable = sorted( + cluster_id + for row in cohorts + if row.get("cohort_status") == "EXECUTABLE" + for cluster_id in row.get("cohort_member_cluster_ids", []) + ) + non_executable = sorted( + cluster_id + for row in cohorts + if row.get("cohort_status") == "NON_EXECUTABLE" + for cluster_id in row.get("cohort_member_cluster_ids", []) + ) + structural = sorted( + cluster_id + for row in cohorts + if row.get("cohort_status") == "STRUCTURAL_ONLY" + for cluster_id in row.get("cohort_member_cluster_ids", []) + ) + return { + "mapping_pass": mapping_pass, + "executable_cluster_ids": executable, + "non_executable_cluster_ids": non_executable, + "structural_cluster_ids": structural, + "passed": mapping_pass and (mode == "STRUCTURAL_FIXTURE" or bool(executable)), + } + + + def _write_fsynced(path: Path, payload: bytes) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + flags = os.O_WRONLY | os.O_CREAT | os.O_EXCL + if hasattr(os, "O_CLOEXEC"): + flags |= os.O_CLOEXEC + descriptor = os.open(path, flags, 0o600) + try: + view = memoryview(payload) + while view: + written = os.write(descriptor, view) + view = view[written:] + os.fsync(descriptor) + finally: + os.close(descriptor) + + + def _fsync_directory(path: Path) -> None: + flags = os.O_RDONLY + if hasattr(os, "O_DIRECTORY"): + flags |= os.O_DIRECTORY + descriptor = os.open(path, flags) + try: + os.fsync(descriptor) + finally: + os.close(descriptor) + + + def _published_tree_matches( + output_dir: Path, + run_binding_digest: str, + *, + stage2_asset_root: str | os.PathLike[str] | None = None, + ) -> bool: + status_path = output_dir / "ingress" / "ingress_status.json" + if not status_path.is_file() or status_path.is_symlink(): + return False + try: + snapshot = open_bounded_snapshot(output_dir, "ingress/ingress_status.json", logical_input_id="published_status") + status_value = load_json_strict(snapshot) + except IngressError: + return False + if not isinstance(status_value, dict): + return False + binding = status_value.get("run_binding_receipt", {}) + if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: + return False + barrier = status_value.get("output_barrier", {}) + if not isinstance(barrier, dict) or barrier.get("written_last") is not True: + return False + expected_artifacts = barrier.get("artifacts", []) + if not isinstance(expected_artifacts, list): + return False + if barrier.get("artifact_set_digest") != canonical_digest(expected_artifacts): + return False + try: + schema_documents = _load_output_schemas( + Path(stage2_asset_root) + if stage2_asset_root is not None + else Path(__file__).resolve().parents[1] + ) + _validate_output_artifact( + "ingress/ingress_status.json", + status_value, + schema_documents, + ) + except IngressError: + return False + for row in expected_artifacts: + if not isinstance(row, dict): + return False + try: + artifact = open_bounded_snapshot(output_dir, row["path"], logical_input_id="published_artifact") + except IngressError: + return False + if artifact.raw_sha256 != row.get("raw_sha256"): + return False + try: + parsed_artifact = load_json_strict(artifact) + _validate_output_artifact(row["path"], parsed_artifact, schema_documents) + except IngressError: + return False + expected_paths = { + row.get("path") + for row in expected_artifacts + if isinstance(row, dict) and isinstance(row.get("path"), str) + } | {"ingress/ingress_status.json"} + observed_paths = { + path.relative_to(output_dir).as_posix() + for path in output_dir.rglob("*") + if path.is_file() and not path.is_symlink() + } + if expected_paths != observed_paths: + return False + return True + + + def _output_schema_id(relative_path: str) -> str: + if relative_path.startswith("context/"): + return CONTEXT_SCHEMA_ID + if relative_path.startswith("review/"): + return REVIEW_SCHEMA_ID + if relative_path.startswith("ingress/"): + return INGRESS_SCHEMA_ID + raise IngressError("OUTPUT_SCHEMA_FAMILY_UNKNOWN", f"no schema family for {relative_path}") + + + def publish_atomically( + output_dir: str | os.PathLike[str], + artifacts: Mapping[str, Any], + *, + run_binding_digest: str, + attempt_id: str, + stage2_asset_root: str | os.PathLike[str] | None = None, + ) -> dict[str, Any]: + """Publish a complete normal or diagnostic tree with one directory rename.""" + + output = Path(output_dir) + if not output.is_absolute(): + output = output.resolve() + if not re.fullmatch(r"[A-Za-z0-9_.-]{1,128}", attempt_id): + raise IngressError("ATTEMPT_ID_INVALID", "attempt_id contains forbidden characters") + if "ingress/ingress_status.json" not in artifacts: + raise IngressError("OUTPUT_BARRIER_MISSING", "ingress_status must be supplied") + safe_paths = {_safe_relative_path(path).as_posix(): value for path, value in artifacts.items()} + if len(safe_paths) != len(artifacts): + raise IngressError("DUPLICATE_OUTPUT_PATH", "duplicate output paths after canonicalization") + schema_root = ( + Path(stage2_asset_root) + if stage2_asset_root is not None + else Path(__file__).resolve().parents[1] + ) + schema_documents = _load_output_schemas(schema_root) + for relative_path, supplied_value in safe_paths.items(): + value = load_json_strict(supplied_value) if isinstance(supplied_value, bytes) else supplied_value + _validate_output_artifact(relative_path, value, schema_documents) + if output.exists(): + if output.is_dir() and _published_tree_matches( + output, + run_binding_digest, + stage2_asset_root=schema_root, + ): + return {"status": "IDEMPOTENT_SUCCESS", "output_dir": str(output), "run_binding_digest": run_binding_digest} + raise IngressError("RUN_TUPLE_CONFLICT", "published output exists with a different binding") + parent = output.parent + parent.mkdir(parents=True, exist_ok=True) + staging_parent = parent / ".staging" + staging_parent.mkdir(parents=True, exist_ok=True) + staging = staging_parent / f"{output.name}.{attempt_id}" + if staging.exists(): + if staging.is_symlink() or staging.parent.resolve() != staging_parent.resolve(): + raise IngressError("STAGING_PATH_UNSAFE", "staging path is not safe") + shutil.rmtree(staging) + staging.mkdir(mode=0o700) + try: + artifact_receipt: list[dict[str, Any]] = [] + for relative_path in sorted(path for path in safe_paths if path != "ingress/ingress_status.json"): + value = safe_paths[relative_path] + payload = value if isinstance(value, bytes) else canonical_json_bytes(value) + _write_fsynced(staging / relative_path, payload) + artifact_receipt.append( + { + "logical_artifact_id": relative_path, + "path": relative_path, + "schema_id": _output_schema_id(relative_path), + "raw_sha256": hashlib.sha256(payload).hexdigest(), + } + ) + status = dict(safe_paths["ingress/ingress_status.json"]) + binding = status.get("run_binding_receipt") + if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: + raise IngressError("RUN_BINDING_RECEIPT_MISMATCH", "status run binding does not match publish binding") + status["output_barrier"] = { + "barrier_id": "S2_00_INGRESS_STATUS_BARRIER", + "barrier_path": "ingress/ingress_status.json", + "publish_semantics": "STATUS_LAST_LOGICAL_COMMIT", + "canonical_output_root": f"stage2_runs/by-binding/{run_binding_digest}/", + "branch": "DIAGNOSTIC" if "ingress/technical_diagnostic.json" in safe_paths else "NORMAL", + "artifacts": artifact_receipt, + "artifact_set_digest": canonical_digest(artifact_receipt), + "non_status_artifacts_read_back_verified": True, + "written_last": True, + } + _validate_output_artifact( + "ingress/ingress_status.json", + status, + schema_documents, + ) + _write_fsynced(staging / "ingress/ingress_status.json", canonical_json_bytes(status)) + directories = sorted((path for path in staging.rglob("*") if path.is_dir()), key=lambda path: len(path.parts), reverse=True) + for directory in directories: + _fsync_directory(directory) + _fsync_directory(staging) + _fsync_directory(staging_parent) + os.rename(staging, output) + _fsync_directory(parent) + except BaseException: + if staging.exists() and staging.parent == staging_parent: + shutil.rmtree(staging) + raise + return {"status": "PUBLISHED", "output_dir": str(output), "run_binding_digest": run_binding_digest} + + + def _load_case_snapshots( + stage1_root: Path, + contracts: Sequence[Mapping[str, Any]], + release_lock: Mapping[str, Any], + ) -> tuple[dict[str, Snapshot], list[dict[str, Any]]]: + snapshots: dict[str, Snapshot] = {} + issues: list[dict[str, Any]] = [] + aggregate = 0 + max_file = int(release_lock.get("limits", {}).get("max_file_bytes", MAX_FILE_BYTES)) + max_run = int(release_lock.get("limits", {}).get("max_run_bytes", MAX_RUN_BYTES)) + for contract in contracts: + logical_id = str(contract["logical_input_id"]) + try: + snapshot = open_bounded_snapshot( + stage1_root, + str(contract["observed_path"]), + logical_input_id=logical_id, + max_bytes=max_file, + ) + except IngressError as exc: + if exc.code != "SOURCE_MISSING": + issues.append(_issue(exc.code, source_refs=[logical_id], message=str(exc))) + continue + aggregate += snapshot.byte_length + if aggregate > max_run: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "aggregate Stage 1 input budget exceeded") + snapshots[logical_id] = snapshot + return snapshots, issues + + + def _load_bound_completion_seal( + stage1_root: Path, + release_lock: Mapping[str, Any], + ) -> tuple[Mapping[str, Any] | None, Snapshot | None]: + dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) + ref = dependency.get("completion_seal_ref", {}) if isinstance(dependency, dict) else {} + if not isinstance(ref, dict): + return None, None + path = ref.get("path") + digest = ref.get("sha256") + if path in {None, "PENDING_SEQUENTIAL_BIND"} or digest in {None, "PENDING_SEQUENTIAL_BIND"}: + return None, None + if not isinstance(path, str) or not isinstance(digest, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", digest) is None: + raise IngressError("COMPLETION_SEAL_UNBOUND", "completion seal ref is not exactly bound") + snapshot = open_bounded_snapshot(stage1_root, path, logical_input_id="stage1_completion_seal") + if snapshot.raw_sha256 != digest.lower(): + raise IngressError("COMPLETION_SEAL_HASH_MISMATCH", "completion seal raw hash differs from release") + document = load_json_strict(snapshot) + if not isinstance(document, dict): + raise IngressError("COMPLETION_SEAL_SHAPE", "completion seal must be an object") + return document, snapshot + + + def _load_stage1_deployment_closure( + stage1_deployment_root: Path, + release_lock: Mapping[str, Any], + ) -> tuple[dict[str, Snapshot], dict[str, Any], list[dict[str, Any]]]: + """Load only the release-enumerated Stage 1 deployment closure.""" + + dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) + locked_rows = dependency.get("concrete_paths", []) if isinstance(dependency, dict) else [] + if not isinstance(locked_rows, list) or not locked_rows: + raise IngressError("STAGE1_DEPENDENCY_LOCK_MISSING", "Stage 1 deployment closure is not enumerated") + expected_count = dependency.get("expected_concrete_path_count") + if expected_count is not None and expected_count != len(locked_rows): + raise IngressError("STAGE1_DEPENDENCY_COUNT_MISMATCH", "Stage 1 dependency row count is not sealed") + snapshots: dict[str, Snapshot] = {} + documents: dict[str, Any] = {} + source_rows: list[dict[str, Any]] = [] + seen_paths: set[str] = set() + aggregate = 0 + max_file = int(release_lock.get("limits", {}).get("max_file_bytes", MAX_FILE_BYTES)) + max_run = int(release_lock.get("limits", {}).get("max_run_bytes", MAX_RUN_BYTES)) + for index, locked in enumerate(locked_rows): + if not isinstance(locked, dict) or not isinstance(locked.get("path"), str): + raise IngressError("STAGE1_DEPENDENCY_ROW_SHAPE", f"invalid dependency row {index}") + path = _safe_relative_path(locked["path"]).as_posix() + if path in seen_paths: + raise IngressError("STAGE1_DEPENDENCY_DUPLICATE_PATH", f"duplicate dependency path: {path}") + seen_paths.add(path) + expected_hash = locked.get("sha256") + if not isinstance(expected_hash, str) or not re.fullmatch(r"[A-Fa-f0-9]{64}", expected_hash): + raise IngressError("STAGE1_DEPENDENCY_UNBOUND", f"dependency hash is not bound: {path}") + lock_id = str(locked.get("lock_id", f"S1-DEPLOY-{index + 1:03d}")) + snapshot = open_bounded_snapshot( + stage1_deployment_root, + path, + logical_input_id=f"deployment:{lock_id}", + max_bytes=max_file, + ) + aggregate += snapshot.byte_length + if aggregate > max_run: + raise IngressError("AGGREGATE_DEPLOYMENT_SIZE_LIMIT", "Stage 1 deployment closure exceeds byte budget") + if snapshot.raw_sha256 != expected_hash: + raise IngressError("STAGE1_DEPENDENCY_HASH_MISMATCH", f"deployment hash mismatch: {path}") + document = load_json_strict( + snapshot, + max_depth=int(release_lock.get("limits", {}).get("max_json_depth", MAX_JSON_DEPTH)), + max_items=int(release_lock.get("limits", {}).get("max_json_items", MAX_JSON_ITEMS)), + ) + snapshots[lock_id] = snapshot + documents[path] = document + schema_id = locked.get("schema_id") + source_rows.append( + { + "logical_input_id": f"deployment:{lock_id}", + "requirement_class": "UPSTREAM_DEPLOYMENT", + "expected_path": path, + "observed_path": path, + "resolution_source": "RELEASE_BOUND_CONTRACT_MANIFEST", + "schema_id": schema_id, + "schema_sha256": None, + "producer_id": None, + "producer_alias_id": None, + "adapter_id": "S2A-UPSTREAM-DEPLOYMENT-V1", + "raw_sha256": snapshot.raw_sha256, + "byte_length": snapshot.byte_length, + "run_identity_ref": {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": path}, + "transaction_identity_ref": {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": path}, + "parse_status": "PASS", + "schema_status": "UNEVALUABLE" if schema_id else "NOT_APPLICABLE", + "seal_status": "PASS", + "scope_technical_disposition": "AVAILABLE", + "impact_scope": "GLOBAL", + "scope_refs": [f"deployment:{lock_id}"], + "source_contract_row_refs": [f"deployment:{lock_id}"], + "reason_codes": [], + "downstream_allowed_actions": [], + "issue_codes": [], + } + ) + return snapshots, documents, source_rows + + + def _signal_source_contract_rows( + signal_all: Mapping[str, Any], + release_lock: Mapping[str, Any], + ) -> list[dict[str, Any]]: + rows: list[dict[str, Any]] = [] + transaction_id = str(signal_all.get("manifest_transaction_id", "MISSING")) + family_contract = next( + ( + row + for row in _release_stage1_source_rows(release_lock) + if row.get("logical_input_id") == "signal_payload_family" + ), + {}, + ) + for file_row in signal_all.get("ordered_file_rows", []): + logical_id = f"signal_file:{int(file_row['manifest_index']):03d}" + hash_status = str(file_row.get("hash_status", "UNEVALUABLE")) + issue_codes = ["SIGNAL_FILE_HASH_MISMATCH"] if hash_status == "FAIL" else [] + rows.append( + { + "logical_input_id": logical_id, + "requirement_class": "SIGNAL_PAYLOAD", + "expected_path": str(file_row["physical_path"]), + "observed_path": str(file_row["physical_path"]), + "resolution_source": "RELEASE_BOUND_CONTRACT_MANIFEST", + "schema_id": None, + "schema_sha256": None, + "producer_id": family_contract.get("producer_id"), + "producer_alias_id": family_contract.get("producer_alias_id"), + "adapter_id": family_contract.get("adapter_id", "S2A-SIGNAL-PAYLOAD-FAMILY-V1"), + "raw_sha256": str(file_row["raw_sha256"]), + "byte_length": int(file_row["byte_length"]), + "run_identity_ref": {"value": None, "disposition": "MISSING", "source_ref": logical_id}, + "transaction_identity_ref": { + "value": transaction_id, + "disposition": "OBSERVED", + "source_ref": "signal_manifest", + }, + "parse_status": "PASS", + "schema_status": "UNEVALUABLE", + "seal_status": hash_status, + "scope_technical_disposition": "AVAILABLE_WITH_ISSUES" if issue_codes else "AVAILABLE", + "impact_scope": "SIGNAL", + "scope_refs": [logical_id], + "source_contract_row_refs": [logical_id], + "reason_codes": issue_codes, + "downstream_allowed_actions": [], + "issue_codes": issue_codes, + } + ) + return rows + + + def _snapshot_state_digest(snapshots: Sequence[Snapshot]) -> str: + """Digest stable source content, not hydration-local filesystem identity.""" + + return canonical_digest( + sorted( + [ + item.logical_input_id, + item.relative_path, + item.raw_sha256, + item.byte_length, + ] + for item in snapshots + ) + ) + + + def _verify_snapshot_state(snapshots: Sequence[Snapshot]) -> str: + rows: list[list[Any]] = [] + for item in snapshots: + path = Path(item.resolved_path) + if path.is_symlink(): + raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source became a symlink: {item.relative_path}") + try: + observed = path.stat(follow_symlinks=False) + except FileNotFoundError as exc: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source disappeared: {item.relative_path}") from exc + identity = (observed.st_dev, observed.st_ino, observed.st_size, observed.st_mtime_ns) + expected = (item.device, item.inode, item.byte_length, item.mtime_ns) + if identity != expected: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source changed: {item.relative_path}") + rows.append( + [ + item.logical_input_id, + item.relative_path, + item.raw_sha256, + item.byte_length, + ] + ) + return canonical_digest(sorted(rows)) + + + def _cluster_inputs(documents: Mapping[str, Any]) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) + les_rows = _array_rows( + documents.get("legal_effect_structures"), + ("structures", "structure_records", "legal_effect_structures", "rows", "items"), + ) + evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) + event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) + members: list[dict[str, Any]] = [] + relations: list[dict[str, Any]] = [] + internal_by_kind_ref: dict[tuple[str, str], str] = {} + + def add_member( + kind: str, + member_ref: str, + logical_id: str, + pointer: str, + raw_value: Any, + ) -> str: + key = (kind, member_ref) + prior = internal_by_kind_ref.get(key) + if prior is not None: + return prior + internal_id = f"{kind}:{member_ref}" + internal_by_kind_ref[key] = internal_id + members.append( + { + "member_id": internal_id, + "member_ref": member_ref, + "member_kind": kind, + "source_refs": [_source_ref(logical_id, pointer, raw_value, stage1_id=member_ref)], + "scope_technical_disposition": "AVAILABLE", + } + ) + return internal_id + + bo_member_by_id: dict[str, str] = {} + for index, row in enumerate(bo_rows): + if not isinstance(row, dict): + continue + bo_id = str(row.get("BO_ID", f"BO-OCCURRENCE-{index}")) + bo_member_by_id[bo_id] = add_member("BO", bo_id, "bo", f"/business_objects/{index}", row) + + fact_member_by_id: dict[str, str] = {} + fact_members_by_bo: dict[str, list[str]] = defaultdict(list) + fact_members_by_evidence: dict[str, list[str]] = defaultdict(list) + fact_members_by_event: dict[str, list[str]] = defaultdict(list) + for index, row in enumerate(ledger_rows): + if not isinstance(row, dict): + continue + fact_id = str(row.get("fact_id", f"F-OCCURRENCE-{index}")) + member_id = add_member("FACT", fact_id, "fact_ledger_base", f"/facts/{index}", row) + fact_member_by_id[fact_id] = member_id + source_bo_id = row.get("source_bo_id") + if source_bo_id is not None: + fact_members_by_bo[str(source_bo_id)].append(member_id) + evidence_refs = row.get("evidence_refs", row.get("evidence_ids", [])) + if isinstance(evidence_refs, list): + for ref in evidence_refs: + fact_members_by_evidence[str(ref)].append(member_id) + event_refs = row.get("event_refs", row.get("event_ids", [])) + if isinstance(event_refs, list): + for ref in event_refs: + fact_members_by_event[str(ref)].append(member_id) + relation_arrays = [ + row.get(key) + for key in ("relations", "explicit_relations", "candidate_relations") + if isinstance(row.get(key), list) + ] + for relation_array in relation_arrays: + for relation_index, relation in enumerate(relation_array): + if not isinstance(relation, dict): + continue + target_ref = relation.get( + "target_fact_id", + relation.get("to_fact_id", relation.get("target_member_id")), + ) + relation_kind = str(relation.get("relation_kind", relation.get("kind", ""))) + if target_ref is None or relation_kind not in CANDIDATE_RELATION_KINDS | {"EXPLICIT_CASE_RELATION"}: + continue + relations.append( + { + "_deferred_source_fact_id": fact_id, + "_deferred_target_fact_id": str(target_ref), + "relation_kind": relation_kind, + "source_refs": [ + _source_ref( + "fact_ledger_base", + f"/facts/{index}/relations/{relation_index}", + relation, + ) + ], + } + ) + + for bo_id, fact_member_ids in sorted(fact_members_by_bo.items()): + bo_member = bo_member_by_id.get(bo_id) + if bo_member is None: + continue + for fact_member in sorted(set(fact_member_ids)): + relations.append( + { + "source_member_id": fact_member, + "target_member_id": bo_member, + "relation_kind": "SAME_BO_ID", + "source_refs": [_source_ref("bo", "", bo_id, stage1_id=bo_id)], + } + ) + + for index, row in enumerate(les_rows): + if not isinstance(row, dict): + continue + structure_ref = str(row.get("structure_id", row.get("legal_effect_structure_id", f"LES-OCCURRENCE-{index}"))) + les_member = add_member( + "LES_STRUCTURE", + structure_ref, + "legal_effect_structures", + f"/structures/{index}", + row, + ) + source_bo_ids = row.get("source_bo_ids", []) + if isinstance(source_bo_ids, list): + for bo_id in source_bo_ids: + bo_member = bo_member_by_id.get(str(bo_id)) + if bo_member is not None: + relations.append( + { + "source_member_id": les_member, + "target_member_id": bo_member, + "relation_kind": "SOURCE_BO_ATTACHMENT", + "source_refs": [ + _source_ref( + "legal_effect_structures", + f"/structures/{index}/source_bo_ids", + source_bo_ids, + ) + ], + } + ) + + for index, row in enumerate(evidence_rows): + if not isinstance(row, dict): + continue + evidence_id = str(row.get("evidence_id", row.get("id", f"EVIDENCE-OCCURRENCE-{index}"))) + evidence_member = add_member("EVIDENCE", evidence_id, "evidence_indexed", f"/items/{index}", row) + for fact_member in sorted(set(fact_members_by_evidence.get(evidence_id, []))): + relations.append( + { + "source_member_id": fact_member, + "target_member_id": evidence_member, + "relation_kind": "SAME_EVIDENCE_REF", + "source_refs": [_source_ref("fact_ledger_base", "", evidence_id, stage1_id=evidence_id)], + } + ) + + for index, row in enumerate(event_rows): + if not isinstance(row, dict): + continue + event_id = str(row.get("event_id", row.get("id", f"EVENT-OCCURRENCE-{index}"))) + event_member = add_member("EVENT", event_id, "evidence_event_candidates", f"/items/{index}", row) + for fact_member in sorted(set(fact_members_by_event.get(event_id, []))): + relations.append( + { + "source_member_id": fact_member, + "target_member_id": event_member, + "relation_kind": "SAME_EVENT_REF", + "source_refs": [_source_ref("fact_ledger_base", "", event_id, stage1_id=event_id)], + } + ) + + resolved_relations: list[dict[str, Any]] = [] + for relation in relations: + if "_deferred_source_fact_id" not in relation: + resolved_relations.append(relation) + continue + source_member = fact_member_by_id.get(str(relation["_deferred_source_fact_id"])) + target_member = fact_member_by_id.get(str(relation["_deferred_target_fact_id"])) + if source_member is None or target_member is None: + continue + resolved_relations.append( + { + "source_member_id": source_member, + "target_member_id": target_member, + "relation_kind": relation["relation_kind"], + "source_refs": relation["source_refs"], + } + ) + return members, resolved_relations + + + def _run_binding_digest( + snapshots: Mapping[str, Snapshot], + release_lock: Mapping[str, Any], + ) -> str: + input_set = [[key, snapshots[key].raw_sha256] for key in sorted(snapshots)] + binding = { + "input_set_digest": canonical_digest(input_set), + "stage2_release_digest": release_lock.get("_release_raw_sha256", release_lock.get("stage2_release_digest")), + "algorithm_digest": ALGORITHM_SEMANTIC_DIGEST, + "release_class": release_lock.get("release_class"), + } + return canonical_digest(binding) + + + def _input_set_digest( + snapshots: Mapping[str, Snapshot], + additional_snapshots: Sequence[Snapshot] = (), + ) -> str: + rows = [ + ["run", key, snapshots[key].relative_path, snapshots[key].raw_sha256] + for key in sorted(snapshots) + ] + rows.extend( + ["closure", item.logical_input_id, item.relative_path, item.raw_sha256] + for item in additional_snapshots + ) + return canonical_digest(sorted(rows, key=canonical_digest)) + + + def _make_run_binding_receipt( + snapshots: Mapping[str, Snapshot], + release_lock: Mapping[str, Any], + *, + request_id: str, + user_context_sha256: str, + workspace_context_sha256: str, + additional_snapshots: Sequence[Snapshot] = (), + ) -> dict[str, Any]: + input_digest = _input_set_digest(snapshots, additional_snapshots) + release_digest = str( + release_lock.get( + "_release_raw_sha256", + release_lock.get( + "release_digest", + release_lock.get("stage2_release_digest", "0" * 64), + ), + ) + ) + algorithm_digest = ALGORITHM_SEMANTIC_DIGEST + binding_digest = canonical_digest( + { + "input_set_digest": input_digest, + "stage2_release_digest": release_digest, + "algorithm_digest": algorithm_digest, + "release_class": release_lock.get("release_class"), + } + ) + return { + "schema_version": "stage2_s2_00_run_binding_receipt.v1.1", + "request_id": request_id, + "run_id": f"S2RUN-{binding_digest}", + "input_set_digest": input_digest, + "stage2_release_digest": release_digest, + "algorithm_digest": algorithm_digest, + "release_class": release_lock.get("release_class"), + "run_binding_digest": binding_digest, + "canonical_output_root": f"stage2_runs/by-binding/{binding_digest}/", + "run_identity_derivation": "RUN_ID_PREFIXED_FROM_RUN_BINDING_DIGEST", + "output_root_derivation": "stage2_runs/by-binding//", + "user_context_sha256": user_context_sha256, + "workspace_context_sha256": workspace_context_sha256, + } + + + def canonical_run_id(run_binding_digest: str) -> str: + """Derive the immutable run identifier from the complete binding digest.""" + + if re.fullmatch(r"[a-f0-9]{64}", run_binding_digest) is None: + raise IngressError( + "RUN_BINDING_DIGEST_INVALID", + "canonical run id requires one lowercase SHA-256 digest", + ) + return f"S2RUN-{run_binding_digest}" + + + def _compact_conservation_checks( + checks: Sequence[Mapping[str, Any]], + ) -> list[dict[str, Any]]: + compact: list[dict[str, Any]] = [] + for check in checks: + check_id = str(check.get("check_id", "UNNAMED_CONSERVATION_CHECK")) + status_value = str(check.get("status", "UNEVALUABLE")) + status = status_value if status_value in {"PASS", "FAIL", "UNEVALUABLE"} else "UNEVALUABLE" + observed_payload = {key: value for key, value in check.items() if key not in {"status"}} + left_digest = canonical_digest(observed_payload) if status != "UNEVALUABLE" else None + right_digest = left_digest if status == "PASS" else canonical_digest([check_id, "EXPECTED"]) if status == "FAIL" else None + left_count = next( + ( + int(check[key]) + for key in ("left_count", "bo_count", "source_bo_ref_count") + if isinstance(check.get(key), int) + ), + None, + ) + right_count = next( + ( + int(check[key]) + for key in ("right_count", "source_bo_ref_count", "bo_count") + if isinstance(check.get(key), int) + ), + None, + ) + compact.append( + { + "check_id": check_id, + "status": status, + "left_counter_digest": left_digest, + "right_counter_digest": right_digest, + "left_count": left_count, + "right_count": right_count, + "source_refs": [check_id], + "issue_codes": [f"{check_id}_FAILED"] if status == "FAIL" else [], + } + ) + return compact + + + def _compact_review_receipt( + normalized_reviews: Mapping[str, Any], + ) -> dict[str, Any]: + normalized = normalized_reviews.get("normalized_occurrences", []) + raw = normalized_reviews.get("raw_occurrences", []) + counts = dict(normalized_reviews.get("partition_counts", {})) + for key in ("SUPPORTED", "CONDITIONAL", "UNRESOLVED", "EXCLUDED", "UNMAPPED"): + counts.setdefault(key, 0) + return { + "schema_version": "stage2_review_normalization_receipt.v1", + "adapter_ids": list(normalized_reviews.get("adapter_ids", [])), + "raw_occurrence_count": len(raw), + "normalized_occurrence_count": len(normalized), + "partition_counts": counts, + "unmapped_occurrence_count": counts["UNMAPPED"], + "conservation_status": str(normalized_reviews.get("conservation_status", "FAIL")), + "ordered_occurrence_refs": [ + str(row.get("review_key", {}).get("value")) + for row in normalized + if row.get("review_key", {}).get("value") + ], + } + + + def _build_issue_ledger( + issues: Sequence[Mapping[str, Any]], + normalized_reviews: Mapping[str, Any], + *, + run_id: str, + input_set_digest: str, + ) -> dict[str, Any]: + issue_rows: list[dict[str, Any]] = [] + seen: set[str] = set() + for index, issue in enumerate(issues): + issue_code = str(issue.get("issue_code", "UNSPECIFIED_ISSUE")) + source_refs = sorted(set(str(item) for item in issue.get("scope_refs", []))) or ["S2_00"] + issue_mint = mint_stage2_id( + "issue", + [issue_code, source_refs, index], + prefix="ISS", + ) + if issue_mint["id"] in seen: + continue + seen.add(issue_mint["id"]) + raw_severity = str(issue.get("severity", "ERROR")).upper() + technical_severity = { + "INFO": "INFO", + "REVIEW": "REVIEW", + "WARNING": "HARD_WARNING", + "HARD_WARNING": "HARD_WARNING", + "ERROR": "BLOCK", + "BLOCK": "BLOCK", + }.get(raw_severity, "REVIEW") + issue_rows.append( + { + "issue_id": issue_mint["id"], + "issue_code": issue_code, + "technical_severity": technical_severity, + "review_partition": "UNRESOLVED", + "impact_scope": str(issue.get("impact_scope", "GLOBAL")), + "scope_refs": source_refs, + "source_contract_row_refs": sorted( + set(str(item) for item in issue.get("source_contract_row_refs", source_refs)) + ), + "reason_codes": sorted(set(str(item) for item in issue.get("reason_codes", [issue_code]))), + "downstream_allowed_actions": sorted( + set(str(item) for item in issue.get("downstream_allowed_actions", [])) + ), + "source_refs": source_refs, + "status": "OPEN", + "upstream_review_key": None, + "raw_item_sha256": None, + } + ) + for occurrence in normalized_reviews.get("normalized_occurrences", []): + if occurrence.get("partition") != "UNMAPPED": + continue + review_key = occurrence["review_key"]["value"] + issue_mint = mint_stage2_id( + "issue", + ["UNMAPPED_REVIEW_STATUS", review_key], + prefix="ISS", + ) + issue_rows.append( + { + "issue_id": issue_mint["id"], + "issue_code": "UNMAPPED_REVIEW_STATUS", + "technical_severity": "REVIEW", + "review_partition": "UNMAPPED", + "impact_scope": "REVIEW_ITEM", + "scope_refs": [review_key], + "source_contract_row_refs": occurrence["source_contract_row_refs"], + "reason_codes": ["UNMAPPED_REVIEW_STATUS"], + "downstream_allowed_actions": occurrence["downstream_allowed_actions"], + "source_refs": occurrence["source_contract_row_refs"], + "status": "OPEN", + "upstream_review_key": review_key, + "raw_item_sha256": occurrence["source_item_raw_sha256"], + } + ) + return { + "schema_version": "stage2_issue_ledger_base.v1", + "producer_id": "S2_00", + "run_id": run_id, + "input_set_digest": input_set_digest, + "mapping_table_version": "S2-REVIEW-MAP-V1", + "issues": sorted(issue_rows, key=lambda row: row["issue_id"]), + "raw_review_occurrence_count": len(normalized_reviews.get("raw_occurrences", [])), + "normalized_review_occurrence_count": len(normalized_reviews.get("normalized_occurrences", [])), + "review_conservation_status": str(normalized_reviews.get("conservation_status", "FAIL")), + } + + + def _route( + source_rows: Sequence[Mapping[str, Any]], + cluster_plan: Mapping[str, Any], + cohort_receipt: Mapping[str, Any], + issues: Sequence[Mapping[str, Any]], + ) -> str: + identity_unavailable = any( + row.get("requirement_class") == "IDENTITY_BACKBONE" + and row.get("scope_technical_disposition") == "UNAVAILABLE" + for row in source_rows + ) + diagnostic_issue_codes = { + "RELEASE_SOURCE_CONTRACT_MISSING", + "RELEASE_SOURCE_PATH_MISMATCH", + "RAW_HASH_MISMATCH", + "SCHEMA_HASH_MISMATCH", + "SCHEMA_ID_MISMATCH", + "RUN_IDENTITY_CONFLICT", + "TRANSACTION_IDENTITY_CONFLICT", + "BO_FACT_CONSERVATION_FAILED", + "FACT_ID_CONSERVATION_FAILED", + "CURRENT_V8_LEDGER_EXTENSION_MISSING", + } + observed_issue_codes = {str(item.get("issue_code")) for item in issues} + if identity_unavailable or observed_issue_codes & diagnostic_issue_codes or not cluster_plan.get("clusters"): + return "TO_S2_40_STATUS_ONLY" + if not cohort_receipt.get("executable_cluster_ids"): + return "TO_S2_40_STATUS_ONLY" + return "TO_S2_10_WITH_ISSUES" if issues else "TO_S2_10" + + + def execute_ingress( + stage1_run_root: str | os.PathLike[str], + release_lock: Mapping[str, Any], + *, + output_dir: str | os.PathLike[str] | None, + attempt_id: str, + run_id: str = "S2-00-REQUEST", + user_context_sha256: str = "0" * 64, + workspace_context_sha256: str = "0" * 64, + stage1_deployment_root: str | os.PathLike[str] | None = None, + stage2_asset_root: str | os.PathLike[str] | None = None, + contract_manifest: Mapping[str, Any] | None = None, + hydration_stability_receipt: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Execute C00→C05→C10→C15 for a canary or production case run.""" + + if release_lock.get("release_class") == "DEV_FIXTURE_RELEASE": + raise IngressError("DEV_FIXTURE_REAL_RUN_FORBIDDEN", "DEV fixture release cannot publish a case run") + if output_dir is None: + raise IngressError("OUTPUT_DIR_REQUIRED", "output_dir is required for canary/production") + if stage1_deployment_root is None: + raise IngressError("STAGE1_DEPLOYMENT_ROOT_REQUIRED", "Stage 1 deployment root is required") + if not re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}", run_id): + raise IngressError("REQUEST_ID_INVALID", "request_id contains forbidden characters") + if not isinstance(hydration_stability_receipt, dict): + raise IngressError( + "HYDRATION_STABILITY_RECEIPT_REQUIRED", + "canary and production ingress require the remote two-pass hydration receipt", + ) + for label, value in ( + ("user_context_sha256", user_context_sha256), + ("workspace_context_sha256", workspace_context_sha256), + ): + if not re.fullmatch(r"[A-Fa-f0-9]{64}", value): + raise IngressError("RUN_CONTEXT_HASH_INVALID", f"{label} must be a SHA-256 digest") + root_arg = Path(stage1_run_root) + deployment_root_arg = Path(stage1_deployment_root) + asset_root_arg = Path(stage2_asset_root) if stage2_asset_root is not None else Path(__file__).resolve().parents[1] + for candidate, code in ( + (root_arg, "SYMLINK_ROOT_REJECTED"), + (deployment_root_arg, "SYMLINK_DEPLOYMENT_ROOT_REJECTED"), + (asset_root_arg, "SYMLINK_ASSET_ROOT_REJECTED"), + ): + if candidate.is_symlink(): + raise IngressError(code, "approved root itself may not be a symlink") + root = root_arg.resolve(strict=True) + deployment_root = deployment_root_arg.resolve(strict=True) + asset_root = ( + asset_root_arg.resolve(strict=True) + if stage2_asset_root is not None + else Path(__file__).resolve().parents[1] + ) + deployment_snapshots, deployment_documents, deployment_rows = _load_stage1_deployment_closure( + deployment_root, + release_lock, + ) + contracts = resolve_stage1_sources(root, contract_manifest) + snapshots, snapshot_issues = _load_case_snapshots(root, contracts, release_lock) + completion_seal, completion_seal_snapshot = _load_bound_completion_seal(root, release_lock) + ingress = validate_ingress_contracts( + snapshots, + contracts, + release_lock, + deployment_snapshots=deployment_snapshots, + deployment_documents=deployment_documents, + completion_seal=completion_seal, + contract_manifest=contract_manifest, + ) + documents = ingress["documents"] + issues = list(snapshot_issues) + list(ingress["issues"]) + signal_all: dict[str, Any] | None = None + if isinstance(documents.get("signal_manifest"), dict): + try: + max_run = int(release_lock.get("limits", {}).get("max_run_bytes", MAX_RUN_BYTES)) + already_loaded = sum(snapshot.byte_length for snapshot in snapshots.values()) + sum( + snapshot.byte_length for snapshot in deployment_snapshots.values() + ) + remaining = max_run - already_loaded + if remaining < 0: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "case and deployment closure exceed aggregate run budget") + signal_all = expand_stage2_signal_all( + root, + documents["signal_manifest"], + max_file_bytes=int(release_lock.get("limits", {}).get("max_file_bytes", MAX_FILE_BYTES)), + max_total_bytes=remaining, + signal_registry=deployment_documents.get("signals/signal_registry.v2.json"), + ) + signal_all = bind_signal_occurrences(signal_all, documents) + issues.extend(signal_all.get("issues", [])) + except IngressError as exc: + issues.append(_issue(exc.code, impact_scope="SIGNAL", source_refs=["signal_manifest"], message=str(exc))) + routing_activation = documents.get("domain_activation_manifest") + if isinstance(routing_activation, dict): + try: + parsed_signal_documents = signal_all.get("_parsed_documents_by_path", {}) if signal_all else {} + signal_activation = parsed_signal_documents.get("domain_activation_manifest.json") + if signal_activation is None: + raise IngressError( + "SG01_SIGNAL_ARTIFACT_MISSING", + "signal ALL does not contain domain_activation_manifest.json", + ) + signal_activation_row = next( + ( + row + for row in signal_all.get("ordered_file_rows", []) + if row.get("file_path") == "domain_activation_manifest.json" + ), + {}, + ) + verify_activation_projection( + routing_activation, + signal_activation, + routing_raw_sha256=snapshots.get("domain_activation_manifest").raw_sha256 if snapshots.get("domain_activation_manifest") else None, + signal_raw_sha256=signal_activation_row.get("raw_sha256"), + ) + except IngressError as exc: + issues.append(_issue(exc.code, impact_scope="SIGNAL", source_refs=["domain_activation_manifest", "signal_sg01_activation"], message=str(exc))) + seal_deployment_snapshots = { + "stage1_domain_registry_index": snapshot + for lock_id, snapshot in deployment_snapshots.items() + if snapshot.relative_path == "domains/_registry_index.json" + } + seals = verify_cross_artifact_seals(documents, snapshots, seal_deployment_snapshots) + issues.extend(seals["issues"]) + review_docs = {key: value for key, value in documents.items() if "review_handoff" in key or "soft_gate_handoff" in key} + normalized_reviews = normalize_review_items(review_docs, release_lock) + issues.extend(normalized_reviews.get("_issues", [])) + conservation = check_conservation( + documents, + signal_all=signal_all, + normalized_reviews=normalized_reviews, + source_snapshots=snapshots, + ) + issues.extend(conservation["issues"]) + signal_snapshots = list(signal_all.get("_payload_snapshots", [])) if signal_all else [] + closure_snapshots = list(deployment_snapshots.values()) + signal_snapshots + all_snapshots = list(snapshots.values()) + closure_snapshots + if completion_seal_snapshot is not None: + all_snapshots.append(completion_seal_snapshot) + snapshot_start_digest = _snapshot_state_digest(all_snapshots) + input_set_digest = _input_set_digest(snapshots, closure_snapshots) + run_binding_receipt = _make_run_binding_receipt( + snapshots, + release_lock, + request_id=run_id, + user_context_sha256=user_context_sha256, + workspace_context_sha256=workspace_context_sha256, + additional_snapshots=closure_snapshots, + ) + run_binding_digest = run_binding_receipt["run_binding_digest"] + run_id = canonical_run_id(run_binding_digest) + context_schema_snapshot = open_bounded_snapshot( + asset_root, + "schemas/context.schema.json", + logical_input_id="stage2_context_schema", + ) + artifact_header = _artifact_header( + schema_sha256=context_schema_snapshot.raw_sha256, + run_id=run_id, + input_set_digest=input_set_digest, + stage2_release_digest=run_binding_receipt["stage2_release_digest"], + algorithm_digest=run_binding_receipt["algorithm_digest"], + release_class=str(release_lock["release_class"]), + ) + domain_configs = { + match.group(1): value + for path, value in deployment_documents.items() + if (match := re.fullmatch(r"domains/([^/]+)/domain_config\.json", path)) is not None + } + routing_payload = ( + _activation_payload(documents["domain_activation_manifest"]) + if isinstance(documents.get("domain_activation_manifest"), dict) + else {} + ) + active_domain_ids = routing_payload.get("active_domain_ids", []) + if isinstance(active_domain_ids, list): + for domain_id in sorted(set(str(value) for value in active_domain_ids)): + if domain_id not in domain_configs: + issues.append( + _issue( + "ACTIVE_DOMAIN_CONFIG_MISSING", + impact_scope="CLUSTER", + source_refs=[f"domain_config:{domain_id}"], + ) + ) + case_context = build_case_context( + documents, + run_binding_digest=run_binding_digest, + issues=issues, + artifact_header=artifact_header, + signal_all=signal_all, + ) + evidence_inventory = build_evidence_inventory( + documents, + run_binding_digest=run_binding_digest, + artifact_header=artifact_header, + ) + object_registry = build_object_registry( + documents, + run_binding_digest=run_binding_digest, + artifact_header=artifact_header, + ) + party_context = build_party_and_title_context( + documents, + run_binding_digest=run_binding_digest, + artifact_header=artifact_header, + ) + slot_crosswalk = build_slot_crosswalk( + documents, + domain_configs, + run_binding_digest=run_binding_digest, + artifact_header=artifact_header, + ) + members, relations = _cluster_inputs(documents) + cluster_plan = compile_cluster_plan(members, relations, artifact_header=artifact_header) + contexts = { + "case_context": case_context, + "evidence_inventory": evidence_inventory, + "object_registry": object_registry, + "party_and_title_context": party_context, + "slot_crosswalk": slot_crosswalk, + "_projection_sources": { + "EVENT": { + str(row.get("event_id", row.get("id"))): row + for row in _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) + if isinstance(row, dict) and (row.get("event_id") is not None or row.get("id") is not None) + }, + "LES_STRUCTURE": { + str(row.get("structure_id", row.get("legal_effect_structure_id"))): row + for row in _array_rows(documents.get("legal_effect_structures"), ("structures", "structure_records", "rows", "items")) + if isinstance(row, dict) and (row.get("structure_id") is not None or row.get("legal_effect_structure_id") is not None) + }, + "SIGNAL_OCCURRENCE": { + str(row.get("occurrence_ref")): row + for row in (signal_all or {}).get("record_occurrences", []) + if row.get("occurrence_ref") is not None + }, + "REVIEW_ITEM": { + str(row.get("review_key", {}).get("value")): row + for row in normalized_reviews.get("normalized_occurrences", []) + if row.get("review_key", {}).get("value") is not None + }, + "ACTIVE_PROFILE": { + str(domain_id): config + for domain_id, config in domain_configs.items() + if domain_id in set(str(value) for value in active_domain_ids) + }, + }, + } + slices = compile_cluster_slices(cluster_plan, contexts) + effective_release_lock = dict(release_lock) + if not isinstance(effective_release_lock.get("bundle"), dict): + issues.append(_issue("BUNDLE_RELEASE_CONTRACT_MISSING", impact_scope="GLOBAL", source_refs=["stage2_release"])) + effective_release_lock["bundle"] = { + "mode": { + "DEV_FIXTURE_RELEASE": "STRUCTURAL_FIXTURE", + "SUBSET_CANARY_RELEASE": "SUBSET_CANARY", + "PRODUCTION_RELEASE": "PRODUCTION", + }.get(str(release_lock.get("release_class")), "STRUCTURAL_FIXTURE"), + "handoff_contract_version": "S2_00_STRUCTURED_CONTEXT_HANDOFF_V2", + "context_cohort_policy_id": "S2-CACHE-STRUCTURED-CONTEXT-HANDOFF-V2", + "s2_10_agent_ref": { + "path": "agent_scripts/Stage_2_S2_10.yml", + "sha256": "0" * 64, + }, + "s2_10_llm_binding_ref": { + "path": "deployment/stage2_s2_10_llm_binding.yml", + "sha256": "0" * 64, + }, + "s2_10_agent_sha256": "0" * 64, + "s2_10_llm_binding_sha256": "0" * 64, + "selected_context_refs": [], + "selected_context_ordered_refs_sha256": canonical_digest([]), + "downstream_hybrid_status": "PENDING_REIMPLEMENTATION_AND_RESEAL", + } + bundle_plan = compile_bundle_plan( + cluster_plan, + slices, + effective_release_lock, + stage2_asset_root=asset_root, + ) + cohort_by_cluster = { + cluster_id: row + for row in bundle_plan.get("cohorts", []) + for cluster_id in row.get("cohort_member_cluster_ids", []) + } + for cluster in cluster_plan.get("clusters", []): + cohort = cohort_by_cluster.get(cluster["cluster_id"]) + if cohort is not None: + cluster["bundle_cohort_id"] = cohort["bundle_cohort_id"] + if cohort["cohort_status"] != "EXECUTABLE" and cluster["cluster_status"] == "EXECUTABLE": + cluster["cluster_status"] = "NON_EXECUTABLE" + for code in cohort.get("reason_codes", []): + issues.append( + _issue( + str(code), + impact_scope="CLUSTER", + source_refs=[cluster["cluster_id"], cohort["bundle_cohort_id"]], + ) + ) + cohort_receipt = validate_bundle_release_cohorts(bundle_plan) + final_executable = cohort_receipt["executable_cluster_ids"] + final_residual = sorted( + set(cluster["cluster_id"] for cluster in cluster_plan.get("clusters", [])) + - set(final_executable) + ) + cluster_plan["executable_cluster_ids"] = final_executable + cluster_plan["residual_review_cluster_ids"] = final_residual + executable_scc_ids = { + scc_id + for cluster in cluster_plan.get("clusters", []) + if cluster["cluster_id"] in set(final_executable) + for scc_id in cluster.get("scc_ids", []) + } + cluster_plan["scheduling_waves"] = [ + [scc_id for scc_id in wave if scc_id in executable_scc_ids] + for wave in cluster_plan.get("scheduling_waves", []) + if any(scc_id in executable_scc_ids for scc_id in wave) + ] + route = _route(ingress["source_contract_rows"], cluster_plan, cohort_receipt, issues) + snapshot_end_digest = _verify_snapshot_state(all_snapshots) + if snapshot_end_digest != snapshot_start_digest: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", "source closure changed after the one-read snapshot") + source_rows = list(ingress["source_contract_rows"]) + if signal_all: + source_rows.extend(_signal_source_contract_rows(signal_all, release_lock)) + source_rows.extend(deployment_rows) + source_rows = sorted(source_rows, key=lambda row: row["logical_input_id"]) + issue_codes = sorted(set(str(item["issue_code"]) for item in issues)) + source_counts = { + "declared": len(source_rows), + "observed": sum(row["raw_sha256"] is not None for row in source_rows), + "available": sum(row["scope_technical_disposition"] == "AVAILABLE" for row in source_rows), + "with_issues": sum(row["scope_technical_disposition"] == "AVAILABLE_WITH_ISSUES" for row in source_rows), + "unavailable": sum(row["scope_technical_disposition"] == "UNAVAILABLE" for row in source_rows), + } + intake_report = { + "schema_version": "stage2_s2_00_intake_report.v1", + "run_id": run_id, + "source_counts": source_counts, + "conservation_checks": _compact_conservation_checks(conservation["checks"]), + "review_normalization_receipt": _compact_review_receipt(normalized_reviews), + "minimum_coherent_package_possible": route != "TO_S2_40_STATUS_ONLY", + "scope_summary": [ + { + "scope_technical_disposition": row["scope_technical_disposition"], + "impact_scope": row["impact_scope"], + "scope_refs": row["scope_refs"], + "source_contract_row_refs": row["source_contract_row_refs"], + "reason_codes": row["reason_codes"], + "downstream_allowed_actions": row["downstream_allowed_actions"], + } + for row in source_rows + ], + "issue_codes": issue_codes, + } + manifest = { + "schema_version": "stage2_s2_00_stage1_input_manifest.v1.1", + "run_id": run_id, + "release_class": release_lock["release_class"], + "source_rows": source_rows, + "source_row_order": [row["logical_input_id"] for row in source_rows], + "signal_all_adapter_id": "S2A-SIGNAL-ALL-V1", + "dual_sg01_adapter_id": "S2A-DUAL-SG01-V1", + "hydration_stability_receipt": dict(hydration_stability_receipt), + "snapshot_start_digest": snapshot_start_digest, + "snapshot_end_digest": snapshot_end_digest, + "snapshot_status": "STABLE", + "input_set_digest": input_set_digest, + } + issue_ledger = _build_issue_ledger( + issues, + normalized_reviews, + run_id=run_id, + input_set_digest=input_set_digest, + ) + status = { + "schema_version": "stage2_s2_00_ingress_status.v1.1", + "run_id": run_id, + "run_binding_receipt": run_binding_receipt, + "route": route, + "executable_cluster_ids": final_executable if route != "TO_S2_40_STATUS_ONLY" else [], + "residual_review_cluster_ids": final_residual, + "issue_codes": issue_codes, + "output_barrier": { + "barrier_id": "S2_00_INGRESS_STATUS_BARRIER", + "barrier_path": "ingress/ingress_status.json", + "publish_semantics": "STATUS_LAST_LOGICAL_COMMIT", + "canonical_output_root": f"stage2_runs/by-binding/{run_binding_digest}/", + "branch": "DIAGNOSTIC" if route == "TO_S2_40_STATUS_ONLY" else "NORMAL", + "artifacts": [], + "artifact_set_digest": canonical_digest([]), + "non_status_artifacts_read_back_verified": True, + "written_last": True, + }, + } + if route == "TO_S2_40_STATUS_ONLY": + artifacts: dict[str, Any] = { + "ingress/stage1_input_manifest.json": manifest, + "ingress/intake_report.json": intake_report, + "ingress/technical_diagnostic.json": { + "schema_version": "stage2_s2_00_technical_diagnostic.v1", + "run_id": run_id, + "route": "TO_S2_40_STATUS_ONLY", + "minimum_coherent_package_possible": False, + "reason_codes": issue_codes or ["MINIMUM_COHERENT_PACKAGE_UNAVAILABLE"], + "source_contract_row_refs": sorted( + set( + ref + for item in issues + for ref in item.get("source_contract_row_refs", []) + ) + ) or ["S2_00"], + "context_published": False, + }, + "review/issue_ledger.base.json": issue_ledger, + "ingress/ingress_status.json": status, + } + else: + artifacts = { + "ingress/stage1_input_manifest.json": manifest, + "ingress/intake_report.json": intake_report, + "context/case_context.json": case_context, + "context/evidence_inventory.json": evidence_inventory, + "context/object_registry.json": object_registry, + "context/party_and_title_context.json": party_context, + "context/slot_crosswalk.json": slot_crosswalk, + "context/cluster_plan.json": cluster_plan, + "context/bundle_plan.json": bundle_plan, + "review/issue_ledger.base.json": issue_ledger, + "ingress/ingress_status.json": status, + } + for cluster_id, slice_body in slices.items(): + if cluster_id in set(final_executable): + artifacts[f"context/cluster_slices/{cluster_id}.json"] = slice_body + publish_receipt = publish_atomically( + output_dir, + artifacts, + run_binding_digest=run_binding_digest, + attempt_id=attempt_id, + stage2_asset_root=asset_root, + ) + return {"route": route, "run_binding_digest": run_binding_digest, "publish": publish_receipt} + + + def _execute_structural_fixture(descriptor_path: Path, release_lock: Mapping[str, Any]) -> dict[str, Any]: + snapshot = open_bounded_snapshot(descriptor_path.parent, descriptor_path.name, logical_input_id="fixture_descriptor") + descriptor = load_json_strict(snapshot) + if not isinstance(descriptor, dict) or descriptor.get("mode") != "STRUCTURAL_FIXTURE": + raise IngressError("FIXTURE_DESCRIPTOR_SHAPE", "fixture descriptor must declare STRUCTURAL_FIXTURE") + if release_lock.get("release_class") != "DEV_FIXTURE_RELEASE": + raise IngressError("FIXTURE_RELEASE_CLASS", "structural fixtures require DEV_FIXTURE_RELEASE") + required = {"fixture_id", "source_locator", "source_hashes", "source_semantics", "expected"} + missing = sorted(required - set(descriptor)) + if missing: + raise IngressError("FIXTURE_DESCRIPTOR_REQUIRED_FIELD", "fixture fields are missing", details={"missing": missing}) + if descriptor.get("source_semantics") not in {"EXPLICIT", "DERIVED"}: + raise IngressError("FIXTURE_SOURCE_SEMANTICS", "source_semantics must be EXPLICIT or DERIVED") + payloads = descriptor.get("fixture_payloads", {}) + source_hashes = descriptor.get("source_hashes", {}) + if not isinstance(payloads, dict) or not isinstance(source_hashes, dict): + raise IngressError("FIXTURE_SOURCE_HASH_SHAPE", "fixture payloads and source hashes must be objects") + for logical_id, payload in payloads.items(): + expected_hash = source_hashes.get(logical_id) + if not isinstance(expected_hash, str) or expected_hash != canonical_digest(payload): + raise IngressError( + "FIXTURE_SOURCE_HASH_MISMATCH", + f"fixture payload hash mismatch for {logical_id}", + ) + return { + "status": "STRUCTURAL_FIXTURE_VALIDATED", + "fixture_id": descriptor["fixture_id"], + "descriptor_sha256": snapshot.raw_sha256, + "canonical_descriptor_sha256": canonical_digest(descriptor), + "published": False, + } + + + def _inline_sha256(value: str, *, code: str) -> str: + if not isinstance(value, str) or re.fullmatch(r"[a-f0-9]{64}", value) is None: + raise IngressError(code, "expected one lowercase SHA-256 digest") + return value + + + def _inline_relative_path(value: str, *, code: str) -> str: + if not isinstance(value, str) or not value or "\x00" in value or "\\" in value: + raise IngressError(code, "logical path is empty or malformed") + if unicodedata.normalize("NFC", value) != value: + raise IngressError(code, "logical path must already be NFC") + path = PurePosixPath(value) + if path.is_absolute() or any(part in {"", ".", ".."} for part in path.parts): + raise IngressError(code, "logical path must be a contained relative path") + rendered = path.as_posix() + if rendered != value: + raise IngressError(code, "logical path is not canonical") + return rendered + + + def _inline_join(root: str, relative: str, *, code: str) -> str: + safe_root = _inline_relative_path(root, code=code) + safe_relative = _inline_relative_path(relative, code=code) + return _inline_relative_path(f"{safe_root}/{safe_relative}", code=code) + + + def _inline_parse_mcp_payload(raw: bytes, expected_id: int) -> Mapping[str, Any]: + """Parse one JSON or SSE JSON-RPC terminal response with an exact ID.""" + + candidates: list[Any] + try: + candidates = [load_json_strict(raw)] + except IngressError: + try: + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise IngressError("MCP_RESPONSE_UTF8", "MCP response is not strict UTF-8") from exc + events: list[bytes] = [] + data_lines: list[str] = [] + for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n"): + if line == "": + if data_lines: + events.append("\n".join(data_lines).encode("utf-8")) + data_lines = [] + continue + if line.startswith(":") or line.startswith("event:") or line.startswith("id:") or line.startswith("retry:"): + continue + if not line.startswith("data:"): + raise IngressError("MCP_SSE_SHAPE", "unexpected non-data SSE line") + payload = line[5:] + if payload.startswith(" "): + payload = payload[1:] + data_lines.append(payload) + if data_lines: + events.append("\n".join(data_lines).encode("utf-8")) + if not events: + raise IngressError("MCP_RESPONSE_SHAPE", "MCP response contains no JSON terminal event") + candidates = [load_json_strict(event) for event in events] + matching = [ + item + for item in candidates + if isinstance(item, dict) and item.get("id") == expected_id + ] + if len(matching) != 1: + raise IngressError( + "MCP_RESPONSE_ID_MISMATCH", + "MCP response must contain exactly one terminal result with the request ID", + ) + response = matching[0] + if response.get("jsonrpc") != "2.0": + raise IngressError("MCP_JSONRPC_VERSION", "MCP response jsonrpc must equal 2.0") + if response.get("error") is not None: + raise IngressError( + "MCP_JSONRPC_ERROR", + "MCP server returned a JSON-RPC error", + details={"rpc_error": response.get("error")}, + ) + if "result" not in response or not isinstance(response["result"], dict): + raise IngressError("MCP_RESULT_SHAPE", "MCP response result must be an object") + return response + + + def _inline_tool_text(result: Mapping[str, Any], tool_name: str) -> str: + if result.get("isError") is True: + content = result.get("content") + rendered = canonical_json_bytes(content).decode("utf-8", errors="replace") if content is not None else "" + lowered = rendered.lower() + code = ( + "LOCALDOCS_NOT_FOUND" + if any(marker in lowered for marker in ("not found", "does not exist", "no such file")) + else "MCP_TOOL_ERROR" + ) + raise IngressError(code, f"localdocs {tool_name} returned isError=true") + content = result.get("content") + if not isinstance(content, list) or len(content) != 1: + raise IngressError("MCP_CONTENT_CARDINALITY", "MCP tool result must contain exactly one content block") + block = content[0] + if not isinstance(block, dict) or block.get("type") != "text" or not isinstance(block.get("text"), str): + raise IngressError("MCP_CONTENT_SHAPE", "MCP tool result must contain one text block") + return block["text"] + + + def _inline_binary_envelope(text: str, logical_path: str) -> bytes: + value = load_json_strict(text) + if isinstance(value, dict) and "results" in value: + results = value.get("results") + if not isinstance(results, list) or len(results) != 1 or not isinstance(results[0], dict): + raise IngressError("LOCALDOCS_RESULT_CARDINALITY", "binary response must contain one result row") + inner: Any = results[0].get("content", results[0].get("text")) + value = load_json_strict(inner) if isinstance(inner, str) else inner + if not isinstance(value, dict) or not isinstance(value.get("content_base64"), str): + raise IngressError("LOCALDOCS_BINARY_ENVELOPE", "binary response lacks content_base64") + try: + payload = base64.b64decode(value["content_base64"].encode("ascii"), validate=True) + except (UnicodeEncodeError, binascii.Error, ValueError) as exc: + raise IngressError("LOCALDOCS_BASE64_INVALID", "binary response is not strict base64") from exc + declared_size = value.get("byte_length", value.get("size")) + if declared_size is not None and (not isinstance(declared_size, int) or declared_size != len(payload)): + raise IngressError("LOCALDOCS_BYTE_LENGTH_MISMATCH", f"binary length mismatch: {logical_path}") + declared_hash = value.get("sha256") + if declared_hash is not None and declared_hash != hashlib.sha256(payload).hexdigest(): + raise IngressError("LOCALDOCS_HASH_MISMATCH", f"binary hash mismatch: {logical_path}") + return payload + + + class _InlineLocaldocs: + """Minimal user/workspace-bound localdocs JSON-RPC client.""" + + def __init__( + self, + user_hash: str, + workspace_hash: str, + *, + client: Any | None = None, + timeout_seconds: int = 60, + ) -> None: + self.user_hash = _inline_sha256(user_hash, code="USER_CONTEXT_HASH_INVALID") + self.workspace_hash = _inline_sha256( + workspace_hash, + code="WORKSPACE_CONTEXT_HASH_INVALID", + ) + if client is None: + try: + import httpx # type: ignore + except ImportError as exc: + raise IngressError("HTTPX_UNAVAILABLE", "Code Executor must supply httpx==0.28.1") from exc + client = httpx.Client(timeout=timeout_seconds) + self.client = client + self.headers = { + "Content-Type": "application/json", + "Accept": "application/json, text/event-stream", + } + self._message_ids = itertools.count(10) + self._initialized = False + self._session_id: str | None = None + + def close(self) -> None: + close = getattr(self.client, "close", None) + if callable(close): + close() + + def _post(self, body: Mapping[str, Any], expected_id: int | None) -> Mapping[str, Any] | None: + try: + response = self.client.post(LOCALDOCS_URL, json=dict(body), headers=dict(self.headers)) + response.raise_for_status() + except Exception as exc: + raise IngressError("MCP_TRANSPORT_ERROR", "localdocs transport failed") from exc + session_id = response.headers.get("mcp-session-id") + if session_id: + if not isinstance(session_id, str) or not session_id.strip(): + raise IngressError("MCP_SESSION_ID_INVALID", "localdocs returned an invalid session ID") + normalized_session_id = session_id.strip() + if self._session_id is None: + if expected_id != 1: + raise IngressError( + "MCP_SESSION_ID_OUTSIDE_INITIALIZE", + "localdocs first bound a session outside initialize", + ) + self._session_id = normalized_session_id + elif normalized_session_id != self._session_id: + raise IngressError( + "MCP_SESSION_ID_CHANGED", + "localdocs changed the initialized session ID", + ) + self.headers["mcp-session-id"] = self._session_id + if expected_id is None: + return None + raw = response.content if isinstance(response.content, bytes) else bytes(response.content) + return _inline_parse_mcp_payload(raw, expected_id) + + def initialize(self) -> None: + response = self._post( + { + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": MCP_PROTOCOL_VERSION, + "capabilities": {}, + "clientInfo": { + "name": INLINE_CLIENT_NAME, + "version": INLINE_CLIENT_VERSION, + "user_id": self.user_hash, + "workspace_id": self.workspace_hash, + }, + }, + }, + 1, + ) + if response is None: + raise IngressError("MCP_INITIALIZE_EMPTY", "localdocs initialize returned no result") + result = response.get("result") + if not isinstance(result, dict) or result.get("protocolVersion") != MCP_PROTOCOL_VERSION: + raise IngressError( + "MCP_PROTOCOL_VERSION_MISMATCH", + "localdocs did not negotiate the requested MCP protocol version", + ) + if self._session_id is None or "mcp-session-id" not in self.headers: + raise IngressError("MCP_SESSION_ID_MISSING", "localdocs initialize did not bind a session ID") + self._post( + {"jsonrpc": "2.0", "method": "notifications/initialized"}, + None, + ) + self._initialized = True + + def call(self, tool_name: str, arguments: Mapping[str, Any]) -> Mapping[str, Any]: + if not self._initialized: + raise IngressError("MCP_NOT_INITIALIZED", "localdocs session is not initialized") + message_id = next(self._message_ids) + response = self._post( + { + "jsonrpc": "2.0", + "id": message_id, + "method": "tools/call", + "params": {"name": tool_name, "arguments": dict(arguments)}, + }, + message_id, + ) + if response is None: + raise IngressError("MCP_TOOL_EMPTY", f"localdocs {tool_name} returned no result") + return response["result"] + + def read_binary(self, logical_path: str) -> bytes: + path = _inline_relative_path(logical_path, code="LOCALDOCS_READ_PATH_INVALID") + result = self.call("read_binary_doc", {"doc_name": path}) + return _inline_binary_envelope(_inline_tool_text(result, "read_binary_doc"), path) + + def read_binary_optional(self, logical_path: str) -> bytes | None: + try: + return self.read_binary(logical_path) + except IngressError as exc: + if exc.code == "LOCALDOCS_NOT_FOUND": + return None + raise + + def write_binary_verified(self, logical_path: str, payload: bytes) -> str: + path = _inline_relative_path(logical_path, code="LOCALDOCS_WRITE_PATH_INVALID") + encoded = base64.b64encode(payload).decode("ascii") + result = self.call( + "write_binary_file", + {"path": path, "content_base64": encoded, "overwrite": True}, + ) + _inline_tool_text(result, "write_binary_file") + observed = self.read_binary(path) + if observed != payload: + raise IngressError("LOCALDOCS_WRITE_READBACK_MISMATCH", f"read-back mismatch: {path}") + return hashlib.sha256(observed).hexdigest() + + + def _inline_validate_request(raw: bytes) -> dict[str, str]: + value = load_json_strict(raw) + if not isinstance(value, dict): + raise IngressError("RUN_REQUEST_SHAPE", "S2_00 request must be an object") + required = { + "schema_version", + "workflow_id", + "request_id", + "attempt_id", + "stage1_run_root_ref", + "stage1_deployment_root_ref", + } + if set(value) != required: + raise IngressError( + "RUN_REQUEST_CLOSED_SHAPE", + "S2_00 request has missing or unknown keys", + details={"missing": sorted(required - set(value)), "extra": sorted(set(value) - required)}, + ) + if value.get("schema_version") != "stage2_s2_00_execution_request.v1": + raise IngressError("RUN_REQUEST_SCHEMA_VERSION", "unsupported S2_00 request schema") + if value.get("workflow_id") != "S2_00": + raise IngressError("WORKFLOW_ID_MISMATCH", "S2_00 request workflow_id must equal S2_00") + request_id = value.get("request_id") + attempt_id = value.get("attempt_id") + if not isinstance(request_id, str) or re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}", request_id) is None: + raise IngressError("REQUEST_ID_INVALID", "request_id contains forbidden characters") + if not isinstance(attempt_id, str) or re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9._-]{0,127}", attempt_id) is None: + raise IngressError("ATTEMPT_ID_INVALID", "attempt_id contains forbidden characters") + run_root = _inline_relative_path( + value.get("stage1_run_root_ref"), + code="STAGE1_RUN_ROOT_REF_INVALID", + ) + deployment_root = _inline_relative_path( + value.get("stage1_deployment_root_ref"), + code="STAGE1_DEPLOYMENT_ROOT_REF_INVALID", + ) + return { + "schema_version": value["schema_version"], + "workflow_id": value["workflow_id"], + "request_id": request_id, + "attempt_id": attempt_id, + "stage1_run_root_ref": run_root, + "stage1_deployment_root_ref": deployment_root, + } + + + def _inline_bound_ref(row: Mapping[str, Any], *, code: str) -> tuple[str, str]: + path = _inline_relative_path(row.get("path"), code=code) + digest = _inline_sha256(str(row.get("sha256", "")), code=f"{code}_HASH") + if row.get("binding_status") not in {None, "BOUND"}: + raise IngressError(code, f"asset is not release-bound: {path}") + return path, digest + + + def _inline_release_materialization_plan( + request: Mapping[str, str], + release: Mapping[str, Any], + first_read: Callable[[str], bytes], + ) -> tuple[list[tuple[str, str, str]], dict[str, str]]: + """Return exact remote->temporary destinations and expected remote hashes.""" + + destinations: list[tuple[str, str, str]] = [] + expected_hashes: dict[str, str] = {} + seen_destinations: set[tuple[str, str]] = set() + + def add( + remote_path: str, + local_family: str, + relative_path: str, + *, + expected_sha256: str | None = None, + ) -> None: + remote = _inline_relative_path(remote_path, code="HYDRATION_REMOTE_PATH_INVALID") + relative = _inline_relative_path(relative_path, code="HYDRATION_LOCAL_PATH_INVALID") + key = (local_family, relative) + if key in seen_destinations: + raise IngressError("HYDRATION_DESTINATION_DUPLICATE", f"duplicate hydration destination: {relative}") + seen_destinations.add(key) + destinations.append((remote, local_family, relative)) + if expected_sha256 is not None: + digest = _inline_sha256(expected_sha256.lower(), code="HYDRATION_EXPECTED_HASH_INVALID") + prior = expected_hashes.get(remote) + if prior is not None and prior != digest: + raise IngressError("HYDRATION_HASH_CONFLICT", f"conflicting expected hash: {remote}") + expected_hashes[remote] = digest + + release_remote = INLINE_STAGE2_RELEASE_PATH + add( + release_remote, + "stage2_asset", + "manifest/stage2_release.json", + expected_sha256=EXPECTED_STAGE2_RELEASE_SHA256, + ) + + source_rows = _release_stage1_source_rows(release) + if not isinstance(source_rows, list): + raise IngressError("RELEASE_STAGE1_SOURCE_SHAPE", "release stage1_sources must be an array") + fixed_rows = [row for row in source_rows if row.get("logical_input_id") != "signal_payload_family"] + if len(fixed_rows) != 16: + raise IngressError("RELEASE_STAGE1_SOURCE_COUNT", "release must enumerate exactly 16 fixed inputs") + signal_manifest_relative: str | None = None + for row in fixed_rows: + if not isinstance(row, dict) or not isinstance(row.get("path"), str): + raise IngressError("RELEASE_STAGE1_SOURCE_ROW", "fixed source row lacks an exact path") + relative = _inline_relative_path(row["path"], code="STAGE1_SOURCE_PATH_INVALID") + remote = _inline_join(request["stage1_run_root_ref"], relative, code="STAGE1_SOURCE_PATH_INVALID") + add(remote, "stage1_run", relative) + first_read(remote) + if row.get("logical_input_id") == "signal_manifest": + signal_manifest_relative = relative + if signal_manifest_relative is None: + raise IngressError("SIGNAL_MANIFEST_RELEASE_ROW_MISSING", "release lacks signal_manifest") + signal_remote = _inline_join( + request["stage1_run_root_ref"], + signal_manifest_relative, + code="SIGNAL_MANIFEST_PATH_INVALID", + ) + signal_manifest = load_json_strict(first_read(signal_remote)) + if not isinstance(signal_manifest, dict) or not isinstance(signal_manifest.get("files"), list): + raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest files must be an array") + for index, row in enumerate(signal_manifest["files"]): + if not isinstance(row, dict) or not isinstance(row.get("path"), str): + raise IngressError("SIGNAL_FILE_ROW_SHAPE", f"invalid signal row: {index}") + relative_payload = _inline_relative_path(row["path"], code="SIGNAL_FILE_PATH_INVALID") + if relative_payload.startswith("signals/"): + raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "signal manifest path includes signals/") + relative = f"signals/{relative_payload}" + remote = _inline_join(request["stage1_run_root_ref"], relative, code="SIGNAL_FILE_PATH_INVALID") + expected = row.get("file_sha256", row.get("sha256", row.get("raw_sha256"))) + add(remote, "stage1_run", relative, expected_sha256=expected if isinstance(expected, str) else None) + first_read(remote) + + dependencies = release.get("dependency_locks") + stage1_dependency = dependencies.get("stage1") if isinstance(dependencies, dict) else None + closure = stage1_dependency.get("concrete_paths") if isinstance(stage1_dependency, dict) else None + if not isinstance(closure, list) or not closure: + raise IngressError("STAGE1_DEPENDENCY_LOCK_MISSING", "Stage 1 deployment closure is absent") + expected_count = stage1_dependency.get("expected_concrete_path_count") + if expected_count is not None and expected_count != len(closure): + raise IngressError("STAGE1_DEPENDENCY_COUNT_MISMATCH", "Stage 1 closure count differs from release") + for row in closure: + if not isinstance(row, dict): + raise IngressError("STAGE1_DEPENDENCY_ROW_SHAPE", "Stage 1 closure row must be an object") + relative, digest = _inline_bound_ref(row, code="STAGE1_DEPENDENCY_UNBOUND") + remote = _inline_join( + request["stage1_deployment_root_ref"], + relative, + code="STAGE1_DEPLOYMENT_PATH_INVALID", + ) + add(remote, "stage1_deployment", relative, expected_sha256=digest) + first_read(remote) + + contract_ref = stage1_dependency.get("contract_manifest_ref", {}) + if isinstance(contract_ref, dict) and isinstance(contract_ref.get("path"), str): + relative, digest = _inline_bound_ref(contract_ref, code="CONTRACT_MANIFEST_UNBOUND") + remote = _inline_join( + request["stage1_deployment_root_ref"], + relative, + code="CONTRACT_MANIFEST_PATH_INVALID", + ) + add(remote, "stage1_deployment", relative, expected_sha256=digest) + first_read(remote) + completion_ref = stage1_dependency.get("completion_seal_ref", {}) + if isinstance(completion_ref, dict) and completion_ref.get("path") not in {None, "PENDING_SEQUENTIAL_BIND"}: + relative, digest = _inline_bound_ref(completion_ref, code="COMPLETION_SEAL_UNBOUND") + remote = _inline_join(request["stage1_run_root_ref"], relative, code="COMPLETION_SEAL_PATH_INVALID") + add(remote, "stage1_run", relative, expected_sha256=digest) + first_read(remote) + + manifest_ref = release.get("module_manifest_ref") + if not isinstance(manifest_ref, dict): + raise IngressError("MODULE_MANIFEST_REF_MISSING", "release lacks module_manifest_ref") + manifest_relative, manifest_digest = _inline_bound_ref(manifest_ref, code="MODULE_MANIFEST_UNBOUND") + manifest_remote = _inline_join(INLINE_STAGE2_ASSET_ROOT, manifest_relative, code="MODULE_MANIFEST_PATH_INVALID") + add(manifest_remote, "stage2_asset", manifest_relative, expected_sha256=manifest_digest) + module_manifest = load_json_strict(first_read(manifest_remote)) + modules = module_manifest.get("modules") if isinstance(module_manifest, dict) else None + if not isinstance(modules, list): + raise IngressError("MODULE_MANIFEST_SHAPE", "module manifest lacks modules array") + found_schema_ids: set[str] = set() + for row in modules: + if not isinstance(row, dict) or row.get("module_id") not in INLINE_SCHEMA_MODULE_IDS: + continue + relative, digest = _inline_bound_ref(row, code="STAGE2_SCHEMA_UNBOUND") + remote = _inline_join(INLINE_STAGE2_ASSET_ROOT, relative, code="STAGE2_SCHEMA_PATH_INVALID") + add(remote, "stage2_asset", relative, expected_sha256=digest) + first_read(remote) + found_schema_ids.add(str(row["module_id"])) + if found_schema_ids != set(INLINE_SCHEMA_MODULE_IDS): + raise IngressError( + "STAGE2_SCHEMA_CLOSURE_INCOMPLETE", + "module manifest does not bind all S2_00 schemas", + details={"missing": sorted(set(INLINE_SCHEMA_MODULE_IDS) - found_schema_ids)}, + ) + + stage2_direct = dependencies.get("stage2_direct") if isinstance(dependencies, dict) else None + if not isinstance(stage2_direct, list): + raise IngressError("STAGE2_DIRECT_LOCK_MISSING", "release lacks Stage 2 direct closure") + for row in stage2_direct: + if not isinstance(row, dict): + raise IngressError("STAGE2_DIRECT_ROW_SHAPE", "Stage 2 direct row must be an object") + relative, digest = _inline_bound_ref(row, code="STAGE2_DIRECT_UNBOUND") + remote = _inline_join(INLINE_STAGE2_ASSET_ROOT, relative, code="STAGE2_DIRECT_PATH_INVALID") + add(remote, "stage2_asset", relative, expected_sha256=digest) + first_read(remote) + return destinations, expected_hashes + + + def _inline_hydrate( + localdocs: _InlineLocaldocs, + temp_root: Path, + ) -> tuple[dict[str, str], Mapping[str, Any], Path, Path, Path, dict[str, Any]]: + first_pass: dict[str, bytes] = {} + first_pass_bytes = 0 + + request_raw = localdocs.read_binary(INLINE_REQUEST_PATH) + if len(request_raw) > MAX_FILE_BYTES: + raise IngressError("SOURCE_SIZE_LIMIT", "S2_00 request exceeds the global pre-parse file limit") + first_pass[INLINE_REQUEST_PATH] = request_raw + first_pass_bytes += len(request_raw) + request = _inline_validate_request(request_raw) + + release_raw = localdocs.read_binary(INLINE_STAGE2_RELEASE_PATH) + if len(release_raw) > MAX_FILE_BYTES: + raise IngressError("SOURCE_SIZE_LIMIT", "Stage 2 release exceeds the global pre-parse file limit") + first_pass[INLINE_STAGE2_RELEASE_PATH] = release_raw + first_pass_bytes += len(release_raw) + if first_pass_bytes > MAX_RUN_BYTES: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "initial request and release exceed the global pre-parse budget") + if EXPECTED_STAGE2_RELEASE_SHA256 == "0" * 64: + raise IngressError( + "EXPECTED_STAGE2_RELEASE_SHA256_UNBOUND", + "build-time Stage 2 release pin has not been bound", + ) + expected_release = _inline_sha256( + EXPECTED_STAGE2_RELEASE_SHA256, + code="EXPECTED_STAGE2_RELEASE_SHA256_INVALID", + ) + if hashlib.sha256(release_raw).hexdigest() != expected_release: + raise IngressError("STAGE2_RELEASE_PIN_MISMATCH", "Stage 2 release bytes differ from the build-time pin") + release = load_json_strict(release_raw) + if not isinstance(release, dict): + raise IngressError("RELEASE_LOCK_SHAPE", "Stage 2 release must be an object") + limits = release.get("limits") if isinstance(release.get("limits"), dict) else {} + max_file = int(limits.get("max_file_bytes", MAX_FILE_BYTES)) + max_total = int(limits.get("max_run_bytes", MAX_RUN_BYTES)) + max_paths = int(limits.get("max_hydration_paths", 1024)) + if max_file <= 0 or max_file > MAX_FILE_BYTES: + raise IngressError("RELEASE_FILE_LIMIT_INVALID", "release max_file_bytes exceeds the build-time ceiling") + if max_total <= 0 or max_total > MAX_RUN_BYTES: + raise IngressError("RELEASE_RUN_LIMIT_INVALID", "release max_run_bytes exceeds the build-time ceiling") + if max_paths <= 0 or max_paths > 1024: + raise IngressError("RELEASE_PATH_LIMIT_INVALID", "release max_hydration_paths exceeds the build-time ceiling") + if len(request_raw) > max_file or len(release_raw) > max_file: + raise IngressError("SOURCE_SIZE_LIMIT", "initial request or release exceeds the release-bound file limit") + if first_pass_bytes > max_total: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "initial request and release exceed the release-bound budget") + + def first_read(path: str) -> bytes: + nonlocal first_pass_bytes + safe = _inline_relative_path(path, code="HYDRATION_REMOTE_PATH_INVALID") + if safe not in first_pass: + payload = localdocs.read_binary(safe) + if len(payload) > max_file: + raise IngressError("SOURCE_SIZE_LIMIT", f"remote source exceeds file limit: {safe}") + first_pass_bytes += len(payload) + if first_pass_bytes > max_total: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "first-pass hydration exceeds release budget") + first_pass[safe] = payload + return first_pass[safe] + + destinations, expected_hashes = _inline_release_materialization_plan( + request, + release, + first_read, + ) + if len(first_pass) > max_paths: + raise IngressError("HYDRATION_PATH_COUNT_LIMIT", "hydration path count exceeds release budget") + for remote, expected in expected_hashes.items(): + observed = hashlib.sha256(first_read(remote)).hexdigest() + if observed != expected: + raise IngressError("HYDRATION_BOUND_HASH_MISMATCH", f"release-bound hash mismatch: {remote}") + + second_pass_bytes = 0 + second_pass: dict[str, bytes] = {} + for remote in sorted(first_pass): + payload = localdocs.read_binary(remote) + second_pass_bytes += len(payload) + if len(payload) > max_file or second_pass_bytes > max_total: + raise IngressError("TWO_PASS_READ_BUDGET", "second-pass hydration exceeds release budget") + if payload != first_pass[remote] or hashlib.sha256(payload).digest() != hashlib.sha256(first_pass[remote]).digest(): + raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"remote source changed between bounded reads: {remote}") + second_pass[remote] = payload + + pass_1_hash_rows = [ + [path, hashlib.sha256(first_pass[path]).hexdigest()] + for path in sorted(first_pass) + ] + pass_2_hash_rows = [ + [path, hashlib.sha256(second_pass[path]).hexdigest()] + for path in sorted(second_pass) + ] + hydration_receipt = { + "schema_version": "stage2_s2_00_two_pass_hydration_receipt.v1", + "transport": "localdocs.read_binary_doc", + "read_policy": "BOUNDED_TWO_PASS_BINARY_RAW_HASH_MAP_EQUALITY", + "read_pass_count": 2, + "max_read_passes": 2, + "pass_1_raw_hash_map_digest": canonical_digest(pass_1_hash_rows), + "pass_2_raw_hash_map_digest": canonical_digest(pass_2_hash_rows), + "source_receipts": [ + { + "logical_input_id": path, + "logical_path": path, + "pass_1_raw_sha256": hashlib.sha256(first_pass[path]).hexdigest(), + "pass_2_raw_sha256": hashlib.sha256(second_pass[path]).hexdigest(), + "pass_1_byte_length": len(first_pass[path]), + "pass_2_byte_length": len(second_pass[path]), + "stability_status": "STABLE", + } + for path in sorted(first_pass) + ], + "stability_status": "STABLE", + } + + roots = { + "stage1_run": temp_root / "stage1_run", + "stage1_deployment": temp_root / "stage1_deployment", + "stage2_asset": temp_root / "stage2_asset", + } + for root in roots.values(): + root.mkdir(parents=True, exist_ok=False) + for remote, family, relative in sorted(destinations): + target = roots[family] / PurePosixPath(relative) + target.parent.mkdir(parents=True, exist_ok=True) + if target.exists(): + raise IngressError("HYDRATION_DESTINATION_EXISTS", f"duplicate materialization: {relative}") + target.write_bytes(first_pass[remote]) + return ( + request, + release, + roots["stage1_run"], + roots["stage1_deployment"], + roots["stage2_asset"], + hydration_receipt, + ) + + + def _inline_output_files(output_root: Path, run_binding_digest: str) -> dict[str, bytes]: + if not output_root.is_dir() or output_root.is_symlink(): + raise IngressError("LOCAL_OUTPUT_TREE_MISSING", "pure core did not produce an output tree") + files: dict[str, bytes] = {} + for path in sorted(output_root.rglob("*")): + if path.is_symlink(): + raise IngressError("LOCAL_OUTPUT_SYMLINK", "pure core output contains a symlink") + if not path.is_file(): + continue + relative = _safe_relative_path(path.relative_to(output_root).as_posix()).as_posix() + files[relative] = path.read_bytes() + status_raw = files.get("ingress/ingress_status.json") + if status_raw is None: + raise IngressError("OUTPUT_BARRIER_MISSING", "pure core output lacks ingress_status") + status = load_json_strict(status_raw) + binding = status.get("run_binding_receipt") if isinstance(status, dict) else None + if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: + raise IngressError("RUN_BINDING_RECEIPT_MISMATCH", "pure core status has a different binding") + barrier = status.get("output_barrier") + rows = barrier.get("artifacts") if isinstance(barrier, dict) else None + if not isinstance(rows, list) or barrier.get("written_last") is not True: + raise IngressError("OUTPUT_BARRIER_INVALID", "pure core output barrier is incomplete") + expected_paths = {"ingress/ingress_status.json"} + for row in rows: + if not isinstance(row, dict) or not isinstance(row.get("path"), str): + raise IngressError("OUTPUT_BARRIER_ROW_SHAPE", "output barrier row is malformed") + relative = _safe_relative_path(row["path"]).as_posix() + payload = files.get(relative) + if payload is None or hashlib.sha256(payload).hexdigest() != row.get("raw_sha256"): + raise IngressError("OUTPUT_BARRIER_HASH_MISMATCH", f"output barrier mismatch: {relative}") + expected_paths.add(relative) + if set(files) != expected_paths: + raise IngressError("OUTPUT_BARRIER_SET_MISMATCH", "output tree differs from its barrier set") + if barrier.get("artifact_set_digest") != canonical_digest(rows): + raise IngressError("OUTPUT_BARRIER_DIGEST_MISMATCH", "output barrier row digest is invalid") + return files + + + def _inline_verify_existing_remote( + localdocs: _InlineLocaldocs, + output_root: str, + files: Mapping[str, bytes], + run_binding_digest: str, + ) -> bool: + status_path = _inline_join(output_root, "ingress/ingress_status.json", code="OUTPUT_PATH_INVALID") + status_raw = localdocs.read_binary_optional(status_path) + if status_raw is None: + return False + status = load_json_strict(status_raw) + binding = status.get("run_binding_receipt") if isinstance(status, dict) else None + if not isinstance(binding, dict) or binding.get("run_binding_digest") != run_binding_digest: + raise IngressError("RUN_ID_BINDING_CONFLICT", "existing remote barrier has a different binding") + if status_raw != files.get("ingress/ingress_status.json"): + raise IngressError("IDEMPOTENT_STATUS_MISMATCH", "existing remote status is not byte-identical") + for relative, expected in sorted(files.items()): + observed = localdocs.read_binary(_inline_join(output_root, relative, code="OUTPUT_PATH_INVALID")) + if observed != expected: + raise IngressError("IDEMPOTENT_ARTIFACT_MISMATCH", f"existing artifact differs: {relative}") + return True + + + def _inline_publish_remote( + localdocs: _InlineLocaldocs, + files: Mapping[str, bytes], + run_binding_digest: str, + ) -> dict[str, Any]: + output_root = f"stage2_runs/by-binding/{run_binding_digest}" + status_raw = files["ingress/ingress_status.json"] + status = load_json_strict(status_raw) + barrier = status.get("output_barrier") if isinstance(status, dict) else None + branch = barrier.get("branch") if isinstance(barrier, dict) else None + if branch not in {"NORMAL", "DIAGNOSTIC"}: + raise IngressError("OUTPUT_BARRIER_BRANCH_INVALID", "pure core output barrier lacks a valid branch") + artifact_rows = [ + { + "logical_artifact_id": relative, + "path": relative, + "schema_id": _output_schema_id(relative), + "raw_sha256": hashlib.sha256(payload).hexdigest(), + "byte_length": len(payload), + } + for relative, payload in sorted(files.items()) + if relative != "ingress/ingress_status.json" + ] + if _inline_verify_existing_remote(localdocs, output_root, files, run_binding_digest): + publication_status = "IDEMPOTENT_SUCCESS" + else: + for relative in sorted(path for path in files if path != "ingress/ingress_status.json"): + localdocs.write_binary_verified( + _inline_join(output_root, relative, code="OUTPUT_PATH_INVALID"), + files[relative], + ) + status_path = _inline_join(output_root, "ingress/ingress_status.json", code="OUTPUT_PATH_INVALID") + localdocs.write_binary_verified(status_path, status_raw) + if localdocs.read_binary(status_path) != status_raw: + raise IngressError("OUTPUT_BARRIER_READBACK_MISMATCH", "remote ingress_status read-back failed") + publication_status = "PUBLISHED_STATUS_LAST" + return { + "schema_version": "stage2_logical_publish_receipt.v1", + "barrier_id": "S2_00_INGRESS_STATUS_BARRIER", + "barrier_path": "ingress/ingress_status.json", + "publish_semantics": "STATUS_LAST_LOGICAL_COMMIT", + "canonical_output_root": f"{output_root}/", + "run_binding_digest": run_binding_digest, + "branch": branch, + "artifacts": artifact_rows, + "artifact_set_digest": canonical_digest(artifact_rows), + "barrier_raw_sha256": hashlib.sha256(status_raw).hexdigest(), + "non_status_artifacts_read_back_verified": True, + "barrier_written_last": True, + "downstream_consumption_allowed": True, + "publication_status": publication_status, + } + + + def _inline_context_is_bound() -> bool: + return ( + re.fullmatch(r"[a-f0-9]{64}", INLINE_USER_HASH) is not None + and re.fullmatch(r"[a-f0-9]{64}", INLINE_WORKSPACE_HASH) is not None + ) + + + def run_inline_mcp(*, client: Any | None = None) -> int: + """Execute one complete MCP Code Executor S2_00 task and emit one JSON receipt.""" + + localdocs: _InlineLocaldocs | None = None + receipt: dict[str, Any] + exit_code = 0 + try: + with contextlib.redirect_stdout(io.StringIO()): + localdocs = _InlineLocaldocs( + INLINE_USER_HASH, + INLINE_WORKSPACE_HASH, + client=client, + ) + try: + localdocs.initialize() + with tempfile.TemporaryDirectory(prefix="liti-s2-00-") as directory: + temp_root = Path(directory) + ( + request, + _release, + stage1_root, + deployment_root, + asset_root, + hydration_receipt, + ) = _inline_hydrate(localdocs, temp_root) + release_lock = load_release_lock(asset_root / "manifest" / "stage2_release.json") + contract_manifest = _load_bound_contract_manifest(deployment_root, release_lock) + local_output = temp_root / "core_output" + result = execute_ingress( + stage1_root, + release_lock, + output_dir=local_output, + attempt_id=request["attempt_id"], + run_id=request["request_id"], + user_context_sha256=INLINE_USER_HASH, + workspace_context_sha256=INLINE_WORKSPACE_HASH, + stage1_deployment_root=deployment_root, + stage2_asset_root=asset_root, + contract_manifest=contract_manifest, + hydration_stability_receipt=hydration_receipt, + ) + run_binding_digest = _inline_sha256( + str(result.get("run_binding_digest", "")), + code="CORE_RUN_BINDING_INVALID", + ) + files = _inline_output_files(local_output, run_binding_digest) + publication = _inline_publish_remote(localdocs, files, run_binding_digest) + receipt = { + "schema_version": "stage2_s2_00_inner_receipt.v1", + "workflow_id": "S2_00", + "ok": True, + "status": ( + "DIAGNOSTIC_PUBLISHED" + if result["route"] == "TO_S2_40_STATUS_ONLY" + else "SUCCEEDED" + ), + "request_id": request["request_id"], + "route": result["route"], + "run_id": canonical_run_id(run_binding_digest), + "run_binding_digest": run_binding_digest, + "expected_release_sha256": EXPECTED_STAGE2_RELEASE_SHA256, + "algorithm_digest": ALGORITHM_SEMANTIC_DIGEST, + "artifact_set_digest": publication["artifact_set_digest"], + "ingress_status_sha256": publication["barrier_raw_sha256"], + "logical_publish_receipt": publication, + } + finally: + if localdocs is not None: + localdocs.close() + except Exception as exc: + error = ( + exc.as_dict() + if isinstance(exc, IngressError) + else {"code": "INLINE_RUNTIME_ERROR", "message": str(exc)} + ) + receipt = { + "schema_version": "stage2_s2_00_inner_receipt.v1", + "workflow_id": "S2_00", + "ok": False, + "status": "FAILED_NO_BARRIER", + "expected_release_sha256": EXPECTED_STAGE2_RELEASE_SHA256, + "error": error, + } + exit_code = 2 + sys.stdout.buffer.write(canonical_json_bytes(receipt)) + return exit_code + + + class _FixedArgvParser(argparse.ArgumentParser): + def error(self, message: str) -> None: + raise IngressError("CLI_ARGUMENT_ERROR", message) + + + def _parser() -> argparse.ArgumentParser: + parser = _FixedArgvParser( + description="S2_00 deterministic Stage 1 ingress", + add_help=False, + ) + parser.add_argument("--workflow-id", required=True) + parser.add_argument("--release-ref", required=True) + parser.add_argument("--run-id", required=True) + parser.add_argument("--attempt-id", required=True) + parser.add_argument("--user-context-sha256", required=True) + parser.add_argument("--workspace-context-sha256", required=True) + parser.add_argument("--stage1-run-root-ref", required=True) + parser.add_argument("--stage1-deployment-root-ref", required=True) + return parser + + + def _project_root_from_runtime() -> Path: + for candidate in Path(__file__).resolve().parents: + if candidate.name == "Case_02_Comparison_Research": + return candidate + raise IngressError("PROJECT_ROOT_NOT_FOUND", "runtime is not located below the canonical project root") + + + def _resolve_contained_ref( + approved_root: Path, + ref: str, + *, + code: str, + ) -> Path: + if not isinstance(ref, str) or not ref or "\x00" in ref: + raise IngressError(code, "workspace reference is empty or malformed") + candidate = Path(ref) + if not candidate.is_absolute(): + candidate = approved_root / candidate + if candidate.is_symlink(): + raise IngressError(code, "workspace reference may not be a symlink") + approved_resolved = approved_root.resolve(strict=True) + try: + lexical_relative = candidate.relative_to(approved_root) + except ValueError: + lexical_relative = None + if lexical_relative is not None: + _assert_no_symlink_components(approved_resolved, PurePosixPath(lexical_relative.as_posix())) + resolved = candidate.resolve(strict=True) + try: + resolved.relative_to(approved_resolved) + except ValueError as exc: + raise IngressError(code, "workspace reference escaped its approved root") from exc + return resolved + + + def _load_bound_contract_manifest( + stage1_deployment_root: Path, + release_lock: Mapping[str, Any], + ) -> Mapping[str, Any] | None: + dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) + ref = dependency.get("contract_manifest_ref", {}) if isinstance(dependency, dict) else {} + if not isinstance(ref, dict) or not isinstance(ref.get("path"), str): + return None + expected_hash = ref.get("sha256") + if not isinstance(expected_hash, str) or not re.fullmatch(r"[A-Fa-f0-9]{64}", expected_hash): + raise IngressError("CONTRACT_MANIFEST_UNBOUND", "contract manifest hash is not bound") + snapshot = open_bounded_snapshot( + stage1_deployment_root, + ref["path"], + logical_input_id="stage1_contract_manifest", + ) + if snapshot.raw_sha256 != expected_hash: + raise IngressError("CONTRACT_MANIFEST_HASH_MISMATCH", "contract manifest raw hash differs from release") + value = load_json_strict(snapshot) + if not isinstance(value, dict): + raise IngressError("CONTRACT_MANIFEST_SHAPE", "contract manifest must be an object") + return value + + + def main(argv: Sequence[str] | None = None) -> int: + """CLI entrypoint. Stdout contains exactly one canonical JSON object.""" + + try: + args = _parser().parse_args(argv) + if args.workflow_id != "S2_00": + raise IngressError("WORKFLOW_ID_MISMATCH", "runtime accepts only workflow_id S2_00") + asset_root = Path(__file__).resolve().parents[1] + project_root = _project_root_from_runtime() + release_path = _resolve_contained_ref(asset_root, args.release_ref, code="RELEASE_REF_OUTSIDE_ROOT") + if not release_path.is_file(): + raise IngressError("RELEASE_REF_NOT_FILE", "release reference must identify one regular file") + stage1_run_root = _resolve_contained_ref( + project_root, + args.stage1_run_root_ref, + code="STAGE1_RUN_ROOT_OUTSIDE_WORKSPACE", + ) + stage1_deployment_root = _resolve_contained_ref( + project_root, + args.stage1_deployment_root_ref, + code="STAGE1_DEPLOYMENT_ROOT_OUTSIDE_WORKSPACE", + ) + if not stage1_run_root.is_dir() or not stage1_deployment_root.is_dir(): + raise IngressError("STAGE1_ROOT_NOT_DIRECTORY", "Stage 1 roots must be directories") + release_lock = load_release_lock(release_path) + contract_manifest = _load_bound_contract_manifest(stage1_deployment_root, release_lock) + output_dir = project_root / "stage2_runs" / args.run_id + result = execute_ingress( + stage1_run_root, + release_lock, + output_dir=output_dir, + attempt_id=args.attempt_id, + run_id=args.run_id, + user_context_sha256=args.user_context_sha256, + workspace_context_sha256=args.workspace_context_sha256, + stage1_deployment_root=stage1_deployment_root, + stage2_asset_root=asset_root, + contract_manifest=contract_manifest, + ) + sys.stdout.buffer.write(canonical_json_bytes({"ok": True, "result": result})) + return 0 + except (IngressError, FileNotFoundError, PermissionError, OSError) as exc: + error = exc.as_dict() if isinstance(exc, IngressError) else {"code": "OS_ERROR", "message": str(exc)} + sys.stdout.buffer.write(canonical_json_bytes({"ok": False, "error": error})) + return 2 + + + if __name__ == "__main__": + raise SystemExit(run_inline_mcp()) + task_procedure: + IN: + nexts: + - Task_S2_00_prepare_request + wait_until: [] + Task_S2_00_prepare_request: + nexts: + - Task_S2_00_deterministic_ingress + wait_until: + - IN + Task_S2_00_deterministic_ingress: + nexts: + - OUT + wait_until: + - Task_S2_00_prepare_request + OUT: + nexts: [] + wait_until: + - Task_S2_00_deterministic_ingress diff --git a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v1.md b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v1.md new file mode 100644 index 00000000..4b246dd9 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v1.md @@ -0,0 +1,260 @@ +# Stage_2_S2_00.yml 개정 전략 v1 + +작성일: 2026-10-02. 상태: **개정 전략 작성 완료 / YAML·실행 자산 개정 및 live 실행 미수행**. + +## 1. 권고안과 개정 범위 + +**Stage 1의 사건 run root와 배포 root를 S2_00 호출에 직접 전달하고, 유일한 `Task_S2_00_deterministic_ingress`가 입력 검증·hydration·C00–C15·결과 발행을 한 번에 수행하도록 개정한다.** `Task_S2_00_prepare_request`, 고정 `stage2_control/s2_00_request.json` 생성·저장·재읽기, 호출자의 `request_id`·`attempt_id` 부여를 실행 경로에서 제거한다. Stage 1 결과물의 파일명·상대경로·원본 bytes는 그대로 사용한다. + +사용자 지시 중 “사건·배포 경로 전달 및 `request_id`·`attempt_id` 부여 방식 폐기”는 **이를 별도의 request 준비 절차로 구성하는 방식을 폐기**한다는 의미로 적용한다. 경로 자체는 직접 인계에 필요한 값이다. ID는 호출자의 필수 입력에서 제거하고, 기존 출력 schema와 downstream 호환에 필요한 내부 값만 S2_00이 처리한다. 출력의 ID 필드까지 전면 삭제하는 것은 본 최소 변경안에 포함하지 않는다. + +개정의 성공 기준은 단일 task 실행 자체가 아니라, Stage 1 원본을 읽어 검증한 뒤 기존 정상 또는 진단 산출물과 마지막 status barrier를 정확히 발행하는 것이다. 기존 구현의 의미 인계 공백 때문에 정상 실행이 성립하지 않는 경우에는 해당 Stage 2 adapter·projection만 보완한다. + +본 문서는 전략서다. 아래 함수 signature, 입력 객체, ID 결정 규칙과 변경 순서는 **개정 제안**이며 현재 구현이나 backend 지원 사실을 뜻하지 않는다. 이번 작성 작업의 산출물은 본 문서와 지정된 Stage 2 `MEMORY.md` 요약뿐이다. + +## 2. 기준 자료와 실제 확인 결과 + +판단 순서는 현행 배포 YAML → 현행 release·배포/빌드 계약 → 현행 분석서 → 과거 YAML·분석서다. IO 표의 역사적 행번호를 현행 YAML의 행번호로 사용하지 않는다. + +| 자료 | 확인 내용과 전략상 역할 | +|---|---| +| [현행 Stage_2_S2_00.yml](Stage_2_S2_00.yml) | `Stage_2_S2_00_v2`, version `1.2.0`; prepare와 ingress의 두 task. 현행 분석의 최우선 기준이다. | +| [Analysis_Stage_2_S2_00.md](Analysis_Stage_2_S2_00.md) | 현행 두-task 구조, 외부 인자 결속 미검증, C00–C15·입출력·운영 경계를 설명한다. | +| [Stage_2_00_IO_info.md](../../../Stage_2_00_IO_info.md) | §3.2의 사건 입력 16개, §3.4의 Stage 1 배포 55개, §3.5의 Stage 2 직접 의존 49개와 출력 목록을 사용한다. 상단 현행 IO 보충과 보존본 기반 본문을 구별한다. | +| [Stage_2_S2_00_outdated_9_09.yml](Stage_2_S2_00_outdated_9_09.yml) | 구형은 단일 ingress task지만 외부 request 파일이 필요하다. 단일 task 구조의 비교 기준이며 그대로 복원할 개정 정본은 아니다. | +| [Stage_2_10_Analysis_v1.md](../../../Stage_2_10_Analysis_v1.md) | 사용자 context의 설명과 달리 실제 제목·본문·정본은 **S2_10 분석**이다. S2_00 산출물을 소비하는 downstream 계약 확인에 사용한다. | +| [Stage_2_00_Analysis_v1.md](../../../Stage_2_00_Analysis_v1.md) | 본문이 구형 S2_00을 분석하고 상단에 현행 보충이 있다. §8의 입력·projection 제한을 교차 확인한다. | +| [stage2_release.json](../manifest/stage2_release.json) | `/stage1_sources`, `/dependency_locks/stage1`, `/dependency_locks/stage2_direct`, `/bundle`이 실제 입력·자산 계약이다. | +| [s2_00_inline_code_receipt.json](../manifest/s2_00_inline_code_receipt.json)·[builder](../offline_build/build_s2_00_inline_projection.py) | authoring 정본은 `Stage_2_S2_00_v.2.yml`이다. main 폴더의 version 없는 동명 YAML을 개정 정본으로 혼동하지 않는다. | + +작성 시 확인한 현행 배포 YAML SHA-256은 `07bc9e235d211f3d4e6886bca2faeeb1f112d525fa29fe111bd7c50c42ff9849`, 구형 `outdated_9_09` SHA-256은 `94c2744d71d222cfb5cc1b2a0d33a6cda20079439823de431eef378a2fa5c943`이다. 현행 배포본과 authoring `Stage_2_S2_00_v.2.yml`은 byte-identical이다. + +현행 ingress inline code와 `outdated_9_09`의 ingress inline code를 비교하면 차이는 `EXPECTED_STAGE2_RELEASE_SHA256` 상수 하나다. 따라서 request 앞단이 추가되었어도 기존 C00–C15의 의미 인계 제한이 해소되었다고 볼 수 없다. 최신 ingress를 기반으로 직접 입력 경계만 바꾸고, 필요한 의미 보완을 제한적으로 추가하는 것이 효율적이다. + +현재 prepare의 `__main__`은 `PREPARE_ARGUMENT_BINDING_UNVERIFIED`로 종료한다. release는 `DEV_FIXTURE_RELEASE / DRAFT_NOT_EXECUTABLE`, bundle은 `STRUCTURAL_FIXTURE`, `selected_context_refs=[]`다. 직접 인계 개정이 이 상태를 자동으로 실행 적격 또는 승인 완료로 바꾸지는 않는다. + +## 3. 목표 실행 흐름과 책임 + +```text +호출자: Stage 1 완료 결과의 사건 root + 그 결과에 대응하는 배포 root +backend: 인증된 user/workspace 세션 + 직접 입력을 executor에 전달 + │ + ▼ +IN → Task_S2_00_deterministic_ingress → OUT + 단일 code-executor.run_code + ① 직접 입력 형태·NFC 상대경로·workspace 결속 검증 + ② 고정 Stage 2 release pin 검증 + ③ 16개 사건 파일 + manifest 신호 + 55개 배포 + 49개 직접 자산 + + module manifest/schema/조건부 release 입력의 exact closure 확정 + ④ 동일 원격 closure 2회 읽기·bytes/hash 대조·임시 hydration + ⑤ C00 → C05 → run binding → C10 → C15 + ⑥ 정상 11+E 또는 진단 5개 출력 검증·발행 + ⑦ ingress_status.json 마지막 write/read-back → stdout receipt + │ + ▼ +외부 orchestrator: barrier·정확한 산출물 집합·route 검증 → S2_10 또는 S2_40 +``` + +C00는 입력 계약·producer·identity·seal·signal ALL 검증, C05는 review 정규화·보존검사, C10은 사건·증거·객체·당사자·slot 문맥, C15는 cluster·slice·bundle·route 구성이다. 이들은 **동일 task 내부의 메모리 처리 단계**로 유지한다. 단계마다 중간 JSON 파일을 새로 만들어 저장·재읽기하는 방식은 도입하지 않는다. + +별도 prepare 성공 뒤 ingress를 시작하는 task 간 게이트는 없어지고, 직접 입력 검증 실패 시 같은 Python 실행 안에서 후속 처리를 중단한다. 외부 실행자의 S2_10/S2_40 dispatch 책임과 status barrier 소비 책임은 유지한다. + +## 4. 최소 직접 인계 계약 + +### 4.1 호출자가 제공할 값 + +직접 입력은 아래 **두 문자열 필드의 닫힌 객체**로 제안한다. 이는 파일 생성 지시가 아니라 호출 데이터 예시다. + +```json +{ + "stage1_run_root_ref": "", + "stage1_deployment_root_ref": "" +} +``` + +| 정보 | 공급·검증 책임 | 최소화 이유 | +|---|---|---| +| `stage1_run_root_ref` | 호출자가 완료된 Stage 1 결과의 실제 root를 공급; S2_00 검증 | 16개 개별 파일 경로를 수동 작성하지 않고 release의 상대경로를 결합한다. | +| `stage1_deployment_root_ref` | 호출자가 해당 Stage 1 결과에 대응하는 실제 배포 root를 공급; S2_00이 55개 잠금 검증 | 사건 자료와 배포 config/schema를 구별한다. 추측한 `Default_Agent` 경로로 대체하지 않는다. | +| user/workspace 인증·hash | 기존 backend/localdocs 세션 결속 | 일반 호출 입력으로 workspace를 임의 지정하지 않는다. | +| Stage 2 자산 root·release pin | 기존 inline 상수·빌드/배포 계약 | 호출자가 release나 output 경로를 자유롭게 바꾸지 않도록 유지한다. | +| `request_id`·`attempt_id` | 호출 입력에서 제거; §4.3에 따른 내부 호환 처리 | 별도 ID 부여 task·request 준비 파일이 필요하지 않다. | + +`W/`는 localdocs 논리 workspace root다. `U=W//`, `D=W//`, `A=W/Default_Agent/Stage_2_Clean/`다. Dropbox 절대경로나 Code Executor의 로컬 mount와 동일하다고 가정하지 않는다. + +두 root는 이미 NFC인 canonical 상대경로여야 한다. 빈 값·절대경로·`..`·역슬래시·NUL·비정규 경로·알 수 없는 필드는 거절한다. Stage 1 완료 여부는 폴더 존재만으로 판정하지 않고 기존 handoff·writer report·identity/seal 검증으로 확인한다. + +### 4.2 executor에 실제 전달하는 방식 + +**직접 인계를 함수 signature에만 적고 현재처럼 실제 진입점에서 인자가 없는 상태로 남겨서는 안 된다.** 개정 계약은 호출 데이터가 유일한 `run_code` 실행의 `run_inline_mcp`에 실제 도달하는 지점까지 포함한다. + +권고 함수 경계는 다음과 같다. 아래는 설계용 signature이며 완성 코드가 아니다. + +```python +run_inline_mcp(direct_handoff, *, client=None) +_inline_validate_direct_handoff(direct_handoff) +_inline_hydrate(localdocs, temp_root, validated_handoff) +``` + +backend가 지원하는 기존 task input binding이 확인되면 두 root를 위 진입점에 넘긴다. 다만 현재 YAML의 `run_code` 파라미터에는 `language`, `requirements`, `network`, `timeout`, `code`만 있고 임의 `arguments`나 환경변수 입력 계약은 확인되지 않았다. `{{stage1_run_root_ref}}` 같은 템플릿 지원을 사실로 가정하지 않는다. + +별도 인자 binding이 없고 backend가 실행 요청의 `code`를 조립하는 기능을 제공한다면, **정적 ingress 본문과 canonical JSON을 base64로 인코딩한 데이터 literal, 고정 진입점 호출**로 한 번의 `run_code` 요청을 구성하는 방식을 대안으로 명세한다. 원시 경로를 Python 코드에 문자열 치환하지 않고 decode 후 동일한 닫힌 입력 검증을 적용한다. 본문·입력 데이터·최종 실행 code hash를 구별하고, 조립기는 임의 코드 변형을 허용하지 않는다. 현재 workflow의 `dynamic_code_allowed:false`와 inline parity 규칙에 대한 이 제한된 데이터 결속 예외도 함께 명시·검증해야 한다. + +두 방식은 실제 backend capability에 따라 **하나만 채택**한다. 이 문서는 어느 방식도 현 플랫폼에서 이미 지원된다고 확정하지 않는다. 구현 착수 시 확인할 외부 정보는 “호출의 두 문자열이 현재 `run_code` 진입점에 어떻게 전달되는가”로 좁힌다. 새로운 request 파일·prepare task를 그 확인의 우회책으로 되살리지 않는다. + +### 4.3 ID는 출력 호환에 필요한 내부 값으로 처리 + +현행 `request_id`는 `run_binding_receipt`와 stdout, S2_10 인계 계약에 존재한다. `attempt_id`는 `publish_atomically`의 임시 staging 이름에 사용된다. 두 값을 전면 삭제하면 ingress·downstream schema와 검증기 변경이 늘어나므로 **외부 부여만 폐기하고 내부 호환 필드는 유지**한다. + +최소 구현안은 core 진입 전 adapter에서 다음과 같이 값을 결정하는 것이다. + +```text +request_id = "S2REQ-" + canonical_digest({ + user_context_sha256, + workspace_context_sha256, + stage1_run_root_ref, + stage1_deployment_root_ref, + expected_stage2_release_sha256, + pass_1_raw_hash_map_digest +}) + +attempt_id = "ATT-" + executor 내부에서 생성한 UUID hex +``` + +`pass_1_raw_hash_map_digest`는 request 파일을 제외한 실제 파일 closure를 두 read-pass에서 검증한 기존 hydration receipt의 값이다. 두 pass가 일치한 뒤 기존 `canonical_digest`로 `request_id`를 결정하므로 추가 원격 read나 새 입력 manifest가 필요하지 않다. UUID에는 Python 표준 라이브러리를 사용한다. 임시 디렉터리마다 격리되므로 원격 attempt 파일도 추가하지 않는다. + +`attempt_id`·현재 시각·랜덤값은 canonical output이나 run binding에 넣지 않는다. `request_id` 결정은 동일 workspace·동일 root·동일 파일 closure·동일 release에서 재현 가능해야 한다. 입력 root가 이동하면 hydration receipt의 logical path도 달라지므로, 경로가 다른 실행까지 output byte 동일성을 보장한다고 확대하지 않는다. + +이 내부 값은 기존 `execute_ingress(..., run_id=request_id, attempt_id=attempt_id, ...)`에 전달할 수 있어 core signature와 출력 ID 필드를 대체로 유지할 수 있다. 최종 `run_id`와 output root는 계속 기존 `run_binding_digest`에서 결정한다. 현행 binding digest의 핵심 재료는 `input_set_digest`, `stage2_release_digest`, `algorithm_digest`, `release_class`이며 request/attempt ID를 추가하지 않는다. + +## 5. Stage 1 원본과 배포 자산 재사용 + +### 5.1 사건 고정 입력 16개 + +다음은 IO 문서 §3.2의 경로를 `U/` 기준으로 묶은 목록이다. 파일 수는 그룹을 전개하면 정확히 16개다. + +| 묶음 | 정확한 상대경로 | +|---|---| +| P1 root 3개 | `evidence_indexed.json`, `evidence_event_candidates.json`, `client_goal.json` | +| P1 routing 2개 | `routing/domain_screening.json`, `routing/domain_activation_manifest.json` | +| P1 gate/handoff 3개 | `quality_gates/B1_evidence_indexed_gate.json`, `quality_gates/B2_event_candidates_gate.json`, `quality_gates/stage1_part1_soft_gate_handoff.json` | +| P2 3개 | `BO.json`, `signals/signal_manifest.json`, `quality_gates/stage1_part2_review_handoff.json` | +| P3 2개 | `legal_effect_structures.json`, `quality_gates/stage1_part3_review_handoff.json` | +| P4 3개 | `Fact_Ledger_base.json`, `stage1_tmp/fact_ledger/fact_ledger_writer_report.json`, `quality_gates/stage1_part4_review_handoff.json` | + +구현의 source-of-truth는 release의 `/stage1_sources`다. 전략 표를 별도의 고정 경로 목록으로 코드에 중복 작성하지 않는다. 원본 JSON을 새 Stage 1 계약으로 변환해 저장하거나 Stage 1에 새 handoff/seal 파일을 요구하지 않는다. 필요한 Stage 2 내부 adapter는 원본 raw hash·pointer를 보존하면서 메모리 객체를 만든다. + +### 5.2 고정 16개 밖의 필수·조건부 입력 + +| 입력군 | 현재 범위 | 개정 원칙 | +|---|---|---| +| signal payload family | `U/signals/` 전체 | `files[]`에서만 전개하고 행 순서·hash·semantic/integrity 구분을 유지한다. 폴더 scan으로 대체하지 않는다. | +| 두 activation | `U/routing/domain_activation_manifest.json` 및 manifest에 포함되는 `U/signals/domain_activation_manifest.json` | 별개 경로의 원본을 유지하고 기존 semantic projection 비교를 수행한다. | +| Stage 1 고정 배포 | `/dependency_locks/stage1/concrete_paths`의 55개 | `D/`의 runtime manifest·registry·domain config·schema를 exact path/hash로 읽는다. 사건 입력과 혼동하지 않는다. | +| Stage 2 직접 의존 | `/dependency_locks/stage2_direct`의 49개 | 49개 전부의 기존 closure 검증을 유지한다. | +| Stage 2 메타데이터·schema | release, module manifest, ingress/context/review schema 3개 | 위 49개와 별도로 기존 hydration에 포함한다. | +| contract manifest·completion seal | release가 유효한 참조를 선택한 경우 | 현재 미설정 참조를 새 Stage 1 필수 파일로 승격하지 않는다. | + +49개 직접 자산은 S2_10 Agent·LLM binding·authority registry/release **4개**, substantive **22개**, crosscut **5개**, special-law **13개**, overlay **5개**다. 각 filename·위치는 [IO 문서 §3.5](../../../Stage_2_00_IO_info.md)와 release가 정의한다. 모든 자산을 읽어 검증하는 것과 그 본문을 모두 LLM에게 전달하는 것은 다른 작업이다. + +이번 개정에서는 49개를 선택된 것만 읽도록 축소하거나 공유 cache를 새로 도입하지 않는다. 그러려면 release closure·cohort·검증 정책까지 바뀐다. 가장 큰 절감은 기존 자산 재사용과 request 전달층 제거에서 얻는다. `selected_context_refs=[]`를 임의로 채우거나 49개 전체를 prompt에 넣어서 정상 handoff를 흉내 내지 않는다. + +## 6. 실제 변경 지점과 최소 영향 범위 + +### 6.1 Agent와 ingress adapter + +| 변경 지점 | 최소 개정 내용 | 재사용하는 부분 | +|---|---|---| +| Agent tasks·DAG | prepare task 제거; `IN → Task_S2_00_deterministic_ingress → OUT`; ingress `wait_until: [IN]` | ingress task 이름·executor·300초·`httpx==0.28.1`·network·비 LLM 실행 | +| task pointer | ingress의 `/Agent/Stages/0/tasks/1/parameters/code`를 `/Agent/Stages/0/tasks/0/parameters/code`로 갱신 | 단일 inline code 실행 원칙 | +| 직접 입력 검사 | `_inline_validate_request(raw)`의 파일/6필드 계약을 두 root의 `_inline_validate_direct_handoff`로 대체 | strict JSON 처리, NFC·상대경로 검증 helper | +| hydration | `_inline_hydrate`가 전달받은 validated handoff 사용; `INLINE_REQUEST_PATH` read·read-pass 등록 제거 | release pin, 크기·개수 상한, exact materialization plan, 두 read-pass, 원본 bytes 복제 | +| 경로 전개 | `_inline_release_materialization_plan`의 인자를 handoff로 변경; 두 root key는 그대로 사용 | 16개·manifest 신호·55개·49개 상대경로 전개 | +| core 호출·receipt | §4.3의 내부 ID를 기존 core에 전달; stdout도 같은 request ID 사용 | C00–C15, run binding, output schema, 정확한 artifact set 검증 | +| 원격 발행 | request read/write 제거 외에는 기존 status-last 발행 유지 | non-status write/read-back, status 마지막 write/read-back, 기존 barrier 전체 bytes 비교 | + +검증과 ingress의 통합은 **하나의 task에서 차례로 처리한다**는 의미다. 실제 Stage 1/자산의 두 read-pass는 변경 중인 원본을 검출하므로 유지한다. 이를 “검증을 한 번만 한다”는 이유로 단일 read로 줄이지 않는다. + +고정 request 경로를 없애면 그 경로의 실행 간 덮어쓰기와 prepare→ingress 사이 직렬화 문제는 사라진다. 그러나 같은 `run_binding_digest`로 동시에 발행하는 문제까지 없어지지는 않는다. 같은 workspace/binding의 publication에는 기존 host 단일 실행 통제를 적용·확인한다. 현재 remote writer는 overwrite와 논리 barrier이므로 이를 원자적 CAS라고 부르지 않는다. 신규 lock JSON을 만드는 방식으로 실행 보장을 대신하지 않는다. + +### 6.2 함께 갱신할 계약·파생물 + +아래는 **향후 YAML 개정 시** 필요한 변경 범위다. 이번 전략 작성에서는 변경하지 않는다. `R/`는 본 `Stage_2_Clean/` 패키지이고 authoring은 Stage 2 main 폴더 기준이다. + +| 파일/묶음 | 필요한 조치 | +|---|---| +| `Stage_2_S2_00_v.2.yml` → `R/agent_scripts/Stage_2_S2_00.yml` | receipt와 builder가 지정한 authoring부터 개정하고 배포본을 재투영한다. 동명 구 authoring으로 덮어쓰지 않는다. | +| `R/runtime/s2_00_ingress.py`, `.txt` | 단일 inline code와 byte parity를 재생성한다. runtime import 자산으로 전환하지 않는다. | +| `R/workflows/S2_00_stage1_ingress_normalize_and_bundle_compile.yml` | 두-task/고정 request 계약을 direct handoff 계약으로 변경하고 task pointer·root/ID 책임을 정정한다. | +| `R/deployment/stage2_code_executor_binding.yml` | S2_00 row의 prepare 계약·fixed request read allowlist 제거; direct input 전달·단일 task hash와 pointer 갱신. 다른 Stage의 request 계약은 유지한다. | +| `R/offline_build/build_s2_00_inline_projection.py`, `.txt` | `EXACTLY_TWO_TASKS_REQUIRED`·prepare 전제·두 pointer/receipt 생성을 단일 ingress 계약으로 변경한다. | +| `R/manifest/s2_00_inline_code_receipt.json` | authoring/projection·단일 code·mirror·DAG hash와 pointer를 다시 기록한다. | +| `R/schemas/deployment.schema.json` | 실제 S2_00 binding의 prepare 필수 규칙을 직접 인계 규칙으로 제한 변경한다. | +| `R/schemas/ingress.schema.json` | 출력·run binding schema는 유지한다. 기존 `execution_request` 정의는 runtime 미사용으로 두어 불필요한 schema/hash 변경을 줄이고, 새 직접 입력의 runtime 닫힌 검증은 adapter에 둔다. | +| `R/runtime/s2_00_prepare_request.py`, `.txt` | 신규 build·manifest·실행 참조에서 제외한다. 역사 파일의 물리 삭제를 개정 전제 작업으로 삼지 않는다. | +| 영향받는 module/parent/child manifests·receipts | 변경 파일을 실제 참조하는 hash·pointer·algorithm 계약만 기존 build 순서로 재봉인한다. unrelated 본문 수정이나 전체 package 재설계는 하지 않는다. | +| 기존 S2_00 시험·mirror | 두-task·request file 전제와 직접 입력/ID 시험을 제한 수정하고 §9의 필요한 실패·보존 검증을 수행한다. | + +공유 executor binding bytes가 바뀌면 다른 Stage가 그 binding hash를 참조할 수 있다. “단일 YAML만 수정”이나 “항상 몇 개 파일만 수정”이라고 미리 확정하지 않고 실제 역참조 closure를 확인한다. parent/module/self-hash의 기존 비순환 규칙을 유지하며, 직접 입력 정책의 의미 변경은 알고리즘 계약/버전에 반영한다. 빌드 산출 hash 변경은 법률 규칙 내용 변경과 구별한다. + +## 7. 기존 목표 달성을 위한 최소 의미 보완 + +직접 인계는 전송 경계의 단순화다. 현행 C00–C15가 완전한 S2_10 문맥을 이미 만든다는 전제는 채택하지 않는다. [구형 분석 §8](../../../Stage_2_00_Analysis_v1.md)의 제한과 현재 inline 비교를 근거로, 다음 사항을 목표 달성 검증에 포함한다. + +| 제한 | 가장 작은 Stage 2 내부 보완 | 완료 증거 | +|---|---|---| +| 00-L2 signal adapter ID 불일치 | release의 `S2A-SIGNAL-ALL-V1`과 core의 family 검사 계약을 하나로 정합화한다. 검증을 제거하거나 모든 문자열을 허용하지 않는다. | 실제 release row에서 family mismatch가 없고 signal ALL hash·개수·activation 검사가 작동한다. | +| 00-L4/L7 raw pointer·slice provenance | 기존 원본 adapter가 실제 배열/wrapper 위치를 결정하고, 기존 source-ref helper로 원본 row의 정확한 pointer·raw-value hash를 projection에 결속한다. | fact/BO/LES/EVENT ref를 원본 bytes에서 dereference해 같은 값·hash를 얻는다. | +| 00-L5 fact·client goal 의미 축약 | 기존 FACT/case payload projection에 법률판단에 필요한 원문 사실·domain effects·목표 지시를 해당 cluster 범위로 전달한다. | 중요한 원문 사실·목표 지시를 S2_10용 문맥에서 원본 참조와 함께 확인한다. | +| 00-L6/L9 signal·review·profile 전달 공백 | 기존 slice projection 배열을 해당 cluster의 명시적 source 관계로 채우고, mapped review의 내용·범위·blocking 정보도 보존한다. | 원본 occurrence와 output occurrence의 대응이 확인되고 unresolved review가 유실되지 않는다. | +| 00-L8 slot skeleton | 원본에 존재하는 slot/fact/evidence 연결만 기존 crosswalk에 반영한다. 근거가 없으면 `UNEVALUABLE`을 유지한다. | 빈 skeleton을 완전한 연결로 오인하지 않고, 원본의 명시적 연결은 누락 없이 인계한다. | + +보완은 기존 adapter·source-ref·projection 함수와 기존 산출물에 집중한다. Stage 1 원본 재작성, 별도 Stage 1 규격 도입, LLM task 추가는 필요하지 않다. 기존 schema로 원문과 출처를 담을 수 없는 경우에만 해당 projection schema와 직접 소비 계약을 함께 제한 변경한다. 확인되지 않은 연결을 추론하여 채우지는 않는다. + +먼저 직접 인계 경계만 개정해 비교 가능한 기준을 만들고, 원본을 보존한 fixture에서 위 공백을 검사한다. **정상 ingress를 가로막거나 필요한 의미를 유실하는 항목은 완료 판정 전에 보완한다.** 기존 review/status-only 기능을 지우거나 release guard를 꺼서 시험을 통과시키지 않는다. S2_10 자체의 adapter/echo/multi-wave 미결 문제는 직접 인계 성공으로 해결된 것으로 기록하지 않는다. + +## 8. 효율적인 개정 순서 + +1. **정본과 입력 계약 고정:** 현행 authoring/배포 hash, 두 root 공급 위치, backend의 실제 직접 입력 전달 방식을 확인한다. Stage 1 결과의 bytes와 경로는 고정한다. +2. **단일 task와 adapter 경계 개정:** prepare/DAG 제거, 직접 입력 validator·hydration 인자·내부 ID·실제 진입점을 함께 변경한다. request 파일과 새 prepare 기능은 만들지 않는다. +3. **원본 재사용·의미 인계 확인:** 16개+신호+55개+49개 closure를 그대로 사용해 C00–C15를 수행하고 §7에서 실제 드러나는 최소 adapter/projection 결함을 보완한다. +4. **계약·projection 일괄 갱신:** authoring을 기준으로 builder·workflow·S2_00 binding·mirror·receipt·영향받는 hash closure를 한 번에 재생성한다. 중간마다 unrelated 자산을 재작성하지 않는다. +5. **필요한 검증 종료:** §9의 정적·offline 결과를 확정하고, 허용된 release와 실제 backend가 준비된 때 동일 입력 경계의 live 검증을 수행한다. 정적 PASS와 live 성공은 따로 기록한다. + +작업량은 executor task **2→1**, 별도 제어 파일 **1→0**, 호출자가 부여하는 ID **2→0**으로 줄어든다. request 저장/read-back·후행 재읽기와 prepare 세션도 제거된다. 총 시간·비용 절감률은 측정하지 않았으며, 16/55/49개와 동적 입력의 핵심 검증 비용은 유지한다. + +## 9. 수용 기준과 검증 계획 + +아래는 향후 구현의 검증 계획이며 이번 작성 작업에서 실행한 시험 결과가 아니다. + +| 검증 면 | 필수 수용 기준 | +|---|---| +| 실행 구조 | S2_00 Agent task와 `run_code`가 정확히 하나; ingress pointer가 `tasks/0`; IN/OUT edge와 workflow·binding·receipt가 일치한다. | +| 실제 직접 입력 | 호출에서 주어진 두 root가 executor 진입점에 그대로 도달한다. 누락·미치환·unknown key·비NFC·절대/탈출 경로는 원본 읽기·output write 전에 거절한다. | +| request 제거 | S2_00 실행에 `stage2_control/s2_00_request.json` read/write가 없고, 호출자가 ID를 공급할 필요가 없다. 기존 파일이 남아 있어도 소비하지 않는다. | +| 원본·closure | Stage 1 원본의 변경은 0; 16개·55개·49개와 manifest 신호·조건부 입력을 release대로 전개한다. 한 파일 hash 불일치·원본 read-pass 변경을 검출한다. | +| ID·재실행 | 동일 직접 입력/closure/release의 내부 request ID가 동일하고 attempt 값은 canonical 결과를 바꾸지 않는다. 기존 output root를 바꾸지 않으며 동일 bytes는 재사용하고 충돌은 거절한다. | +| 문맥 보존 | §7의 실제 source pointer를 dereference하고 사실·목표·신호·review·명시적 slot 관계의 필요한 내용이 후속 slice에서 확인된다. 원본 미상은 미상으로 남는다. | +| 정상 route | `TO_S2_10`/`TO_S2_10_WITH_ISSUES`; 아래 고정 11개+실행 가능 slice E개가 schema·artifact set/hash 검증을 통과한다. | +| 진단 route | `TO_S2_40_STATUS_ONLY`; 아래 5개만 발행하고 정상 context를 만들지 않는다. 직접 입력 오류가 항상 진단 5파일을 만든다고 가정하지 않는다. | +| 발행 실패·동시 실행 | non-status write/read-back 실패 시 유효한 새 barrier를 게시하지 않는다. 기존 barrier와 잔여 파일을 실제 확인하고 동일 binding 동시 writer 통제를 검증한다. | +| 빌드·release | authoring↔projection↔inline/mirror parity, task pointer와 참조 hash closure, algorithm 의미 계약이 일치한다. 기존 DEV guard와 pending admission을 보존한다. | +| live | 직접 입력 전달·인증 세션·closure 읽기·실제 발행·downstream barrier 소비를 실환경에서 확인한다. offline 함수 성공만으로 live 완료를 선언하지 않는다. | + +정상 고정 출력 11개는 `ingress/stage1_input_manifest.json`, `ingress/intake_report.json`, `review/issue_ledger.base.json`, `ingress/ingress_status.json`의 공통 4개와 `context/case_context.json`, `context/evidence_inventory.json`, `context/object_registry.json`, `context/party_and_title_context.json`, `context/slot_crosswalk.json`, `context/cluster_plan.json`, `context/bundle_plan.json`의 7개다. 가변 출력은 `context/cluster_slices/.json`이다. + +진단 5개는 공통 4개와 `ingress/technical_diagnostic.json`이다. 모든 최종 파일은 `W/stage2_runs/by-binding//`에 저장하고 `ingress/ingress_status.json`을 마지막에 발행한다. stdout receipt와 임시 hydration 파일은 이 출력 개수에 포함하지 않는다. + +단일 task, 직접 입력 validator, 원본 closure, C00–C15 의미 보존, 분기 출력과 barrier가 함께 충족되어야 S2_00 개정 목표 달성으로 판정한다. 법률적 판단의 타당성, S2_10 플랫폼/model/legal admission 또는 전체 Stage 2 production readiness는 별도의 확인 대상이다. + +## 10. 결정·검증 기록 + +| 결정 | 근거 | 효과 | +|---|---|---| +| 직접 인계와 ingress를 유일 task에 통합 | 사용자 개정 방향; 현행 prepare 인자 결속 미검증 | 별도 제어 파일과 task 사이 전달을 제거한다. | +| 두 root만 호출 데이터로 사용 | 기존 release가 16/55/49 및 신호 경로를 정의 | 개별 파일 경로·ID 준비 작업을 줄이고 Stage 1 변경을 피한다. | +| ID 출력 필드는 내부 호환 유지 | 현행 run binding/stdout 및 S2_10 계약 | downstream 전면 변경을 피한다. | +| 두 read-pass·직접 자산 49개 유지 | 현행 hydration·release closure 계약 | 전송층 단순화가 무결성 검증 축소로 이어지지 않도록 한다. | +| 필요한 의미 보완은 Stage 2 adapter/projection에 한정 | 구형 분석 §8; 현행/구형 ingress code 차이는 release pin 하나 | 입력 통합과 실제 목표 달성을 구별하면서 원본을 보존한다. | + +문서 작성 중 현행/구형 YAML, 두 분석 종류, IO 표, release의 55/49 경로 수, authoring 정본 및 inline code 차이를 정적으로 대조했다. 저장 후 확인에서 로컬 출처 링크 11개, Markdown fence 8개, UTF-8·종결 LF, release와 일치하는 사건 입력 16개 및 MEMORY 요약 위치·유일성이 통과했다. 참조한 원본 문서·YAML·자산 15개의 SHA-256은 작업 전후 동일했다. 이는 문서 정적 확인 결과다. YAML 개정, builder 실행·재봉인, 회귀시험, MCP/backend/live 실행, 배포·commit은 이번 작업에서 수행하지 않았다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v2.md b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v2.md new file mode 100644 index 00000000..18938ea0 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v2.md @@ -0,0 +1,196 @@ +# Stage_2_S2_00.yml 개정 전략 v2 + +작성일: 2026-10-02. 상태: **전략서 작성 완료 / YAML 개정·실행 시험 미수행**. + +## 1. 목표·산출물·범위 + +**Stage 1 사건·배포 root를 S2_00의 단일 ingress task에 직접 전달하고, Stage 1 원본을 그대로 읽어 검증·보존·정규화·묶음 구성·발행하는 것으로 S2_00을 완결한다. 별도의 `request_id`는 외부 입력, 내부 생성, 출력 모두에서 제거한다.** + +이번 작업의 산출물은 이 전략서와 Stage 2 `MEMORY.md`의 작업 요약뿐이다. 실제 YAML·schema·runtime·manifest·builder·배포 자산은 수정하지 않는다. 향후 구현의 유일한 개정 대상도 이 폴더의 [Stage_2_S2_00.yml](Stage_2_S2_00.yml)이다. 다른 authoring 파일부터 수정하고 이 YAML을 재생성하는 v1의 절차는 이번 범위에 적용하지 않는다. + +`Stage_2_S2_10.yml`, `Stage_2_S2_20.yml`, `Stage_2_S2_30.yml`, `Stage_2_S2_40.yml`의 자산 사용 schema, 인계 필드, route, 출력 수, admission·hash 호환 요구는 설계 기준과 수용 기준에서 제외한다. 해당 파일을 분석하거나 개정하는 작업도 포함하지 않는다. S2_00의 자체 입력·출력 계약과 검증은 개정 YAML의 metadata 및 inline code 안에 둔다. + +S2_00의 현행 기능 목표인 C00 입력 검증, C05 보존·정규화, C10 claim-neutral cluster 구성, C15 묶음·결과 발행은 유지한다. 후속 Agent에 맞추기 위한 구조·필드의 보존은 목표에서 제거한다. 이 문서의 함수 signature, 출력 구조와 변경 순서는 구현 제안이며 현 플랫폼의 지원 사실이나 실행 결과가 아니다. + +## 2. 확인한 기준과 v1에서 폐기할 전제 + +| 기준 자료 | 현재 확인 내용 | v2에서의 사용 | +|---|---|---| +| [현행 S2_00 YAML](Stage_2_S2_00.yml) | Agent `Stage_2_S2_00_v2`, version `1.2.0`; prepare·ingress 두 task; inline C00–C15 | S2_00 내부 변경 위치와 보존할 처리 기능의 기준 | +| [현행 S2_00 분석서](Analysis_Stage_2_S2_00.md) | 직접 인자 결속 미검증, 고정 request 파일, 정상/진단 발행, DEV release 제한 | 현재 구현과 미확인 운영 경계 구별 | +| [개정 전략 v1](stage_2_s2_00_revision_strategy_v1.md) | 내부 request ID 유지, 기존 downstream 출력/schema와 49개 자산 유지, 여러 파일 재봉인 계획 | 사용자 원칙에 맞지 않는 전제를 식별하는 비교 자료 | +| [현행 release](../manifest/stage2_release.json) | Stage 1 고정 사건 입력 16개와 동적 signal family; Stage 1 배포 잠금 55개; Stage 2 직접 잠금 49개; `DEV_FIXTURE_RELEASE` | Stage 1 경로·adapter·검증 근거의 확인 자료. Stage 2 전체 의존 목록을 v2에 자동 승계하지 않음 | + +v1의 다음 결정은 v2에서 대체한다. + +- 내부 `request_id = "S2REQ-" + digest(...)` 생성과 stdout·receipt 출력은 전면 폐기한다. +- `attempt_id` UUID를 생성해 기존 core signature에 맞추는 방식도 폐기한다. 임시 작업 격리는 실행별 임시 디렉터리로 해결한다. +- 기존 `run_binding_receipt`, `S2RUN-...` 생성, `by-binding/` 출력 경로를 유지할 호환 의무를 제거한다. 무결성 hash는 실제 검증에 필요한 용도로만 사용한다. +- 기존 정상 11+E개·진단 5개 출력, S2_10용 bundle schema, S2_40용 status-only route를 수용 기준으로 삼지 않는다. +- Stage 2 직접 자산 49개 일괄 읽기와 다른 Stage의 hash closure 재봉인은 개정 절차에서 제외한다. + +## 3. 절대 원칙과 완료 기준 + +1. 호출 데이터는 `stage1_run_root_ref`, `stage1_deployment_root_ref` 두 값이다. 사건 자료와 배포 자료의 실제 위치를 추정하거나 새 request 파일로 중계하지 않는다. +2. Stage 1 원본의 bytes·파일명·상대경로·원문 의미를 보존한다. 기존 사건/transaction 식별자가 원본에 있으면 출처 검증에 그대로 사용하며 새 요청 식별자로 재포장하지 않는다. +3. prepare task, 고정 request 파일 read/write, 별도의 ID 발급·저장·출력을 제거한다. `request_id`를 이름만 바꾼 `handoff_id`·`execution_id`도 만들지 않는다. +4. S2_00은 비 LLM deterministic task 하나로 수행한다. 입력 본문 전체를 prompt에 복사하거나 LLM task를 추가하지 않는다. +5. 검증은 S2_00의 실제 입력·연산·출력에 필요한 것으로 한정한다. 원본의 무결성·출처·필수 내용·blocking review를 생략하여 단순화하지 않는다. +6. 정상 결과는 원본을 추적할 수 있고 미해결 사항을 보존해야 한다. 입력 전달 성공만으로 S2_00의 기능 목표 달성을 선언하지 않는다. + +완료는 **두 root의 실제 인자 결속 → 원본 읽기와 필요한 검증 → 의미 보존·정규화·cluster/bundle 구성 → 자체 출력 검증 → 마지막 status 발행**이 한 task에서 성립하는 것으로 판정한다. 기존 S2_10~40에서 읽을 수 있는지는 판정 대상이 아니다. + +## 4. 직접 입력과 실행 경계 + +호출 데이터의 닫힌 형태는 다음과 같다. 이는 파일로 저장할 request가 아니라 실행 인자다. + +```json +{ + "stage1_run_root_ref": "", + "stage1_deployment_root_ref": "<그 결과에 대응하는 Stage 1 배포의 workspace 상대 root>" +} +``` + +`W/`는 인증된 localdocs 논리 workspace root, `U=W//`, `D=W//`다. Dropbox 절대경로나 executor 임시 mount와 동일한 것으로 취급하지 않는다. 인증 정보는 기존 backend 세션을 사용하며 호출자가 임의 workspace를 지정하게 하지 않는다. + +권고 경계는 다음과 같다. + +```python +run_inline_mcp(stage1_run_root_ref, stage1_deployment_root_ref, *, client=None) +validate_direct_roots(stage1_run_root_ref, stage1_deployment_root_ref) +hydrate_stage1(localdocs, temp_root, validated_roots) +execute_ingress(stage1_root, stage1_deployment_root, *, output_dir, source_checks) +publish_result(localdocs, output_root, files, status) +``` + +별도의 request 객체를 core에 요구하지 않는다. core의 `run_id="S2-00-REQUEST"`, `attempt_id` 필수 인자와 ID 정규식 검사는 삭제하고 필요한 경로·검증 객체만 전달한다. 내부 함수 분리는 유지할 수 있으나 추가 executor task로 나누지 않는다. + +두 root는 canonical NFC 상대경로인지 검사한다. 누락·빈 값·미치환 템플릿·절대경로·`..`·NUL·경로 탈출·알 수 없는 입력 필드는 원본 읽기나 출력 쓰기 전에 거절한다. Stage 1 완료 여부는 폴더 존재만으로 판단하지 않고 기존 writer report·gate/handoff와 원본 간 결속을 확인한다. 기존 완료/seal 정보가 없으면 없다고 기록하며 새 Stage 1 파일을 요구하지 않는다. + +실제 backend가 이 두 문자열을 `run_code` 진입점에 어떻게 전달하는지는 아직 미확인이다. 지원되는 기존 task input binding을 우선 확인한다. 실행 code 조립을 지원하는 경우에만 canonical JSON의 인코딩된 데이터 literal과 고정 호출로 두 값을 전달하는 제한된 대안을 검토한다. 원시 문자열을 Python 코드에 직접 치환하지 않는다. `{{prev...}}`의 cross-run 지원, 임의 `arguments` 파라미터, 환경변수 주입 기능을 가정하지 않는다. + +실제 전달 방식이 YAML 한 파일 개정으로 표현되지 않으면 **직접 입력 live 연결 미확인**으로 남긴다. 이를 이유로 prepare/request 파일을 되살리거나 backend·다른 자산의 개정을 전략 범위에 추가하지 않는다. 함수와 offline fixture 검증은 이 운영 미확인 사항과 구별하여 수행할 수 있다. + +## 5. Stage 1 원본 재사용과 필요한 자산만 읽기 + +현행 고정 사건 입력 16개를 그대로 사용한다. 아래 경로는 모두 `U/` 기준이다. + +| 묶음 | 원본 상대경로 | +|---|---| +| P1 3개 | `evidence_indexed.json`, `evidence_event_candidates.json`, `client_goal.json` | +| P1 routing 2개 | `routing/domain_screening.json`, `routing/domain_activation_manifest.json` | +| P1 gate/handoff 3개 | `quality_gates/B1_evidence_indexed_gate.json`, `quality_gates/B2_event_candidates_gate.json`, `quality_gates/stage1_part1_soft_gate_handoff.json` | +| P2 3개 | `BO.json`, `signals/signal_manifest.json`, `quality_gates/stage1_part2_review_handoff.json` | +| P3 2개 | `legal_effect_structures.json`, `quality_gates/stage1_part3_review_handoff.json` | +| P4 3개 | `Fact_Ledger_base.json`, `stage1_tmp/fact_ledger/fact_ledger_writer_report.json`, `quality_gates/stage1_part4_review_handoff.json` | + +signal payload는 기존 `signal_manifest.files[]`의 경로·hash·조건을 따라 읽는다. 폴더 전체 scan이나 추측한 파일 목록으로 대체하지 않는다. routing activation과 signal activation은 서로 다른 원본으로 취급하며 실제 의미 일치 여부를 확인한다. + +Stage 1 배포의 현행 55개 잠금 목록은 검증 근거의 기준선이다. v2에서는 각 자산이 사건 JSON 해석, schema 검증, producer/alias·domain 연결, 원본 identity 검증 중 어디에서 필요한지 S2_00 내부에서 추적한다. 필요 자산과 그 검증 의존은 유지하고, 미사용임이 확인된 자산만 읽기 대상에서 제외한다. 검토 전부터 특정 감소 개수를 약속하지 않는다. Stage 1 원본이나 배포 파일을 고쳐 개정에 맞추지 않는다. + +Stage 2의 현행 49개 직접 의존은 후속 작업을 위한 결속까지 포함하므로 전체를 유지하지 않는다. S2_10 YAML·LLM binding과 후속 authority/retrieval 선택용 자산은 S2_00 의존에서 제거한다. registry/config도 C05~C10에서 실제 사용하는 내용만 남기고, 남기는 이유·경로·검증 hash를 YAML 내부의 단일 자산 표에 둔다. 필요한 비실행 데이터는 읽기 전용으로 참조할 수 있으나 그 파일을 개정하거나 새 외부 schema를 만들지 않는다. 문서 안의 법률 규칙을 새로 작성하는 작업도 포함하지 않는다. + +자체 입력 adapter·출력 validator·경로 목록의 정본은 개정 YAML 안에서 하나로 관리한다. 기존 전체 release·module·downstream schema를 그대로 validator에 주입해 제거한 필드를 다시 요구하게 해서는 안 된다. 외부 참조가 있는 기존 schema는 실제 S2_00에 필요한 규칙과 참조 의존만 inline 자체 계약에 반영한다. 이것은 후속 호환 검사 제거이며 Stage 1 원본 검증의 면제는 아니다. + +## 6. C00–C15의 자체 기능 목표 + +| 논리 단계 | 유지할 작업 | 단순화·보완 방향 | +|---|---|---| +| C00 | 사건·배포 원본 읽기, 필수 입력·hash·producer/transaction·gate/review 확인 | 직접 root로 전개. request/ID 검증과 downstream 자산 admission은 제거. hash 근거가 없는 입력은 무근거 상태를 보존 | +| C05 | 증거·사실·객체·당사자·목표·LES·signal의 원본 보존과 필요한 정규화 | 원문을 덮어쓰지 않고 필요한 view만 생성. source path·정확한 JSON pointer·raw hash로 원문 연결 | +| C10 | 원본의 명시적 관계에 따른 claim-neutral cluster와 작업 범위 구성 | 확인된 fact/evidence/BO/LES/signal 관계만 사용. 빈 slot skeleton이나 근거 없는 관계를 채우지 않음 | +| C15 | cluster별 입력 묶음·미해결 항목·결과 검증과 발행 | 후속 Agent prompt·dispatch schema 대신 자체 bundle 표현. S2_10/40 route를 자체 처리 상태로 대체 | + +adapter가 wrapper/배열 위치를 잘못 해석하여 원문 pointer가 틀리거나 사실·client goal·signal·review를 축약해 의미를 잃는 문제는 S2_00 inline code 안에서 보완한다. 원문 사실, 목표 제약, 반대 자료, blocking/unresolved review, 기존 명시적 연결을 보존한다. `UNEVALUABLE`·미상은 원본 근거 없이 해결된 상태로 변경하지 않는다. + +원본의 fact/evidence/object/transaction ID는 의미와 참조를 구성하는 기존 값이므로 유지한다. cluster의 구조상 식별값이 필요한 경우에도 해당 묶음 참조 용도에 한정한다. 이러한 자료 식별값을 요청 ID나 실행 ID 발급의 근거로 사용하지 않는다. + +## 7. ID 없는 출력·재실행·임시 작업 + +최종 출력 root는 원본 사건 root를 재사용하여 다음처럼 결정한다. + +```text +O = W/stage2_runs/from-stage1//s2_00/ +``` + +검증된 상대 root의 경로 구성을 그대로 아래에 결합한다. 새 요청 ID, UUID, `run_binding_digest`로 출력 폴더를 만들지 않는다. `O/`가 `U/` 또는 `D/`와 겹치거나 어느 한쪽 안에 놓이는 경우는 쓰기 전에 거절한다. Stage 1 원본 트리는 읽기 전용으로 유지한다. + +권고 정상 출력은 다음 5개이며 cluster 수가 크거나 부분 읽기가 필요한 경우에만 가변 slice를 분리한다. 기존 11+E개 파일을 모두 보존하기 위한 빈 파일은 만들지 않는다. + +| 파일 | 필요한 내용 | +|---|---| +| `ingress/stage1_input_manifest.json` | 두 root, 읽은 원본 상대경로·raw hash·검증 근거, 기존 원본 identity가 있으면 그 값 | +| `ingress/intake_report.json` | 필수 입력·무결성·gate 확인 결과와 누락/미검증 항목 | +| `review/issue_ledger.base.json` | 원본 review의 내용·범위·blocking 여부·출처와 S2_00 처리 중 발견한 문제 | +| `context/case_context.json` | 정규화 view·원본 참조, 증거/객체/당사자/LES/사실/목표/signal 관계, cluster 목록과 bundle별 원본 참조. 분리 slice 사용 시 해당 경로 | +| `ingress/ingress_status.json` | `READY`, `READY_WITH_ISSUES`, `BLOCKED` 중 상태, 두 root, algorithm version, 실제 발행 파일 경로·hash. 마지막에 발행 | + +가변 파일은 필요한 경우에만 `context/cluster_slices/.json`으로 분리한다. ``는 자료 묶음 참조이며 요청/실행 ID가 아니다. 한 내용은 context 또는 slice 중 한 곳에 담고 다른 곳에서는 참조하여 중복을 줄인다. 원본 전체 사본을 새 context에 복제하지 않는다. + +안전하게 root를 확정하고 입력 검사를 수행한 뒤 발생한 차단은 `BLOCKED`로 기록한다. 이때 정상 context/slice는 발행하지 않으며 intake·issue ledger·technical diagnostic과 마지막 status를 발행한다. 입력 목록을 확정할 수 있을 때만 manifest도 발행한다. 입력 인자·인증 실패처럼 출력 위치를 안전하게 확정하지 못한 오류는 stdout 오류로 종료하고 결과 파일 발행을 강제하지 않는다. + +모든 출력의 닫힌 필드·상태별 required/forbidden artifact 규칙은 inline validator에서 정의한다. 기존 `run_binding_receipt`와 출력 schema의 `request_id`·신규 `run_id` 필수 조건은 함께 제거한다. stdout은 처리 성공 여부·상태·출력 경로·오류·마지막 status hash 정도만 담는다. 별도 ID receipt는 만들지 않는다. + +원본 raw hash와 산출물 hash는 변조·내용 동일성 확인을 위해 유지한다. 이 hash를 접두사와 결합하여 요청 식별자로 재출력하지 않는다. 동일 `O/`가 이미 있으면 기존 입력 경로·hash, 배포 근거, algorithm version과 발행 파일을 확인하고 일치하는 완료 결과만 재사용한다. 내용·버전이 다르거나 잔여 부분 파일이 충돌하면 덮어쓰지 않고 명시적으로 종료한다. 이번 전략은 여러 개정 버전의 결과를 같은 사건 root 아래 동시에 보관하는 새 체계를 도입하지 않는다. + +실행별 `TemporaryDirectory` 안에서 hydration·로컬 결과·staging을 처리한다. 임시 디렉터리의 고유한 이름은 도구 내부 자원이며 호출 필드, canonical output, stdout에 ID로 노출하지 않는다. `publish_atomically(..., attempt_id=...)`의 의존도 함께 제거한다. + +원격 발행은 비 status 파일 write/read-back 이후 status를 마지막에 기록하는 논리적 완료 경계로 유지한다. 임시 디렉터리가 원격 동시 writer 문제까지 해결한다고 주장하지 않는다. 같은 `O/`의 동시 writer는 기존 host 직렬화 기능을 사용할 수 있는지 확인하며, 미지원·미확인 상태에서는 동시 발행 지원을 수용 완료로 기록하지 않는다. 새 lock JSON·lock ID·CAS 지원을 가정하지 않는다. + +## 8. YAML 내부의 실제 변경 위치와 순서 + +1. **범위 고정:** 현행 배포 YAML을 기준으로 두 root, S2_00 자체 목표·출력, 필수 검증과 필요한 읽기 전용 자산을 확정한다. 다른 Stage와의 호환 검토를 하지 않는다. +2. **단일 실행 구조:** `Task_S2_00_prepare_request`를 제거한다. `task_procedure`는 `IN → Task_S2_00_deterministic_ingress → OUT`으로 바꾸고 ingress pointer를 `tasks/0`으로 정리한다. +3. **직접 입력 진입점:** `run_inline_mcp`·`_inline_hydrate`·`_inline_release_materialization_plan`을 두 root 입력으로 변경한다. `_inline_validate_request`, request 파일의 두 read-pass, prepare 관련 상수·입출력 설명을 제거한다. 실제 원본의 두 read-pass 안정성 확인은 필요한 입력 집합에 대해 유지한다. +4. **ID 의존 제거:** `execute_ingress`의 request용 `run_id`와 `attempt_id`, `_make_run_binding_receipt`의 요청/실행 ID 생성, `canonical_run_id`, CLI ID 옵션, staging/remote publisher·stdout의 ID 의존을 제거한다. 빈 문자열·고정 가짜 ID로 기존 signature를 통과시키지 않는다. +5. **자체 계약·자산 의존:** Agent description·metadata와 inline code에서 전체 Stage 2 workflow/binding/schema를 런타임 필수 계약으로 연결하는 부분을 S2_00 자체 계약으로 정리한다. C00–C10에 실제 필요한 읽기 전용 원본/schema/config 검증은 남긴다. 다른 Stage 자산을 유지하려고 전체 49개 closure를 복원하지 않는다. +6. **의미 보존과 출력:** adapter·pointer·review 처리의 실제 공백을 보완하고 C15를 자체 bundle·상태·출력 root로 바꾼다. 출력 validator, 파일 수·artifact-set 검사, publish/read-back, stdout을 함께 정합화한다. +7. **S2_00 한정 검증:** 아래 수용 기준으로 inline 함수를 시험하고 실제 전달 가능한 환경에서 단일 task를 확인한다. 다른 Stage 시험과 패키지 전체 재봉인은 개정 범위에 넣지 않는다. + +향후 수정 파일은 해당 `Stage_2_S2_00.yml` 하나다. `Stage_2_S2_00_v.2.yml`, runtime mirror, 외부 workflow/binding/schema, manifest, builder는 함께 수정하지 않는다. 그 결과 기존 package builder의 parity나 parent release 봉인을 충족한다고 주장할 수 없다. 기존 builder를 돌려 이 YAML을 다시 덮어쓰는 것도 하지 않는다. 전체 패키지의 배포 승인과 S2_00 자체 기능 구현은 구별한다. + +## 9. 실행 허용 조건과 미확인 사항 + +현재 release는 `DEV_FIXTURE_RELEASE`이며 현행 core에는 실제 사건 발행을 금지하는 `DEV_FIXTURE_REAL_RUN_FORBIDDEN` 방어가 있다. 후속 Agent와 무관한 인증·원본 신뢰 근거·S2_00 실행 허용 조건은 보존한다. 기존 전체 Stage 2 admission을 제거하면서 이를 임의의 `active:true`나 새 self-certified receipt로 대체하지 않는다. + +YAML 한 파일 안에서 기존 근거로 독립 검증 가능한 S2_00 실행 조건을 명시한다. 현재 DEV 상태만으로는 offline fixture 검증까지만 가능하다. 실제 사건 발행에 필요한 승인·backend 기능이 외부 패키지 변경을 요구한다면 그 부분은 **현재 범위에서 live 미완료**로 기록한다. 다른 Stage의 준비 상태를 S2_00 내용 설계의 제약으로 다시 가져오거나, guard를 꺼 실제 사건 시험을 통과시키는 방법은 채택하지 않는다. + +미확인 사항은 두 root의 실제 backend 결속, 인증 치환, 필요한 Stage 1 신뢰 근거의 완비, 같은 출력 root의 writer 직렬화, S2_00 독립 실행 허용 조건이다. 전략 작성이나 offline 함수 성공만으로 어느 것도 해결되었다고 기록하지 않는다. + +## 10. 수용 기준과 검증 계획 + +아래는 향후 YAML 구현의 검증 계획이며 이번 작업에서 실행한 시험 결과가 아니다. 검증용 fixture·임시 출력은 임시 디렉터리를 사용하며 저장소에 다른 시험 자산을 새로 만들 필요는 없다. + +| 검증 면 | 수용 기준 | +|---|---| +| 개정 범위 | 실제 수정 YAML은 S2_00 배포본 하나. S2_10~40·authoring·외부 자산 수정 없음 | +| 구조 | Agent stage 하나, ingress task·`run_code` 하나. DAG와 `tasks/0` pointer 일치 | +| 직접 전달 | 호출한 두 root가 실행 진입점에 그대로 도달. 경로 누락·미치환·탈출·출력/원본 중첩은 read/write 전에 차단 | +| 군더더기 제거 | prepare·고정 request 파일 read/write 0; `request_id` 입력·생성·출력 0; request를 대체하는 새 ID 0; `attempt_id` 필수 인자·출력 0 | +| 원본·검증 | Stage 1 원본 변경 0. 고정 16개·동적 signal 및 필요한 배포/schema 검증 수행. hash 불일치, manifest 탈출, 두 read-pass 변경 검출 | +| 자산 범위 | 실제 S2_00 사용 근거가 있는 자산만 읽음. 후속 YAML·LLM binding·후속 schema/route 계약을 요구하지 않음 | +| 내용 보존 | source pointer dereference 결과와 hash 일치. 사실·목표 제약·증거·신호·review·명시적 관계 누락 없음. 미상·blocking 유지 | +| 정상·차단 | 자체 상태와 실제 artifact 목록 일치. 차단 결과를 정상 context로 표현하지 않음. 원본에 없는 관계를 생성하지 않음 | +| 재실행·발행 | 동일 완료 결과만 재사용. 불일치·부분 충돌은 거절. 비 status write/read-back 실패 시 새 정상 status를 발행하지 않음. 기존 status 존재는 별도 확인 | +| 임시·동시 실행 | 임시 작업 디렉터리는 분리되고 출력에 임시 ID 없음. 원격 동시 발행 보장은 실증 전 미확인으로 기록 | +| live | 허용된 환경에서 root 결속·인증·원본 읽기·결과 write/read-back·마지막 status 확인. offline PASS와 구별 | + +검증은 먼저 경로·ID 제거·정상/차단·원본 보존·pointer·재실행·발행 실패를 다루는 한 차례의 종합 offline 검증으로 수행한다. 수정이나 새 실패가 생겼을 때 해당 영향만 재확인한다. S2_10~40과의 E2E·호환 시험, 법률 판단의 타당성 승인, 전체 패키지 production readiness는 수용 기준에 넣지 않는다. + +## 11. 효율성과 남는 비용 + +계획상 executor task는 2→1, 고정 request 파일은 1→0, 별도 request ID 생성·출력은 0으로 줄인다. prepare 호출, request 저장·read-back·재읽기, ID 호환 처리, 후속 자산의 불필요한 읽기를 제거하는 것이 절감 지점이다. 이미 Stage 1 결과를 직접 읽는 방식에 추가 ID 처리를 붙이는 것이 더 효율적이라고 주장하지 않는다. + +S2_00은 비 LLM 작업이므로 ID 제거 자체가 큰 모델 토큰 절감을 보장하지 않는다. inline code의 전달량·billing과 필요한 원본 검증 비용은 남는다. 속도·비용·토큰 절감률은 실제 실행 요청 크기, 읽기 횟수, task 시간과 청구 값을 측정하기 전까지 미측정으로 둔다. + +## 12. 작성 작업의 결정·검증 기록 + +| 결정 | 근거 | 결과 | +|---|---|---| +| 두 Stage 1 root 직접 전달과 request ID 전면 제거 | 이번 사용자 절대 원칙 | v1 §4.3의 내부 호환 ID 정책 폐기 | +| S2_00 YAML 한 파일을 향후 개정 대상으로 한정 | 사용자 범위 제한 | downstream 호환·외부 자산 재봉인 계획 제외 | +| 원본 의미·검증 유지, 자체 출력 축소 | S2_00의 C00–C15 기능 목표 | 전송 단순화와 기능 완결을 함께 검증 | +| 기존 root·자료 참조·hash 재사용 | 추가 요청 식별자를 만들지 않는 원칙 | 원본 보존, 출처 확인, 재실행 충돌 판별 가능 | +| backend·live 허용 조건 미확인 유지 | 현행 S2_00 진입점과 DEV guard | 전략·offline·live·패키지 승인 상태 구별 | + +이번에는 이 전략서 생성과 지정 MEMORY 요약만 수행했다. 문서의 정확한 저장 경로·UTF-8·종결 LF, fence 6개·로컬 링크 5개, 사용자 원칙 반영과 요약 위치를 정적으로 확인했다. 작업 전 기준선 610개 파일 중 MEMORY를 제외한 기존 609개는 SHA-256이 동일했고, MEMORY도 요약 삽입 이외의 기존 내용은 보존했다. 새 파일은 이 전략서 하나다. YAML 개정, inline 실행, builder·release 재봉인, MCP/backend/live 시험, 배포·commit은 수행하지 않았다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v3.md b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v3.md new file mode 100644 index 00000000..4bcd1935 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v3.md @@ -0,0 +1,221 @@ +# Stage_2_S2_00.yml 개정 전략 v3 + +작성일: 2026-10-02. 상태: **전략서 작성 완료 / YAML 개정·실행 시험 미수행**. + +## 1. 목표·산출물·범위 + +**Stage 1 사건·배포 root를 S2_00의 단일 ingress task에 직접 전달하고, Stage 1 결과물은 backend Python 실행 환경이 지원하는 `{{prev.###}}` 참조로 그대로 사용한다. `###`는 Stage 1의 결과물 파일이다. 원본을 검증·보존·정규화·묶음 구성·발행하는 것으로 S2_00을 완결하며, 별도의 `request_id`는 외부 입력, 내부 생성, 출력 모두에서 제거한다.** + +이번 작업의 산출물은 이 전략서와 Stage 2 `MEMORY.md`의 작업 요약뿐이다. 실제 YAML·schema·runtime·manifest·builder·배포 자산은 수정하지 않는다. 향후 구현의 유일한 개정 대상도 이 폴더의 [Stage_2_S2_00.yml](Stage_2_S2_00.yml)이다. 다른 authoring 파일부터 수정하고 이 YAML을 재생성하는 v1의 절차는 이번 범위에 적용하지 않는다. + +`Stage_2_S2_10.yml`, `Stage_2_S2_20.yml`, `Stage_2_S2_30.yml`, `Stage_2_S2_40.yml`의 자산 사용 schema, 인계 필드, route, 출력 수, admission·hash 호환 요구는 설계 기준과 수용 기준에서 제외한다. 해당 파일을 분석하거나 개정하는 작업도 포함하지 않는다. S2_00의 자체 입력·출력 계약과 검증은 개정 YAML의 metadata 및 inline code 안에 둔다. + +S2_00의 현행 기능 목표인 C00 입력 검증, C05 보존·정규화, C10 claim-neutral cluster 구성, C15 묶음·결과 발행은 유지한다. 후속 Agent에 맞추기 위한 구조·필드의 보존은 목표에서 제거한다. `{{prev.###}}` 사용을 backend Python 실행 환경이 허용한다는 점은 이번 사용자 추가 설명에 따른 설계 전제로 채택한다. 이 문서의 함수 signature, 자체 출력 구조와 변경 순서는 구현 제안이며 이번 작업에서 실제 YAML을 개정하거나 backend 실행을 시험한 결과는 아니다. + +## 2. 개정 기준과 v2에서 달라지는 전제 + +| 기준 자료 | 기준 내용 | v3에서의 사용 | +|---|---|---| +| [현행 S2_00 YAML](Stage_2_S2_00.yml) | Agent `Stage_2_S2_00_v2`, version `1.2.0`; prepare·ingress 두 task; inline C00–C15 | S2_00 내부 변경 위치와 보존할 처리 기능의 기준 | +| [현행 S2_00 분석서](Analysis_Stage_2_S2_00.md) | v2 작성 시 기준선: 기존 직접 인자 결속 미검증, 고정 request 파일, 정상/진단 발행, DEV release 제한 | 기존 구현의 기록. 결과물 참조 지원 여부는 이번 사용자 설명을 우선 적용 | +| [개정 전략 v2](stage_2_s2_00_revision_strategy_v2.md) | root 직접 전달·ID 제거·S2_00 한정 개정; backend 결속은 미확인으로 기술 | 본 개정의 원본. 직접 전달 원칙을 유지하고 `{{prev.###}}` 지원 전제를 반영 | +| 사용자 추가 설명 | Stage 1 결과물 파일을 `{{prev.###}}`로 사용할 수 있고 backend Python 실행 환경이 이를 허용 | 결과물 접근 방식의 설계 근거. 지원 여부 재조사를 개정 선행 조건으로 두지 않음 | +| [현행 release](../manifest/stage2_release.json) | Stage 1 고정 사건 입력 16개와 동적 signal family; Stage 1 배포 잠금 55개; Stage 2 직접 잠금 49개; `DEV_FIXTURE_RELEASE` | Stage 1 경로·adapter·검증 근거의 확인 자료. Stage 2 전체 의존 목록을 v3에 자동 승계하지 않음 | + +기존 YAML·release의 설명은 v2에 기록된 기준선을 승계한다. 이번 문서 개정의 변경 근거는 사용자 추가 설명이며, 이를 별도의 live 실증 결과로 표시하지 않는다. + +v2의 직접 전달·ID 제거·S2_00 한정 원칙은 유지한다. v2 §4의 backend 전달 방식 탐색, code 조립 대안, 직접 입력 지원 여부를 미확인으로 두는 전제는 `{{prev.###}}`의 지원을 채택하는 것으로 대체한다. v1에서 이미 폐기한 다음 전제도 되살리지 않는다. + +- 내부 `request_id = "S2REQ-" + digest(...)` 생성과 stdout·receipt 출력은 전면 폐기한다. +- `attempt_id` UUID를 생성해 기존 core signature에 맞추는 방식도 폐기한다. 임시 작업 격리는 실행별 임시 디렉터리로 해결한다. +- 기존 `run_binding_receipt`, `S2RUN-...` 생성, `by-binding/` 출력 경로를 유지할 호환 의무를 제거한다. 무결성 hash는 실제 검증에 필요한 용도로만 사용한다. +- 기존 정상 11+E개·진단 5개 출력, S2_10용 bundle schema, S2_40용 status-only route를 수용 기준으로 삼지 않는다. +- Stage 2 직접 자산 49개 일괄 읽기와 다른 Stage의 hash closure 재봉인은 개정 절차에서 제외한다. + +## 3. 절대 원칙과 완료 기준 + +1. 사건·배포 위치는 `stage1_run_root_ref`, `stage1_deployment_root_ref`로 직접 전달하고, Stage 1 결과물은 `{{prev.###}}`로 직접 사용한다. 기존 실행 정보와 결과 참조를 활용하며 호출자에게 별도 request나 결과 사본을 만들게 하지 않는다. +2. Stage 1 원본의 bytes·파일명·상대경로·원문 의미를 보존한다. 기존 사건/transaction 식별자가 원본에 있으면 출처 검증에 그대로 사용하며 새 요청 식별자로 재포장하지 않는다. +3. prepare task, 고정 request 파일 read/write, 별도의 ID 발급·저장·출력을 제거한다. `request_id`를 이름만 바꾼 `handoff_id`·`execution_id`도 만들지 않는다. +4. S2_00은 비 LLM deterministic task 하나로 수행한다. 입력 본문 전체를 prompt에 복사하거나 LLM task를 추가하지 않는다. +5. 검증은 S2_00의 실제 입력·연산·출력에 필요한 것으로 한정한다. 원본의 무결성·출처·필수 내용·blocking review를 생략하여 단순화하지 않는다. +6. 정상 결과는 원본을 추적할 수 있고 미해결 사항을 보존해야 한다. 입력 전달 성공만으로 S2_00의 기능 목표 달성을 선언하지 않는다. + +완료는 **두 root 직접 전달·`{{prev.###}}` 결과 참조의 실제 해석 → 원본 사용과 필요한 검증 → 의미 보존·정규화·cluster/bundle 구성 → 자체 출력 검증 → 마지막 status 발행**이 한 task에서 성립하는 것으로 판정한다. 기존 S2_10~40에서 읽을 수 있는지는 판정 대상이 아니다. + +## 4. `{{prev.###}}` 결과 사용과 직접 실행 경계 + +### 4.1 backend 지원을 채택하는 설계 전제 + +Stage 2 S2_00은 Stage 1 결과물을 아래 형식으로 사용할 수 있다. + +```text +{{prev.###}} +### = Stage 1의 결과물 파일 +``` + +**이 사용은 backend의 Python code 실행 환경이 허용한다.** 이는 사용자 추가 설명을 반영한 설계 전제다. 따라서 v2의 “`prev` 지원을 가정하지 않는다”, “backend 전달 방식을 먼저 확인한다”, “별도 code 조립 대안을 검토한다”는 접근은 적용하지 않는다. 지원 여부 확인을 개정 착수의 장애나 prepare task 부활의 이유로 삼지 않는다. + +단순 파일명에 적용하면 `{{prev.evidence_indexed.json}}`, `{{prev.BO.json}}`, `{{prev.Fact_Ledger_base.json}}`과 같은 형태다. 이 예시는 사용자가 제시한 형식에 실제 파일명을 대입한 것이다. `###`를 task 이름·새 ID로 바꾸지 않는다. 하위 폴더 파일과 동적 signal의 경우 실제 Stage 1 결과물에 등록된 파일 참조를 그대로 사용하고, 문서에서 임의의 별칭이나 새로운 참조 문법을 만들지 않는다. + +`prev` 표현식을 일반 Python 문법으로 직접 평가하는 것은 아니다. backend가 해당 참조를 해석하여 Python 실행에 제공하는 결과물을 ingress에서 사용하는 경계다. 이 문서의 예시만으로 `prev`라는 Python 객체·신규 resolver API·임의 `run_code.arguments` 파라미터를 구현 사실로 가정하지 않는다. + +### 4.2 두 root와 파일 참조의 역할 + +두 root는 원본 위치·배포 검증·출처 기록·출력 위치 결정에 그대로 사용한다. 위치 정보의 형태는 다음과 같다. 이는 파일로 저장할 request가 아니라 기존 실행 정보의 직접 전달 값이다. + +```json +{ + "stage1_run_root_ref": "", + "stage1_deployment_root_ref": "<그 결과에 대응하는 Stage 1 배포의 workspace 상대 root>" +} +``` + +`{{prev.###}}`는 이 위치 정보와 함께 실제 Stage 1 결과물 파일을 사용하는 방식이다. 파일 참조 사용을 위해 별도 요청 ID를 발급하거나 caller에게 파일 목록·본문을 새 request로 재작성하게 하지 않는다. backend에서 제공된 원본 결과를 직접 사용하고, root에서 파일을 다시 찾는 별도 중계 절차로 되돌리지 않는다. + +`W/`는 인증된 localdocs 논리 workspace root, `U=W//`, `D=W//`다. Dropbox 절대경로나 executor 임시 mount와 동일한 것으로 취급하지 않는다. 인증 정보는 기존 backend 세션을 사용하며 호출자가 임의 workspace를 지정하게 하지 않는다. 배포 자산은 `D/`에서 필요한 항목만 읽기 전용으로 검증한다. + +### 4.3 단일 Python ingress의 내부 경계 + +권고 내부 함수 경계는 다음과 같다. `stage1_results`는 backend가 `{{prev.###}}`로 제공한 결과를 ingress 안에서 다루는 메모리 값이며 별도 request 파일·새 호출자 준비 산출물이 아니다. + +```python +run_inline_mcp(stage1_run_root_ref, stage1_deployment_root_ref, *, stage1_results, client=None) +validate_direct_inputs(stage1_run_root_ref, stage1_deployment_root_ref, stage1_results) +hydrate_stage1(localdocs, temp_root, validated_roots, stage1_results) +execute_ingress(stage1_root, stage1_deployment_root, *, output_dir, source_checks) +publish_result(localdocs, output_root, files, status) +``` + +backend에서 제공되는 실제 결과 표현에 따라 기존 adapter를 연결한다. 파일 경로/참조가 제공되면 해당 원본 파일을 사용하고, 파일 bytes나 파싱된 결과가 제공되면 그 값을 재사용한다. 표현 형식을 문서에서 하나로 단정하지 않는다. 파싱된 객체만으로 원본 raw bytes hash가 검증되었다고 주장하지 않으며, 검증에 필요한 원본 bytes가 없을 때만 기존 원본 경로에서 보완한다. 이미 제공된 원본을 불필요하게 다시 읽거나 직렬화하여 사본을 만들지 않는다. + +별도의 request 객체를 core에 요구하지 않는다. core의 `run_id="S2-00-REQUEST"`, `attempt_id` 필수 인자와 ID 정규식 검사는 삭제하고 필요한 경로·검증 객체만 전달한다. 내부 함수 분리는 유지할 수 있으나 추가 executor task로 나누지 않는다. 원문 결과에 따옴표·개행이 포함되어도 데이터로 보존하며 원문을 Python 소스 문자열에 직접 끼워 넣지 않는다. + +두 root는 canonical NFC 상대경로인지 검사한다. 누락·빈 값·절대경로·`..`·NUL·경로 탈출·알 수 없는 입력 필드는 원본 읽기나 출력 쓰기 전에 거절한다. 지원되는 `{{prev.###}}` 표현식 자체를 오류로 판단하지 않는다. backend 해석 이후에도 필수 참조가 미치환 상태이거나 결과물이 누락·잘못 대응된 경우에는 원본으로 사용하지 않고 진단한다. Stage 1 완료 여부는 폴더 존재만으로 판단하지 않고 기존 writer report·gate/handoff와 실제 제공된 결과 간 결속을 확인한다. 기존 완료/seal 정보가 없으면 없다고 기록하며 새 Stage 1 파일을 요구하지 않는다. + +참조 지원은 채택한 전제이며, 각 결과물의 정확한 결속·원본 동일성·한 task 내 발행 성공은 향후 구현 검증 항목이다. 이번 전략서 작성은 그 실행 검증을 수행하지 않는다. prepare/request 파일·새 ID·backend 개정을 추가하지 않고 기존 실행 환경 안에서 S2_00 YAML만 개정한다. + +## 5. Stage 1 원본 재사용과 필요한 자산만 읽기 + +현행 고정 사건 입력 16개를 `{{prev.###}}`로 그대로 사용한다. 아래 경로는 원본의 `U/` 기준 상대경로이며, 실제 backend 파일 참조와 원본 출처를 대응하는 기준이다. 이 표를 새 request 파일이나 결과 사본으로 저장하지 않는다. + +| 묶음 | 원본 상대경로 | +|---|---| +| P1 3개 | `evidence_indexed.json`, `evidence_event_candidates.json`, `client_goal.json` | +| P1 routing 2개 | `routing/domain_screening.json`, `routing/domain_activation_manifest.json` | +| P1 gate/handoff 3개 | `quality_gates/B1_evidence_indexed_gate.json`, `quality_gates/B2_event_candidates_gate.json`, `quality_gates/stage1_part1_soft_gate_handoff.json` | +| P2 3개 | `BO.json`, `signals/signal_manifest.json`, `quality_gates/stage1_part2_review_handoff.json` | +| P3 2개 | `legal_effect_structures.json`, `quality_gates/stage1_part3_review_handoff.json` | +| P4 3개 | `Fact_Ledger_base.json`, `stage1_tmp/fact_ledger/fact_ledger_writer_report.json`, `quality_gates/stage1_part4_review_handoff.json` | + +signal manifest 자체를 Stage 1 결과 참조로 사용하고, payload는 기존 `signal_manifest.files[]`의 경로·hash·조건에 해당하는 실제 결과물 참조를 따라 사용한다. 필요한 원본 bytes 접근만 보완한다. 폴더 전체 scan이나 추측한 파일 목록으로 대체하지 않는다. routing activation과 signal activation은 서로 다른 원본으로 취급하며 실제 의미 일치 여부를 확인한다. + +Stage 1 배포의 현행 55개 잠금 목록은 검증 근거의 기준선이다. v3에서는 각 자산이 사건 JSON 해석, schema 검증, producer/alias·domain 연결, 원본 identity 검증 중 어디에서 필요한지 S2_00 내부에서 추적한다. 필요 자산과 그 검증 의존은 유지하고, 미사용임이 확인된 자산만 읽기 대상에서 제외한다. 검토 전부터 특정 감소 개수를 약속하지 않는다. Stage 1 원본이나 배포 파일을 고쳐 개정에 맞추지 않는다. + +Stage 2의 현행 49개 직접 의존은 후속 작업을 위한 결속까지 포함하므로 전체를 유지하지 않는다. S2_10 YAML·LLM binding과 후속 authority/retrieval 선택용 자산은 S2_00 의존에서 제거한다. registry/config도 C05~C10에서 실제 사용하는 내용만 남기고, 남기는 이유·경로·검증 hash를 YAML 내부의 단일 자산 표에 둔다. 필요한 비실행 데이터는 읽기 전용으로 참조할 수 있으나 그 파일을 개정하거나 새 외부 schema를 만들지 않는다. 문서 안의 법률 규칙을 새로 작성하는 작업도 포함하지 않는다. + +자체 입력 adapter·출력 validator·경로 목록의 정본은 개정 YAML 안에서 하나로 관리한다. 기존 전체 release·module·downstream schema를 그대로 validator에 주입해 제거한 필드를 다시 요구하게 해서는 안 된다. 외부 참조가 있는 기존 schema는 실제 S2_00에 필요한 규칙과 참조 의존만 inline 자체 계약에 반영한다. 이것은 후속 호환 검사 제거이며 Stage 1 원본 검증의 면제는 아니다. + +## 6. C00–C15의 자체 기능 목표 + +| 논리 단계 | 유지할 작업 | 단순화·보완 방향 | +|---|---|---| +| C00 | Stage 1 결과 참조·배포 원본, 필수 입력·hash·producer/transaction·gate/review 확인 | 결과물은 `{{prev.###}}`로 사용하고 두 root로 출처·배포를 확인. request/ID 검증과 downstream 자산 admission은 제거. hash 근거가 없는 입력은 무근거 상태를 보존 | +| C05 | 증거·사실·객체·당사자·목표·LES·signal의 원본 보존과 필요한 정규화 | 원문을 덮어쓰지 않고 필요한 view만 생성. source path·정확한 JSON pointer·raw hash로 원문 연결 | +| C10 | 원본의 명시적 관계에 따른 claim-neutral cluster와 작업 범위 구성 | 확인된 fact/evidence/BO/LES/signal 관계만 사용. 빈 slot skeleton이나 근거 없는 관계를 채우지 않음 | +| C15 | cluster별 입력 묶음·미해결 항목·결과 검증과 발행 | 후속 Agent prompt·dispatch schema 대신 자체 bundle 표현. S2_10/40 route를 자체 처리 상태로 대체 | + +adapter가 wrapper/배열 위치를 잘못 해석하여 원문 pointer가 틀리거나 사실·client goal·signal·review를 축약해 의미를 잃는 문제는 S2_00 inline code 안에서 보완한다. 원문 사실, 목표 제약, 반대 자료, blocking/unresolved review, 기존 명시적 연결을 보존한다. `UNEVALUABLE`·미상은 원본 근거 없이 해결된 상태로 변경하지 않는다. + +원본의 fact/evidence/object/transaction ID는 의미와 참조를 구성하는 기존 값이므로 유지한다. cluster의 구조상 식별값이 필요한 경우에도 해당 묶음 참조 용도에 한정한다. 이러한 자료 식별값을 요청 ID나 실행 ID 발급의 근거로 사용하지 않는다. + +## 7. ID 없는 출력·재실행·임시 작업 + +최종 출력 root는 원본 사건 root를 재사용하여 다음처럼 결정한다. + +```text +O = W/stage2_runs/from-stage1//s2_00/ +``` + +검증된 상대 root의 경로 구성을 그대로 아래에 결합한다. 새 요청 ID, UUID, `run_binding_digest`로 출력 폴더를 만들지 않는다. `O/`가 `U/` 또는 `D/`와 겹치거나 어느 한쪽 안에 놓이는 경우는 쓰기 전에 거절한다. Stage 1 원본 트리는 읽기 전용으로 유지한다. + +권고 정상 출력은 다음 5개이며 cluster 수가 크거나 부분 읽기가 필요한 경우에만 가변 slice를 분리한다. 기존 11+E개 파일을 모두 보존하기 위한 빈 파일은 만들지 않는다. + +| 파일 | 필요한 내용 | +|---|---| +| `ingress/stage1_input_manifest.json` | 두 root, 읽은 원본 상대경로·raw hash·검증 근거, 기존 원본 identity가 있으면 그 값 | +| `ingress/intake_report.json` | 필수 입력·무결성·gate 확인 결과와 누락/미검증 항목 | +| `review/issue_ledger.base.json` | 원본 review의 내용·범위·blocking 여부·출처와 S2_00 처리 중 발견한 문제 | +| `context/case_context.json` | 정규화 view·원본 참조, 증거/객체/당사자/LES/사실/목표/signal 관계, cluster 목록과 bundle별 원본 참조. 분리 slice 사용 시 해당 경로 | +| `ingress/ingress_status.json` | `READY`, `READY_WITH_ISSUES`, `BLOCKED` 중 상태, 두 root, algorithm version, 실제 발행 파일 경로·hash. 마지막에 발행 | + +가변 파일은 필요한 경우에만 `context/cluster_slices/.json`으로 분리한다. ``는 자료 묶음 참조이며 요청/실행 ID가 아니다. 한 내용은 context 또는 slice 중 한 곳에 담고 다른 곳에서는 참조하여 중복을 줄인다. 원본 전체 사본을 새 context에 복제하지 않는다. + +안전하게 root를 확정하고 입력 검사를 수행한 뒤 발생한 차단은 `BLOCKED`로 기록한다. 이때 정상 context/slice는 발행하지 않으며 intake·issue ledger·technical diagnostic과 마지막 status를 발행한다. 입력 목록을 확정할 수 있을 때만 manifest도 발행한다. 입력 인자·인증 실패처럼 출력 위치를 안전하게 확정하지 못한 오류는 stdout 오류로 종료하고 결과 파일 발행을 강제하지 않는다. + +모든 출력의 닫힌 필드·상태별 required/forbidden artifact 규칙은 inline validator에서 정의한다. 기존 `run_binding_receipt`와 출력 schema의 `request_id`·신규 `run_id` 필수 조건은 함께 제거한다. stdout은 처리 성공 여부·상태·출력 경로·오류·마지막 status hash 정도만 담는다. 별도 ID receipt는 만들지 않는다. + +원본 raw hash와 산출물 hash는 변조·내용 동일성 확인을 위해 유지한다. 이 hash를 접두사와 결합하여 요청 식별자로 재출력하지 않는다. 동일 `O/`가 이미 있으면 기존 입력 경로·hash, 배포 근거, algorithm version과 발행 파일을 확인하고 일치하는 완료 결과만 재사용한다. 내용·버전이 다르거나 잔여 부분 파일이 충돌하면 덮어쓰지 않고 명시적으로 종료한다. 이번 전략은 여러 개정 버전의 결과를 같은 사건 root 아래 동시에 보관하는 새 체계를 도입하지 않는다. + +실행별 `TemporaryDirectory` 안에서 hydration·로컬 결과·staging을 처리한다. 임시 디렉터리의 고유한 이름은 도구 내부 자원이며 호출 필드, canonical output, stdout에 ID로 노출하지 않는다. `publish_atomically(..., attempt_id=...)`의 의존도 함께 제거한다. + +원격 발행은 비 status 파일 write/read-back 이후 status를 마지막에 기록하는 논리적 완료 경계로 유지한다. 임시 디렉터리가 원격 동시 writer 문제까지 해결한다고 주장하지 않는다. 같은 `O/`의 동시 writer는 기존 host 직렬화 기능을 사용할 수 있는지 확인하며, 미지원·미확인 상태에서는 동시 발행 지원을 수용 완료로 기록하지 않는다. 새 lock JSON·lock ID·CAS 지원을 가정하지 않는다. + +## 8. YAML 내부의 실제 변경 위치와 순서 + +1. **범위 고정:** 두 root 직접 전달과 backend가 허용하는 `{{prev.###}}` 결과 사용을 전제로 S2_00 자체 목표·출력, 필수 검증과 필요한 읽기 전용 자산을 확정한다. 다른 Stage와의 호환 검토를 하지 않는다. +2. **단일 실행 구조:** `Task_S2_00_prepare_request`를 제거한다. `task_procedure`는 `IN → Task_S2_00_deterministic_ingress → OUT`으로 바꾸고 ingress pointer를 `tasks/0`으로 정리한다. +3. **직접 결과 사용 진입점:** `run_inline_mcp`·`_inline_hydrate`·`_inline_release_materialization_plan`을 두 root와 backend가 제공한 `{{prev.###}}` 결과를 사용하도록 변경한다. 파일 참조와 원본 source pointer를 대응시킨다. `_inline_validate_request`, request 파일의 두 read-pass, prepare 관련 상수·입출력 설명을 제거한다. 실제 원본의 두 read-pass 안정성 확인은 필요한 입력 집합에 대해 유지하되 결과 참조의 해석을 두 번 수행하는 것으로 대체하지 않는다. +4. **ID 의존 제거:** `execute_ingress`의 request용 `run_id`와 `attempt_id`, `_make_run_binding_receipt`의 요청/실행 ID 생성, `canonical_run_id`, CLI ID 옵션, staging/remote publisher·stdout의 ID 의존을 제거한다. 빈 문자열·고정 가짜 ID로 기존 signature를 통과시키지 않는다. +5. **자체 계약·자산 의존:** Agent description·metadata와 inline code에서 전체 Stage 2 workflow/binding/schema를 런타임 필수 계약으로 연결하는 부분을 S2_00 자체 계약으로 정리한다. C00–C10에 실제 필요한 읽기 전용 원본/schema/config 검증은 남긴다. 다른 Stage 자산을 유지하려고 전체 49개 closure를 복원하지 않는다. +6. **의미 보존과 출력:** adapter·pointer·review 처리의 실제 공백을 보완하고 C15를 자체 bundle·상태·출력 root로 바꾼다. 출력 validator, 파일 수·artifact-set 검사, publish/read-back, stdout을 함께 정합화한다. +7. **S2_00 한정 검증:** 아래 수용 기준으로 inline 함수를 시험하고 실제 전달 가능한 환경에서 단일 task를 확인한다. 다른 Stage 시험과 패키지 전체 재봉인은 개정 범위에 넣지 않는다. + +향후 수정 파일은 해당 `Stage_2_S2_00.yml` 하나다. `Stage_2_S2_00_v.2.yml`, runtime mirror, 외부 workflow/binding/schema, manifest, builder는 함께 수정하지 않는다. 그 결과 기존 package builder의 parity나 parent release 봉인을 충족한다고 주장할 수 없다. 기존 builder를 돌려 이 YAML을 다시 덮어쓰는 것도 하지 않는다. 전체 패키지의 배포 승인과 S2_00 자체 기능 구현은 구별한다. + +## 9. 실행 허용 조건과 미확인 사항 + +현재 release는 `DEV_FIXTURE_RELEASE`이며 현행 core에는 실제 사건 발행을 금지하는 `DEV_FIXTURE_REAL_RUN_FORBIDDEN` 방어가 있다. 후속 Agent와 무관한 인증·원본 신뢰 근거·S2_00 실행 허용 조건은 보존한다. 기존 전체 Stage 2 admission을 제거하면서 이를 임의의 `active:true`나 새 self-certified receipt로 대체하지 않는다. + +YAML 한 파일 안에서 기존 근거로 독립 검증 가능한 S2_00 실행 조건을 명시한다. 현재 DEV 상태만으로는 offline fixture 검증까지만 가능하다. 실제 사건 발행에 필요한 승인·참조 지원 이외의 실행 조건이 외부 패키지 변경을 요구한다면 그 부분은 **현재 범위에서 live 미완료**로 기록한다. 다른 Stage의 준비 상태를 S2_00 내용 설계의 제약으로 다시 가져오거나, guard를 꺼 실제 사건 시험을 통과시키는 방법은 채택하지 않는다. + +`{{prev.###}}` 결과물 사용을 backend Python 실행 환경이 허용한다는 점은 사용자 설명으로 채택했으므로 미지원·지원 여부 미확인으로 기록하지 않는다. 이번 작성에서 실행 확인하지 않은 항목은 실제 파일 참조·두 root의 올바른 대응, 인증 치환, 필요한 Stage 1 신뢰 근거의 완비, 같은 출력 root의 writer 직렬화, S2_00 독립 실행 허용 조건이다. 이는 지원 기능의 존재와 그 기능을 사용한 개정 YAML의 실행 검증을 구별한 것이다. + +## 10. 수용 기준과 검증 계획 + +아래는 향후 YAML 구현의 검증 계획이며 이번 작업에서 실행한 시험 결과가 아니다. 검증용 fixture·임시 출력은 임시 디렉터리를 사용하며 저장소에 다른 시험 자산을 새로 만들 필요는 없다. + +| 검증 면 | 수용 기준 | +|---|---| +| 개정 범위 | 실제 수정 YAML은 S2_00 배포본 하나. S2_10~40·authoring·외부 자산 수정 없음 | +| 구조 | Agent stage 하나, ingress task·`run_code` 하나. DAG와 `tasks/0` pointer 일치 | +| 직접 전달·결과 참조 | 두 root가 실행 진입점에 그대로 도달하고 `{{prev.###}}`가 의도한 Stage 1 결과물로 해석됨. 지원 표현식은 허용하며 해석 후 필수 미치환·누락·잘못된 파일 대응·경로 탈출·출력/원본 중첩은 사용·쓰기 전에 차단 | +| 군더더기 제거 | prepare·고정 request 파일 read/write 0; `request_id` 입력·생성·출력 0; request를 대체하는 새 ID 0; `attempt_id` 필수 인자·출력 0 | +| 원본·검증 | Stage 1 원본 변경 0. 참조로 제공된 고정 16개·동적 signal과 원본 경로/pointer의 대응, 필요한 배포/schema 검증 수행. 원문 따옴표·개행 보존, raw hash 불일치·manifest 탈출·필요한 두 read-pass 변경 검출 | +| 자산 범위 | 실제 S2_00 사용 근거가 있는 자산만 읽음. 후속 YAML·LLM binding·후속 schema/route 계약을 요구하지 않음 | +| 내용 보존 | source pointer dereference 결과와 hash 일치. 사실·목표 제약·증거·신호·review·명시적 관계 누락 없음. 미상·blocking 유지 | +| 정상·차단 | 자체 상태와 실제 artifact 목록 일치. 차단 결과를 정상 context로 표현하지 않음. 원본에 없는 관계를 생성하지 않음 | +| 재실행·발행 | 동일 완료 결과만 재사용. 불일치·부분 충돌은 거절. 비 status write/read-back 실패 시 새 정상 status를 발행하지 않음. 기존 status 존재는 별도 확인 | +| 임시·동시 실행 | 임시 작업 디렉터리는 분리되고 출력에 임시 ID 없음. 원격 동시 발행 보장은 실증 전 미확인으로 기록 | +| live | 기존 backend Python 실행 환경에서 `{{prev.###}}` 결과 사용·두 root 대응·인증·원본 검증·결과 write/read-back·마지막 status 확인. 채택된 지원 전제, offline PASS, 실제 실행 결과를 구별 | + +검증은 먼저 경로·결과 참조 대응·참조 누락/미치환·ID 제거·정상/차단·원본 보존·pointer·재실행·발행 실패를 다루는 한 차례의 종합 offline 검증으로 수행한다. 수정이나 새 실패가 생겼을 때 해당 영향만 재확인한다. S2_10~40과의 E2E·호환 시험, 법률 판단의 타당성 승인, 전체 패키지 production readiness는 수용 기준에 넣지 않는다. + +## 11. 효율성과 남는 비용 + +계획상 executor task는 2→1, 고정 request 파일은 1→0, 별도 request ID 생성·출력은 0으로 줄인다. prepare 호출, request 저장·read-back·재읽기, ID 호환 처리, 후속 자산의 불필요한 읽기를 제거하고 `{{prev.###}}`로 이미 제공된 Stage 1 결과를 재사용하는 것이 절감 지점이다. 이미 Stage 1 결과를 직접 읽는 방식에 추가 ID 처리를 붙이는 것이 더 효율적이라고 주장하지 않는다. + +S2_00은 비 LLM 작업이므로 ID 제거 자체가 큰 모델 토큰 절감을 보장하지 않는다. inline code의 전달량·billing과 필요한 원본 검증 비용은 남는다. 속도·비용·토큰 절감률은 실제 실행 요청 크기, 읽기 횟수, task 시간과 청구 값을 측정하기 전까지 미측정으로 둔다. + +## 12. 작성 작업의 결정·검증 기록 + +| 결정 | 근거 | 결과 | +|---|---|---| +| 두 Stage 1 root 직접 전달과 request ID 전면 제거 | 이번 사용자 절대 원칙 | v1 §4.3의 내부 호환 ID 정책 폐기 | +| S2_00 YAML 한 파일을 향후 개정 대상으로 한정 | 사용자 범위 제한 | downstream 호환·외부 자산 재봉인 계획 제외 | +| 원본 의미·검증 유지, 자체 출력 축소 | S2_00의 C00–C15 기능 목표 | 전송 단순화와 기능 완결을 함께 검증 | +| 기존 root·자료 참조·hash 재사용 | 추가 요청 식별자를 만들지 않는 원칙 | 원본 보존, 출처 확인, 재실행 충돌 판별 가능 | +| `{{prev.###}}`와 backend Python 실행 허용 채택 | 이번 사용자 추가 설명 | v2의 지원 여부 미확인·전달 방식 탐색 전제를 대체 | +| 참조·root 대응과 live 허용 조건은 구현 시 검증 | 실제 YAML 개정·실행 시험은 이번 범위 밖 | 지원 전제·전략·offline·live·패키지 승인 상태 구별 | + +이번에는 v2를 보존하고 이 v3 전략서를 생성하며 지정 MEMORY에 요약을 삽입했다. v2의 직접 전달·ID 제거·S2_00 한정 원칙을 유지하고, `{{prev.###}}` 사용 및 backend Python 실행 환경의 허용을 목표·입력 경계·원본 사용·변경 순서·수용 기준·상태 기록에 반영했다. 저장 후 UTF-8·종결 LF·fence 8개·로컬 링크 5개, 이전 미확인 전제의 제거와 관련 절 간 정합성, 요약 위치를 정적으로 확인했다. 작업 전 기존 파일 611개 중 MEMORY를 제외한 610개의 SHA-256이 동일하며, MEMORY도 이번 요약 삽입 외의 내용은 보존했다. 새 파일은 이 전략서 하나다. 실제 YAML 개정, inline 실행, builder·release 재봉인, MCP/backend/live 시험, 배포·commit은 수행하지 않았다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/MEMORY.md b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/MEMORY.md index 244a2b7a..4acdee55 100644 --- a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/MEMORY.md +++ b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/MEMORY.md @@ -16,6 +16,37 @@ - 하위 에이전트(subagents)를 사용하여 핵심 테마와 교훈을 식별하고 MEMORY.md에 섹션으로 저장하라. - 향후 세션 시작 시 MEMORY.md의 이전 섹션을 참조하라. +## 2026-10-02 — S2_00 YAML v3 직접 인계 구현 + +한 줄 요약: 전략 v3·지정 SKILL·Code Executor notebook에 따라 S2_00을 단일 deterministic ingress로 개정했다. Stage 1 사건·배포 root와 `{{prev.###}}` 결과물을 직접 재사용하며 별도 request/attempt/run ID와 준비 task를 제거한다. + +- C00–C15의 원본 계약·seal·hash·보존 검증, 원본 pointer 정규화, claim-neutral cluster·의존 관계·review/signal 보존을 유지한다. 필요한 Stage 1 배포 의존만 읽고 자체 5개 산출물을 검증한 뒤 status를 마지막에 기록하며, 완료본은 byte 일치 시 재사용하고 불완전·충돌 출력은 덮어쓰지 않는다. 기존 중첩 Counter 비교 오류와 원본 배열 pointer 불일치도 수정했다. +- 원본은 `Default_Agent/Stage_2_Clean/agent_scripts/Stage_2_S2_00_10_02.yml`, 개정 복제본은 본 폴더 `Stage_2_S2_00_v.3.yml`로 byte-identical 저장했다. 오프라인·모의 시험 39건, YAML/DAG·Python 3.11/3.12 문법, Stage 1 승인 의존 55개 hash와 선택적 의존 로딩을 확인했다. 실제 backend 치환·MCP 실행·동일 root 동시 writer·배포 성공은 검증하지 않았다. +- 기존 `DEV_FIXTURE_RELEASE`의 실제 사건 발행 차단은 유지한다. S2_10~40·전략서·SKILL·notebook·schema·mirror·release는 변경하지 않았으며 builder/재봉인/live 실행/commit도 수행하지 않았다. 단일 YAML 구현 검증을 기존 패키지 전체 정합성이나 운영 승인으로 해석하지 않는 것이 핵심이다. + +## 2026-10-02 — S2_00 Stage 1 결과 참조 반영 전략 v3 + +한 줄 요약: v2를 보존하고 `Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v3.md`를 생성하여 `{{prev.###}}`(`###`는 Stage 1 결과물 파일)를 통한 원본 사용과 backend Python 실행 환경의 허용을 사용자 설명에 따른 설계 전제로 반영했다. + +- Stage 1 사건·배포 root 직접 전달, request ID 전면 제거, S2_00 YAML 한 파일 개정·S2_10~40 계약 제외 원칙은 유지한다. v2의 참조 지원 여부 미확인·전달 방식 탐색 전제를 대체하고 결과 참조를 목표·입력 경계·원본 사용·변경 순서·수용 기준에 연결했다. +- 지원되는 기능의 존재와 개정 YAML에서의 실제 파일/root 대응·원본 동일성·발행 성공 검증을 구별한다. 문서 구조·로컬 링크 5개·미확인 전제 제거·요약 위치와 기존 610개 파일의 hash 보존을 정적으로 확인했다. MEMORY 기존 내용도 삽입 외에는 보존했다. 이번 작성은 v3 전략서·본 요약만 변경했으며 YAML/자산 개정·backend/live 시험·배포·commit은 수행하지 않았다. + +## 2026-10-02 — S2_00 한정·request ID 제거 개정 전략 v2 + +한 줄 요약: `Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v2.md`를 작성하여 Stage 1 사건·배포 root 직접 전달, 별도 `request_id` 입력·내부 생성·출력 제거, S2_00 배포 YAML 한 파일 개정을 원칙으로 확정했다. + +- v1의 내부 호환 ID·downstream schema/route·49개 자산 일괄 유지·다중 파일 재봉인 전제를 대체한다. S2_10~40의 자산 사용 계약을 무시하고 C00–C15 자체 기능, 필요한 원본 검증·의미 보존·자체 출력과 마지막 status만 완료 기준으로 삼는다. +- 기존 사건 root·원본 식별값·검증 hash를 재사용하고 임시 디렉터리로 작업을 격리한다. 직접 전달에 ID 처리를 추가하는 것은 효율 향상 수단이 아니며, 불필요한 전달층·의존 읽기를 제거하는 것이 핵심 교훈이다. backend 인자 결속·원격 동시 writer·DEV guard에 따른 live 제한과 기존 패키지 봉인 미충족 가능성은 남긴다. +- 이번 작성은 전략서·본 요약만 변경했다. 저장 경로·로컬 링크 5개·UTF-8/fence·요약 위치와 기존 609개 파일의 hash 보존을 정적으로 확인했다. MEMORY 기존 내용도 삽입 외에는 보존했다. 실제 YAML/자산 개정·실행 시험·builder·재봉인·배포·commit은 수행하지 않았다. v1은 과거 전략으로 보존한다. + +## 2026-10-02 — S2_00 직접 인계 개정 전략 v1 + +한 줄 요약: `Default_Agent/Stage_2_Clean/agent_scripts/stage_2_s2_00_revision_strategy_v1.md`에 두 Stage 1 root를 직접 받아 단일 ingress task에서 검증·C00–C15·발행을 수행하고, prepare·고정 request 파일·외부 ID 부여를 없애는 최소 변경 전략을 작성했다. + +- Stage 1 원본과 16개+manifest 신호·배포 55개·Stage 2 직접 자산 49개, 두 read-pass·기존 출력/barrier를 재사용하며 ID는 내부 호환 값으로 처리한다. 입력 전달 단순화와 기존 slice·pointer·review 의미 보완을 구별하는 것이 핵심 교훈이다. +- 지정된 `Stage_2_10_Analysis_v1.md`는 실제 S2_10 분석이므로 구형 S2_00 YAML·`Stage_2_00_Analysis_v1.md`와 교차 확인했다. backend 직접 입력 binding과 DEV release 제한은 미해결로 명시했다. 전략서·본 요약만 작성했으며 YAML/자산 개정·재봉인·시험/live 실행·배포·commit은 수행하지 않았다. +- 문서 링크·16개 입력 목록·구조와 요약 위치를 정적으로 확인했고 참조 원본 15개의 SHA-256은 작업 전후 동일했다. + ## 2026-10-01 — 현행 S2_00 배포 YAML 분석 한 줄 요약: `Default_Agent/Stage_2_Clean/agent_scripts/Analysis_Stage_2_S2_00.md`에 현행 두-task DAG, C00–C15의 단일 ingress 실행 경계, 고정·동적 입력과 분기별 출력, 배포 자산·후속 route를 기록했다. prepare 함수의 request 저장 구현과 배포 YAML의 직접 실행 가능성을 구분하는 것이 핵심이다. 현재 네 외부 인자 결속·실패 차단·workspace 직렬화는 미검증이고, 직접 실행은 fail closed하며 DEV fixture release는 실제 사건 발행을 금지한다. diff --git a/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Stage_2_S2_00_v.3.yml b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Stage_2_S2_00_v.3.yml new file mode 100644 index 00000000..593248b1 --- /dev/null +++ b/Case_02_Comparison_Research/YAML_Prompts/2. Stage_2/Stage_2_S2_00_v.3.yml @@ -0,0 +1,3057 @@ +Agent: + name: Stage_2_S2_00_v3 + version: 3.0.0 + description: Stage 1 사건·배포 root와 {{prev.###}} 결과물 참조를 직접 받아 C00 검증, C05 원본 보존·정규화, C10 claim-neutral cluster, + C15 묶음·status-last 발행을 단일 비 LLM task로 수행한다. + metadata: + workflow_id: S2_00 + execution_class: NON-LLM-DETERMINISTIC + execution_authority: MCP_CODE_EXECUTOR_INLINE + implementation_status: IMPLEMENTED_OFFLINE_VERIFIED_LIVE_NOT_RUN + algorithm_version: s2_00_direct_ingress/3.0.0 + input_contract: + stage1_run_root_ref: '{{prev.stage1_run_root_ref}}' + stage1_deployment_root_ref: '{{prev.stage1_deployment_root_ref}}' + stage1_result_reference: '{{prev.###}}; ### = Stage 1 결과물 파일' + source_contract_authority: parameters.code::SOURCE_POLICY + output_contract: + root: stage2_runs/from-stage1//s2_00/ + states: + - READY + - READY_WITH_ISSUES + - BLOCKED + normal_artifact_count: 5 + status_last: ingress/ingress_status.json + publication_semantics: STATUS_LAST_LOGICAL_COMMIT; SAME_ROOT_CONCURRENT_WRITERS_UNVERIFIED + execution_admission: DEV_FIXTURE_RELEASE + standalone_contract: true + Stages: + - name: S2_00 + description: 두 root 직접 전달과 Stage 1 결과물 참조 사용. 별도 준비 task 없이 한 run_code에서 원본·배포 검증, 의미 보존, cluster/bundle 구성, + 자체 출력 검증 및 마지막 status 발행. + prevs: [] + nexts: [] + tools: + mcpServers: + localdocs: + type: streamable-http + url: http://mcp-localdocs:8012/mcp + code-executor: + type: streamable-http + url: https://code-executor.mcp.eroomai.com/mcp + tasks: + - task_name: Task_S2_00_deterministic_ingress + description: backend Python 환경이 제공하는 {{prev.###}} 원본 결과물을 재사용한다. 필요한 배포·원본 검증만 수행하고 자체 5개 정상 산출물 또는 차단 진단을 + 발행한다. 기존 DEV release의 실제 사건 발행 금지를 유지한다. + mcp: code-executor + tool_name: run_code + parameters: + language: python + requirements: httpx==0.28.1 + network: agent-network + timeout: 300 + code: | + #!/usr/bin/env python3 + """S2_00 direct Stage 1 ingress; one deterministic Code Executor task. + + Stage 1 results are supplied through the backend's prev file references. + This module owns its contract; no downstream Agent or output schema is loaded. + MCP transport follows the required Code Executor notebook and SKILL guide. + """ + from __future__ import annotations + import base64 + import binascii + from collections import Counter, defaultdict + import contextlib + from dataclasses import dataclass + import hashlib + import io + import itertools + import json + import math + import os + from pathlib import Path, PurePosixPath + import posixpath + import re + import stat + import sys + import tempfile + import unicodedata + from typing import Any, Callable, Iterable, Mapping, MutableMapping, Sequence + + ALGORITHM_VERSION = "s2_00_direct_ingress/3.0.0" + LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp" + MCP_PROTOCOL_VERSION = "2025-03-26" + INLINE_CLIENT_NAME = "liti-stage2-s2-00-direct" + INLINE_CLIENT_VERSION = "3.0.0" + INLINE_USER_HASH = r"""{{__user_hash__}}""" + INLINE_WORKSPACE_HASH = r"""{{__workspace_hash__}}""" + RAW_RUN_ROOT = r"""{{prev.stage1_run_root_ref}}""" + RAW_DEPLOYMENT_ROOT = r"""{{prev.stage1_deployment_root_ref}}""" + + + MAX_FILE_BYTES = 32 * 1024 * 1024 + + + MAX_RUN_BYTES = 256 * 1024 * 1024 + + + MAX_JSON_DEPTH = 96 + + + MAX_JSON_ITEMS = 1_000_000 + + + SEMANTIC_SIGNAL_KINDS = frozenset({"canonical", "domain_signal"}) + + + HARD_RELATION_KINDS = frozenset( + { + "SAME_BO_ID", + "SOURCE_BO_ATTACHMENT", + "SAME_EVIDENCE_REF", + "SAME_EVENT_REF", + "EXPLICIT_CASE_RELATION", + } + ) + + + CANDIDATE_RELATION_KINDS = frozenset( + {"claim_precondition", "accessory_of", "incompatible_with", "EXPLICIT_DEPENDENCY"} + ) + + + P1_DIGEST_KEYS = { + "evidence_indexed_sha256": "evidence_indexed", + "evidence_event_candidates_sha256": "evidence_event_candidates", + "b1_gate_sha256": "b1_evidence_indexed_gate", + "b2_gate_sha256": "b2_event_candidates_gate", + "screening_sha256": "domain_screening", + "activation_manifest_sha256": "domain_activation_manifest", + "registry_index_sha256": "stage1_domain_registry_index", + } + + + REQUIREMENT_CLASS_ENUM = { + "identity_backbone": "IDENTITY_BACKBONE", + "routing_profile_backbone": "ROUTING_PROFILE_BACKBONE", + "evidence_scope": "EVIDENCE_EVENT_SCOPE", + "event_scope": "EVIDENCE_EVENT_SCOPE", + "integrity_corroborator": "INTEGRITY_CORROBORATOR", + "optimization_context": "OPTIMIZATION_CONTEXT", + } + + + ADAPTER_IDS = { + "evidence_indexed": "S2A-EVIDENCE-V3-ENVELOPE-V1", + "evidence_event_candidates": "S2A-EVENTS-V1-ENVELOPE-V1", + "client_goal": "S2A-CLIENT-GOAL-V8-V1", + "domain_screening": "S2A-DOMAIN-SCREENING-V1", + "domain_activation_manifest": "S2A-DUAL-SG01-V1", + "b1_evidence_indexed_gate": "S2A-B1-GATE-V1", + "b2_event_candidates_gate": "S2A-B2-GATE-V1", + "stage1_part1_soft_gate_handoff": "S2A-P1-HANDOFF-FLAT-V1", + "bo": "S2A-BO-V8-LIST-V1", + "signal_manifest": "S2A-SIGNAL-ALL-V1", + "stage1_part2_review_handoff": "S2A-P2-HANDOFF-FLAT-V1", + "legal_effect_structures": "S2A-LES-CURRENT-V8-V1", + "stage1_part3_review_handoff": "S2A-P3-HANDOFF-WRAPPED-V1", + "fact_ledger_base": "S2A-FACT-LEDGER-CURRENT-V8-V1", + "fact_ledger_writer_report": "S2A-FACT-LEDGER-WRITER-REPORT-V1", + "stage1_part4_review_handoff": "S2A-P4-HANDOFF-WRAPPED-V1", + } + + + SG01_PROJECTION_FIELDS: tuple[str, ...] = ( + "schema_version", + "signal_id", + "status", + "registry_version", + "registry_index_sha256", + "screening_sha256", + "domain_entries", + "active_domain_ids", + "supporting_domain_ids", + "monitor_domain_ids", + "expected_runnable_domain_ids", + "required_calculation_domains", + "unrouted_material", + "conservation_gate", + "fail_open_policy", + "review_items", + "contract_guards", + ) + + + SG01_SET_FIELDS = frozenset( + { + "active_domain_ids", + "supporting_domain_ids", + "monitor_domain_ids", + "expected_runnable_domain_ids", + "required_calculation_domains", + } + ) + + + _RAW_VALUE_UNSET = object() + + + DEFAULT_SOURCE_CONTRACTS: tuple[dict[str, Any], ...] = ( + {"logical_input_id": "evidence_indexed", "path": "evidence_indexed.json", "criticality": "evidence_scope"}, + {"logical_input_id": "evidence_event_candidates", "path": "evidence_event_candidates.json", "criticality": "event_scope"}, + {"logical_input_id": "client_goal", "path": "client_goal.json", "criticality": "optimization_context"}, + {"logical_input_id": "domain_screening", "path": "routing/domain_screening.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "domain_activation_manifest", "path": "routing/domain_activation_manifest.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "b1_evidence_indexed_gate", "path": "quality_gates/B1_evidence_indexed_gate.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "b2_event_candidates_gate", "path": "quality_gates/B2_event_candidates_gate.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "stage1_part1_soft_gate_handoff", "path": "quality_gates/stage1_part1_soft_gate_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "bo", "path": "BO.json", "criticality": "identity_backbone"}, + {"logical_input_id": "signal_manifest", "path": "signals/signal_manifest.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "stage1_part2_review_handoff", "path": "quality_gates/stage1_part2_review_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "legal_effect_structures", "path": "legal_effect_structures.json", "criticality": "routing_profile_backbone"}, + {"logical_input_id": "stage1_part3_review_handoff", "path": "quality_gates/stage1_part3_review_handoff.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "fact_ledger_base", "path": "Fact_Ledger_base.json", "criticality": "identity_backbone"}, + {"logical_input_id": "fact_ledger_writer_report", "path": "stage1_tmp/fact_ledger/fact_ledger_writer_report.json", "criticality": "integrity_corroborator"}, + {"logical_input_id": "stage1_part4_review_handoff", "path": "quality_gates/stage1_part4_review_handoff.json", "criticality": "integrity_corroborator"}, + ) + + + class IngressError(RuntimeError): + """A machine-readable deterministic ingress failure.""" + + def __init__( + self, + code: str, + message: str, + *, + logical_input_id: str | None = None, + details: Mapping[str, Any] | None = None, + ) -> None: + super().__init__(message) + self.code = code + self.logical_input_id = logical_input_id + self.details = dict(details or {}) + + def as_dict(self) -> dict[str, Any]: + result: dict[str, Any] = {"code": self.code, "message": str(self)} + if self.logical_input_id is not None: + result["logical_input_id"] = self.logical_input_id + if self.details: + result["details"] = self.details + return result + + + @dataclass(frozen=True, slots=True) + class Snapshot: + logical_input_id: str + relative_path: str + resolved_path: str + raw: bytes + raw_sha256: str + byte_length: int + device: int + inode: int + mtime_ns: int + + + def _reject_constant(value: str) -> None: + raise ValueError(f"non-finite JSON number is forbidden: {value}") + + + def _pairs_without_duplicates(pairs: Sequence[tuple[str, Any]]) -> dict[str, Any]: + result: dict[str, Any] = {} + for key, value in pairs: + if key in result: + raise ValueError(f"duplicate JSON key: {key}") + result[key] = value + return result + + + def _walk_json_limits(value: Any, *, max_depth: int, max_items: int) -> int: + count = 0 + stack: list[tuple[Any, int]] = [(value, 1)] + while stack: + current, depth = stack.pop() + if depth > max_depth: + raise IngressError("JSON_DEPTH_LIMIT", "JSON nesting depth exceeded") + if isinstance(current, dict): + count += len(current) + stack.extend((item, depth + 1) for item in current.values()) + elif isinstance(current, list): + count += len(current) + stack.extend((item, depth + 1) for item in current) + if count > max_items: + raise IngressError("JSON_ITEM_LIMIT", "JSON aggregate item limit exceeded") + return count + + + def load_json_strict( + source: Snapshot | bytes | bytearray | memoryview | str, + *, + max_depth: int = MAX_JSON_DEPTH, + max_items: int = MAX_JSON_ITEMS, + ) -> Any: + """Parse one UTF-8 JSON value, rejecting duplicate keys and non-finite numbers.""" + + if isinstance(source, Snapshot): + raw = source.raw + elif isinstance(source, str): + raw = source.encode("utf-8") + else: + raw = bytes(source) + try: + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise IngressError("INVALID_UTF8", "JSON source is not strict UTF-8") from exc + try: + value = json.loads( + text, + object_pairs_hook=_pairs_without_duplicates, + parse_constant=_reject_constant, + ) + except (json.JSONDecodeError, ValueError) as exc: + message = str(exc) + code = "DUPLICATE_JSON_KEY" if "duplicate JSON key" in message else "STRICT_JSON_PARSE_FAILED" + raise IngressError(code, message) from exc + _walk_json_limits(value, max_depth=max_depth, max_items=max_items) + return value + + + def canonical_json_bytes(value: Any) -> bytes: + """Return the project canonical parsed representation without normalizing strings.""" + + def reject_nonfinite(item: Any) -> None: + if isinstance(item, float) and not math.isfinite(item): + raise IngressError("NON_FINITE_NUMBER", "NaN and Infinity are forbidden") + if isinstance(item, dict): + for nested in item.values(): + reject_nonfinite(nested) + elif isinstance(item, (list, tuple)): + for nested in item: + reject_nonfinite(nested) + + reject_nonfinite(value) + try: + rendered = json.dumps( + value, + ensure_ascii=False, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ) + except (TypeError, ValueError) as exc: + raise IngressError("CANONICAL_SERIALIZATION_FAILED", str(exc)) from exc + return (rendered + "\n").encode("utf-8") + + + def canonical_digest(value: Any) -> str: + return hashlib.sha256(canonical_json_bytes(value)).hexdigest() + + + class _SchemaViolation(ValueError): + """Internal deterministic JSON Schema validation failure.""" + + + def _json_equal(left: Any, right: Any) -> bool: + try: + return canonical_json_bytes(left) == canonical_json_bytes(right) + except IngressError: + return False + + + def _schema_pointer(document: Mapping[str, Any], fragment: str) -> Mapping[str, Any]: + if fragment in {"", "#"}: + return document + pointer = fragment[1:] if fragment.startswith("#") else fragment + if not pointer.startswith("/"): + raise _SchemaViolation(f"unsupported schema fragment: {fragment}") + current: Any = document + for token in pointer[1:].split("/"): + key = token.replace("~1", "/").replace("~0", "~") + if not isinstance(current, dict) or key not in current: + raise _SchemaViolation(f"unresolved schema pointer: {fragment}") + current = current[key] + if not isinstance(current, dict): + raise _SchemaViolation(f"schema pointer is not an object: {fragment}") + return current + + + def _schema_type_matches(value: Any, expected: str) -> bool: + return { + "object": isinstance(value, dict), + "array": isinstance(value, list), + "string": isinstance(value, str), + "integer": isinstance(value, int) and not isinstance(value, bool), + "number": isinstance(value, (int, float)) and not isinstance(value, bool), + "boolean": isinstance(value, bool), + "null": value is None, + }.get(expected, False) + + + def _validate_schema_node( + value: Any, + schema: Mapping[str, Any], + *, + root_schema: Mapping[str, Any], + schema_documents: Mapping[str, Mapping[str, Any]], + instance_path: str, + ) -> None: + reference = schema.get("$ref") + if isinstance(reference, str): + if reference.startswith("#"): + target_root = root_schema + fragment = reference + else: + name, separator, tail = reference.partition("#") + target_root = schema_documents.get(name) + if target_root is None: + raise _SchemaViolation(f"{instance_path}: external schema ref is not release-local: {reference}") + fragment = f"#{tail}" if separator else "#" + _validate_schema_node( + value, + _schema_pointer(target_root, fragment), + root_schema=target_root, + schema_documents=schema_documents, + instance_path=instance_path, + ) + return + if "const" in schema and not _json_equal(value, schema["const"]): + raise _SchemaViolation(f"{instance_path}: const mismatch") + if "enum" in schema and not any(_json_equal(value, candidate) for candidate in schema["enum"]): + raise _SchemaViolation(f"{instance_path}: enum mismatch") + forbidden = schema.get("not") + if isinstance(forbidden, dict) and _schema_branch_matches( + value, + forbidden, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ): + raise _SchemaViolation(f"{instance_path}: forbidden schema branch matched") + expected_type = schema.get("type") + if expected_type is not None: + alternatives = [expected_type] if isinstance(expected_type, str) else list(expected_type) + if not any(_schema_type_matches(value, item) for item in alternatives): + raise _SchemaViolation(f"{instance_path}: expected type {alternatives}") + for keyword in ("oneOf", "anyOf"): + branches = schema.get(keyword) + if isinstance(branches, list): + matches = 0 + for branch in branches: + try: + _validate_schema_node( + value, + branch, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + matches += 1 + except _SchemaViolation: + continue + required_matches = 1 if keyword == "oneOf" else None + if (required_matches is not None and matches != required_matches) or (keyword == "anyOf" and matches == 0): + raise _SchemaViolation(f"{instance_path}: {keyword} matched {matches} branches") + all_of = schema.get("allOf") + if isinstance(all_of, list): + for branch in all_of: + _validate_schema_node( + value, + branch, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + condition = schema.get("if") + if isinstance(condition, dict): + condition_matches = True + try: + _validate_schema_node( + value, + condition, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + except _SchemaViolation: + condition_matches = False + selected = schema.get("then" if condition_matches else "else") + if isinstance(selected, dict): + _validate_schema_node( + value, + selected, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + if isinstance(value, dict): + minimum_properties = schema.get("minProperties") + maximum_properties = schema.get("maxProperties") + if isinstance(minimum_properties, int) and len(value) < minimum_properties: + raise _SchemaViolation(f"{instance_path}: minProperties {minimum_properties}") + if isinstance(maximum_properties, int) and len(value) > maximum_properties: + raise _SchemaViolation(f"{instance_path}: maxProperties {maximum_properties}") + required = schema.get("required", []) + if isinstance(required, list): + missing = [key for key in required if key not in value] + if missing: + raise _SchemaViolation(f"{instance_path}: missing required keys {missing}") + properties = schema.get("properties", {}) + if isinstance(properties, dict): + pattern_properties = schema.get("patternProperties", {}) + matched_by_pattern: set[str] = set() + if isinstance(pattern_properties, dict): + for key, child_value in value.items(): + for pattern_text, child_schema in pattern_properties.items(): + if re.search(pattern_text, key) is not None and isinstance(child_schema, dict): + matched_by_pattern.add(key) + _validate_schema_node( + child_value, + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + extras = sorted(set(value) - set(properties) - matched_by_pattern) + additional = schema.get("additionalProperties") + if additional is False: + if extras: + raise _SchemaViolation(f"{instance_path}: additional properties {extras}") + elif isinstance(additional, dict): + for key in extras: + _validate_schema_node( + value[key], + additional, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + for key, child_schema in properties.items(): + if key in value and isinstance(child_schema, dict): + _validate_schema_node( + value[key], + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{key}", + ) + if isinstance(value, list): + minimum = schema.get("minItems") + maximum = schema.get("maxItems") + if isinstance(minimum, int) and len(value) < minimum: + raise _SchemaViolation(f"{instance_path}: minItems {minimum}") + if isinstance(maximum, int) and len(value) > maximum: + raise _SchemaViolation(f"{instance_path}: maxItems {maximum}") + if schema.get("uniqueItems") is True: + digests = [canonical_digest(item) for item in value] + if len(digests) != len(set(digests)): + raise _SchemaViolation(f"{instance_path}: duplicate array items") + prefix_items = schema.get("prefixItems") + if isinstance(prefix_items, list): + for index, child_schema in enumerate(prefix_items): + if index < len(value) and isinstance(child_schema, dict): + _validate_schema_node( + value[index], + child_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + item_schema = schema.get("items") + if item_schema is False and isinstance(prefix_items, list) and len(value) > len(prefix_items): + raise _SchemaViolation(f"{instance_path}: additional array items are forbidden") + if isinstance(item_schema, dict): + start = len(prefix_items) if isinstance(prefix_items, list) else 0 + for index, item in enumerate(value[start:], start=start): + _validate_schema_node( + item, + item_schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + contains = schema.get("contains") + if isinstance(contains, dict): + if not any( + _schema_branch_matches( + item, + contains, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=f"{instance_path}/{index}", + ) + for index, item in enumerate(value) + ): + raise _SchemaViolation(f"{instance_path}: contains did not match") + if isinstance(value, str): + min_length = schema.get("minLength") + if isinstance(min_length, int) and len(value) < min_length: + raise _SchemaViolation(f"{instance_path}: minLength {min_length}") + max_length = schema.get("maxLength") + if isinstance(max_length, int) and len(value) > max_length: + raise _SchemaViolation(f"{instance_path}: maxLength {max_length}") + pattern = schema.get("pattern") + if isinstance(pattern, str) and re.search(pattern, value) is None: + raise _SchemaViolation(f"{instance_path}: pattern mismatch") + if isinstance(value, (int, float)) and not isinstance(value, bool): + minimum = schema.get("minimum") + if isinstance(minimum, (int, float)) and value < minimum: + raise _SchemaViolation(f"{instance_path}: minimum {minimum}") + maximum = schema.get("maximum") + if isinstance(maximum, (int, float)) and value > maximum: + raise _SchemaViolation(f"{instance_path}: maximum {maximum}") + + + def _schema_branch_matches( + value: Any, + schema: Mapping[str, Any], + *, + root_schema: Mapping[str, Any], + schema_documents: Mapping[str, Mapping[str, Any]], + instance_path: str, + ) -> bool: + try: + _validate_schema_node( + value, + schema, + root_schema=root_schema, + schema_documents=schema_documents, + instance_path=instance_path, + ) + return True + except _SchemaViolation: + return False + + + def _safe_relative_path(relative_path: str) -> PurePosixPath: + if not isinstance(relative_path, str) or not relative_path: + raise IngressError("INVALID_SOURCE_PATH", "source path must be a non-empty string") + if "\x00" in relative_path or "\\" in relative_path: + raise IngressError("INVALID_SOURCE_PATH", "NUL and backslash are forbidden in logical paths") + logical = PurePosixPath(relative_path) + if logical.is_absolute() or any(part in {"", ".", ".."} for part in logical.parts): + raise IngressError("PATH_TRAVERSAL", f"unsafe relative path: {relative_path}") + return logical + + + def _assert_no_symlink_components(root: Path, logical: PurePosixPath) -> None: + current = root + for part in logical.parts: + current = current / part + try: + current_stat = current.lstat() + except FileNotFoundError: + return + if stat.S_ISLNK(current_stat.st_mode): + raise IngressError("SYMLINK_ESCAPE", f"symlink component rejected: {logical}") + + + def open_bounded_snapshot( + approved_root: str | os.PathLike[str], + relative_path: str, + *, + logical_input_id: str = "anonymous", + max_bytes: int = MAX_FILE_BYTES, + require_single_link: bool = True, + ) -> Snapshot: + """Read one regular file once from one descriptor and verify post-read identity.""" + + root_arg = Path(approved_root) + if root_arg.is_symlink(): + raise IngressError("SYMLINK_ROOT_REJECTED", "approved root itself may not be a symlink") + try: + root = root_arg.resolve(strict=True) + except FileNotFoundError as exc: + raise IngressError("APPROVED_ROOT_MISSING", "approved root does not exist") from exc + if not root.is_dir(): + raise IngressError("APPROVED_ROOT_NOT_DIRECTORY", "approved root must be a directory") + logical = _safe_relative_path(relative_path) + _assert_no_symlink_components(root, logical) + candidate = root.joinpath(*logical.parts) + try: + resolved = candidate.resolve(strict=True) + except FileNotFoundError as exc: + raise IngressError("SOURCE_MISSING", f"source is missing: {relative_path}", logical_input_id=logical_input_id) from exc + try: + resolved.relative_to(root) + except ValueError as exc: + raise IngressError("PATH_ESCAPE", f"resolved source escaped approved root: {relative_path}") from exc + flags = os.O_RDONLY + if hasattr(os, "O_CLOEXEC"): + flags |= os.O_CLOEXEC + if hasattr(os, "O_NOFOLLOW"): + flags |= os.O_NOFOLLOW + try: + descriptor = os.open(candidate, flags) + except OSError as exc: + raise IngressError("SOURCE_OPEN_FAILED", f"unable to open source: {relative_path}") from exc + try: + before = os.fstat(descriptor) + if not stat.S_ISREG(before.st_mode): + raise IngressError("NON_REGULAR_SOURCE", f"source is not a regular file: {relative_path}") + if require_single_link and before.st_nlink != 1: + raise IngressError("HARDLINK_POLICY_VIOLATION", f"source link count is {before.st_nlink}") + if before.st_size > max_bytes: + raise IngressError("SOURCE_SIZE_LIMIT", f"source exceeds {max_bytes} bytes") + chunks: list[bytes] = [] + total = 0 + while True: + chunk = os.read(descriptor, min(1024 * 1024, max_bytes + 1 - total)) + if not chunk: + break + chunks.append(chunk) + total += len(chunk) + if total > max_bytes: + raise IngressError("SOURCE_SIZE_LIMIT", f"source exceeds {max_bytes} bytes") + after = os.fstat(descriptor) + finally: + os.close(descriptor) + try: + path_after = candidate.stat(follow_symlinks=False) + except FileNotFoundError as exc: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", "source disappeared after snapshot") from exc + identity_before = (before.st_dev, before.st_ino, before.st_size, before.st_mtime_ns) + identity_after = (after.st_dev, after.st_ino, after.st_size, after.st_mtime_ns) + path_identity = (path_after.st_dev, path_after.st_ino, path_after.st_size, path_after.st_mtime_ns) + if identity_before != identity_after or identity_after != path_identity: + raise IngressError("SOURCE_SNAPSHOT_CHANGED", f"source changed during snapshot: {relative_path}") + raw = b"".join(chunks) + return Snapshot( + logical_input_id=logical_input_id, + relative_path=logical.as_posix(), + resolved_path=str(resolved), + raw=raw, + raw_sha256=hashlib.sha256(raw).hexdigest(), + byte_length=len(raw), + device=after.st_dev, + inode=after.st_ino, + mtime_ns=after.st_mtime_ns, + ) + + + def resolve_stage1_sources( + stage1_run_root: str | os.PathLike[str], + contract_manifest: Mapping[str, Any] | None = None, + ) -> list[dict[str, Any]]: + """Resolve only approved logical kinds; a relocation manifest cannot invent kinds.""" + + root = Path(stage1_run_root).resolve(strict=True) + if not root.is_dir(): + raise IngressError("STAGE1_ROOT_NOT_DIRECTORY", "Stage 1 run root must be a directory") + contracts = [dict(row) for row in DEFAULT_SOURCE_CONTRACTS] + overrides = dict((contract_manifest or {}).get("path_overrides", {})) + approved_ids = {row["logical_input_id"] for row in contracts} + invented = sorted(set(overrides) - approved_ids) + if invented: + raise IngressError("UNAPPROVED_LOGICAL_KIND", "relocation manifest invented logical kinds", details={"ids": invented}) + seen_paths: set[str] = set() + for row in contracts: + path = overrides.get(row["logical_input_id"], row["path"]) + safe = _safe_relative_path(path).as_posix() + if safe in seen_paths: + raise IngressError("DUPLICATE_LOGICAL_MAPPING", f"duplicate physical mapping: {safe}") + seen_paths.add(safe) + row["expected_path"] = row.pop("path") + row["observed_path"] = safe + row["resolution_source"] = ( + "RELEASE_BOUND_CONTRACT_MANIFEST" + if row["logical_input_id"] in overrides + else "DEFAULT_EXACT_PATH" + ) + return contracts + + + def _issue( + code: str, + *, + impact_scope: str = "GLOBAL", + source_refs: Sequence[str] = (), + severity: str = "ERROR", + message: str | None = None, + ) -> dict[str, Any]: + return { + "issue_code": code, + "severity": severity, + "impact_scope": impact_scope, + "scope_refs": sorted(set(source_refs)), + "source_contract_row_refs": sorted(set(source_refs)), + "reason_codes": [code], + "downstream_allowed_actions": [], + "message": message or code, + } + + + def _shape_required(value: Any, keys: Sequence[str]) -> list[str]: + if not isinstance(value, dict): + return list(keys) + return [key for key in keys if key not in value] + + + def _json_pointer_value(document: Any, pointer: str | None) -> tuple[bool, Any]: + if pointer in {None, ""}: + return (pointer == "", document) + if not isinstance(pointer, str) or not pointer.startswith("/"): + return False, None + current = document + for raw_token in pointer[1:].split("/"): + token = raw_token.replace("~1", "/").replace("~0", "~") + if isinstance(current, dict) and token in current: + current = current[token] + elif isinstance(current, list) and token.isdigit() and int(token) < len(current): + current = current[int(token)] + else: + return False, None + return True, current + + + def _release_stage1_source_rows(release_lock: Mapping[str, Any]) -> list[Mapping[str, Any]]: + rows = release_lock.get("stage1_sources") + if not isinstance(rows, list): + dependency = release_lock.get("dependency_locks", {}).get("stage1", {}) + rows = dependency.get("stage1_sources") if isinstance(dependency, dict) else None + return [row for row in rows if isinstance(row, dict)] if isinstance(rows, list) else [] + + + def _adapter_decision(release_lock: Mapping[str, Any], adapter_id: str) -> Mapping[str, Any] | None: + for row in release_lock.get("adapter_decisions", []): + if isinstance(row, dict) and row.get("adapter_id") == adapter_id and isinstance(row.get("decision"), dict): + return row["decision"] + return None + + + def _closed_adapter_shape_errors( + document: Any, + *, + logical_id: str, + adapter_id: str, + required_keys: Sequence[str], + release_lock: Mapping[str, Any], + ) -> list[str]: + errors: list[str] = [] + if required_keys: + errors.extend(f"missing root key {key}" for key in _shape_required(document, required_keys)) + decision = _adapter_decision(release_lock, adapter_id) + if decision is not None: + root_shape = decision.get("root_shape") + if root_shape == "ARRAY" and not isinstance(document, list): + errors.append("root must be an array") + elif root_shape == "OBJECT_ENVELOPE" and not isinstance(document, dict): + errors.append("root must be an object envelope") + if isinstance(document, dict): + errors.extend( + f"missing root key {key}" + for key in _shape_required(document, decision.get("required_root_fields", [])) + ) + if isinstance(document, list): + required_item_fields = decision.get("required_item_fields", decision.get("required_row_fields", [])) + if isinstance(required_item_fields, list): + for index, item in enumerate(document): + for key in _shape_required(item, required_item_fields): + errors.append(f"row {index} missing {key}") + if decision is None: + fallback_required: dict[str, tuple[str, ...]] = { + "evidence_indexed": ("schema_contract_version", "items"), + "evidence_event_candidates": ("schema_version", "items"), + "domain_activation_manifest": SG01_PROJECTION_FIELDS, + "signal_manifest": ("downstream_read_sets", "files"), + "legal_effect_structures": ("schema_version", "structure_records"), + "fact_ledger_writer_report": ( + "schema_version", + "row_count", + "gate_firings", + "domain_effect_coverage", + "calculation_readiness", + "blocked_review_items", + "conservation", + "final_sha256", + ), + } + fallback = fallback_required.get(logical_id, ()) + if fallback: + errors.extend(f"missing root key {key}" for key in _shape_required(document, fallback)) + if logical_id in {"bo", "fact_ledger_base"} and not isinstance(document, list): + errors.append("root must be an array") + return sorted(set(errors)) + + + def _schema_document_index(deployment_documents: Mapping[str, Any]) -> dict[str, Mapping[str, Any]]: + result: dict[str, Mapping[str, Any]] = {} + for path, document in deployment_documents.items(): + if not isinstance(document, dict): + continue + result[path] = document + result[PurePosixPath(path).name] = document + schema_id = document.get("$id") + if isinstance(schema_id, str): + result[schema_id] = document + return result + + + def _source_hash_index(document: Mapping[str, Any] | None) -> dict[str, str]: + result: dict[str, str] = {} + if not isinstance(document, dict): + return result + candidate_arrays: list[Any] = [] + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(document.get(key), list): + candidate_arrays.append(document[key]) + for wrapper in ("completion_seal", "manifest", "payload", "data"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(nested.get(key), list): + candidate_arrays.append(nested[key]) + for rows in candidate_arrays: + for row in rows: + if not isinstance(row, dict): + continue + digest = row.get("raw_sha256", row.get("sha256")) + if not isinstance(digest, str) or re.fullmatch(r"[A-Fa-f0-9]{64}", digest) is None: + continue + for key in ("logical_input_id", "path", "observed_path", "logical_id"): + identifier = row.get(key) + if isinstance(identifier, str) and identifier: + result[identifier] = digest.lower() + return result + + + def _source_producer_index(document: Mapping[str, Any] | None) -> dict[str, str]: + """Index producer evidence carried by a bounded completion/manifest row.""" + + result: dict[str, str] = {} + if not isinstance(document, dict): + return result + candidate_arrays: list[Any] = [] + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(document.get(key), list): + candidate_arrays.append(document[key]) + for wrapper in ("completion_seal", "manifest", "payload", "data"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("source_rows", "sources", "artifacts", "files", "entries"): + if isinstance(nested.get(key), list): + candidate_arrays.append(nested[key]) + for rows in candidate_arrays: + for row in rows: + if not isinstance(row, dict): + continue + producer = next( + ( + row.get(key) + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by") + if isinstance(row.get(key), str) and row.get(key) + ), + None, + ) + if not isinstance(producer, str): + continue + for key in ("logical_input_id", "path", "observed_path", "logical_id"): + identifier = row.get(key) + if isinstance(identifier, str) and identifier: + result[identifier] = producer + return result + + + def _producer_value(document: Any) -> str | None: + if not isinstance(document, dict): + return None + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): + value = document.get(key) + if isinstance(value, str) and value: + return value + for wrapper in ("metadata", "meta", "handoff", "payload"): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("producer_id", "created_by", "writer_id", "writer", "finalized_by"): + value = nested.get(key) + if isinstance(value, str) and value: + return value + # P3/P4 are closed one-key wrappers in the Stage 1 v8 handoff contract. + for wrapper in ( + "stage1_part3_review_handoff", + "stage1_part4_review_handoff", + ): + nested = document.get(wrapper) + if isinstance(nested, dict): + for key in ("created_by", "finalized_by"): + value = nested.get(key) + if isinstance(value, str) and value: + return value + return None + + + def _producer_matches( + observed: str, + expected: str, + alias_id: str | None, + release_lock: Mapping[str, Any], + ) -> bool: + if observed == expected: + return True + if alias_id is None: + return False + decision = _adapter_decision(release_lock, alias_id) + if decision is None or decision.get("bidirectional_match_allowed") is not True: + return False + pair = {decision.get("schema_writer_id"), decision.get("orchestration_producer_id")} + return {observed, expected} == pair + + + def _identity_ref(document: Any, pointer: str | None, logical_id: str) -> dict[str, Any]: + if pointer is None: + return {"value": None, "disposition": "NOT_APPLICABLE", "source_ref": logical_id} + found, value = _json_pointer_value(document, pointer) + if not found or value is None: + return {"value": None, "disposition": "MISSING", "source_ref": f"{logical_id}#{pointer}"} + return {"value": str(value), "disposition": "OBSERVED", "source_ref": f"{logical_id}#{pointer}"} + + + def validate_ingress_contracts( + snapshots: Mapping[str, Snapshot], + contracts: Sequence[Mapping[str, Any]], + release_lock: Mapping[str, Any], + *, + deployment_snapshots: Mapping[str, Snapshot] | None = None, + deployment_documents: Mapping[str, Any] | None = None, + completion_seal: Mapping[str, Any] | None = None, + contract_manifest: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Strictly parse sources and verify release-bound schema, producer, identity, and seal rows.""" + + documents: dict[str, Any] = {} + rows: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + deployment_snapshots = deployment_snapshots or {} + deployment_documents = deployment_documents or {} + deployment_by_path = {snapshot.relative_path: snapshot for snapshot in deployment_snapshots.values()} + schema_documents = _schema_document_index(deployment_documents) + release_source_rows = _release_stage1_source_rows(release_lock) + release_ids = [str(row.get("logical_input_id")) for row in release_source_rows] + duplicate_release_ids = sorted(key for key, count in Counter(release_ids).items() if count > 1) + if duplicate_release_ids: + raise IngressError( + "RELEASE_SOURCE_CONTRACT_DUPLICATE", + "release stage1_sources contains duplicate logical_input_id rows", + details={"logical_input_ids": duplicate_release_ids}, + ) + expected_fixed = { + str(row["logical_input_id"]): str(row["path"]) + for row in DEFAULT_SOURCE_CONTRACTS + } + expected_release_ids = set(expected_fixed) | {"signal_payload_family"} + observed_release_ids = set(release_ids) + if observed_release_ids != expected_release_ids: + raise IngressError( + "RELEASE_SOURCE_CONTRACT_SET_MISMATCH", + "release stage1_sources must be the exact 16 fixed inputs plus signal_payload_family", + details={ + "missing": sorted(expected_release_ids - observed_release_ids), + "extra": sorted(observed_release_ids - expected_release_ids), + }, + ) + release_rows = {str(row.get("logical_input_id")): row for row in release_source_rows} + for logical_id, expected_path in expected_fixed.items(): + release_row = release_rows[logical_id] + if release_row.get("path") != expected_path or release_row.get("path_rule") not in {None, ""}: + raise IngressError( + "RELEASE_SOURCE_FIXED_PATH_MISMATCH", + f"fixed source path contract mismatch: {logical_id}", + ) + signal_family = release_rows["signal_payload_family"] + if ( + signal_family.get("path") is not None + or signal_family.get("path_rule") != "signals/" + or signal_family.get("adapter_id") != "S2A-SIGNAL-ALL-V1" + or signal_family.get("raw_hash_source") != "MANIFEST_ROW" + ): + raise IngressError( + "SIGNAL_PAYLOAD_FAMILY_CONTRACT_MISMATCH", + "signal_payload_family must use the approved manifest-expanded path contract", + ) + completion_hashes = _source_hash_index(completion_seal) + manifest_hashes = _source_hash_index(contract_manifest) + completion_producers = _source_producer_index(completion_seal) + manifest_producers = _source_producer_index(contract_manifest) + for contract in contracts: + logical_id = str(contract["logical_input_id"]) + snapshot = snapshots.get(logical_id) + release_row = release_rows.get(logical_id) + contract_missing = release_row is None + release_row = release_row or {} + alias_value = release_row.get("producer_alias", release_row.get("producer_alias_id")) + alias_id = str(alias_value) if isinstance(alias_value, str) else None + schema_ref = release_row.get("schema_ref") if isinstance(release_row.get("schema_ref"), dict) else None + row = { + "logical_input_id": logical_id, + "requirement_class": REQUIREMENT_CLASS_ENUM.get( + str(contract.get("criticality")), + "INTEGRITY_CORROBORATOR", + ), + "expected_path": contract.get("expected_path"), + "observed_path": contract.get("observed_path"), + "resolution_source": contract.get("resolution_source"), + "schema_id": schema_ref.get("$id") if schema_ref else release_row.get("schema_id"), + "schema_sha256": schema_ref.get("sha256") if schema_ref else release_row.get("schema_sha256"), + "producer_id": release_row.get("producer_id"), + "producer_alias_id": alias_id, + "adapter_id": release_row.get("adapter_id", ADAPTER_IDS.get(logical_id, "S2A-UNBOUND-V1")), + "run_identity_ref": release_row.get("run_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), + "transaction_identity_ref": release_row.get("transaction_identity_ref", {"value": None, "disposition": "MISSING", "source_ref": logical_id}), + "scope_refs": [logical_id], + "source_contract_row_refs": [logical_id], + "reason_codes": [], + "downstream_allowed_actions": [], + "issue_codes": [], + } + if contract_missing: + code = "RELEASE_SOURCE_CONTRACT_MISSING" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + declared_path = release_row.get("path") + if isinstance(declared_path, str) and declared_path != contract.get("expected_path"): + code = "RELEASE_SOURCE_PATH_MISMATCH" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if snapshot is None: + row.update( + { + "raw_sha256": None, + "byte_length": 0, + "parse_status": "NOT_OBSERVED", + "schema_status": "UNEVALUABLE", + "seal_status": "UNEVALUABLE", + "scope_technical_disposition": "UNAVAILABLE", + "impact_scope": "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER", + } + ) + code = "SOURCE_MISSING" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope=row["impact_scope"], source_refs=[logical_id])) + rows.append(row) + continue + row["raw_sha256"] = snapshot.raw_sha256 + row["byte_length"] = snapshot.byte_length + try: + document = load_json_strict( + snapshot, + max_depth=int(release_lock.get("limits", {}).get("max_json_depth", MAX_JSON_DEPTH)), + max_items=int(release_lock.get("limits", {}).get("max_json_items", MAX_JSON_ITEMS)), + ) + documents[logical_id] = document + row["parse_status"] = "PASS" + except IngressError as exc: + row["parse_status"] = "FAIL" + row["schema_status"] = "UNEVALUABLE" + row["seal_status"] = "UNEVALUABLE" + row["scope_technical_disposition"] = "UNAVAILABLE" + row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" + row["reason_codes"].append(exc.code) + row["issue_codes"].append(exc.code) + issues.append(_issue(exc.code, impact_scope=row["impact_scope"], source_refs=[logical_id], message=str(exc))) + rows.append(row) + continue + expected_adapter = ADAPTER_IDS.get(logical_id) + if expected_adapter is not None and release_row.get("adapter_id") not in {None, expected_adapter}: + code = "ADAPTER_ID_MISMATCH" + row["schema_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if schema_ref is not None: + schema_path = schema_ref.get("path") + schema_snapshot = deployment_by_path.get(schema_path) if isinstance(schema_path, str) else None + schema_document = deployment_documents.get(schema_path) if isinstance(schema_path, str) else None + expected_schema_hash = schema_ref.get("sha256") + expected_schema_id = schema_ref.get("$id") + if schema_snapshot is None or not isinstance(schema_document, dict): + schema_error = "SCHEMA_REF_NOT_IN_BOUNDED_DEPLOYMENT" + elif not isinstance(expected_schema_hash, str) or schema_snapshot.raw_sha256 != expected_schema_hash.lower(): + schema_error = "SCHEMA_HASH_MISMATCH" + elif expected_schema_id is not None and schema_document.get("$id") != expected_schema_id: + schema_error = "SCHEMA_ID_MISMATCH" + else: + schema_error = None + try: + _validate_schema_node( + document, + schema_document, + root_schema=schema_document, + schema_documents=schema_documents, + instance_path=logical_id, + ) + except _SchemaViolation as exc: + schema_error = "SOURCE_SCHEMA_VALIDATION_FAILED" + issues.append( + _issue( + schema_error, + impact_scope="CLUSTER", + source_refs=[logical_id], + message=str(exc), + ) + ) + if schema_error is not None: + row["schema_status"] = "FAIL" + row["reason_codes"].append(schema_error) + row["issue_codes"].append(schema_error) + if schema_error != "SOURCE_SCHEMA_VALIDATION_FAILED": + issues.append(_issue(schema_error, impact_scope="GLOBAL", source_refs=[logical_id])) + else: + row["schema_status"] = "PASS" + else: + adapter_errors = _closed_adapter_shape_errors( + document, + logical_id=logical_id, + adapter_id=str(row["adapter_id"]), + required_keys=release_row.get("required_keys", []), + release_lock=release_lock, + ) + if contract_missing: + row["schema_status"] = "UNEVALUABLE" + elif adapter_errors: + code = "ADAPTER_REQUIRED_KEY_MISSING" + row["schema_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append( + _issue( + code, + impact_scope="CLUSTER", + source_refs=[logical_id], + message="; ".join(adapter_errors), + ) + ) + else: + row["schema_status"] = "PASS" + expected_producer = release_row.get("producer_id") + document_producer = _producer_value(document) + sealed_producer = ( + completion_producers.get(logical_id) + or completion_producers.get(str(contract.get("observed_path"))) + or manifest_producers.get(logical_id) + or manifest_producers.get(str(contract.get("observed_path"))) + ) + if ( + document_producer is not None + and sealed_producer is not None + and document_producer != sealed_producer + ): + code = "PRODUCER_EVIDENCE_CONFLICT" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + observed_producer = document_producer or sealed_producer + if isinstance(expected_producer, str): + if observed_producer is None: + code = "PRODUCER_ID_UNEVALUABLE" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="CLUSTER", source_refs=[logical_id])) + elif not _producer_matches(observed_producer, expected_producer, alias_id, release_lock): + code = "PRODUCER_ID_MISMATCH" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + row["run_identity_ref"] = _identity_ref(document, release_row.get("run_identity_pointer"), logical_id) + row["transaction_identity_ref"] = _identity_ref( + document, + release_row.get("transaction_identity_pointer"), + logical_id, + ) + raw_hash_source = str(release_row.get("raw_hash_source", "NONE")) + if raw_hash_source in {"CASE_RUN_COMPLETION_SEAL", "COMPLETION_SEAL", "COMPLETION_SEAL_ROW"}: + expected_hash = completion_hashes.get(logical_id) or completion_hashes.get(str(contract.get("observed_path"))) + elif raw_hash_source in {"CONTRACT_MANIFEST", "CONTRACT_MANIFEST_ROW", "MANIFEST_ROW"}: + expected_hash = manifest_hashes.get(logical_id) or manifest_hashes.get(str(contract.get("observed_path"))) + elif raw_hash_source in {"COMPLETION_SEAL_OR_CONTRACT_MANIFEST", "SEALED_ROW"}: + expected_hash = ( + completion_hashes.get(logical_id) + or completion_hashes.get(str(contract.get("observed_path"))) + or manifest_hashes.get(logical_id) + or manifest_hashes.get(str(contract.get("observed_path"))) + ) + elif raw_hash_source in {"UNAVAILABLE_DEV", "NONE"}: + expected_hash = None + else: + expected_hash = None + code = "RAW_HASH_SOURCE_UNAPPROVED" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + if expected_hash is not None and expected_hash != snapshot.raw_sha256: + code = "RAW_HASH_MISMATCH" + row["seal_status"] = "FAIL" + row["reason_codes"].append(code) + row["issue_codes"].append(code) + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=[logical_id])) + else: + row["seal_status"] = "PASS" if expected_hash else "UNEVALUABLE" + row["scope_technical_disposition"] = ( + "UNAVAILABLE" + if contract_missing or any(code in row["issue_codes"] for code in {"RAW_HASH_MISMATCH", "SCHEMA_HASH_MISMATCH", "SCHEMA_ID_MISMATCH"}) + else "AVAILABLE" + if not row["issue_codes"] + else "AVAILABLE_WITH_ISSUES" + ) + row["impact_scope"] = "GLOBAL" if contract.get("criticality") == "identity_backbone" else "CLUSTER" + rows.append(row) + for identity_kind, field in ( + ("RUN", "run_identity_ref"), + ("TRANSACTION", "transaction_identity_ref"), + ): + observed_values = { + str(row[field]["value"]) + for row in rows + if row[field].get("disposition") == "OBSERVED" and row[field].get("value") is not None + } + if len(observed_values) > 1: + code = f"{identity_kind}_IDENTITY_CONFLICT" + issues.append(_issue(code, impact_scope="GLOBAL", source_refs=sorted(observed_values))) + for row in rows: + if row[field].get("disposition") == "OBSERVED": + row["issue_codes"] = sorted(set(row["issue_codes"] + [code])) + row["reason_codes"] = sorted(set(row["reason_codes"] + [code])) + row["scope_technical_disposition"] = "UNAVAILABLE" + return {"documents": documents, "source_contract_rows": rows, "issues": issues} + + + def _records_from_signal_document(document: Any) -> list[Any]: + if isinstance(document, list): + return list(document) + if isinstance(document, dict): + for key in ("signals", "records", "items"): + value = document.get(key) + if isinstance(value, list): + return list(value) + return [document] + return [document] + + + def _record_signal_id(record: Any) -> str | None: + if not isinstance(record, dict): + return None + value = record.get("signal_id") + if isinstance(value, str) and value: + return value + for wrapper in ("domain_activation_manifest", "payload", "data"): + nested = record.get(wrapper) + if isinstance(nested, dict) and isinstance(nested.get("signal_id"), str): + return nested["signal_id"] + return None + + + def expand_stage2_signal_all( + stage1_run_root: str | os.PathLike[str], + signal_manifest: Mapping[str, Any], + *, + max_file_bytes: int = MAX_FILE_BYTES, + max_total_bytes: int = MAX_RUN_BYTES, + signal_registry: Mapping[str, Any] | None = None, + ) -> dict[str, Any]: + """Expand Stage 2 ALL while separating semantic and integrity-only universes.""" + + downstream = signal_manifest.get("downstream_read_sets", {}) + stage2 = downstream.get("stage2", []) if isinstance(downstream, dict) else [] + if stage2 != ["ALL"]: + raise IngressError("SIGNAL_ALL_CONTRACT", "downstream_read_sets.stage2 must equal ['ALL']") + files = signal_manifest.get("files") + if not isinstance(files, list): + raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest files must be an array") + transaction_id = str(signal_manifest.get("manifest_transaction_id", signal_manifest.get("transaction_id", "MISSING"))) + file_rows: list[dict[str, Any]] = [] + semantic_rows: list[dict[str, Any]] = [] + integrity_rows: list[dict[str, Any]] = [] + occurrences: list[dict[str, Any]] = [] + payload_snapshots: list[Snapshot] = [] + issues: list[dict[str, Any]] = [] + path_counter: Counter[str] = Counter() + parsed_documents: dict[str, Any] = {} + aggregate_bytes = 0 + registry_entries = { + str(row.get("file")): row + for row in (signal_registry or {}).get("entries", []) + if isinstance(row, dict) and isinstance(row.get("file"), str) + } + compatibility_files = { + str(path) + for path in (signal_registry or {}).get("compatibility_views", []) + if isinstance(path, str) + } + domain_envelope_schema = (signal_registry or {}).get("domain_envelope") + observed_registry_files: set[str] = set() + for index, entry in enumerate(files): + if not isinstance(entry, dict) or not isinstance(entry.get("path"), str): + raise IngressError("SIGNAL_FILE_ROW_SHAPE", f"invalid signal file row at index {index}") + relative_payload = _safe_relative_path(entry["path"]).as_posix() + if relative_payload.startswith("signals/"): + raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "manifest file path must not include signals/ prefix") + physical = f"signals/{relative_payload}" + snapshot = open_bounded_snapshot( + stage1_run_root, + physical, + logical_input_id=f"signal_file:{index}", + max_bytes=max_file_bytes, + ) + document = load_json_strict(snapshot) + payload_snapshots.append(snapshot) + aggregate_bytes += snapshot.byte_length + if aggregate_bytes > max_total_bytes: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "signal ALL payloads exceed remaining run byte budget") + parsed_documents[relative_payload] = document + kind = entry.get("kind", "canonical") + if kind not in SEMANTIC_SIGNAL_KINDS | {"compatibility_view"}: + raise IngressError("SIGNAL_KIND_UNAPPROVED", f"unapproved signal file kind: {kind}") + semantic = kind in SEMANTIC_SIGNAL_KINDS + expected_hash = entry.get( + "file_sha256", entry.get("sha256", entry.get("raw_sha256")) + ) + row = { + "manifest_index": index, + "file_path": relative_payload, + "physical_path": physical, + "kind": kind, + "raw_sha256": snapshot.raw_sha256, + "byte_length": snapshot.byte_length, + "semantic": semantic, + "manifest_declared_record_count": entry.get("record_count"), + } + if expected_hash is not None and expected_hash != snapshot.raw_sha256: + row["hash_status"] = "FAIL" + issues.append(_issue("SIGNAL_FILE_HASH_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + else: + row["hash_status"] = "PASS" if expected_hash else "UNEVALUABLE" + records = _records_from_signal_document(document) + row["observed_record_count"] = len(records) + declared_count = entry.get("record_count") + if isinstance(declared_count, int) and declared_count != len(records): + row["record_count_status"] = "FAIL" + issues.append(_issue("SIGNAL_RECORD_COUNT_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + else: + row["record_count_status"] = "PASS" if isinstance(declared_count, int) else "UNEVALUABLE" + registry_row = registry_entries.get(relative_payload) + if kind == "canonical": + if signal_registry is not None and registry_row is None: + issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif registry_row is not None: + observed_registry_files.add(relative_payload) + declared_schema = entry.get("schema", entry.get("schema_path")) + if declared_schema is not None and declared_schema != registry_row.get("schema"): + issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif kind == "compatibility_view": + if relative_payload in registry_entries: + issues.append(_issue("SIGNAL_COMPATIBILITY_SUBSTITUTION", impact_scope="SIGNAL", source_refs=[physical])) + if signal_registry is not None and relative_payload not in compatibility_files: + issues.append(_issue("SIGNAL_REGISTRY_COVERAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + elif kind == "domain_signal": + declared_schema = entry.get("schema", entry.get("schema_path")) + if signal_registry is not None and declared_schema not in {None, domain_envelope_schema}: + issues.append(_issue("SIGNAL_SCHEMA_LINEAGE_MISMATCH", impact_scope="SIGNAL", source_refs=[physical])) + file_rows.append(row) + path_counter[relative_payload] += 1 + if semantic: + semantic_rows.append(row) + for record_ordinal, record in enumerate(records): + signal_id = _record_signal_id(record) + occurrence_key = [transaction_id, relative_payload, record_ordinal, signal_id] + occurrences.append( + { + "occurrence_key": occurrence_key, + "occurrence_ref": f"SIGO-{canonical_digest(occurrence_key)[:24]}", + "manifest_transaction_id": transaction_id, + "file_path": relative_payload, + "record_ordinal": record_ordinal, + "signal_id": signal_id, + "disposition": "UNMAPPED" if signal_id is None else "UNUSED", + "binding_refs": [], + "raw_record_sha256": canonical_digest(record), + "record": record, + } + ) + else: + integrity_rows.append(row) + duplicates = sorted(path for path, count in path_counter.items() if count > 1) + if duplicates: + issues.append(_issue("SIGNAL_ALL_DUPLICATE_FILE_ROW", impact_scope="SIGNAL", source_refs=duplicates)) + manifest_counter = Counter((i, row["file_path"], row["kind"]) for i, row in enumerate(file_rows)) + partition_counter = Counter((row["manifest_index"], row["file_path"], row["kind"]) for row in semantic_rows + integrity_rows) + missing_registry_files = sorted(set(registry_entries) - observed_registry_files) if signal_registry is not None else [] + if missing_registry_files: + issues.append( + _issue( + "SIGNAL_REGISTRY_COVERAGE_MISMATCH", + impact_scope="SIGNAL", + source_refs=[f"signals/{path}" for path in missing_registry_files], + ) + ) + file_conservation = ( + manifest_counter == partition_counter + and not duplicates + and not missing_registry_files + and not any(row["hash_status"] == "FAIL" or row["record_count_status"] == "FAIL" for row in file_rows) + ) + record_counter = Counter(tuple(row["occurrence_key"]) for row in occurrences) + partitioned_record_counter = Counter( + tuple(row["occurrence_key"]) + for row in occurrences + if row["disposition"] in {"USED", "UNUSED", "UNMAPPED"} + ) + record_conservation = record_counter == partitioned_record_counter + return { + "manifest_transaction_id": transaction_id, + "ordered_file_rows": file_rows, + "semantic_file_rows": semantic_rows, + "integrity_only_file_rows": integrity_rows, + "record_occurrences": occurrences, + "used_record_occurrences": [], + "unused_record_occurrences": [row for row in occurrences if row["disposition"] == "UNUSED"], + "unmapped_record_occurrences": [row for row in occurrences if row["disposition"] == "UNMAPPED"], + "_parsed_documents_by_path": parsed_documents, + "_payload_snapshots": payload_snapshots, + "file_conservation_pass": file_conservation, + "record_conservation_pass": record_conservation, + "aggregate_payload_bytes": aggregate_bytes, + "issues": issues, + } + + + def _collect_values_for_keys(value: Any, keys: frozenset[str]) -> set[str]: + result: set[str] = set() + stack = [value] + while stack: + current = stack.pop() + if isinstance(current, dict): + for key, child in current.items(): + if key in keys: + if isinstance(child, list): + result.update(str(item) for item in child if item is not None) + elif child is not None: + result.add(str(child)) + stack.append(child) + elif isinstance(current, list): + stack.extend(current) + return result + + + def bind_signal_occurrences(signal_all: MutableMapping[str, Any], documents: Mapping[str, Any]) -> dict[str, Any]: + """Bind each semantic signal occurrence to explicit Stage 1 references without deduplication.""" + + explicit_signal_ids = _collect_values_for_keys( + documents, + frozenset({"signal_id", "signal_ids", "signal_refs", "emitted_signal_ids", "required_signal_ids"}), + ) + known_refs = { + "fact_id": _collect_values_for_keys(documents.get("fact_ledger_base"), frozenset({"fact_id"})), + "source_bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "source_bo_id", "source_bo_ids"})), + "bo_id": _collect_values_for_keys(documents, frozenset({"BO_ID", "bo_id"})), + "structure_id": _collect_values_for_keys(documents.get("legal_effect_structures"), frozenset({"structure_id"})), + "domain_id": _collect_values_for_keys(documents, frozenset({"domain_id", "domain_ids", "active_domain_ids"})), + "evidence_id": _collect_values_for_keys(documents.get("evidence_indexed"), frozenset({"evidence_id", "id"})), + "event_id": _collect_values_for_keys(documents.get("evidence_event_candidates"), frozenset({"event_id", "id"})), + } + link_keys = { + "fact_id": ("fact_id", "fact_ids"), + "source_bo_id": ("source_bo_id", "source_bo_ids"), + "bo_id": ("bo_id", "bo_ids"), + "structure_id": ("structure_id", "structure_ids"), + "domain_id": ("domain_id", "domain_ids"), + "evidence_id": ("evidence_id", "evidence_ids"), + "event_id": ("event_id", "event_ids"), + } + for occurrence in signal_all.get("record_occurrences", []): + signal_id = occurrence.get("signal_id") + record = occurrence.get("record") + bindings: set[str] = set() + if isinstance(signal_id, str) and signal_id in explicit_signal_ids: + bindings.add(f"signal_id:{signal_id}") + for ref_kind, candidate_keys in link_keys.items(): + observed = _collect_values_for_keys(record, frozenset(candidate_keys)) + for ref in sorted(observed & known_refs[ref_kind]): + bindings.add(f"{ref_kind}:{ref}") + if not isinstance(signal_id, str) or not signal_id: + occurrence["disposition"] = "UNMAPPED" + elif bindings: + occurrence["disposition"] = "USED" + else: + occurrence["disposition"] = "UNUSED" + occurrence["binding_refs"] = sorted(bindings) + for disposition, key in ( + ("USED", "used_record_occurrences"), + ("UNUSED", "unused_record_occurrences"), + ("UNMAPPED", "unmapped_record_occurrences"), + ): + signal_all[key] = [ + row for row in signal_all.get("record_occurrences", []) if row.get("disposition") == disposition + ] + source_counter = Counter(tuple(row["occurrence_key"]) for row in signal_all.get("record_occurrences", [])) + partition_counter = Counter( + tuple(row["occurrence_key"]) + for key in ("used_record_occurrences", "unused_record_occurrences", "unmapped_record_occurrences") + for row in signal_all[key] + ) + signal_all["record_conservation_pass"] = source_counter == partition_counter + return dict(signal_all) + + + def _activation_payload(value: Mapping[str, Any]) -> Mapping[str, Any]: + for key in ("domain_activation_manifest", "activation", "payload", "data"): + nested = value.get(key) + if isinstance(nested, dict) and any(field in nested for field in SG01_PROJECTION_FIELDS): + return nested + return value + + + def verify_activation_projection( + routing_activation: Mapping[str, Any], + signal_activation: Mapping[str, Any], + *, + routing_raw_sha256: str | None = None, + signal_raw_sha256: str | None = None, + ) -> dict[str, Any]: + """Compare approved semantic SG-01 projection while retaining both raw hashes.""" + + left = _activation_payload(routing_activation) + right = _activation_payload(signal_activation) + missing_left = [field for field in SG01_PROJECTION_FIELDS if field not in left] + missing_right = [field for field in SG01_PROJECTION_FIELDS if field not in right] + if missing_left or missing_right: + raise IngressError( + "SG01_PROJECTION_SHAPE", + "both activation artifacts must expose the complete approved 17-field projection", + details={"routing_missing": missing_left, "signal_missing": missing_right}, + ) + + def project(value: Mapping[str, Any]) -> dict[str, Any]: + result: dict[str, Any] = {} + for field in SG01_PROJECTION_FIELDS: + child = value[field] + if field in SG01_SET_FIELDS: + if not isinstance(child, list): + raise IngressError("SG01_PROJECTION_SHAPE", f"{field} must be an array") + child = sorted({canonical_json_bytes(item): item for item in child}.values(), key=canonical_json_bytes) + result[field] = child + return result + + left_projection = project(left) + right_projection = project(right) + if left_projection != right_projection: + raise IngressError( + "SG01_SEMANTIC_DRIFT", + "routing activation and signal SG-01 semantic projections differ", + details={"routing_projection": left_projection, "signal_projection": right_projection}, + ) + return { + "status": "PASS", + "projection": left_projection, + "projection_sha256": canonical_digest(left_projection), + "routing_raw_sha256": routing_raw_sha256, + "signal_raw_sha256": signal_raw_sha256, + "compared_keys": list(SG01_PROJECTION_FIELDS), + } + + + def verify_cross_artifact_seals( + documents: Mapping[str, Any], + snapshots: Mapping[str, Snapshot], + deployment_snapshots: Mapping[str, Snapshot] | None = None, + ) -> dict[str, Any]: + """Recompute the P1 guard and current-v8 producer invariants.""" + + checks: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + deployment_snapshots = deployment_snapshots or {} + p1 = documents.get("stage1_part1_soft_gate_handoff") + if isinstance(p1, dict): + digest_guard = p1.get("digest_guard") + if not isinstance(digest_guard, dict): + issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff"])) + digest_guard = {} + elif any(key not in digest_guard for key in P1_DIGEST_KEYS): + issues.append(_issue("P1_SEVEN_KEY_MISSING", source_refs=["stage1_part1_soft_gate_handoff#digest_guard"])) + for digest_key, logical_id in P1_DIGEST_KEYS.items(): + source = snapshots.get(logical_id) or deployment_snapshots.get(logical_id) + observed = source.raw_sha256 if source else None + expected = digest_guard.get(digest_key) + passed = expected is not None and observed is not None and expected == observed + checks.append({"check_id": f"P1:{digest_key}", "status": "PASS" if passed else "UNEVALUABLE" if source is None else "FAIL"}) + if expected is not None and observed is not None and not passed: + issues.append(_issue("P1_DIGEST_MISMATCH", source_refs=[logical_id])) + else: + issues.append(_issue("P1_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part1_soft_gate_handoff"])) + p2 = documents.get("stage1_part2_review_handoff") + if p2 is not None and not isinstance(p2, dict): + issues.append(_issue("P2_HANDOFF_NOT_FLAT_OBJECT", source_refs=["stage1_part2_review_handoff"])) + for stage in (3, 4): + logical = f"stage1_part{stage}_review_handoff" + value = documents.get(logical) + if value is not None: + wrapper_present = isinstance(value, dict) and isinstance(value.get(logical), dict) + if not wrapper_present: + issues.append(_issue(f"P{stage}_WRAPPER_MISSING", source_refs=[logical])) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + for index, row in enumerate(ledger_rows): + if not isinstance(row, dict) or "domain_effects" not in row or "calculation_requests" not in row: + issues.append(_issue("CURRENT_V8_LEDGER_EXTENSION_MISSING", impact_scope="FACT", source_refs=[f"fact_ledger_base#/{index}"])) + return {"checks": checks, "issues": issues, "passed": not any(item["severity"] == "ERROR" for item in issues)} + + + def check_conservation( + documents: Mapping[str, Any], + *, + signal_all: Mapping[str, Any] | None = None, + normalized_reviews: Mapping[str, Any] | None = None, + source_snapshots: Mapping[str, Snapshot] | None = None, + ) -> dict[str, Any]: + """Independently compute core set, cardinality, and multiset invariants.""" + + checks: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + source_snapshots = source_snapshots or {} + + def add_check( + check_id: str, + passed: bool | None, + left: Sequence[Any] | Counter[Any] | None, + right: Sequence[Any] | Counter[Any] | None, + *, + issue_code: str, + impact_scope: str, + source_refs: Sequence[str], + details: Mapping[str, Any] | None = None, + ) -> None: + left_counter = left if isinstance(left, Counter) else Counter(canonical_digest(value) for value in (left or [])) + right_counter = right if isinstance(right, Counter) else Counter(canonical_digest(value) for value in (right or [])) + row: dict[str, Any] = { + "check_id": check_id, + "status": "UNEVALUABLE" if passed is None else "PASS" if passed else "FAIL", + "left_count": sum(left_counter.values()) if left is not None else None, + "right_count": sum(right_counter.values()) if right is not None else None, + "left_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in left_counter.items())) if left is not None else None, + "right_counter_digest": canonical_digest(sorted((canonical_digest(key), count) for key, count in right_counter.items())) if right is not None else None, + } + if details: + row.update(details) + checks.append(row) + if passed is False: + issues.append(_issue(issue_code, impact_scope=impact_scope, source_refs=source_refs)) + + bo_rows = _array_rows(documents.get("bo"), ("business_objects", "BO", "rows", "items")) + ledger_rows = _array_rows(documents.get("fact_ledger_base"), ("facts", "fact_ledger", "rows", "items")) + bo_ids = [str(row["BO_ID"]) for row in bo_rows if isinstance(row, dict) and row.get("BO_ID") is not None] + source_bo_ids = [ + str(row["source_bo_id"]) + for row in ledger_rows + if isinstance(row, dict) and row.get("source_bo_id") is not None + ] + missing_bo_id_rows = [index for index, row in enumerate(bo_rows) if not isinstance(row, dict) or row.get("BO_ID") is None] + missing_source_bo_rows = [ + index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("source_bo_id") is None + ] + bo_pass = ( + not missing_bo_id_rows + and not missing_source_bo_rows + and Counter(bo_ids) == Counter(source_bo_ids) + ) + add_check( + "BO_FACT_MULTISET", + bo_pass, + bo_ids, + source_bo_ids, + issue_code="BO_FACT_CONSERVATION_FAILED", + impact_scope="FACT", + source_refs=["bo", "fact_ledger_base"], + details={ + "missing_bo_id_rows": missing_bo_id_rows, + "missing_source_bo_id_rows": missing_source_bo_rows, + "duplicate_bo_ids": sorted(key for key, count in Counter(bo_ids).items() if count > 1), + "dangling_source_bo_ids": sorted(set(source_bo_ids) - set(bo_ids)), + }, + ) + missing_fact_id_rows = [ + index for index, row in enumerate(ledger_rows) if not isinstance(row, dict) or row.get("fact_id") is None + ] + fact_ids = [str(row["fact_id"]) for row in ledger_rows if isinstance(row, dict) and row.get("fact_id") is not None] + expected_fact_ids = [f"F-{index:03d}" for index in range(1, len(ledger_rows) + 1)] + fact_pass = not missing_fact_id_rows and fact_ids == expected_fact_ids and len(fact_ids) == len(set(fact_ids)) + add_check( + "FACT_ID_SEQUENCE", + fact_pass, + fact_ids, + expected_fact_ids, + issue_code="FACT_ID_CONSERVATION_FAILED", + impact_scope="FACT", + source_refs=["fact_ledger_base"], + details={"missing_fact_id_rows": missing_fact_id_rows, "observed": fact_ids}, + ) + extension_missing = [ + index + for index, row in enumerate(ledger_rows) + if not isinstance(row, dict) + or not isinstance(row.get("domain_effects"), dict) + or not isinstance(row.get("calculation_requests"), list) + ] + add_check( + "CURRENT_V8_LEDGER_EXTENSIONS", + not extension_missing, + list(range(len(ledger_rows))), + [index for index in range(len(ledger_rows)) if index not in extension_missing], + issue_code="CURRENT_V8_LEDGER_EXTENSION_MISSING", + impact_scope="FACT", + source_refs=["fact_ledger_base"], + details={"missing_row_indices": extension_missing}, + ) + les_rows = _array_rows( + documents.get("legal_effect_structures"), + ("structures", "structure_records", "legal_effect_structures", "rows", "items"), + ) + dangling_les: list[str] = [] + les_ids: list[str] = [] + for row in les_rows: + if not isinstance(row, dict): + continue + structure_id = row.get("structure_id", row.get("legal_effect_structure_id")) + if structure_id is not None: + les_ids.append(str(structure_id)) + refs = row.get("source_bo_ids", []) + if isinstance(refs, list): + dangling_les.extend(str(ref) for ref in refs if ref not in set(bo_ids)) + duplicate_les_ids = sorted(key for key, count in Counter(les_ids).items() if count > 1) + les_pass = not dangling_les and not duplicate_les_ids and len(les_ids) == len(les_rows) + add_check( + "LES_BO_JOIN", + les_pass, + [str(row.get("structure_id", row.get("legal_effect_structure_id"))) for row in les_rows if isinstance(row, dict)], + les_ids, + issue_code="LES_BO_JOIN_FAILED", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures", "bo"], + details={"dangling_refs": sorted(dangling_les), "duplicate_structure_ids": duplicate_les_ids}, + ) + declared_les_count = None + les_document = documents.get("legal_effect_structures") + if isinstance(les_document, dict): + for key in ("declared_structure_count", "structure_count", "record_count"): + if isinstance(les_document.get(key), int): + declared_les_count = int(les_document[key]) + break + declared_les_pass = None if declared_les_count is None else declared_les_count == len(les_rows) + add_check( + "LES_DECLARED_ACTUAL_COUNT", + declared_les_pass, + [None] * declared_les_count if declared_les_count is not None else None, + [None] * len(les_rows), + issue_code="LES_DECLARED_COUNT_MISMATCH", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures"], + ) + actual_domain_index: dict[str, list[str]] = defaultdict(list) + actual_bo_index: dict[str, list[str]] = defaultdict(list) + ledger_structure_refs: list[tuple[str, str, str]] = [] + ledger_type_refs: list[tuple[str, str, str]] = [] + actual_structure_refs: list[tuple[str, str, str]] = [] + actual_type_refs: list[tuple[str, str, str]] = [] + route_count_errors: list[str] = [] + for row in les_rows: + if not isinstance(row, dict): + continue + structure_id = str(row.get("structure_id", row.get("legal_effect_structure_id", "MISSING"))) + domain_id = str(row.get("domain_id", "MISSING")) + type_id = str(row.get("type_id", row.get("type", "MISSING"))) + actual_domain_index[domain_id].append(structure_id) + source_ids = row.get("source_bo_ids", []) + if isinstance(source_ids, list): + for bo_id in source_ids: + actual_bo_index[str(bo_id)].append(structure_id) + actual_structure_refs.append((str(bo_id), domain_id, structure_id)) + actual_type_refs.append((str(bo_id), domain_id, type_id)) + routes = row.get("routes", []) + if isinstance(routes, list) and row.get("route_count", len(routes)) != len(routes): + route_count_errors.append(structure_id) + for row in ledger_rows: + if not isinstance(row, dict): + continue + bo_id = str(row.get("source_bo_id", "MISSING")) + effects = row.get("domain_effects", {}) + if not isinstance(effects, dict): + continue + for domain_id, effect in effects.items(): + if not isinstance(effect, dict): + continue + for structure_id in effect.get("structure_ids", []) if isinstance(effect.get("structure_ids"), list) else []: + ledger_structure_refs.append((bo_id, str(domain_id), str(structure_id))) + for type_id in effect.get("type_ids", []) if isinstance(effect.get("type_ids"), list) else []: + ledger_type_refs.append((bo_id, str(domain_id), str(type_id))) + structure_index = les_document.get("structure_index", {}) if isinstance(les_document, dict) else {} + index_present = isinstance(structure_index, dict) and bool(structure_index) + index_ok = True + if index_present: + declared_by_domain = structure_index.get("by_domain_id", {}) + declared_by_bo = structure_index.get("by_bo_id", {}) + index_ok = ( + isinstance(declared_by_domain, dict) + and isinstance(declared_by_bo, dict) + and {str(key): Counter(map(str, value)) for key, value in declared_by_domain.items() if isinstance(value, list)} + == {key: Counter(value) for key, value in actual_domain_index.items()} + and {str(key): Counter(map(str, value)) for key, value in declared_by_bo.items() if isinstance(value, list)} + == {key: Counter(value) for key, value in actual_bo_index.items()} + ) + reverse_ok = ( + (not ledger_structure_refs or Counter(ledger_structure_refs) == Counter(actual_structure_refs)) + and (not ledger_type_refs or Counter(ledger_type_refs) == Counter(actual_type_refs)) + and not route_count_errors + and index_ok + ) + add_check( + "LES_REVERSE_INDEX", + reverse_ok, + ledger_structure_refs + ledger_type_refs, + actual_structure_refs + actual_type_refs, + issue_code="LES_REVERSE_INDEX_MISMATCH", + impact_scope="CLUSTER", + source_refs=["legal_effect_structures", "fact_ledger_base"], + details={"index_present": index_present, "route_count_errors": route_count_errors}, + ) + evidence_rows = _array_rows(documents.get("evidence_indexed"), ("evidence", "evidence_items", "rows", "items")) + event_rows = _array_rows(documents.get("evidence_event_candidates"), ("events", "event_candidates", "rows", "items")) + evidence_ids = [ + str(row.get("evidence_id", row.get("id"))) + for row in evidence_rows + if isinstance(row, dict) and (row.get("evidence_id") is not None or row.get("id") is not None) + ] + event_ids = [ + str(row.get("event_id", row.get("id"))) + for row in event_rows + if isinstance(row, dict) and (row.get("event_id") is not None or row.get("id") is not None) + ] + fact_evidence_refs: list[str] = [] + fact_event_refs: list[str] = [] + event_evidence_refs: list[str] = [] + for row in ledger_rows: + if not isinstance(row, dict): + continue + evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) + event_values = row.get("event_refs", row.get("event_ids", [])) + if isinstance(evidence_values, list): + fact_evidence_refs.extend(str(ref) for ref in evidence_values) + if isinstance(event_values, list): + fact_event_refs.extend(str(ref) for ref in event_values) + for row in event_rows: + if not isinstance(row, dict): + continue + evidence_values = row.get("evidence_refs", row.get("evidence_ids", [])) + if isinstance(evidence_values, list): + event_evidence_refs.extend(str(ref) for ref in evidence_values) + evidence_failures = sorted( + set(fact_evidence_refs + event_evidence_refs) - set(evidence_ids) + ) + duplicate_evidence_ids = sorted(key for key, count in Counter(evidence_ids).items() if count > 1) + evidence_pass = not evidence_failures and not duplicate_evidence_ids + add_check( + "EVIDENCE_REFERENCE_CONSERVATION", + evidence_pass, + fact_evidence_refs + event_evidence_refs, + evidence_ids, + issue_code="EVIDENCE_REFERENCE_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_indexed", "fact_ledger_base", "evidence_event_candidates"], + details={"dangling_refs": evidence_failures, "duplicate_evidence_ids": duplicate_evidence_ids}, + ) + event_failures = sorted(set(fact_event_refs) - set(event_ids)) + duplicate_event_ids = sorted(key for key, count in Counter(event_ids).items() if count > 1) + event_pass = not event_failures and not duplicate_event_ids + add_check( + "EVENT_REFERENCE_CONSERVATION", + event_pass, + fact_event_refs, + event_ids, + issue_code="EVENT_REFERENCE_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_event_candidates", "fact_ledger_base"], + details={"dangling_refs": event_failures, "duplicate_event_ids": duplicate_event_ids}, + ) + disposition_rows = [row.get("disposition") for row in event_rows if isinstance(row, dict) and "disposition" in row] + b2_gate = documents.get("b2_event_candidates_gate") + declared_dispositions = None + if isinstance(b2_gate, dict): + declared_dispositions = b2_gate.get("event_disposition_counts") + if declared_dispositions is None and isinstance(b2_gate.get("summary"), dict): + declared_dispositions = b2_gate["summary"].get("event_disposition_counts") + if isinstance(declared_dispositions, dict): + disposition_expected = Counter( + {str(key): int(value) for key, value in declared_dispositions.items() if isinstance(value, int)} + ) + disposition_actual = Counter(str(value) for value in disposition_rows) + disposition_pass: bool | None = disposition_actual == disposition_expected + elif disposition_rows: + disposition_expected = Counter(str(value) for value in disposition_rows) + disposition_actual = Counter(str(value) for value in disposition_rows) + disposition_pass = all(isinstance(value, str) and value for value in disposition_rows) + else: + disposition_expected = Counter() + disposition_actual = Counter() + disposition_pass = None + add_check( + "EVENT_DISPOSITION_CONSERVATION", + disposition_pass, + disposition_actual, + disposition_expected, + issue_code="EVENT_DISPOSITION_CONSERVATION_FAILED", + impact_scope="EVIDENCE", + source_refs=["evidence_event_candidates", "b2_event_candidates_gate"], + ) + writer_report = documents.get("fact_ledger_writer_report") + if isinstance(writer_report, dict): + observed_domain_coverage = Counter( + str(domain_id) + for row in ledger_rows + if isinstance(row, dict) and isinstance(row.get("domain_effects"), dict) + for domain_id in row["domain_effects"] + ) + declared_domain_coverage = Counter( + {str(key): int(value) for key, value in writer_report.get("domain_effect_coverage", {}).items() if isinstance(value, int)} + ) + observed_readiness = Counter( + str(request.get("operand_state")) + for row in ledger_rows + if isinstance(row, dict) and isinstance(row.get("calculation_requests"), list) + for request in row["calculation_requests"] + if isinstance(request, dict) + ) + declared_readiness = Counter( + {str(key): int(value) for key, value in writer_report.get("calculation_readiness", {}).items() if isinstance(value, int)} + ) + ledger_snapshot = source_snapshots.get("fact_ledger_base") + final_hash = writer_report.get("final_sha256") + writer_pass = ( + writer_report.get("row_count") == len(ledger_rows) + and declared_domain_coverage == observed_domain_coverage + and declared_readiness == observed_readiness + and (ledger_snapshot is None or final_hash == ledger_snapshot.raw_sha256) + ) + add_check( + "FACT_LEDGER_WRITER_REPORT_CONNECTION", + writer_pass, + [len(ledger_rows), observed_domain_coverage, observed_readiness, ledger_snapshot.raw_sha256 if ledger_snapshot else None], + [writer_report.get("row_count"), declared_domain_coverage, declared_readiness, final_hash], + issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", + impact_scope="FACT", + source_refs=["fact_ledger_base", "fact_ledger_writer_report"], + ) + else: + add_check( + "FACT_LEDGER_WRITER_REPORT_CONNECTION", + None, + None, + None, + issue_code="FACT_LEDGER_WRITER_REPORT_MISMATCH", + impact_scope="FACT", + source_refs=["fact_ledger_base", "fact_ledger_writer_report"], + ) + if signal_all is not None: + file_pass = bool(signal_all.get("file_conservation_pass")) + record_pass = bool(signal_all.get("record_conservation_pass")) + checks.append({"check_id": "SIGNAL_FILE_ROW_CONSERVATION", "status": "PASS" if file_pass else "FAIL"}) + checks.append({"check_id": "SIGNAL_RECORD_OCCURRENCE_CONSERVATION", "status": "PASS" if record_pass else "FAIL"}) + issues.extend(signal_all.get("issues", [])) + if not file_pass: + issues.append(_issue("SIGNAL_FILE_CONSERVATION_FAILED", impact_scope="SIGNAL")) + if not record_pass: + issues.append(_issue("SIGNAL_RECORD_CONSERVATION_FAILED", impact_scope="SIGNAL")) + if normalized_reviews is not None: + review_pass = normalized_reviews.get("conservation_status") == "PASS" + checks.append({"check_id": "REVIEW_OCCURRENCE_CONSERVATION", "status": "PASS" if review_pass else "FAIL"}) + if not review_pass: + issues.append(_issue("REVIEW_CONSERVATION_FAILED", impact_scope="REVIEW_ITEM")) + issues.extend(normalized_reviews.get("_issues", [])) + return {"checks": checks, "issues": issues, "passed": not any(check["status"] == "FAIL" for check in checks)} + + + def _source_ref( + logical_id: str, + pointer: str, + raw_value: Any = _RAW_VALUE_UNSET, + *, + stage1_id: str | None = None, + ) -> dict[str, Any]: + """Build a truthful RFC 6901 provenance row without pointer narrowing.""" + + row: dict[str, Any] = { + "logical_artifact_id": logical_id, + "json_pointer": pointer, + "raw_value_sha256": canonical_digest( + [logical_id, pointer] + if raw_value is _RAW_VALUE_UNSET + else raw_value + ), + "source_contract_row_ref": logical_id, + } + if stage1_id is not None: + row["stage1_id"] = stage1_id + return row + + + def _tarjan_scc(nodes: Sequence[str], edges: Sequence[tuple[str, str]]) -> list[list[str]]: + adjacency: dict[str, list[str]] = {node: [] for node in nodes} + for source, target in edges: + adjacency.setdefault(source, []).append(target) + adjacency.setdefault(target, []) + for value in adjacency.values(): + value.sort() + index = 0 + stack: list[str] = [] + on_stack: set[str] = set() + indices: dict[str, int] = {} + lowlink: dict[str, int] = {} + components: list[list[str]] = [] + + def visit(node: str) -> None: + nonlocal index + indices[node] = index + lowlink[node] = index + index += 1 + stack.append(node) + on_stack.add(node) + for neighbor in adjacency[node]: + if neighbor not in indices: + visit(neighbor) + lowlink[node] = min(lowlink[node], lowlink[neighbor]) + elif neighbor in on_stack: + lowlink[node] = min(lowlink[node], indices[neighbor]) + if lowlink[node] == indices[node]: + component: list[str] = [] + while True: + member = stack.pop() + on_stack.remove(member) + component.append(member) + if member == node: + break + components.append(sorted(component)) + + for node in sorted(adjacency): + if node not in indices: + visit(node) + return sorted(components, key=lambda component: component[0]) + + + def _inline_sha256(value: str, *, code: str) -> str: + if not isinstance(value, str) or re.fullmatch(r"[a-f0-9]{64}", value) is None: + raise IngressError(code, "expected one lowercase SHA-256 digest") + return value + + + def _inline_relative_path(value: str, *, code: str) -> str: + if not isinstance(value, str) or not value or "\x00" in value or "\\" in value: + raise IngressError(code, "logical path is empty or malformed") + if unicodedata.normalize("NFC", value) != value: + raise IngressError(code, "logical path must already be NFC") + path = PurePosixPath(value) + if path.is_absolute() or any(part in {"", ".", ".."} for part in path.parts): + raise IngressError(code, "logical path must be a contained relative path") + rendered = path.as_posix() + if rendered != value: + raise IngressError(code, "logical path is not canonical") + return rendered + + + def _inline_parse_mcp_payload(raw: bytes, expected_id: int) -> Mapping[str, Any]: + """Parse one JSON or SSE JSON-RPC terminal response with an exact ID.""" + + candidates: list[Any] + try: + candidates = [load_json_strict(raw)] + except IngressError: + try: + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise IngressError("MCP_RESPONSE_UTF8", "MCP response is not strict UTF-8") from exc + events: list[bytes] = [] + data_lines: list[str] = [] + for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n"): + if line == "": + if data_lines: + events.append("\n".join(data_lines).encode("utf-8")) + data_lines = [] + continue + if line.startswith(":") or line.startswith("event:") or line.startswith("id:") or line.startswith("retry:"): + continue + if not line.startswith("data:"): + raise IngressError("MCP_SSE_SHAPE", "unexpected non-data SSE line") + payload = line[5:] + if payload.startswith(" "): + payload = payload[1:] + data_lines.append(payload) + if data_lines: + events.append("\n".join(data_lines).encode("utf-8")) + if not events: + raise IngressError("MCP_RESPONSE_SHAPE", "MCP response contains no JSON terminal event") + candidates = [load_json_strict(event) for event in events] + matching = [ + item + for item in candidates + if isinstance(item, dict) and item.get("id") == expected_id + ] + if len(matching) != 1: + raise IngressError( + "MCP_RESPONSE_ID_MISMATCH", + "MCP response must contain exactly one terminal result with the JSON-RPC message ID", + ) + response = matching[0] + if response.get("jsonrpc") != "2.0": + raise IngressError("MCP_JSONRPC_VERSION", "MCP response jsonrpc must equal 2.0") + if response.get("error") is not None: + raise IngressError( + "MCP_JSONRPC_ERROR", + "MCP server returned a JSON-RPC error", + details={"rpc_error": response.get("error")}, + ) + if "result" not in response or not isinstance(response["result"], dict): + raise IngressError("MCP_RESULT_SHAPE", "MCP response result must be an object") + return response + + + def _inline_tool_text(result: Mapping[str, Any], tool_name: str) -> str: + if result.get("isError") is True: + content = result.get("content") + rendered = canonical_json_bytes(content).decode("utf-8", errors="replace") if content is not None else "" + lowered = rendered.lower() + code = ( + "LOCALDOCS_NOT_FOUND" + if any(marker in lowered for marker in ("not found", "does not exist", "no such file")) + else "MCP_TOOL_ERROR" + ) + raise IngressError(code, f"localdocs {tool_name} returned isError=true") + content = result.get("content") + if not isinstance(content, list) or len(content) != 1: + raise IngressError("MCP_CONTENT_CARDINALITY", "MCP tool result must contain exactly one content block") + block = content[0] + if not isinstance(block, dict) or block.get("type") != "text" or not isinstance(block.get("text"), str): + raise IngressError("MCP_CONTENT_SHAPE", "MCP tool result must contain one text block") + return block["text"] + + + def _inline_binary_envelope(text: str, logical_path: str) -> bytes: + value = load_json_strict(text) + if isinstance(value, dict) and "results" in value: + results = value.get("results") + if not isinstance(results, list) or len(results) != 1 or not isinstance(results[0], dict): + raise IngressError("LOCALDOCS_RESULT_CARDINALITY", "binary response must contain one result row") + inner: Any = results[0].get("content", results[0].get("text")) + value = load_json_strict(inner) if isinstance(inner, str) else inner + if not isinstance(value, dict) or not isinstance(value.get("content_base64"), str): + raise IngressError("LOCALDOCS_BINARY_ENVELOPE", "binary response lacks content_base64") + try: + payload = base64.b64decode(value["content_base64"].encode("ascii"), validate=True) + except (UnicodeEncodeError, binascii.Error, ValueError) as exc: + raise IngressError("LOCALDOCS_BASE64_INVALID", "binary response is not strict base64") from exc + declared_size = value.get("byte_length", value.get("size")) + if declared_size is not None and (not isinstance(declared_size, int) or declared_size != len(payload)): + raise IngressError("LOCALDOCS_BYTE_LENGTH_MISMATCH", f"binary length mismatch: {logical_path}") + declared_hash = value.get("sha256") + if declared_hash is not None and declared_hash != hashlib.sha256(payload).hexdigest(): + raise IngressError("LOCALDOCS_HASH_MISMATCH", f"binary hash mismatch: {logical_path}") + return payload + + + class _InlineLocaldocs: + """Minimal user/workspace-bound localdocs JSON-RPC client.""" + + def __init__( + self, + user_hash: str, + workspace_hash: str, + *, + client: Any | None = None, + timeout_seconds: int = 60, + ) -> None: + self.user_hash = _inline_sha256(user_hash, code="USER_CONTEXT_HASH_INVALID") + self.workspace_hash = _inline_sha256( + workspace_hash, + code="WORKSPACE_CONTEXT_HASH_INVALID", + ) + if client is None: + try: + import httpx # type: ignore + except ImportError as exc: + raise IngressError("HTTPX_UNAVAILABLE", "Code Executor must supply httpx==0.28.1") from exc + client = httpx.Client(timeout=timeout_seconds) + self.client = client + self.headers = { + "Content-Type": "application/json", + "Accept": "application/json, text/event-stream", + } + self._message_ids = itertools.count(10) + self._initialized = False + self._session_id: str | None = None + + def close(self) -> None: + close = getattr(self.client, "close", None) + if callable(close): + close() + + def _post(self, body: Mapping[str, Any], expected_id: int | None) -> Mapping[str, Any] | None: + try: + response = self.client.post(LOCALDOCS_URL, json=dict(body), headers=dict(self.headers)) + response.raise_for_status() + except Exception as exc: + raise IngressError("MCP_TRANSPORT_ERROR", "localdocs transport failed") from exc + session_id = response.headers.get("mcp-session-id") + if session_id: + if not isinstance(session_id, str) or not session_id.strip(): + raise IngressError("MCP_SESSION_ID_INVALID", "localdocs returned an invalid session ID") + normalized_session_id = session_id.strip() + if self._session_id is None: + if expected_id != 1: + raise IngressError( + "MCP_SESSION_ID_OUTSIDE_INITIALIZE", + "localdocs first bound a session outside initialize", + ) + self._session_id = normalized_session_id + elif normalized_session_id != self._session_id: + raise IngressError( + "MCP_SESSION_ID_CHANGED", + "localdocs changed the initialized session ID", + ) + self.headers["mcp-session-id"] = self._session_id + if expected_id is None: + return None + raw = response.content if isinstance(response.content, bytes) else bytes(response.content) + return _inline_parse_mcp_payload(raw, expected_id) + + def initialize(self) -> None: + response = self._post( + { + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": MCP_PROTOCOL_VERSION, + "capabilities": {}, + "clientInfo": { + "name": INLINE_CLIENT_NAME, + "version": INLINE_CLIENT_VERSION, + "user_id": self.user_hash, + "workspace_id": self.workspace_hash, + }, + }, + }, + 1, + ) + if response is None: + raise IngressError("MCP_INITIALIZE_EMPTY", "localdocs initialize returned no result") + result = response.get("result") + if not isinstance(result, dict) or result.get("protocolVersion") != MCP_PROTOCOL_VERSION: + raise IngressError( + "MCP_PROTOCOL_VERSION_MISMATCH", + "localdocs did not negotiate the requested MCP protocol version", + ) + if self._session_id is None or "mcp-session-id" not in self.headers: + raise IngressError("MCP_SESSION_ID_MISSING", "localdocs initialize did not bind a session ID") + self._post( + {"jsonrpc": "2.0", "method": "notifications/initialized"}, + None, + ) + self._initialized = True + + def call(self, tool_name: str, arguments: Mapping[str, Any]) -> Mapping[str, Any]: + if not self._initialized: + raise IngressError("MCP_NOT_INITIALIZED", "localdocs session is not initialized") + message_id = next(self._message_ids) + response = self._post( + { + "jsonrpc": "2.0", + "id": message_id, + "method": "tools/call", + "params": {"name": tool_name, "arguments": dict(arguments)}, + }, + message_id, + ) + if response is None: + raise IngressError("MCP_TOOL_EMPTY", f"localdocs {tool_name} returned no result") + return response["result"] + + def read_binary(self, logical_path: str) -> bytes: + path = _inline_relative_path(logical_path, code="LOCALDOCS_READ_PATH_INVALID") + result = self.call("read_binary_doc", {"doc_name": path}) + return _inline_binary_envelope(_inline_tool_text(result, "read_binary_doc"), path) + + def read_binary_optional(self, logical_path: str) -> bytes | None: + try: + return self.read_binary(logical_path) + except IngressError as exc: + if exc.code == "LOCALDOCS_NOT_FOUND": + return None + raise + + def write_binary_verified(self, logical_path: str, payload: bytes, *, overwrite: bool = False) -> str: + path = _inline_relative_path(logical_path, code="LOCALDOCS_WRITE_PATH_INVALID") + encoded = base64.b64encode(payload).decode("ascii") + result = self.call( + "write_binary_file", + {"path": path, "content_base64": encoded, "overwrite": overwrite}, + ) + _inline_tool_text(result, "write_binary_file") + observed = self.read_binary(path) + if observed != payload: + raise IngressError("LOCALDOCS_WRITE_READBACK_MISMATCH", f"read-back mismatch: {path}") + return hashlib.sha256(observed).hexdigest() + + + SOURCE_POLICY = load_json_strict(r'''{"stage1_sources":[{"adapter_id":"S2A-EVIDENCE-V3-ENVELOPE-V1","logical_input_id":"evidence_indexed","path":"evidence_indexed.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B1_quality_gate_evidence_indexed","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_contract_version","items"],"requirement_class":"EVIDENCE_EVENT_SCOPE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-EVENTS-V1-ENVELOPE-V1","logical_input_id":"evidence_event_candidates","path":"evidence_event_candidates.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B2_quality_gate_event_candidates","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version","items"],"requirement_class":"EVIDENCE_EVENT_SCOPE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-CLIENT-GOAL-V8-V1","logical_input_id":"client_goal","path":"client_goal.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_A_client_goal","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["primary_goal","constraints","parties"],"requirement_class":"OPTIMIZATION_CONTEXT","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-DOMAIN-SCREENING-V1","logical_input_id":"domain_screening","path":"routing/domain_screening.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_A0_domain_screener_02","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["domain_screening"],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-DUAL-SG01-V1","logical_input_id":"domain_activation_manifest","path":"routing/domain_activation_manifest.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_D0_domain_activation_gate","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["domain_activation_manifest"],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/s5/domain_activation_manifest.schema.json","path":"signals/schemas/domain_activation_manifest.schema.json","sha256":"013a6ebd230ebe46dda665af9f6c4448b267444b44e7b8f701f2fae80a2ee92a"},"transaction_identity_pointer":null},{"adapter_id":"S2A-B1-GATE-V1","logical_input_id":"b1_evidence_indexed_gate","path":"quality_gates/B1_evidence_indexed_gate.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B12_gate_audit_finalizer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_contract_version","gate_id","overall_severity","hard_gate_findings","review_findings","stage2_auto_progression_allowed"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-B2-GATE-V1","logical_input_id":"b2_event_candidates_gate","path":"quality_gates/B2_event_candidates_gate.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B12_gate_audit_finalizer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_contract_version","gate_id","overall_severity","hard_gate_findings","review_findings","stage2_auto_progression_allowed"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-P1-HANDOFF-FLAT-V1","logical_input_id":"stage1_part1_soft_gate_handoff","path":"quality_gates/stage1_part1_soft_gate_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_B2_SHA256_soft_gate_handoff_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version","handoff_status","review_items","stage2_auto_progression_allowed","hard_gate_summary","review_item_conservation","digest_guard"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-BO-V8-LIST-V1","logical_input_id":"bo","path":"BO.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_BO_F0_final_bo_compiler_gate_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":[],"requirement_class":"IDENTITY_BACKBONE","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-SIGNAL-ALL-V1","logical_input_id":"signal_manifest","path":"signals/signal_manifest.json","path_rule":null,"producer_alias_id":"PA-SG-COMPILER-001","producer_id":"Task_C_BO_S0_signal_bundle_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["files","downstream_read_sets"],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/s5/signal_manifest.schema.json","path":"signals/schemas/signal_manifest.schema.json","sha256":"5e72084780b82b29582c9ffcf48f3e4894d7c0b152e5ce8df394583c07dde681"},"transaction_identity_pointer":"/transaction_id"},{"adapter_id":"S2A-P2-HANDOFF-FLAT-V1","logical_input_id":"stage1_part2_review_handoff","path":"quality_gates/stage1_part2_review_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_BO_F0_final_bo_compiler_gate_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version","status","review_items"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-LES-CURRENT-V8-V1","logical_input_id":"legal_effect_structures","path":"legal_effect_structures.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_LE_L2_final_structure_index_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":[],"requirement_class":"ROUTING_PROFILE_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/part3/legal_effect_structures.schema.json","path":"platform/schemas/legal_effect_structures.schema.json","sha256":"fc962e8ae39f9bede64ba017297eded6413689204a065e00c3b3bdca8f1854df"},"transaction_identity_pointer":"/signal_manifest_transaction_id"},{"adapter_id":"S2A-P3-HANDOFF-WRAPPED-V1","logical_input_id":"stage1_part3_review_handoff","path":"quality_gates/stage1_part3_review_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_LE_L2_final_structure_index_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["stage1_part3_review_handoff"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-FACT-LEDGER-CURRENT-V8-V1","logical_input_id":"fact_ledger_base","path":"Fact_Ledger_base.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_FL_F2_final_fact_ledger_gate_and_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":[],"requirement_class":"IDENTITY_BACKBONE","run_identity_pointer":null,"schema_ref":{"$id":"https://schemas.liti-agent.local/stage1/part4/fact_ledger_base.schema.json","path":"platform/schemas/fact_ledger_base.schema.json","sha256":"b3f0e79ecb4c2f720f3e07e89154aadbd2327e4129cc703569fb5635240d2fe8"},"transaction_identity_pointer":null},{"adapter_id":"S2A-FACT-LEDGER-WRITER-REPORT-V1","logical_input_id":"fact_ledger_writer_report","path":"stage1_tmp/fact_ledger/fact_ledger_writer_report.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_FL_F2_final_fact_ledger_gate_and_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["schema_version"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-P4-HANDOFF-WRAPPED-V1","logical_input_id":"stage1_part4_review_handoff","path":"quality_gates/stage1_part4_review_handoff.json","path_rule":null,"producer_alias_id":null,"producer_id":"Task_C_FL_F2_final_fact_ledger_gate_and_writer","raw_hash_source":"UNAVAILABLE_DEV","required_keys":["stage1_part4_review_handoff"],"requirement_class":"INTEGRITY_CORROBORATOR","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null},{"adapter_id":"S2A-SIGNAL-ALL-V1","logical_input_id":"signal_payload_family","path":null,"path_rule":"signals/","producer_alias_id":"PA-SG-COMPILER-001","producer_id":"Task_C_BO_S0_signal_bundle_writer","raw_hash_source":"MANIFEST_ROW","required_keys":[],"requirement_class":"SIGNAL_PAYLOAD","run_identity_pointer":null,"schema_ref":null,"transaction_identity_pointer":null}],"dependency_locks":{"stage1":{"closure_scope":"REFERENCED_55_ONLY_NOT_FULL_STAGE1_RUNTIME_RELEASE","closure_snapshot_date":"2026-08-29","concrete_paths":[{"binding_status":"BOUND","lock_id":"S1-DEPLOY-001","path":"runtime_manifest.json","schema_id":"stage1_runtime_manifest.v1","sha256":"8964593a64a9b1bc90122054bb09eb3911827a06ed62dab0d6b7c745e7e18f54","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-002","path":"domains/_registry_index.json","schema_id":null,"sha256":"9f177ebf8860e20e05483967a2037f3baa09c2ac92c69ddeb260c04ca31ebf39","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-003","path":"signals/signal_registry.v2.json","schema_id":"signal_registry.v2","sha256":"4392b40da458102f8dd11b40b40ae3f694b7b5911849b050e2e4118c569e5ab0","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-004","path":"domains/E-00/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"5919f7ea1d7be02666b0c48aa6a66445e6d454fc2fbb21d7fe5154b0a1e68f6f","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-005","path":"domains/E-01/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"b557e92cd1b093bf31792dbcf5b62cab8ad064a65c4421e79f141e06b4cc2192","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-006","path":"domains/E-02/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"be407c980c28226a15406f85b5861b04a4e19a13870513ac6626349fc05ac434","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-007","path":"domains/E-03/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"7d3814f9b50cd5b33ef65a4eb778693552b3685bd369e765e9ac032734ebe23e","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-008","path":"domains/E-04/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"95a8c600cdce5a687f766788af0f763ee1b6a895e6ed80934afd28fe9a107e25","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-009","path":"domains/E-05/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"e50541018f47aa27de2f8b13ec3fa52210cf8456a6feed8356af78c1f1da144a","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-010","path":"domains/E-06/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"1e2bee36cb3c24dd37fc3beb3cf70236d531126c4f62ee97b5b42e55f4b0745c","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-011","path":"domains/E-07/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"22ac562084b1ce231b7257d099c18b6a4619defa2fc42504d590e0bcc494c5f8","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-012","path":"domains/E-08/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"4d545306d42120e8552dd953d4336ef6de828ea827779944ba73acfda3d3a8bb","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-013","path":"domains/E-09/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"7ef7340750094efeb372c397eb3134e21d62dda6988b1f6fa0a197b9963868e0","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-014","path":"domains/E-10/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"e8d4f45fa76ea9e09333256dd4ea36cd3dd963bf60c04814a2cb8dc90d152f0a","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-015","path":"domains/E-11/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"eb78d0188a0a2400307b1c34c8f8703c54cd86dd06b942c1709c44a8630a68e1","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-016","path":"domains/E-12/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"f336accdc6de10cdcc28c1190328054bca402fb77a2a9859d59fbaf5e84dd170","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-017","path":"domains/E-13/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"e27e2e3855b5868a3ec12c7093b872434702f2e465b73c0bfc948b516aa0fc35","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-018","path":"domains/E-14/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"a273148cc17f07d90cda500fa5cb7df30c9cd253f7b86048cf4f63495d36a156","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-019","path":"domains/E-15/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"7f37edddc101a08ed8a0e25f3a2c638e72571edc212ac91d96c1e33c51202a69","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-020","path":"domains/E-16/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"ae9af46ee31b6ef0dafedd35ccd7959a941d67d1a3dcc70e0b13647896873323","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-021","path":"domains/E-17/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"5c61f4486bdc968e4b30734b3c045404f0a711f47ea3abbe6c5c64652fb7f68c","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-022","path":"domains/E-18/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"74ff76148d175929bdeeeced77e9a9922d29b51ad3d00ae6711c43c55717692c","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-023","path":"domains/E-19/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"a8578f54a3fead3bbd35c62d7199b0f8aafb77f5d409a87236a55d2550fbfd37","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-024","path":"domains/E-20/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"29ee14cfe7789f004e6b6978d5360cf1bebe33bf88715df0cb47257993a11d00","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-025","path":"domains/E-21/domain_config.json","schema_id":"stage1_domain_config.v2","sha256":"4e1684a843d9e0c5af82f45332ad85abe94eda0c3ff0d578235d892aee39b908","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-026","path":"domains/EC-00/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"fe74de112b73289485dcead7e0fc7d270c794b3cf8a29ee00fab1eb64ba13861","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-027","path":"domains/X1/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"ad2fee7d206018f9a1f66e5fdf40dd67b686f6938099bad1ffc5d538db14ac57","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-028","path":"domains/X2/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"eba7d4671546bd66f1350d144ae0884f8beffb0dffdc147b45b9d5292676d46f","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-029","path":"domains/X3/domain_config.json","schema_id":"stage1_domain_config.v1","sha256":"8b67a638ae4a86aca3a2216974242b11ec39790162c9f366edfa91b02c3d270a","source_manifest":"domains/_registry_index.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-030","path":"platform/schemas/client_goal_domain_profiles.schema.json","schema_id":null,"sha256":"ae2bfe0d754a09cbae16b2c15bf1518fc23f9e1bda8fa1f5f949606c8e42c010","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-031","path":"platform/schemas/domain_fanout_plan.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s1/domain_fanout_plan.schema.json","sha256":"3b0948613a5996028b9c030a99f0b51d682f6035e019557756b1a15d43971113","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-032","path":"platform/schemas/domain_seed_output.schema.v3.json","schema_id":"https://schemas.liti-agent.local/stage1/s1/domain_seed_output.schema.v3.json","sha256":"992acf05dbccb34c65ead4e8c592f424e3b91672dc109cbd1bfa76a0a71a13c9","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-033","path":"platform/schemas/domain_slice.schema.v2.json","schema_id":"https://schemas.liti-agent.local/stage1/s1/domain_slice.schema.v2.json","sha256":"212a405088e7cf7ba2c65528a1c716938c946df7fe3bae3256b613051ed31aa3","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-034","path":"platform/schemas/fact_exception_pack.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part4/fact_exception_pack.schema.json","sha256":"4eba7e51ed46a99e3704bc2333169749f4a16935de26a8c8027c1cac98ea58cf","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-035","path":"platform/schemas/fact_ledger_base.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part4/fact_ledger_base.schema.json","sha256":"b3f0e79ecb4c2f720f3e07e89154aadbd2327e4129cc703569fb5635240d2fe8","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-036","path":"platform/schemas/fact_ledger_candidate_bundle.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part4/fact_ledger_candidate_bundle.schema.json","sha256":"4e481504fb795b2be510680a8fa88124f5763a8124462f7870be4125ed9a7730","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-037","path":"platform/schemas/legal_effect_structures.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part3/legal_effect_structures.schema.json","sha256":"fc962e8ae39f9bede64ba017297eded6413689204a065e00c3b3bdca8f1854df","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-038","path":"platform/schemas/structure_seed_bundle.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/part3/structure_seed_bundle.schema.json","sha256":"b7af9e422b6ac3876cffea39ec4f617eea76a631a57dfdfcd57d3785a83c667a","source_manifest":"runtime_manifest.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-039","path":"signals/_common/evidence_slot_status.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/evidence_slot_status.schema.json","sha256":"292b03960b187cef668b8635a8d7539fde7c31f0c20d01af52c4f6ff8519d7b1","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-040","path":"signals/_common/signal_item.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/signal_item.schema.json","sha256":"de8695f98041c06cf50c0d8d2ebc31e7b3c518ca9d39a27da940438704c58bb1","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-041","path":"signals/schemas/domain_activation_manifest.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/domain_activation_manifest.schema.json","sha256":"013a6ebd230ebe46dda665af9f6c4448b267444b44e7b8f701f2fae80a2ee92a","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-042","path":"signals/schemas/procedural_posture_relief_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/procedural_posture_relief_signals.schema.json","sha256":"fefb4317ad63088919b61777c71fe75ee6aa507b9f599dcf455d2af63dfc5e0d","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-043","path":"signals/schemas/party_capacity_standing_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/party_capacity_standing_signals.schema.json","sha256":"66de89ac53964166f6caabd50cbc03eb82dede0acf702d5e6d825c1d82ef81d0","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-044","path":"signals/schemas/governing_law_version_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/governing_law_version_signals.schema.json","sha256":"13a3f62f03356090d2cb24de2da0ba217928dfe8eb3c111d0f5e87c7df3119ee","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-045","path":"signals/schemas/legal_relation_lifecycle_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/legal_relation_lifecycle_signals.schema.json","sha256":"420613a5900c4360487b89b978efedde58f5ddc61644130e4b9e63ef8ab33d8b","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-046","path":"signals/schemas/timeline_notice_condition_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/timeline_notice_condition_signals.schema.json","sha256":"99c66208524155cea6bbd5e24fd26998cc9b653c89b24b569c793e36f1623d35","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-047","path":"signals/schemas/asset_right_state_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/asset_right_state_signals.schema.json","sha256":"fc34fbb3d33a284c3d57f3c278cbda8b3555ef26ee2f06b803fd2410ebce38b6","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-048","path":"signals/schemas/liability_causation_damage_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/liability_causation_damage_signals.schema.json","sha256":"34102cb8eeda80773eb62a5ee61e3d714bf90424ed5350dcac4b7bf873a72c5a","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-049","path":"signals/schemas/defense_exception_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/defense_exception_signals.schema.json","sha256":"010148c15e60e4d112b142f80b1723c06e34ba22b3edefae9e4371f2353b053e","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-050","path":"signals/schemas/evidence_proof_conflict_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/evidence_proof_conflict_signals.schema.json","sha256":"c419f568e28c06c629bc715aff7b0737b77e9c4871c91d4fae8f6ecf04196390","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-051","path":"signals/schemas/calculation_requirements.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/calculation_requirements.schema.json","sha256":"7fdb5ef0f50d7af22ac417abc4022cd238f5ab0dc942866420616729a9e3571f","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-052","path":"signals/schemas/remedy_enforcement_signals.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/remedy_enforcement_signals.schema.json","sha256":"999e1969b983748f209e9b5239f7edd0ec43bc642d9ea8fd7edbf34f9ce653f3","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-053","path":"signals/schemas/legal_effect_routes.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/legal_effect_routes.schema.json","sha256":"c24cb740c370aa2477199a0225be8291164ef5c787601fd962a370c642cc3cc0","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-054","path":"signals/schemas/domain_signal_envelope.schema.v2.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/domain_signal_envelope.schema.v2.json","sha256":"1483d6c5f98083f59172feff9b7c15b44d3ed789db5b6172d0de05f06e9d3fbc","source_manifest":"signals/signal_registry.v2.json"},{"binding_status":"BOUND","lock_id":"S1-DEPLOY-055","path":"signals/schemas/signal_manifest.schema.json","schema_id":"https://schemas.liti-agent.local/stage1/s5/signal_manifest.schema.json","sha256":"5e72084780b82b29582c9ffcf48f3e4894d7c0b152e5ce8df394583c07dde681","source_manifest":"signals/signal_registry.v2.json"}],"contract_manifest_ref":{"mode":"CONDITIONAL_RELOCATION_ONLY","path":null,"sha256":null,"status":"NOT_REQUIRED_DEFAULT_PATHS"},"expected_concrete_path_count":55,"full_stage1_runtime_release_status":"STAGE1_NOT_RELEASE_READY"}},"adapter_decisions":[{"adapter_id":"S2A-SIGNAL-ALL-V1","decision":{"file_conservation_equation":"semantic_file_rows + integrity_only_file_rows = Counter(signal_manifest.files[])","global_signal_id_uniqueness_assumed":false,"integrity_only_kinds":["compatibility_view"],"manifest_selector":"/downstream_read_sets/stage2","physical_path_rule":"U/signals/","record_conservation_equation":"used_record_occurrences + unused_record_occurrences + unmapped_record_occurrences = records_from_semantic_files","record_occurrence_key":["manifest_transaction_id","file_path","record_ordinal","signal_id"],"row_order":"PRESERVE_MANIFEST_ORDER","row_source":"/files","semantic_kinds":["canonical","domain_signal"],"sentinel":["ALL"]}},{"adapter_id":"S2A-DUAL-SG01-V1","decision":{"comparison":"PARSED_CANONICAL_PROJECTION_EQUAL","payload_root":"/domain_activation_manifest","projection_json_pointers":["/schema_version","/signal_id","/status","/registry_version","/registry_index_sha256","/screening_sha256","/domain_entries","/active_domain_ids","/supporting_domain_ids","/monitor_domain_ids","/expected_runnable_domain_ids","/required_calculation_domains","/unrouted_material","/conservation_gate","/fail_open_policy","/review_items","/contract_guards"],"raw_hash_policy":"PRESERVE_AND_VERIFY_SEPARATELY","routing_path":"routing/domain_activation_manifest.json","set_semantics_json_pointers":["/active_domain_ids","/supporting_domain_ids","/monitor_domain_ids","/expected_runnable_domain_ids","/required_calculation_domains"],"signal_path":"signals/domain_activation_manifest.json"}},{"adapter_id":"S2A-P1-HANDOFF-FLAT-V1","decision":{"count_field_required":false,"logical_input_id":"P1_REVIEW_HANDOFF","p1_digest_keys":["evidence_indexed_sha256","evidence_event_candidates_sha256","b1_gate_sha256","b2_gate_sha256","screening_sha256","activation_manifest_sha256","registry_index_sha256"],"review_items_json_pointer":"/review_items","schema_version":"stage1_part1_soft_gate_handoff.v1","seal_sources":["routing/domain_screening.json","routing/domain_activation_manifest.json","domains/_registry_index.json"],"source_stage":"P1","status_json_pointer":"/handoff_status","wrapper_json_pointer":""}},{"adapter_id":"S2A-P2-HANDOFF-FLAT-V1","decision":{"count_field_required":false,"logical_input_id":"P2_REVIEW_HANDOFF","review_items_json_pointer":"/review_items","schema_version":"stage1_part2_review_handoff.v1","seal_sources":["BO.json","signals/signal_manifest.json"],"source_stage":"P2","status_json_pointer":"/status","wrapper_json_pointer":""}},{"adapter_id":"S2A-P3-HANDOFF-WRAPPED-V1","decision":{"count_field_required":true,"logical_input_id":"P3_REVIEW_HANDOFF","review_items_json_pointer":"/review_items","schema_version":"stage1_part3_review_handoff.v1","seal_sources":["legal_effect_structures.json","validation_assets/routing/part3_receipt.json"],"source_stage":"P3","status_json_pointer":"/status","wrapper_json_pointer":"/stage1_part3_review_handoff"}},{"adapter_id":"S2A-P4-HANDOFF-WRAPPED-V1","decision":{"count_field_required":true,"logical_input_id":"P4_REVIEW_HANDOFF","review_items_json_pointer":"/review_items","schema_version":"stage1_part4_review_handoff.v1","seal_sources":["Fact_Ledger_base.json","validation_assets/routing/part4_receipt.json","stage1_tmp/fact_ledger/fact_ledger_writer_report.json"],"source_stage":"P4","status_json_pointer":"/status","wrapper_json_pointer":"/stage1_part4_review_handoff"}},{"adapter_id":"S2-REVIEW-MAP-V1","decision":{"aggregate_handoff_status_never_resolves_item":true,"handoff_status_mappings":[{"source_stage":"P1","source_value":"READY_NO_REVIEW","technical_disposition":"AVAILABLE"},{"source_stage":"P1","source_value":"READY_WITH_REVIEW","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P1","source_value":"BLOCKED","technical_disposition":"UNAVAILABLE"},{"source_stage":"P2","source_value":"PENDING_FINALIZE","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P2","source_value":"FINALIZED","technical_disposition":"AVAILABLE"},{"source_stage":"P3","source_value":"OPEN","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P3","source_value":"FINALIZED","technical_disposition":"AVAILABLE"},{"source_stage":"P4","source_value":"OPEN","technical_disposition":"AVAILABLE_WITH_ISSUES"},{"source_stage":"P4","source_value":"FINALIZED","technical_disposition":"AVAILABLE"}],"mappings":[{"mapping_id":"S2RM-001","normalized_partition":"SUPPORTED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"SUPPORTED","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-002","normalized_partition":"CONDITIONAL","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"CONDITIONAL","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-003","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"UNRESOLVED","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-004","normalized_partition":"EXCLUDED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"EXCLUDED","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-005","normalized_partition":"SUPPORTED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"observed","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-006","normalized_partition":"CONDITIONAL","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"inferred","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-007","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"contested","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-008","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"missing_required","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-009","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"review","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-010","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_STATUS","source_stage":"ANY","source_value":"NO_SUPPORT","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-011","normalized_partition":"CONDITIONAL","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"info","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-012","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"review","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-013","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"SOFT_WARNING","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-014","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"hard_warning","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-015","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"HARD_WARNING","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-016","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"block","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"},{"mapping_id":"S2RM-017","normalized_partition":"UNRESOLVED","resolution_inference_allowed":false,"source_field_kind":"REVIEW_ITEM_SEVERITY","source_stage":"ANY","source_value":"BLOCK","unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"}],"normalized_partitions":["SUPPORTED","CONDITIONAL","UNRESOLVED","EXCLUDED","UNMAPPED"],"resolution_inference_allowed":false,"unknown_value_policy":"MAP_TO_UNMAPPED_AND_ISSUE"}},{"adapter_id":"S2A-BO-V8-LIST-V1","decision":{"logical_input_id":"BO","open_source_fields_policy":"PRESERVE_UNMODELED_FIELDS_WITH_RAW_HASH","producer_id":"Task_C_BO_F0_final_bo_compiler_gate_writer","required_item_fields":["BO_ID","id","BOType","ActionType","JuristicAct","Action","Reason","PriorAct","ReasonRefs","Legal_Keywords","core_field_base","amount","EvidenceTitles","Evidence","source_evidence_indexes","provenance","downstream_seed_refs","extensions"],"required_root_fields":[],"root_shape":"ARRAY","schema_contract_version":null}},{"adapter_id":"S2A-EVIDENCE-V3-ENVELOPE-V1","decision":{"logical_input_id":"EVIDENCE_INDEXED","open_source_fields_policy":"PRESERVE_UNMODELED_FIELDS_WITH_RAW_HASH","producer_id":"Task_B1_quality_gate_evidence_indexed","required_root_fields":["schema_contract_version","items"],"root_shape":"OBJECT_ENVELOPE","schema_contract_version":"evidence_indexed.v3"}},{"adapter_id":"S2A-EVENTS-V1-ENVELOPE-V1","decision":{"logical_input_id":"EVIDENCE_EVENT_CANDIDATES","open_source_fields_policy":"PRESERVE_UNMODELED_FIELDS_WITH_RAW_HASH","producer_id":"Task_B2_quality_gate_event_candidates","required_root_fields":["schema_version","items"],"root_shape":"OBJECT_ENVELOPE","schema_contract_version":"evidence_event_candidates.v1"}},{"adapter_id":"S2A-DOMAIN-CONFIG-V1","decision":{"accepted_schema_version":"stage1_domain_config.v1","depends_on_legal_dependency_allowed":false,"rebuttal_slot_synthesis_allowed":false,"required_slot_fields":["element_slots","opposing_fact_slots","defense_map","calculation_bindings","emits_signals"],"undeclared_slot_policy":"PRESERVE_AS_PROPOSED_NEW_SLOT_ISSUE"}},{"adapter_id":"S2A-DOMAIN-CONFIG-V2","decision":{"accepted_schema_version":"stage1_domain_config.v2","depends_on_legal_dependency_allowed":false,"rebuttal_slot_synthesis_allowed":false,"required_slot_fields":["element_slots","opposing_fact_slots","defense_map","calculation_bindings","emits_signals"],"undeclared_slot_policy":"PRESERVE_AS_PROPOSED_NEW_SLOT_ISSUE"}},{"adapter_id":"S2A-FACT-LEDGER-CURRENT-V8-V1","decision":{"bo_source_bo_id_multiset_equality_required":true,"fact_id_pattern":"^F-[0-9]{3,}$","legacy_adapter_status":"DISABLED_NO_APPROVED_ADAPTER","producer_generation":"CURRENT_V8","required_row_fields":["fact_id","source_bo_id","domain_effects","calculation_requests"],"root_shape":"ARRAY"}},{"adapter_id":"PA-SG-COMPILER-001","decision":{"bidirectional_match_allowed":true,"global_alias_allowed":false,"orchestration_producer_id":"Task_C_BO_S0_signal_bundle_writer","schema_writer_id":"Task_C_BO_S0_canonical_signal_compiler","scope":"STAGE1_PART2_SIGNAL_TRANSACTION_ONLY"}}],"release_class":"DEV_FIXTURE_RELEASE","limits":{"max_file_bytes":33554432,"max_run_bytes":268435456,"max_json_depth":96,"max_json_items":1000000}}''') + + + RAW_STAGE1_RESULTS = { + 'evidence_indexed': r"""{{prev.evidence_indexed.json}}""", + 'evidence_event_candidates': r"""{{prev.evidence_event_candidates.json}}""", + 'client_goal': r"""{{prev.client_goal.json}}""", + 'domain_screening': r"""{{prev.routing/domain_screening.json}}""", + 'domain_activation_manifest': r"""{{prev.routing/domain_activation_manifest.json}}""", + 'b1_evidence_indexed_gate': r"""{{prev.quality_gates/B1_evidence_indexed_gate.json}}""", + 'b2_event_candidates_gate': r"""{{prev.quality_gates/B2_event_candidates_gate.json}}""", + 'stage1_part1_soft_gate_handoff': r"""{{prev.quality_gates/stage1_part1_soft_gate_handoff.json}}""", + 'bo': r"""{{prev.BO.json}}""", + 'signal_manifest': r"""{{prev.signals/signal_manifest.json}}""", + 'stage1_part2_review_handoff': r"""{{prev.quality_gates/stage1_part2_review_handoff.json}}""", + 'legal_effect_structures': r"""{{prev.legal_effect_structures.json}}""", + 'stage1_part3_review_handoff': r"""{{prev.quality_gates/stage1_part3_review_handoff.json}}""", + 'fact_ledger_base': r"""{{prev.Fact_Ledger_base.json}}""", + 'fact_ledger_writer_report': r"""{{prev.stage1_tmp/fact_ledger/fact_ledger_writer_report.json}}""", + 'stage1_part4_review_handoff': r"""{{prev.quality_gates/stage1_part4_review_handoff.json}}""", + } + + + STATUS_PATH = "ingress/ingress_status.json" + NORMAL_PATHS = frozenset({"ingress/stage1_input_manifest.json", "ingress/intake_report.json", "review/issue_ledger.base.json", "context/case_context.json", STATUS_PATH}) + BLOCKED_PATHS = frozenset({"ingress/stage1_input_manifest.json", "ingress/intake_report.json", "review/issue_ledger.base.json", "ingress/technical_diagnostic.json", STATUS_PATH}) + ROW_KEYS = { + "bo": ("business_objects", "BO", "rows", "items"), + "fact_ledger_base": ("facts", "fact_ledger", "rows", "items"), + "legal_effect_structures": ("structures", "structure_records", "legal_effect_structures", "rows", "items"), + "evidence_indexed": ("evidence", "evidence_items", "rows", "items"), + "evidence_event_candidates": ("events", "event_candidates", "rows", "items"), + } + WRAPPER_KEYS = ("payload", "data", "fact_ledger_base", "Fact_Ledger_base", "legal_effect_structures") + REVIEW_ARRAY_KEYS = frozenset({"review_items", "review_queue", "blocked_review_items", "unresolved_review_items", "review_findings", "hard_gate_findings"}) + + + def _pointer_token(value: str) -> str: + return value.replace("~", "~0").replace("/", "~1") + + + def _row_locations(document: Any, keys: Sequence[str], pointer: str = "") -> list[tuple[str, Any]]: + if isinstance(document, list): + return [(f"{pointer}/{i}", row) for i, row in enumerate(document)] + if not isinstance(document, dict): + raise IngressError("SOURCE_ROWS_SHAPE", "record source must be an array or approved envelope") + arrays = [(key, document[key]) for key in keys if isinstance(document.get(key), list)] + if len(arrays) > 1: + raise IngressError("SOURCE_ROWS_AMBIGUOUS", "multiple record arrays in one source envelope") + if arrays: + key, rows = arrays[0] + return [(f"{pointer}/{_pointer_token(key)}/{i}", row) for i, row in enumerate(rows)] + nested = [key for key in WRAPPER_KEYS if isinstance(document.get(key), dict)] + if len(nested) != 1: + raise IngressError("SOURCE_ROWS_SHAPE", "approved record array is missing or ambiguous") + key = nested[0] + return _row_locations(document[key], keys, f"{pointer}/{_pointer_token(key)}") + + + def _array_rows(document: Any, keys: Sequence[str]) -> list[Any]: + if document is None: + return [] + return [row for _, row in _row_locations(document, keys)] + + + def _json_value(raw: Any) -> Any: + if not isinstance(raw, (str, bytes)): + return raw + if isinstance(raw, str) and re.fullmatch(r"\s*\{\{[^{}]+\}\}\s*", raw): + raise IngressError("PREV_REFERENCE_UNRESOLVED", "required backend result reference was not resolved") + try: + return load_json_strict(raw) + except IngressError: + if isinstance(raw, str) and raw.strip() and not raw.lstrip().startswith(("{", "[", '"')): + return raw.strip() + raise + + + def validate_direct_roots(run_root: Any, deployment_root: Any) -> dict[str, str]: + result = { + "stage1_run_root_ref": _inline_relative_path(_json_value(run_root), code="STAGE1_RUN_ROOT_INVALID"), + "stage1_deployment_root_ref": _inline_relative_path(_json_value(deployment_root), code="STAGE1_DEPLOYMENT_ROOT_INVALID"), + } + output = f"stage2_runs/from-stage1/{result['stage1_run_root_ref']}/s2_00" + out = PurePosixPath(output) + for value in result.values(): + original = PurePosixPath(value) + if out == original or original in out.parents or out in original.parents: + raise IngressError("OUTPUT_SOURCE_OVERLAP", "output and source roots must be disjoint") + result["output_root"] = output + return result + + + def _previous_source(value: Any, expected_path: str) -> tuple[bytes | None, Any | None]: + """Return exact bytes when supplied; otherwise retain parsed value for comparison.""" + if isinstance(value, bytes): + load_json_strict(value) + return value, None + parsed = _json_value(value) + if isinstance(parsed, dict) and "content_base64" in parsed: + return _inline_binary_envelope(canonical_json_bytes(parsed).decode(), expected_path), None + if isinstance(parsed, dict) and set(parsed).issubset({"path", "content", "text", "name", "doc_name", "sha256", "byte_length"}): + supplied_path = parsed.get("path", parsed.get("doc_name", parsed.get("name"))) + if supplied_path is not None and supplied_path != expected_path: + raise IngressError("PREV_SOURCE_PATH_MISMATCH", "backend result names a different source file") + content = parsed.get("content", parsed.get("text")) + if isinstance(content, str): + raw = content.encode("utf-8") + load_json_strict(raw) + return raw, None + if content is not None: + return None, content + if supplied_path is not None: + return None, None + if isinstance(parsed, str): + if parsed != expected_path: + raise IngressError("PREV_SOURCE_PATH_MISMATCH", "backend result names a different source file") + return None, None + if isinstance(parsed, (dict, list)): + return None, parsed + raise IngressError("PREV_SOURCE_SHAPE", "backend result must provide source JSON, raw bytes, or its exact path") + + + def _copy_to_temp(root: Path, path: str, raw: bytes) -> None: + safe = _safe_relative_path(path) + target = root.joinpath(*safe.parts) + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(raw) + + + def _walk_values(value: Any, pointer: str = "") -> Iterable[tuple[str, Any]]: + yield pointer, value + if isinstance(value, dict): + for key, item in value.items(): + yield from _walk_values(item, f"{pointer}/{_pointer_token(key)}") + elif isinstance(value, list): + for index, item in enumerate(value): + yield from _walk_values(item, f"{pointer}/{index}") + + + def _schema_dependencies(document: Mapping[str, Any], current_path: str, locks: Mapping[str, Any]) -> set[str]: + dependencies = set() + for _, item in _walk_values(document): + if not isinstance(item, dict) or not isinstance(item.get("$ref"), str): + continue + ref = item["$ref"].split("#", 1)[0] + if not ref: + continue + candidates = [path for path, row in locks.items() if row.get("schema_id") == ref] + if not candidates and "://" not in ref: + relative = posixpath.normpath(posixpath.join(posixpath.dirname(current_path), ref)) + if relative in locks: + candidates = [relative] + elif ref in locks: + candidates = [ref] + if not candidates: + candidates = [path for path in locks if PurePosixPath(path).name == PurePosixPath(ref).name] + if len(candidates) != 1: + raise IngressError("SCHEMA_DEPENDENCY_UNBOUND", "schema reference is not uniquely bound to Stage 1 deployment") + dependencies.add(candidates[0]) + return dependencies + + + def hydrate_stage1(localdocs: _InlineLocaldocs, temp_root: Path, roots: Mapping[str, str], stage1_results: Mapping[str, Any], policy: Mapping[str, Any]) -> dict[str, Any]: + """Reuse prev results; fetch only missing raw bytes and needed upstream dependencies.""" + stage1_root = temp_root / "stage1" + deployment_root = temp_root / "deployment" + stage1_root.mkdir(); deployment_root.mkdir() + observed: dict[str, bytes] = {} + documents: dict[str, Any] = {} + source_snapshots: dict[str, Snapshot] = {} + issues = [] + total = 0 + def remember(path: str, raw: bytes) -> None: + nonlocal total + if path in observed: + if observed[path] != raw: + raise IngressError("SOURCE_PATH_CONTENT_CONFLICT", "one source path has conflicting results") + return + if len(raw) > MAX_FILE_BYTES: + raise IngressError("SOURCE_SIZE_LIMIT", "input exceeds per-file byte limit") + total += len(raw) + if total > MAX_RUN_BYTES: + raise IngressError("AGGREGATE_RUN_SIZE_LIMIT", "input set exceeds byte limit") + observed[path] = raw + for contract in DEFAULT_SOURCE_CONTRACTS: + logical = contract["logical_input_id"] + relative = contract["path"] + logical_path = f"{roots['stage1_run_root_ref']}/{relative}" + if logical not in stage1_results: + raise IngressError("PREV_SOURCE_MISSING", "required Stage 1 result reference is missing", logical_input_id=logical) + exact, parsed = _previous_source(stage1_results[logical], logical_path) + raw = exact if exact is not None else localdocs.read_binary_optional(logical_path) + if raw is None: + issues.append(_issue("SOURCE_MISSING", source_refs=[logical])) + continue + value = load_json_strict(raw) + if parsed is not None and not _json_equal(value, parsed): + raise IngressError("PREV_SOURCE_CONTENT_MISMATCH", "backend result differs from the original file", logical_input_id=logical) + remember(logical_path, raw) + _copy_to_temp(stage1_root, relative, raw) + documents[logical] = value + source_snapshots[logical] = open_bounded_snapshot(stage1_root, relative, logical_input_id=logical) + manifest = documents.get("signal_manifest") + if isinstance(manifest, dict): + files = manifest.get("files") + if not isinstance(files, list): + raise IngressError("SIGNAL_FILES_SHAPE", "signal manifest must contain its actual files array") + for index, row in enumerate(files): + if not isinstance(row, dict) or not isinstance(row.get("path"), str): + raise IngressError("SIGNAL_FILE_ROW_SHAPE", "signal manifest row is malformed") + relative = _safe_relative_path(row["path"]).as_posix() + if relative.startswith("signals/"): + raise IngressError("SIGNAL_PATH_PREFIX_FORBIDDEN", "signal row path must not repeat signals/") + relative = f"signals/{relative}" + path = f"{roots['stage1_run_root_ref']}/{relative}" + # A dynamic file is an existing Stage 1 path selected by its manifest. + # Provided raw results can be reused; no new result-list request is created. + provided = stage1_results.get(relative) + if provided is not None: + exact, parsed = _previous_source(provided, path) + else: + exact, parsed = None, None + raw = exact if exact is not None else localdocs.read_binary(path) + if parsed is not None and not _json_equal(load_json_strict(raw), parsed): + raise IngressError("PREV_SOURCE_CONTENT_MISMATCH", "dynamic result differs from its source") + remember(path, raw) + _copy_to_temp(stage1_root, relative, raw) + locks = {row["path"]: row for row in policy["dependency_locks"]["stage1"]["concrete_paths"]} + if len(locks) != len(policy["dependency_locks"]["stage1"]["concrete_paths"]): + raise IngressError("STAGE1_DEPENDENCY_DUPLICATE_PATH", "upstream dependency table contains duplicate paths") + deployment_snapshots: dict[str, Snapshot] = {} + deployment_documents: dict[str, Any] = {} + needed = {"domains/_registry_index.json", "signals/signal_registry.v2.json"} + needed.update(row["schema_ref"]["path"] for row in policy["stage1_sources"] if isinstance(row.get("schema_ref"), dict)) + activation = documents.get("domain_activation_manifest") + payload = _activation_payload(activation) if isinstance(activation, dict) else {} + for domain in payload.get("active_domain_ids", []): + needed.add(f"domains/{_safe_relative_path(str(domain)).as_posix()}/domain_config.json") + while needed: + relative = min(needed); needed.remove(relative) + if relative in deployment_documents: + continue + row = locks.get(relative) + if row is None: + raise IngressError("STAGE1_DEPENDENCY_UNBOUND", "required upstream dependency is not pinned") + expected = _inline_sha256(row.get("sha256"), code="STAGE1_DEPENDENCY_UNBOUND") + path = f"{roots['stage1_deployment_root_ref']}/{relative}" + raw = localdocs.read_binary(path) + if hashlib.sha256(raw).hexdigest() != expected: + raise IngressError("STAGE1_DEPENDENCY_HASH_MISMATCH", "upstream deployment file differs from its pin") + remember(path, raw) + value = load_json_strict(raw) + _copy_to_temp(deployment_root, relative, raw) + deployment_snapshots[relative] = open_bounded_snapshot(deployment_root, relative, logical_input_id=f"deployment:{relative}") + deployment_documents[relative] = value + if isinstance(value, dict): + needed.update(_schema_dependencies(value, relative, locks) - deployment_documents.keys()) + if relative == "signals/signal_registry.v2.json" and isinstance(value, dict): + for entry in value.get("entries", []): + if isinstance(entry, dict) and isinstance(entry.get("schema"), str): + schema = entry["schema"] + needed.add(schema if schema.startswith("signals/") else f"signals/{schema}") + envelope = value.get("domain_envelope") + if isinstance(envelope, str): + needed.add(envelope if envelope.startswith("signals/") else f"signals/{envelope}") + return {"stage1_root": stage1_root, "deployment_root": deployment_root, "snapshots": source_snapshots, "documents": documents, "deployment_snapshots": deployment_snapshots, "deployment_documents": deployment_documents, "observed": observed, "issues": issues} + + + def verify_remote_stability(localdocs: _InlineLocaldocs, observed: Mapping[str, bytes]) -> None: + for path, expected in sorted(observed.items()): + if localdocs.read_binary(path) != expected: + raise IngressError("HYDRATION_SOURCE_CHANGED", "source differs from the first read/result reference") + + + def _provenance(logical: str, pointer: str, documents: Mapping[str, Any]) -> dict[str, Any]: + found, value = _json_pointer_value(documents[logical], pointer) + if not found: + raise IngressError("SOURCE_POINTER_INVALID", "projection pointer does not address the original") + return _source_ref(logical, pointer, value) + + + def normalize_review_items(review_documents: Mapping[str, Any], release_lock: Mapping[str, Any] | None = None) -> dict[str, Any]: + """Preserve every review/gate occurrence, its content, exact pointer, and blocking state.""" + mapping = _adapter_decision(release_lock or {}, "S2-REVIEW-MAP-V1") or {} + table = {(row.get("source_stage", "ANY"), row.get("source_field_kind"), str(row.get("source_value"))): row.get("normalized_partition") for row in mapping.get("mappings", [])} + rows = [] + partitions = Counter() + adapter_issues = [] + for stage_number in range(1, 5): + logical = 'stage1_part1_soft_gate_handoff' if stage_number == 1 else f'stage1_part{stage_number}_review_handoff' + if logical not in review_documents: + continue + adapter = f'S2A-P{stage_number}-HANDOFF-' + ('FLAT-V1' if stage_number < 3 else 'WRAPPED-V1') + decision = _adapter_decision(release_lock or {}, adapter) + if not isinstance(decision, dict): + adapter_issues.append(_issue('HANDOFF_ADAPTER_CONTRACT_MISSING', source_refs=[logical])) + continue + found, wrapper = _json_pointer_value(review_documents[logical], decision.get('wrapper_json_pointer')) + if not found or not isinstance(wrapper, dict): + adapter_issues.append(_issue(f'P{stage_number}_WRAPPER_MISSING', source_refs=[logical])) + continue + if wrapper.get('schema_version') != decision.get('schema_version'): + adapter_issues.append(_issue(f'P{stage_number}_HANDOFF_SCHEMA_VERSION_MISMATCH', source_refs=[logical])) + found, handoff_items = _json_pointer_value(wrapper, decision.get('review_items_json_pointer')) + if not found or not isinstance(handoff_items, list): + adapter_issues.append(_issue(f'P{stage_number}_REVIEW_ITEMS_SHAPE', source_refs=[logical])) + elif decision.get('count_field_required') is True and wrapper.get('review_item_count') != len(handoff_items): + adapter_issues.append(_issue('REVIEW_CONSERVATION_FAILED', source_refs=[logical])) + for logical, document in sorted(review_documents.items()): + stage_match = re.search(r"part([1-4])", logical) + stage = f"P{stage_match.group(1)}" if stage_match else "ANY" + for pointer, value in _walk_values(document): + if not isinstance(value, dict): + continue + for key in sorted(REVIEW_ARRAY_KEYS): + items = value.get(key) + if not isinstance(items, list): + continue + for index, item in enumerate(items): + item_pointer = f"{pointer}/{_pointer_token(key)}/{index}" + raw_status = item.get("status") if isinstance(item, dict) else None + raw_severity = item.get("severity") if isinstance(item, dict) else None + kind = "REVIEW_ITEM_STATUS" if raw_status is not None else "REVIEW_ITEM_SEVERITY" + raw_value = str(raw_status if raw_status is not None else raw_severity) + partition = table.get((stage, kind, raw_value), table.get(("ANY", kind, raw_value), "UNMAPPED")) + explicit_block = key == "blocked_review_items" or isinstance(item, dict) and (item.get("blocking") is True or item.get("blocked") is True or str(item.get("status", "")).upper() == "BLOCKED" or str(item.get("severity", "")).upper() in {"BLOCKING", "CRITICAL", "FATAL"}) + row = {"review_ref": f"{logical}#{item_pointer}", "source_ref": _provenance(logical, item_pointer, review_documents), "source_status_raw": raw_status, "source_severity_raw": raw_severity, "partition": partition, "blocking": bool(explicit_block), "content": item} + rows.append(row); partitions[partition] += 1 + # Count source occurrences independently; duplicates remain distinct by pointer. + expected = sum(len(v[k]) for doc in review_documents.values() for _, v in _walk_values(doc) if isinstance(v, dict) for k in REVIEW_ARRAY_KEYS if isinstance(v.get(k), list)) + return {"normalized_occurrences": rows, "partition_counts": dict(partitions), "conservation_status": "PASS" if expected == len(rows) and len({r['review_ref'] for r in rows}) == expected and not any(x["issue_code"] == "REVIEW_CONSERVATION_FAILED" for x in adapter_issues) else "FAIL", "_issues": adapter_issues} + + + def _project_content(value: Any) -> Any: + if not isinstance(value, dict): + return value + # Envelope/protocol metadata remains reachable through provenance instead of copying files. + return {key: item for key, item in value.items() if key not in {"schema_version", "schema_contract_version", "producer_id", "created_by", "finalized_by", "metadata", "meta"}} + + + def compile_case_context(documents: Mapping[str, Any], signal_all: Mapping[str, Any], reviews: Mapping[str, Any], deployment_documents: Mapping[str, Any]) -> dict[str, Any]: + """Normalize original records once and group only explicit source relationships.""" + members = [] + lookup = {} + identities = {"bo": ("BO", ("BO_ID",)), "fact_ledger_base": ("FACT", ("fact_id",)), "legal_effect_structures": ("LES", ("structure_id", "legal_effect_structure_id")), "evidence_indexed": ("EVIDENCE", ("evidence_id", "id")), "evidence_event_candidates": ("EVENT", ("event_id", "id"))} + raw_rows = {} + for logical, keys in ROW_KEYS.items(): + for pointer, value in _row_locations(documents[logical], keys): + if not isinstance(value, dict): + raise IngressError("SOURCE_RECORD_SHAPE", "original record must be an object") + kind, id_keys = identities[logical] + identifier = next((str(value[k]) for k in id_keys if value.get(k) is not None), None) + ref = f"{logical}#{pointer}" + if identifier is not None: + if (kind, identifier) in lookup: + raise IngressError("SOURCE_RECORD_ID_DUPLICATE", "original record ID occurs more than once") + lookup[(kind, identifier)] = ref + member = {"member_ref": ref, "kind": kind, "stage1_id": identifier, "source_ref": _provenance(logical, pointer, documents), "field_refs": {key: _provenance(logical, f"{pointer}/{_pointer_token(key)}", documents) for key in value}, "projection": _project_content(value)} + members.append(member); raw_rows[ref] = (logical, pointer, value) + relationships = []; candidates = []; unresolved = [] + parent = {m['member_ref']: m['member_ref'] for m in members} + def find(ref): + while parent[ref] != ref: + parent[ref] = parent[parent[ref]]; ref = parent[ref] + return ref + def join(a,b): + a,b=find(a),find(b) + if a!=b:parent[max(a,b)]=min(a,b) + def edge(source, kind, identifier, relation, pointer, *, hard=True): + logical, _, _ = raw_rows[source] + target = lookup.get((kind, str(identifier))) + row = {"from_ref": source, "to_ref": target, "target_stage1_id": str(identifier), "relation_kind": relation, "source_ref": _provenance(logical, pointer, documents), "hard_join_allowed": hard, "disposition": "OBSERVED" if target else "UNEVALUABLE"} + if target is None: + unresolved.append(row) + elif hard: + relationships.append(row); join(source,target) + else: + candidates.append(row) + for member in members: + ref=member['member_ref']; logical,pointer,row=raw_rows[ref] + if member['kind']=='FACT': + if row.get('source_bo_id') is not None:edge(ref,'BO',row['source_bo_id'],'SAME_BO_ID',f"{pointer}/source_bo_id") + for keys,kind,relation in [(('evidence_refs','evidence_ids'),'EVIDENCE','SAME_EVIDENCE_REF'),(('event_refs','event_ids'),'EVENT','SAME_EVENT_REF')]: + key=next((k for k in keys if isinstance(row.get(k),list)),None) + if key: + for index,identifier in enumerate(row[key]):edge(ref,kind,identifier,relation,f"{pointer}/{key}/{index}") + for key in ('relations','explicit_relations','candidate_relations'): + for index,item in enumerate(row.get(key,[]) if isinstance(row.get(key),list) else []): + if not isinstance(item,dict):continue + target=item.get('target_fact_id',item.get('to_fact_id')) + relation=str(item.get('relation_kind',item.get('kind','UNCLASSIFIED'))) + if target is not None:edge(ref,'FACT',target,relation,f"{pointer}/{key}/{index}",hard=relation=='EXPLICIT_CASE_RELATION') + elif member['kind']=='LES': + for index,identifier in enumerate(row.get('source_bo_ids',[]) if isinstance(row.get('source_bo_ids'),list) else []):edge(ref,'BO',identifier,'SOURCE_BO_ATTACHMENT',f"{pointer}/source_bo_ids/{index}") + elif member['kind']=='EVENT': + key=next((k for k in ('evidence_refs','evidence_ids') if isinstance(row.get(k),list)),None) + if key: + for index,identifier in enumerate(row[key]):edge(ref,'EVIDENCE',identifier,'SAME_EVIDENCE_REF',f"{pointer}/{key}/{index}") + member_by_ref = {row['member_ref']: row for row in members} + grouped=defaultdict(list) + for ref in sorted(parent):grouped[find(ref)].append(ref) + clusters=[]; membership={} + for index,refs in enumerate(sorted(grouped.values(),key=lambda v:v[0]),1): + cluster_ref=f"CL-{index:03d}" + clusters.append({'cluster_ref':cluster_ref,'member_refs':refs,'source_refs':[member_by_ref[ref]['source_ref'] for ref in refs]}) + for ref in refs:membership[ref]=cluster_ref + cluster_edges=sorted({(membership[r['from_ref']],membership[r['to_ref']]) for r in candidates if r['relation_kind'] in CANDIDATE_RELATION_KINDS and membership[r['from_ref']]!=membership[r['to_ref']]}) + sccs=_tarjan_scc([c['cluster_ref'] for c in clusters],cluster_edges) + component={ref:index for index,group in enumerate(sccs) for ref in group} + indegree={i:0 for i in range(len(sccs))}; adjacency=defaultdict(set) + for left,right in cluster_edges: + a,b=component[left],component[right] + if a!=b and b not in adjacency[a]:adjacency[a].add(b); indegree[b]+=1 + ready=sorted(i for i in indegree if indegree[i]==0); waves=[] + while ready: + waves.append([sccs[i] for i in ready]); upcoming=[] + for i in ready: + for j in sorted(adjacency[i]): + indegree[j]-=1 + if indegree[j]==0:upcoming.append(j) + ready=sorted(set(upcoming)) + signal_refs=[] + for occurrence in signal_all.get('record_occurrences',[]): + logical=f"signal:{occurrence['file_path']}" + document=documents[logical] + locations=_record_locations_for_signal(document) + ordinal=occurrence['record_ordinal'] + pointer,value=locations[ordinal] + signal_refs.append({'source_ref':_provenance(logical,pointer,documents),'signal_id':occurrence['signal_id'],'disposition':occurrence['disposition'],'binding_refs':occurrence.get('binding_refs',[]),'projection':_project_content(value)}) + for cluster in clusters: + member_set=set(cluster['member_refs']) + cluster_members = [member_by_ref[ref] for ref in cluster['member_refs']] + bound_ids={f"{m['kind']}:{m['stage1_id']}" for m in cluster_members if m['stage1_id'] is not None} + selected=[] + for index,row in enumerate(signal_refs): + tokens={t.replace('fact_id:','FACT:').replace('source_bo_id:','BO:').replace('bo_id:','BO:').replace('evidence_id:','EVIDENCE:').replace('event_id:','EVENT:') for t in row['binding_refs']} + if tokens & bound_ids:selected.append(index) + cluster['signal_indexes']=selected + cluster['review_refs']=[r['review_ref'] for r in reviews['normalized_occurrences'] if any(str(m['stage1_id']) in _collect_values_for_keys(r['content'], {'fact_id','fact_ids','BO_ID','bo_id','bo_ids','source_bo_id','source_bo_ids','evidence_id','evidence_ids','event_id','event_ids'}) for m in cluster_members if m['stage1_id'] is not None)] + cluster['bundle']={'member_refs':cluster['member_refs'],'signal_indexes':selected,'review_refs':cluster['review_refs']} + slot_links=[]; party_object_refs=[] + for logical,document in documents.items(): + if logical.startswith('deployment:'):continue + for pointer,value in _walk_values(document): + if not isinstance(value,dict):continue + if any(k in value for k in ('slot_id','slot_ref','evidence_slot_id')): + slot_links.append({'source_ref':_provenance(logical,pointer,documents),'projection':_project_content(value),'disposition':'OBSERVED'}) + for key in ('parties','party_refs','object_refs','objects','title_refs'): + if isinstance(value.get(key),(list,dict)): + party_object_refs.append({'kind':key,'source_ref':_provenance(logical,f"{pointer}/{key}",documents)}) + return {'source_documents':[_provenance(logical,'',documents) for logical in sorted(documents) if not logical.startswith('deployment:')], 'members':members,'relationships':relationships,'candidate_dependencies':candidates,'unresolved_relationships':unresolved,'clusters':clusters,'scheduling_waves':waves,'client_goal':{'source_ref':_provenance('client_goal','',documents),'projection':_project_content(documents['client_goal'])},'routing':{'source_ref':_provenance('domain_activation_manifest','',documents),'projection':_activation_payload(documents['domain_activation_manifest'])},'signals':signal_refs,'global_review_refs':[r['review_ref'] for r in reviews['normalized_occurrences']],'object_and_party_refs':party_object_refs,'slot_links':slot_links,'slot_link_status':'OBSERVED' if slot_links else 'UNEVALUABLE','active_profiles':[{'path':path,'sha256':canonical_digest(value),'profile':value} for path,value in sorted(deployment_documents.items()) if re.fullmatch(r'domains/[^/]+/domain_config\.json',path)]} + + + def _record_locations_for_signal(document: Any) -> list[tuple[str, Any]]: + rows=_records_from_signal_document(document) + if isinstance(document,list):return [(f'/{i}',v) for i,v in enumerate(document)] + if not isinstance(document,dict):return [] + if rows == [document]:return [('',document)] + candidates=[(p,v) for p,v in _walk_values(document) if isinstance(v,list) and v==rows] + if len(candidates)!=1: + raise IngressError('SIGNAL_RECORD_POINTER_AMBIGUOUS','signal record array cannot be located uniquely') + p,v=candidates[0] + return [(f'{p}/{i}',item) for i,item in enumerate(v)] + + + def _validate_provenance(value: Any, documents: Mapping[str, Any]) -> None: + for _,row in _walk_values(value): + if not isinstance(row,dict) or not {'logical_artifact_id','json_pointer','raw_value_sha256'}.issubset(row):continue + logical=row['logical_artifact_id'] + if logical not in documents:raise IngressError('SOURCE_REF_UNKNOWN','output refers to an unknown source') + found,raw=_json_pointer_value(documents[logical],row['json_pointer']) + if not found or canonical_digest(raw)!=row['raw_value_sha256']: + raise IngressError('SOURCE_REF_HASH_MISMATCH','output provenance does not match original content') + + + def _clean_issues(issues: Sequence[Mapping[str, Any]]) -> list[dict[str, Any]]: + rows=[]; seen=set() + for row in issues: + cleaned={k:row[k] for k in ('issue_code','severity','impact_scope','scope_refs','source_refs','message') if k in row} + key=canonical_digest(cleaned) + if key not in seen:seen.add(key); rows.append(cleaned) + return sorted(rows,key=canonical_digest) + + + def execute_ingress(hydrated: Mapping[str, Any], roots: Mapping[str, str], *, policy: Mapping[str, Any] = SOURCE_POLICY, fixture: bool = False) -> dict[str, Any]: + """Pure C00-C15 core. Fixture evaluation never enables remote publication.""" + if not fixture and policy.get('release_class')=='DEV_FIXTURE_RELEASE': + raise IngressError('DEV_FIXTURE_REAL_RUN_FORBIDDEN','DEV fixture admission cannot publish a real case') + snapshots=hydrated['snapshots']; deployment=hydrated['deployment_documents']; dep_snapshots=hydrated['deployment_snapshots'] + contracts=resolve_stage1_sources(hydrated['stage1_root']) + ingress=validate_ingress_contracts(snapshots,contracts,policy,deployment_snapshots=dep_snapshots,deployment_documents=deployment) + documents=ingress['documents']; issues=list(hydrated['issues'])+ingress['issues']; checks=[] + signal_all={}; reviews={'normalized_occurrences':[],'partition_counts':{},'conservation_status':'PASS','_issues':[]} + try: + if set(documents)!={r['logical_input_id'] for r in DEFAULT_SOURCE_CONTRACTS}: + raise IngressError('SOURCE_SET_INCOMPLETE','required Stage 1 sources are unavailable') + signal_all=expand_stage2_signal_all(hydrated['stage1_root'],documents['signal_manifest'],signal_registry=deployment.get('signals/signal_registry.v2.json')) + signal_all=bind_signal_occurrences(signal_all,documents) + issues.extend(signal_all['issues']) + for row in signal_all['ordered_file_rows']: + logical=f"signal:{row['file_path']}" + document=signal_all['_parsed_documents_by_path'][row['file_path']] + documents[logical]=document + manifest_row=documents['signal_manifest']['files'][row['manifest_index']] + schema_path=manifest_row.get('schema',manifest_row.get('schema_path')) + if isinstance(schema_path,str): + if not schema_path.startswith('signals/'):schema_path=f'signals/{schema_path}' + schema=deployment.get(schema_path) + if not isinstance(schema,dict):raise IngressError('SIGNAL_SCHEMA_UNBOUND','signal schema is not in the selected upstream closure') + try:_validate_schema_node(document,schema,root_schema=schema,schema_documents=_schema_document_index(deployment),instance_path=logical) + except _SchemaViolation as exc:raise IngressError('SIGNAL_SCHEMA_VALIDATION_FAILED',str(exc)) from exc + activation=signal_all['_parsed_documents_by_path'].get('domain_activation_manifest.json') + if activation is None:raise IngressError('SG01_SIGNAL_ARTIFACT_MISSING','signal ALL lacks domain activation') + verify_activation_projection(documents['domain_activation_manifest'],activation) + seals=verify_cross_artifact_seals(documents,snapshots,{'stage1_domain_registry_index':dep_snapshots['domains/_registry_index.json']} if 'domains/_registry_index.json' in dep_snapshots else {}) + checks.extend(seals['checks']); issues.extend(seals['issues']) + reviews=normalize_review_items(documents,policy) + conserved=check_conservation(documents,signal_all=signal_all,normalized_reviews=reviews,source_snapshots=snapshots) + checks.extend(conserved['checks']); issues.extend(conserved['issues']) + for logical,doc in documents.items(): + if logical.startswith('signal:'):continue + for pointer,value in _walk_values(doc): + if not isinstance(value,dict):continue + if value.get('stage2_auto_progression_allowed') is False or value.get('blocking') is True or value.get('blocked') is True or str(value.get('status',value.get('handoff_status',''))).upper()=='BLOCKED': + issues.append(_issue('UPSTREAM_BLOCKING_GATE',source_refs=[f'{logical}#{pointer}'])) + if any(r['blocking'] for r in reviews['normalized_occurrences']):issues.append(_issue('UPSTREAM_BLOCKING_REVIEW')) + except (IngressError,_SchemaViolation) as exc: + code=exc.code if isinstance(exc,IngressError) else 'SOURCE_SCHEMA_VALIDATION_FAILED' + issues.append(_issue(code,message=str(exc))) + issues=_clean_issues(issues) + serious=any(row.get('severity')=='ERROR' and row.get('issue_code') not in {'PRODUCER_ID_UNEVALUABLE','UNMAPPED_REVIEW_STATUS'} for row in issues) + if any(row.get('parse_status')!='PASS' or row.get('schema_status')=='FAIL' or row.get('seal_status')=='FAIL' for row in ingress['source_contract_rows']):serious=True + if any(c.get('status')=='FAIL' for c in checks):serious=True + status='BLOCKED' if serious else 'READY_WITH_ISSUES' if issues or any(r['partition'] in {'UNRESOLVED','CONDITIONAL','UNMAPPED'} for r in reviews['normalized_occurrences']) or any(r.get('seal_status')=='UNEVALUABLE' for r in ingress['source_contract_rows']) else 'READY' + context=None + if status!='BLOCKED': + try: + context=compile_case_context(documents,signal_all,reviews,deployment) + if not context['clusters']: + raise IngressError('NO_COHERENT_CLUSTER', 'no source records form a usable case context') + if context['unresolved_relationships']: + issues=_clean_issues(issues+[_issue('RELATION_TARGET_UNEVALUABLE',severity='WARNING')]); status='READY_WITH_ISSUES' + _validate_provenance(context,documents) + except IngressError as exc: + issues=_clean_issues(issues+[_issue(exc.code,message=str(exc))]); status='BLOCKED'; context=None + _validate_provenance(reviews['normalized_occurrences'],documents) + header={'schema_version':'stage2_s2_00_direct.v3','algorithm_version':ALGORITHM_VERSION,'stage1_run_root_ref':roots['stage1_run_root_ref'],'stage1_deployment_root_ref':roots['stage1_deployment_root_ref']} + manifest_rows=[{'logical_input_id':row['logical_input_id'],'path':snapshots[row['logical_input_id']].relative_path if row['logical_input_id'] in snapshots else row.get('expected_path'),'raw_sha256':row.get('raw_sha256'),'byte_length':row.get('byte_length'),'parse_status':row.get('parse_status'),'schema_status':row.get('schema_status'),'seal_status':row.get('seal_status'),'run_identity_ref':row.get('run_identity_ref'),'transaction_identity_ref':row.get('transaction_identity_ref')} for row in ingress['source_contract_rows']] + for row in signal_all.get('ordered_file_rows',[]):manifest_rows.append({'logical_input_id':f"signal:{row['file_path']}",'path':row['physical_path'],'raw_sha256':row['raw_sha256'],'byte_length':row['byte_length'],'hash_status':row['hash_status'],'record_count_status':row['record_count_status']}) + deployment_rows=[{'path':path,'raw_sha256':snap.raw_sha256,'byte_length':snap.byte_length} for path,snap in sorted(dep_snapshots.items())] + source_hashes={path:hashlib.sha256(raw).hexdigest() for path,raw in sorted(hydrated['observed'].items())} + files={ + 'ingress/stage1_input_manifest.json':{**header,'sources':manifest_rows,'deployment_sources':deployment_rows}, + 'ingress/intake_report.json':{**header,'checks':checks,'issues':issues,'source_contract_rows':[{k:v for k,v in row.items() if k!='downstream_allowed_actions'} for row in ingress['source_contract_rows']]}, + 'review/issue_ledger.base.json':{**header,'review_items':reviews['normalized_occurrences'],'partition_counts':reviews['partition_counts'],'conservation_status':reviews['conservation_status'],'issues':issues}, + } + if status=='BLOCKED':files['ingress/technical_diagnostic.json']={**header,'status':status,'issues':issues,'checks':checks} + else:files['context/case_context.json']={**header,**context} + serialized={path:canonical_json_bytes(value)+b'\n' for path,value in files.items()} + artifact_rows=[{'path':path,'raw_sha256':hashlib.sha256(raw).hexdigest(),'byte_length':len(raw)} for path,raw in sorted(serialized.items())] + files[STATUS_PATH]={**header,'status':status,'output_root':roots['output_root'],'source_hashes':source_hashes,'artifacts':artifact_rows,'written_last':True,'publication_semantics':'STATUS_LAST_LOGICAL_COMMIT'} + serialized[STATUS_PATH]=canonical_json_bytes(files[STATUS_PATH])+b'\n' + validate_output_files(serialized,roots) + return {'status':status,'files':serialized,'documents':documents} + + + def validate_output_files(files: Mapping[str, bytes], roots: Mapping[str, str]) -> dict[str, Any]: + status=load_json_strict(files.get(STATUS_PATH,b'')) + allowed=NORMAL_PATHS if status.get('status') in {'READY','READY_WITH_ISSUES'} else BLOCKED_PATHS if status.get('status')=='BLOCKED' else frozenset() + if set(files)!=allowed:raise IngressError('OUTPUT_ARTIFACT_SET_INVALID','output set differs from its processing state') + common={'schema_version','algorithm_version','stage1_run_root_ref','stage1_deployment_root_ref'} + fields={ + 'ingress/stage1_input_manifest.json':{'sources','deployment_sources'}, + 'ingress/intake_report.json':{'checks','issues','source_contract_rows'}, + 'review/issue_ledger.base.json':{'review_items','partition_counts','conservation_status','issues'}, + 'context/case_context.json':{'source_documents','members','relationships','candidate_dependencies','unresolved_relationships','clusters','scheduling_waves','client_goal','routing','signals','global_review_refs','object_and_party_refs','slot_links','slot_link_status','active_profiles'}, + 'ingress/technical_diagnostic.json':{'status','issues','checks'}, + STATUS_PATH:{'status','output_root','source_hashes','artifacts','written_last','publication_semantics'}, + } + for path,raw in files.items(): + value=load_json_strict(raw) + if not isinstance(value,dict) or set(value)!=common|fields[path]:raise IngressError('OUTPUT_CLOSED_SCHEMA_INVALID','output fields do not match the inline contract') + if value['algorithm_version']!=ALGORITHM_VERSION or value['schema_version']!='stage2_s2_00_direct.v3':raise IngressError('OUTPUT_VERSION_INVALID','output algorithm/schema version differs') + if any(value[key]!=roots[key] for key in ('stage1_run_root_ref','stage1_deployment_root_ref')):raise IngressError('OUTPUT_SOURCE_BINDING_INVALID','output roots differ from inputs') + if status['output_root']!=roots['output_root'] or status['written_last'] is not True or status['publication_semantics']!='STATUS_LAST_LOGICAL_COMMIT':raise IngressError('OUTPUT_STATUS_INVALID','status does not identify the logical completion boundary') + rows=status['artifacts'] + if not isinstance(rows,list) or len(rows)!=len(files)-1 or {r.get('path') for r in rows}!=set(files)-{STATUS_PATH}:raise IngressError('OUTPUT_STATUS_SET_INVALID','status inventory differs from actual outputs') + for row in rows: + raw=files[row['path']] + if set(row)!={'path','raw_sha256','byte_length'} or row['raw_sha256']!=hashlib.sha256(raw).hexdigest() or row['byte_length']!=len(raw):raise IngressError('OUTPUT_STATUS_HASH_INVALID','status inventory does not match output bytes') + return status + + + def publish_result(localdocs: _InlineLocaldocs, roots: Mapping[str,str], files: Mapping[str,bytes]) -> dict[str,Any]: + """No overwrite, exact completed-result reuse, and status-last publication.""" + status=validate_output_files(files,roots); output=roots['output_root'] + existing=localdocs.read_binary_optional(f'{output}/{STATUS_PATH}') + if existing is not None: + if existing!=files[STATUS_PATH]:raise IngressError('EXISTING_OUTPUT_CONFLICT','existing completed output differs in source, version, status, or inventory') + for relative,raw in sorted(files.items()): + if localdocs.read_binary(f'{output}/{relative}')!=raw:raise IngressError('EXISTING_OUTPUT_CORRUPT','existing artifact differs from completed status') + publication='REUSED_COMPLETED_OUTPUT' + else: + for relative in sorted(NORMAL_PATHS|BLOCKED_PATHS): + if relative!=STATUS_PATH and localdocs.read_binary_optional(f'{output}/{relative}') is not None:raise IngressError('PARTIAL_OUTPUT_CONFLICT','unfinished output requires explicit recovery; no overwrite') + for relative in sorted(set(files)-{STATUS_PATH}):localdocs.write_binary_verified(f'{output}/{relative}',files[relative],overwrite=False) + localdocs.write_binary_verified(f'{output}/{STATUS_PATH}',files[STATUS_PATH],overwrite=False) + publication='PUBLISHED_STATUS_LAST' + return {'ok':status['status']!='BLOCKED','status':status['status'],'output_root':output,'publication':publication,'ingress_status_sha256':hashlib.sha256(files[STATUS_PATH]).hexdigest()} + + + def run_inline_mcp(run_root: Any = RAW_RUN_ROOT, deployment_root: Any = RAW_DEPLOYMENT_ROOT, *, stage1_results: Mapping[str,Any] | None = None, client: Any | None = None) -> int: + localdocs=None + try: + roots=validate_direct_roots(run_root,deployment_root) + supplied=RAW_STAGE1_RESULTS if stage1_results is None else stage1_results + # Validate all required references before making a remote call. + for row in DEFAULT_SOURCE_CONTRACTS: + logical=row['logical_input_id'] + if logical not in supplied:raise IngressError('PREV_SOURCE_MISSING','required backend result is absent',logical_input_id=logical) + _previous_source(supplied[logical],f"{roots['stage1_run_root_ref']}/{row['path']}") + if SOURCE_POLICY['release_class']=='DEV_FIXTURE_RELEASE':raise IngressError('DEV_FIXTURE_REAL_RUN_FORBIDDEN','DEV fixture admission cannot publish a real case') + localdocs=_InlineLocaldocs(INLINE_USER_HASH,INLINE_WORKSPACE_HASH,client=client) + localdocs.initialize() + with tempfile.TemporaryDirectory(prefix='liti-s2-00-') as directory: + hydrated=hydrate_stage1(localdocs,Path(directory),roots,supplied,SOURCE_POLICY) + result=execute_ingress(hydrated,roots,policy=SOURCE_POLICY) + verify_remote_stability(localdocs,hydrated['observed']) + receipt=publish_result(localdocs,roots,result['files']) + print(json.dumps(receipt,ensure_ascii=False,separators=(',',':'))) + return 0 if receipt['ok'] else 2 + except Exception as exc: + error=exc.as_dict() if isinstance(exc,IngressError) else {'code':'S2_00_RUNTIME_ERROR','message':str(exc)} + # Failure does not assert that a status already written remotely is absent. + print(json.dumps({'ok':False,'status':'FAILED','error':error},ensure_ascii=False,separators=(',',':'))) + return 2 + finally: + if localdocs is not None:localdocs.close() + + + if __name__ == '__main__': + raise SystemExit(run_inline_mcp()) + task_procedure: + IN: + nexts: + - Task_S2_00_deterministic_ingress + wait_until: [] + Task_S2_00_deterministic_ingress: + nexts: + - OUT + wait_until: + - IN + OUT: + nexts: [] + wait_until: + - Task_S2_00_deterministic_ingress