Files
Liti-agent-Development/0. YAML_Updated/dag_script_claim_identification.yml

537 lines
26 KiB
YAML

Agent:
name: Law-aid_Claim_Identification_Agent
description: 청구권 식별 에이전트 - Fact Ledger + BO → 청구권 식별 뷰 생성 → LLM 청구권 추론
version: 0.1
Stages:
# ══════════════════════════════════════════════════════════════════
# Stage: 청구권 식별
#
# Workflow (task_procedure DAG):
# Task_A (Python: claim_identification_view.json 생성)
# ↓
# Task_B (GPT-5.4: 청구권식별 추론 → claims_identified.json)
# ══════════════════════════════════════════════════════════════════
- name: stage_청구권식별
description: Fact_Ledger_base.json + BO.json → claim_identification_view.json → claims_identified.json
llm_provider: openai
llm_model: gpt-4o-2024-08-06
tools:
mcpServers:
localdocs:
type: streamable-http
url: "http://mcp-localdocs:8012/mcp"
description: Get the content of local documents
code-executor:
type: streamable-http
url: https://code-executor.mcp.eroomai.com/mcp
description: Run scripts of programming languages
headers:
Authorization: Bearer rR8OXqWrVZA1gFEo8oWfBkw2XgpWoGrrspw5ObsxTCM=
prevs: []
nexts: []
task_procedure:
IN:
nexts: [Task_A]
wait_until: []
Task_A:
nexts: [Task_B]
wait_until: [IN]
Task_B:
nexts: [OUT]
wait_until: [Task_A]
OUT:
nexts: []
wait_until: [Task_B]
tasks:
# ────────────────────────────────────────────────────
# Task_A: Python → claim_identification_view.json
# IN: Fact_Ledger_base.json
# BO.json
# OUT: claim_identification_view.json
# ────────────────────────────────────────────────────
- task_name: Task_A
mcp: code-executor
tool_name: run_code
parameters:
language: python
requirements: "httpx"
network: "agent-network"
timeout: 120
code: |
#!/usr/bin/env python3
from __future__ import annotations
import json
import sys
from dataclasses import asdict, dataclass, field
from typing import Any, Literal
import httpx
# =============================================================================
# MCP helpers (localdocs)
# =============================================================================
LOCALDOCS_URL = "http://mcp-localdocs:8012/mcp"
_MCP_HEADERS = {
"Content-Type": "application/json",
"Accept": "application/json, text/event-stream",
}
_mcp_client = httpx.Client(timeout=60)
def _init_mcp_session():
r = _mcp_client.post(LOCALDOCS_URL, json={
"jsonrpc": "2.0", "id": 1, "method": "initialize",
"params": {
"protocolVersion": "2025-03-26",
"capabilities": {},
"clientInfo": {"name": "claim-identification", "version": "1.0"},
},
}, headers=_MCP_HEADERS)
sid = r.headers.get("mcp-session-id")
if sid:
_MCP_HEADERS["mcp-session-id"] = sid
_mcp_client.post(LOCALDOCS_URL, json={
"jsonrpc": "2.0", "method": "notifications/initialized"
}, headers=_MCP_HEADERS)
print(f"[MCP] Connected to localdocs (session: {sid})", file=sys.stderr)
def _parse_sse(text):
for line in text.strip().split("\n"):
if line.startswith("data: "):
return json.loads(line[6:])
try:
return json.loads(text)
except Exception:
return None
def _call_tool(name, arguments, msg_id=10):
r = _mcp_client.post(LOCALDOCS_URL, json={
"jsonrpc": "2.0", "id": msg_id,
"method": "tools/call",
"params": {"name": name, "arguments": arguments},
}, headers=_MCP_HEADERS)
result = _parse_sse(r.text)
if result and "result" in result:
return result
print(f"[MCP] {name} error: {json.dumps(result)[:300]}", file=sys.stderr)
return result
def _read_json_doc(doc_name):
result = _call_tool("read_docs", {"doc_names": [doc_name]})
if not (result and "result" in result):
raise RuntimeError(f"Failed to read {doc_name}: {json.dumps(result)[:300]}")
content = result["result"].get("content", [])
text = content[0].get("text", "") if content else ""
try:
parsed = json.loads(text)
except json.JSONDecodeError:
raise RuntimeError(f"Invalid JSON envelope for {doc_name}: {text[:300]}")
if isinstance(parsed, dict) and "results" in parsed:
results_list = parsed["results"]
if not results_list:
raise RuntimeError(f"No results for {doc_name} - file may not exist")
inner = results_list[0].get("content", "")
if not inner:
raise RuntimeError(f"Empty content for {doc_name} - file may not exist at this path")
if isinstance(inner, str):
try:
return json.loads(inner)
except json.JSONDecodeError:
raise RuntimeError(f"Content of {doc_name} is not valid JSON: {inner[:300]}")
return inner
return parsed
def _write_doc(doc_name, content):
_call_tool("write_file", {"path": doc_name, "content": content}, msg_id=20)
print(f"[MCP] Written: {doc_name}", file=sys.stderr)
# =============================================================================
# Data classes
# =============================================================================
Credibility = Literal["high", "medium", "low"]
@dataclass
class EvidenceStrength:
has_direct_evidence: bool | None = None
has_corroboration: bool | None = None
authentication_status_summary: str | None = None
@dataclass
class ClaimIdentificationItem:
fact_id: str
bo_id: str | None
prior_bo_id: str | None
event_date: str | None
event_class: str | None
action_text: str | None
actors: list[str] = field(default_factory=list)
object_spec: str | None = None
amount: str | None = None
claim_event_type: str | None = None
claim_nature_tags: list[str] = field(default_factory=list)
claim_creditor: str | None = None
claim_debtor: str | None = None
claim_status_tags: list[str] = field(default_factory=list)
legal_keywords: list[str] = field(default_factory=list)
evidence_refs: list[str] = field(default_factory=list)
credibility: Credibility | None = None
evidence_strength: EvidenceStrength = field(default_factory=EvidenceStrength)
@dataclass
class ClaimIdentificationView:
schema_version: str
items: list[ClaimIdentificationItem]
# =============================================================================
# Build helpers
# =============================================================================
def _uniq(values: list[str | None]) -> list[str]:
seen: set[str] = set()
out: list[str] = []
for value in values:
if not value:
continue
if value in seen:
continue
seen.add(value)
out.append(value)
return out
def _strip_event_class(value: str | None) -> str | None:
if not value:
return None
return value.split("(")[0].strip()
def _clean_auth_status(value: str | None) -> str | None:
if value is None:
return None
value = value.strip()
return value or None
def _summarize_auth_status(evidence: list[dict[str, Any]]) -> str | None:
statuses = _uniq([_clean_auth_status(e.get("authentication_status")) for e in evidence])
if not statuses:
return None
if len(statuses) == 1:
return statuses[0]
return "혼재"
def _build_evidence_strength(evidence: list[dict[str, Any]]) -> EvidenceStrength:
if not evidence:
return EvidenceStrength()
return EvidenceStrength(
has_direct_evidence=any(e.get("content_relevance") == "직접" for e in evidence),
has_corroboration=any(e.get("corroboration") == "보강존재" for e in evidence),
authentication_status_summary=_summarize_auth_status(evidence),
)
def _normalize_evidence_refs(
fact_refs: list[str] | None,
bo_evidence: list[dict[str, Any]],
) -> list[str]:
refs = [ref for ref in (fact_refs or []) if ref and ref != "증거공백"]
if refs:
return _uniq(refs)
return _uniq([e.get("evidence_index") for e in bo_evidence])
def _coerce_list(values: Any) -> list[str]:
if not isinstance(values, list):
return []
return [str(value) for value in values if value not in (None, "")]
def _build_item(
fact: dict[str, Any],
bo_index: dict[str, dict[str, Any]],
) -> ClaimIdentificationItem:
bo_id = fact.get("source_bo_id")
bo = bo_index.get(bo_id, {})
legal_calc = fact.get("legal_calculation_object") or {}
evidence = bo.get("Evidence") or []
actors = _uniq(
_coerce_list(fact.get("parties"))
+ [bo.get("Performer"), bo.get("Subject")]
)
return ClaimIdentificationItem(
fact_id=str(fact["fact_id"]),
bo_id=bo_id,
prior_bo_id=bo.get("PriorAct"),
event_date=fact.get("date") or bo.get("BehaviorTime"),
event_class=_strip_event_class(fact.get("type") or bo.get("ActionType")),
action_text=fact.get("action") or bo.get("Action"),
actors=actors,
object_spec=fact.get("object_spec"),
amount=fact.get("amount"),
claim_event_type=legal_calc.get("claim_event_type"),
claim_nature_tags=_coerce_list(legal_calc.get("claim_nature_tags")),
claim_creditor=legal_calc.get("claim_creditor"),
claim_debtor=legal_calc.get("claim_debtor"),
claim_status_tags=_coerce_list(legal_calc.get("claim_status_tags")),
legal_keywords=_coerce_list(bo.get("Legal_Keywords")),
evidence_refs=_normalize_evidence_refs(fact.get("evidence_refs"), evidence),
credibility=fact.get("credibility"),
evidence_strength=_build_evidence_strength(evidence),
)
# =============================================================================
# Main
# =============================================================================
_init_mcp_session()
fact_ledger = _read_json_doc("Fact_Ledger_base.json")
bo_list = _read_json_doc("BO.json")
if not isinstance(fact_ledger, list):
raise ValueError("Fact_Ledger_base.json must be a JSON array.")
if not isinstance(bo_list, list):
raise ValueError("BO.json must be a JSON array.")
bo_index = {
str(bo["id"]): bo
for bo in bo_list
if isinstance(bo, dict) and bo.get("id") is not None
}
items = [
_build_item(fact, bo_index)
for fact in fact_ledger
if isinstance(fact, dict) and fact.get("fact_id") is not None
]
items.sort(key=lambda item: ((item.event_date or "9999-99-99"), item.fact_id))
view = ClaimIdentificationView(
schema_version="claim_identification_view.v1",
items=items,
)
output = json.dumps(asdict(view), ensure_ascii=False, indent=2)
_write_doc("claim_identification_view.json", output)
print(f"Wrote claim_identification_view.json ({len(items)} items)")
# ────────────────────────────────────────────────────
# Task_B: GPT-5.4 청구권식별 추론 → claims_identified.json
# IN: claim_identification_view.json (← Task_A)
# Task_C_result.md
# client_goal.json
# Task_A_result.md
# 프롬프트_청구권식별_v2.txt (system prompt)
# OUT: claims_identified.json
# ────────────────────────────────────────────────────
- task_name: Task_B
llm_provider: openai
llm_model: gpt-5.4
llm_reasoning: medium
llm_verbosity: low
llm_endpoint: responses
use_tools: [localdocs]
depends_on: [Task_A]
prompts:
- role: user
content: |
# Checklist
[] 1. Preflight: list_docs 사용하여 claim_identification_view.json, Task_C_result.md, client_goal.json, Task_A_result.md 확인하고 read_docs 사용하여 claim_identification_view.json, Task_C_result.md, client_goal.json, Task_A_result.md 읽는다.
[] 2. write_file 사용하여 claims_identified.json 생성한다.
[] 3. list_docs 사용하여 claims_identified.json 파일을 확인한다. 절대 claims_identified.json을 읽지(read) 않는다.
[] 4. Terminate
<role>
당신은 대한민국 민사소송 원고대리 실무를 지원하는 LLM이다.
임무는 입력 자료만 근거로 원고가 소장에서 주장할 수 있는 본안 청구권을 식별하고, 다음 단계의 사건 종류 분류에 바로 사용할 최소 JSON만 출력하는 것이다.
</role>
<source_order>
- 읽기 순서: claim_identification_view.json -> Task_C_result.md -> client_goal.json -> Task_A_result.md
- 법적 존재, 권리발생 사실, 요건 충족, 현재 행사 가능성은 claim_identification_view.json을 우선 판단 근거로 삼는다.
- Task_C_result.md는 쟁점 정리, 누락 점검, 피고 범위 점검에만 보조 사용한다.
- client_goal.json과 Task_A_result.md는 소송 목적, 피고 우선순위, 병합 여부, 보전 필요성 판단에만 사용한다.
- claim_identification_view.json과 충돌하면 나머지 문서는 모두 후순위다.
</source_order>
<reading_scope>
- claim_identification_view.json: items 배열만 읽고 다음 필드만 사용
fact_id, prior_bo_id, event_date, event_class, action_text, actors, object_spec, amount, claim_event_type, claim_nature_tags, claim_creditor, claim_debtor, claim_status_tags, legal_keywords, evidence_refs, credibility, evidence_strength
- Task_C_result.md: `사건개요 파악` 이하만 읽음
- client_goal.json: primary_goal, constraints, summary_key_incidents, parties.defendants만 사용
- Task_A_result.md: `고객의사 확인` 이하만 읽음
</reading_scope>
<grounding_rules>
- Hallucination 금지. 입력에 없는 사건 고유 사실을 추가하지 말라.
- claim_identification_view.json에 없는 사실을 다른 문서만으로 보충하여 새로운 청구권을 만들지 말라.
- 다른 문서는 청구권 생성 근거가 아니라 보조 정리 근거다.
- 불명확하면 추정하지 말고 conditional_claims 또는 excluded_items로 보낸다.
</grounding_rules>
<claim_rules>
- 청구권은 법원이 주문으로 선고할 수 있는 본안 구제단위여야 한다.
- 청구권과 쟁점, 항변, 입증문제, 보전처분, 배경사실을 구별하라.
- 악의, 고의, 과실, 시효, 동시이행항변, 상계, 무효, 취소사유, 상당가액 변제 항변, 무자력, 상속 가능성은 원칙적으로 청구권이 아니다.
- 가압류, 가처분, 집행곤란, 상속 가능성은 related_measures에만 적는다.
- 이자, 지연손해금, 원상회복 부수항목은 독립 청구권으로 쪼개지 말고 본안 청구권에 포함된 것으로 본다.
- client_goal.json의 "constraints"를 적용하여 소를 제기할 경제적 실효가 없는 피고가 존재한다면 그 대상을 처음부터 피고에서 제외한다.
</claim_rules>
<claim_chain_rules>
- claim_identification_view.json의 item들만으로 claim chain을 구성한다.
- 아래 요소가 이어지면 같은 chain으로 묶는다:
- prior_bo_id 연결
- 같은 claim_creditor 또는 같은 원고 후보
- 같은 claim_debtor 또는 같은 피고 후보
- 같은 object_spec, 같은 금전채권, 같은 담보, 같은 재산처분
- claim_status_tags가 발생 -> 이행/불이행 -> 잔존/연체처럼 이어짐
- 같은 사실관계라도 피고가 다르거나 주문 형태가 다르면 별개 chain으로 분리한다.
</claim_chain_rules>
<classification_rules>
- 각 chain마다 아래 4문턱을 점검한다:
1. 권리발생 사실이 있는가
2. 원고가 그 권리의 주체인가
3. 피고가 그 의무의 상대방 또는 침해자인가
4. 현재 행사 가능한 상태인가
- 4개 모두 충족: identified_claims
- 청구권 골격은 있으나 1개 이상 불명확: conditional_claims
- 청구권이 아니라 쟁점/항변이거나 핵심요소 부족: excluded_items
- 증거 강도 보정:
- credibility=high: 원칙적으로 직접 근거
- credibility=medium: evidence_strength.has_direct_evidence 또는 has_corroboration이 있으면 유지 가능
- credibility=low이고 직접 증거도 없으면 원칙적으로 conditional_claims 또는 excluded_items
</classification_rules>
<claim_title_rules>
- claim_title은 다음 단계의 사건 종류 분류기가 바로 읽을 수 있는 짧은 명사형으로 쓴다.
- 법적 성질이 드러나야 한다.
- 쟁점명, 설명문, 당사자명 나열은 금지한다.
- 예시:
- 대여금 청구
- 대출금 청구
- 연대보증채무 이행청구
- 구상금 청구
- 사해행위취소 및 원상회복청구
- 말소등기청구
- 손해배상청구
- 부당이득반환청구
- 원상회복 방식이 불명확하면 `사해행위취소 및 원상회복청구`로 두고 `RESTORATION_METHOD_REVIEW`를 붙인다.
</claim_title_rules>
<field_rules>
- claim_statement: 반드시 1문장. 형식은 `원고 X는 피고 Y를 상대로 Z 청구를 할 수 있다.`
- claim_statement에는 금액, 상세 법리, 세부 원상회복 방식, 장식어를 넣지 말라. 사건 종류 식별에 필요한 핵심 명칭만 남겨라.
- source_fact_ids: 각 청구권마다 가장 결정적인 fact_id만 2개 이상 5개 이하.
- source_fact_ids에는 claim_identification_view.json에 존재하는 id만 넣는다.
- review_flags: 짧은 코드형 문자열만 사용하고 0개 이상 3개 이하로 제한한다.
- 권장 review_flags:
EVIDENCE_GAP, PARTY_ID_RECHECK, DEFENDANT_SCOPE_REVIEW, AMOUNT_RECHECK, CURRENT_ENFORCEABILITY_REVIEW, BAD_FAITH_REVIEW, RESTORATION_METHOD_REVIEW, LIMITATION_REVIEW, JOINDER_REVIEW, PRESERVATION_RECOMMENDED
</field_rules>
<related_measures_rules>
- related_measures에는 본안 외 조치만 짧게 적는다.
- 예: 보전처분 필요, 특정 피고에 대한 집행보전 필요
</related_measures_rules>
<output_contract>
- 반드시 JSON 객체 하나만 출력한다.
- 설명문, 서문, 해설, 마크다운을 출력하지 말라.
- 아래 키만, 아래 순서로 출력한다:
1. identified_claims
2. conditional_claims
3. excluded_items
4. related_measures
- 각 배열이 비어도 반드시 빈 배열로 출력한다.
</output_contract>
<output_schema>
{
"identified_claims": [
{
"claim_id": "C-001",
"claim_title": "string",
"plaintiffs": ["string"],
"defendants": ["string"],
"claim_statement": "원고 X는 피고 Y를 상대로 Z 청구를 할 수 있다.",
"source_fact_ids": ["F-001", "F-002"],
"review_flags": ["AMOUNT_RECHECK"]
}
],
"conditional_claims": [
{
"claim_id": "CC-001",
"claim_title": "string",
"possible_plaintiffs": ["string"],
"possible_defendants": ["string"],
"claim_statement": "원고 X는 피고 Y를 상대로 Z 청구를 검토할 수 있다.",
"source_fact_ids": ["F-001", "F-002"],
"review_flags": ["EVIDENCE_GAP"]
}
],
"excluded_items": [
{
"item": "string",
"reason": "쟁점 또는 근거 부족"
}
],
"related_measures": [
"보전 필요 조치가 있으면 짧게 기재"
]
}
</output_schema>
<sorting_rules>
- identified_claims는 권리발생 시점이 빠른 순으로 정렬한다.
- 같은 시점이면 client_goal.json과 Task_A_result.md에서 드러나는 소송 우선순위가 높은 순으로 정렬한다.
- conditional_claims도 같은 방식으로 정렬한다.
</sorting_rules>
<completeness_contract>
- 과업은 아래 4개 배열이 모두 채워지거나 빈 배열로 확정될 때까지 완료된 것이 아니다.
- identified_claims, conditional_claims, excluded_items, related_measures를 모두 검토하라.
- 청구권 후보를 누락하지 말되, 근거가 부족하면 억지로 identified_claims에 올리지 말라.
</completeness_contract>
<missing_context_gating>
- required context가 없으면 추정하지 말라.
- claim_identification_view.json만으로 확정되지 않는 요소는 conditional_claims 또는 excluded_items로 처리하라.
</missing_context_gating>
<verification_loop>
- finalizing 전 아래를 내부적으로 점검하라:
- correctness: 각 항목이 정말 청구권인지
- grounding: claim_identification_view.json에 없는 사실을 쓰지 않았는지
- source discipline: 전략 자료만으로 청구권을 만들지 않았는지
- formatting: JSON만 출력하는지, 키 순서가 맞는지
- field limits: claim_statement는 1문장인지, source_fact_ids는 2~5개인지, review_flags는 코드형인지
</verification_loop>
<output>
- claims_identified.json 으로 결과물을 저장한다.
</output>