This file is a merged representation of the entire codebase, combined into a single document by Repomix. This section contains a summary of this file. This file contains a packed representation of the entire repository's contents. It is designed to be easily consumable by AI systems for analysis, code review, or other automated processes. The content is organized as follows: 1. This summary section 2. Repository information 3. Directory structure 4. Repository files (if enabled) 5. Multiple file entries, each consisting of: - File path as an attribute - Full contents of the file - This file should be treated as read-only. Any changes should be made to the original repository files, not this packed version. - When processing this file, use the file path to distinguish between different files in the repository. - Be aware that this file may contain sensitive information. Handle it with the same level of security as you would the original repository. - Some files may have been excluded based on .gitignore rules and Repomix's configuration - Binary files are not included in this packed representation. Please refer to the Repository Structure section for a complete list of file paths, including binary files - Files matching patterns in .gitignore are excluded - Files matching default ignore patterns are excluded - Files are sorted by Git change count (files with more changes are at the bottom) agentbackend/ .gitea/ workflows/ ci-dev.yml ci.yml examples/ web_client.html websocket_client.py mcp_docparser/ __init__.py json_utils.py mcp_docparser.py PROMPT.md requirements.txt mcp-localdocs/ (견적서)SI20251203_01_이룸에이아이파트너스_수냉킷_(주)삼인HNT.html (견적서)SI20251203_01_이룸에이아이파트너스_수냉킷_(주)삼인HNT.json (견적서)SI20251203_01_이룸에이아이파트너스_수냉킷_(주)삼인HNT.pdf evidence_등기사항전부증명서_용인시.html evidence_등기사항전부증명서_용인시.json evidence_등기사항전부증명서_용인시.pdf src/ __init__.py agent.py server.py web-client/ src/ components/ AgentExecution.vue AgentSettings.vue FileUpload.vue TaskProcedureGraph.vue services/ sse.ts websocket.ts types/ agent.ts file.ts App.vue main.ts style.css .dockerignore .env.production .gitignore Dockerfile index.html nginx.conf package.json README.md tsconfig.json tsconfig.node.json vite.config.ts .dockerignore .gitignore .gitmodules dev_start.sh dev_stop.sh docker-compose.dev.yml docker-compose.yml Dockerfile Dockerfile.mcp-docparser Dockerfile.mcp-localdocs Dockerfile.mcp-redis Dockerfile.mcp-weaviate FIX_ORPHANED_TOOL_CALL_ID.md main.py mcp_localdocs_server.py mcp_redis_server.py mcp_weaviate_server.py mcp-read_pdf pyproject.toml README.docker.md README.md README.redis.md requirements.mcp-redis.txt requirements.mcp-weaviate.txt requirements.txt test_mcp_tools.py test_redis_mcp.py weaviate_upload_all.py yaml2json.ipynb Data_for_Agent/ case_kinds.json client_meeting.md evidence_all.json MCP_도구사용_Weaviate.txt MCP_도구사용_에러처리.txt Rule_Set_법리키워드.md Rule_Set_사실관계키워드.md Rule_Set_조문요건요소키워드.md Weaviate_DB_Structure_updated.md 적격피고자제외조건.json llm_bridge/ examples/ mcp_config_examples.py README.md src/ llm_bridge/ __init__.py base.py bridge.py client.py format_converters.py models.py types.py llm_bridge.egg-info/ dependency_links.txt PKG-INFO requires.txt SOURCES.txt top_level.txt __init__.py .gitignore CHANGELOG_orphaned_tool_fix.md INSTALL.md MANIFEST.in pyproject.toml README.md requirements-dev.txt requirements-extras.txt requirements.txt setup.py test_install.py outsourcing/ .claude/ settings.json .gitea/ workflows/ ci.yaml examples/ integration_example.py stage_config.yaml test_orchestrator.py src/ mcp_gitea_orchestrator/ __init__.py auth.py dag_executor.py server.py __init__.py web-client/ src/ components/ ExecutionMonitor.vue TaskProcedureGraph.vue services/ sse.ts types/ execution.ts App.vue main.ts vite-env.d.ts Dockerfile index.html nginx.conf package.json tsconfig.json tsconfig.node.json vite.config.ts .env.example .gitignore .gitmodules ARCHITECTURE.md DEPLOYMENT.md docker-compose.yml Dockerfile Dockerfile.executor PROJECT_STRUCTURE.md PROJECT_SUMMARY.md pyproject.toml QUICKSTART.md README.md requirements.txt run_server.sh test_deployed_server.ipynb test_orchestrator.ipynb Results_from_Agent/ BO.json Case_Dashboard.html case_research_report.md client_goal.json evidence_indexed.json Fact_Ledger.json queries_case_research.json Queries_facts_satisfying_the_legal_elements.json search_prompt.md search_results_facts_satisfying_legal_elements.json 소장.md 청구전작업.md 청구항변반박전략.md legal_agent_v03_parallel_2_C_mod.yaml This section contains the contents of the repository's files. name: CI/CD (Dev) on: push: branches: [dev] pull_request: branches: [dev] jobs: # 테스트 Job test: runs-on: ubuntu-latest steps: - name: Checkout uses: actions/checkout@v4 - name: Checkout submodules run: | git config --global url."https://${{ secrets.GIT_TOKEN }}@git.eroomai.com/".insteadOf "https://git.eroomai.com/" git submodule update --init --recursive - name: Set up Python uses: actions/setup-python@v5 with: python-version: "3.13" - name: Install dependencies run: | python -m pip install --upgrade pip pip install ./llm_bridge pip install -r requirements.txt pip install pytest pytest-asyncio - name: Run tests env: GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }} OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} LOCALDOCS_MCP_URL: ${{ secrets.LOCALDOCS_MCP_URL }} run: | pytest tests/ -v || echo "No tests found or tests skipped" # 배포 Job (Dev 환경) deploy: runs-on: dev-deploy needs: test if: github.event_name == 'push' steps: - name: Deploy to Dev Server uses: appleboy/ssh-action@v1.0.3 env: GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }} OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} PHOENIX_BASE_URL: ${{ secrets.PHOENIX_BASE_URL }} with: host: ${{ secrets.DEV_DEPLOY_HOST }} port: ${{ secrets.DEV_DEPLOY_SSH_PORT }} username: ${{ secrets.DEV_DEPLOY_USER }} key: ${{ secrets.DEV_DEPLOY_SSH_KEY }} envs: GOOGLE_API_KEY,OPENAI_API_KEY,ANTHROPIC_API_KEY,PHOENIX_BASE_URL timeout: 600s command_timeout: 30m script: | cd ${{ secrets.DEV_DEPLOY_PATH }} # Create .env file for docker compose cat > .env << EOF GOOGLE_API_KEY=$GOOGLE_API_KEY OPENAI_API_KEY=$OPENAI_API_KEY ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY PHOENIX_ENABLED=true PHOENIX_BASE_URL=$PHOENIX_BASE_URL PHOENIX_PROJECT_NAME=legal_POC_dev EOF # Reset and pull latest code git fetch origin dev git reset --hard origin/dev git submodule update --init --recursive # Ensure llm_bridge is on main branch cd llm_bridge && git fetch origin main && git checkout main && git pull origin main && cd .. # Stop backend and web-client containers echo "Stopping backend and web-client containers..." docker compose stop backend web-client || true sleep 5 docker compose rm -f backend web-client || true # Wait for containers to be fully removed echo "Waiting for container removal..." for i in $(seq 1 12); do if ! docker compose ps backend web-client --status running -q 2>/dev/null | grep -q .; then echo "Containers removed." break fi echo "Waiting... ($i/12)" sleep 5 done # Build backend and web-client without cache echo "Building backend and web-client containers..." docker compose build --no-cache backend web-client # Start backend and web-client services echo "Starting backend and web-client containers..." docker compose up -d backend web-client # Clean up old images docker image prune -f # Health check sleep 10 curl -f http://localhost:8800/ || exit 1 name: CI/CD (Main) on: push: branches: [main] pull_request: branches: [main] jobs: # 테스트 Job - main-deploy 레이블을 가진 runner-base-1에서 실행 test: runs-on: main-deploy steps: - name: Checkout uses: actions/checkout@v4 - name: Checkout submodules run: | git config --global url."https://${{ secrets.GIT_TOKEN }}@git.eroomai.com/".insteadOf "https://git.eroomai.com/" git submodule update --init --recursive - name: Set up Python uses: actions/setup-python@v5 with: python-version: "3.13" - name: Install dependencies run: | python -m pip install --upgrade pip pip install ./llm_bridge pip install -r requirements.txt pip install pytest pytest-asyncio - name: Run tests env: GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }} OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} LOCALDOCS_MCP_URL: ${{ secrets.LOCALDOCS_MCP_URL }} run: | pytest tests/ -v || echo "No tests found or tests skipped" # 배포 Job (서버에서 직접 빌드) deploy: runs-on: main-deploy needs: test if: github.event_name == 'push' && github.ref == 'refs/heads/main' steps: - name: Deploy to Server uses: appleboy/ssh-action@v1.0.3 env: GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }} OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} PHOENIX_BASE_URL: ${{ secrets.PHOENIX_BASE_URL }} with: host: ${{ secrets.DEPLOY_HOST }} port: ${{ secrets.DEPLOY_SSH_PORT }} username: ${{ secrets.DEPLOY_USER }} key: ${{ secrets.DEPLOY_SSH_KEY }} envs: GOOGLE_API_KEY,OPENAI_API_KEY,ANTHROPIC_API_KEY,PHOENIX_BASE_URL script: | cd ${{ secrets.DEPLOY_PATH }} # Create .env file for docker compose cat > .env << EOF GOOGLE_API_KEY=$GOOGLE_API_KEY OPENAI_API_KEY=$OPENAI_API_KEY ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY PHOENIX_ENABLED=true PHOENIX_BASE_URL=$PHOENIX_BASE_URL EOF # Reset and pull latest code git fetch origin main git reset --hard origin/main git submodule update --init --recursive # Ensure llm_bridge is on main branch cd llm_bridge && git fetch origin main && git checkout main && git pull origin main && cd .. # Stop all containers with retry echo "Stopping containers..." docker compose down --timeout 120 --remove-orphans || true # Wait for all containers to be fully removed echo "Waiting for containers to be fully removed..." for i in $(seq 1 30); do REMAINING=$(docker compose ps -q 2>/dev/null | wc -l) if [ "$REMAINING" -eq 0 ]; then echo "All containers removed." break fi echo "Waiting... ($i/30) - $REMAINING containers remaining" sleep 5 done # Force cleanup any remaining containers docker compose kill 2>/dev/null || true docker compose rm -f 2>/dev/null || true # Wait for Docker to fully release resources sleep 10 # Build without cache to ensure changes are applied docker compose build --no-cache # Start all services with environment variables docker compose up -d # Clean up old images docker image prune -f # Health check sleep 10 curl -f http://localhost:8800/ || exit 1 Agent WebSocket Client

🤖 Agent WebSocket Client

Interactive Agent Execution with Stage Confirmation

● Disconnected

Event Log

""" WebSocket 클라이언트 예제 에이전트를 단계별로 실행하고 각 단계 후 사용자 확인을 받습니다. """ import asyncio import json import websockets async def run_agent_with_confirmation(): """WebSocket을 통해 에이전트를 단계별로 실행하는 예제""" agent_name = "TestAgent" uri = f"ws://localhost:8000/ws/agent/{agent_name}" async with websockets.connect(uri) as websocket: print(f"✓ Connected to agent: {agent_name}\n") # 1. 에이전트 실행 시작 start_message = { "action": "start", "user_input": "판례" } print(f"→ Sending start command with input: {start_message['user_input']}") await websocket.send(json.dumps(start_message)) current_stage_index = 0 while True: # 서버로부터 메시지 수신 response = await websocket.recv() data = json.loads(response) event_type = data.get("type") if event_type == "session_started": print(f"\n✓ Session started: {data.get('session_id')}") print(f" Total stages: {data.get('total_stages')}\n") elif event_type == "stage_start": print(f"\n{'='*60}") print(f"▶ Starting Stage {data.get('stage_index') + 1}/{data.get('total_stages')}: {data.get('stage_name')}") print(f" Description: {data.get('description')}") print(f"{'='*60}\n") elif event_type == "stage_complete": stage_name = data.get("stage_name") stage_index = data.get("stage_index") result = data.get("result") print(f"\n✓ Stage '{stage_name}' completed!\n") print(f"Result (first 500 chars):") print(f"{'-'*60}") print(result[:500] if result else "No result") print(f"{'-'*60}\n") # 사용자에게 다음 액션 입력 받기 print("Choose an action:") print(" [c] Confirm - Proceed to next stage") print(" [r] Retry - Re-run current stage") print(" [d] Deny - Stop execution") while True: action_input = input("\nYour choice (c/r/d): ").strip().lower() if action_input == 'c': action = "confirm" break elif action_input == 'r': action = "retry" break elif action_input == 'd': action = "deny" break else: print("Invalid input. Please enter 'c', 'r', or 'd'.") # 선택한 액션을 서버로 전송 action_message = { "action": action, "stage_index": stage_index } print(f"\n→ Sending action: {action}\n") await websocket.send(json.dumps(action_message)) current_stage_index = stage_index elif event_type == "stage_retry": print(f"\n↻ Retrying stage: {data.get('stage_name')}...\n") elif event_type == "stage_error": print(f"\n✗ Error in stage '{data.get('stage_name')}':") print(f" {data.get('error')}\n") # 에러 발생 시에도 재시도 또는 중단 선택 가능 print("Choose an action:") print(" [r] Retry - Re-run current stage") print(" [d] Deny - Stop execution") while True: action_input = input("\nYour choice (r/d): ").strip().lower() if action_input == 'r': action = "retry" break elif action_input == 'd': action = "deny" break else: print("Invalid input. Please enter 'r' or 'd'.") action_message = { "action": action, "stage_index": data.get("stage_index") } await websocket.send(json.dumps(action_message)) elif event_type == "execution_complete": print(f"\n{'='*60}") print("✓ All stages completed successfully!") print(f"{'='*60}\n") print("Final outputs:") for stage_name, output in data.get("final_outputs", {}).items(): print(f"\n{stage_name}:") print(f" {output[:200]}..." if len(output) > 200 else f" {output}") break elif event_type == "execution_stopped": print(f"\n⊗ Execution stopped by user at stage {data.get('stage_index')}") break elif event_type == "error": print(f"\n✗ Error: {data.get('error')}") break else: print(f"\nReceived unknown event type: {event_type}") print(f"Data: {data}") async def simple_test(): """간단한 자동 테스트 (모든 단계 자동 확인)""" agent_name = "TestAgent" uri = f"ws://localhost:8000/ws/agent/{agent_name}" async with websockets.connect(uri) as websocket: print(f"✓ Connected to agent: {agent_name}\n") # 시작 await websocket.send(json.dumps({ "action": "start", "user_input": "판례" })) while True: response = await websocket.recv() data = json.loads(response) event_type = data.get("type") print(f"Event: {event_type}") if event_type == "stage_complete": # 자동으로 확인 await websocket.send(json.dumps({ "action": "confirm", "stage_index": data.get("stage_index") })) elif event_type in ["execution_complete", "execution_stopped", "error"]: print(f"Final event: {data}") break if __name__ == "__main__": print("Agent WebSocket Client") print("=" * 60) print("\nSelect mode:") print("1. Interactive mode (manual confirmation for each stage)") print("2. Automatic mode (auto-confirm all stages)") choice = input("\nYour choice (1/2): ").strip() try: if choice == "1": asyncio.run(run_agent_with_confirmation()) elif choice == "2": asyncio.run(simple_test()) else: print("Invalid choice. Exiting.") except KeyboardInterrupt: print("\n\nInterrupted by user. Exiting...") except Exception as e: print(f"\nError: {e}") # mcp_docparser package from .json_utils import safe_json_dumps, safe_json_loads, fix_and_validate_json, extract_json_from_text, sanitize_json_string """ JSON utility functions for sanitizing and validating JSON content. Handles common JSON syntax errors that may occur from LLM outputs. """ import json import re from typing import Any, Union def sanitize_json_string(json_str: str) -> str: """ Sanitize a JSON string by fixing common syntax errors. Handles: - Trailing commas in arrays and objects - Single quotes instead of double quotes - Unescaped control characters - Missing quotes around keys - Comments (// and /* */) - Trailing text after valid JSON Args: json_str: The potentially malformed JSON string Returns: A sanitized JSON string """ if not json_str or not json_str.strip(): return json_str text = json_str.strip() # Remove markdown code blocks if present if text.startswith("```json"): text = text[7:] if text.startswith("```"): text = text[3:] if text.endswith("```"): text = text[:-3] text = text.strip() # Remove single-line comments (// ...) but NOT inside strings # Only remove // comments that are outside of quoted strings def remove_comments_outside_strings(json_text): result = [] in_string = False escape_next = False i = 0 while i < len(json_text): char = json_text[i] if escape_next: result.append(char) escape_next = False i += 1 continue if char == '\\' and in_string: result.append(char) escape_next = True i += 1 continue if char == '"': in_string = not in_string result.append(char) i += 1 continue # Check for // comment outside string if not in_string and char == '/' and i + 1 < len(json_text) and json_text[i + 1] == '/': # Skip until end of line while i < len(json_text) and json_text[i] != '\n': i += 1 continue # Check for /* */ comment outside string if not in_string and char == '/' and i + 1 < len(json_text) and json_text[i + 1] == '*': i += 2 while i + 1 < len(json_text) and not (json_text[i] == '*' and json_text[i + 1] == '/'): i += 1 i += 2 # Skip */ continue result.append(char) i += 1 return ''.join(result) text = remove_comments_outside_strings(text) # Fix trailing commas before closing brackets/braces # This handles: [1, 2, 3,] or {"a": 1,} text = re.sub(r',(\s*[\]\}])', r'\1', text) # Fix unescaped newlines within strings (but not between elements) # This is tricky - we need to be careful not to break valid JSON # Fix control characters that should be escaped in strings # Replace actual tab/newline in strings with escaped versions def escape_control_chars_in_strings(match): s = match.group(0) # Escape unescaped control characters s = s.replace('\t', '\\t') # Handle newlines - but only unescaped ones # Replace actual newline with \n escape sequence s = s.replace('\r\n', '\\n') s = s.replace('\r', '\\n') s = s.replace('\n', '\\n') return s # Find strings and escape control chars within them # This regex matches JSON strings (handling escaped quotes) text = re.sub(r'"(?:[^"\\]|\\.)*"', escape_control_chars_in_strings, text) return text def extract_json_from_text(text: str) -> str: """ Extract valid JSON from text that may contain extra content. Handles cases where LLM outputs JSON with surrounding text like: "Here is the JSON: [...]" or "[...] Hope this helps!" Args: text: Text that may contain JSON Returns: The extracted JSON string, or original text if no JSON found """ if not text or not text.strip(): return text text = text.strip() # Remove markdown code blocks if text.startswith("```json"): text = text[7:] elif text.startswith("```"): text = text[3:] if text.endswith("```"): text = text[:-3] text = text.strip() # Try to find JSON array or object boundaries # Look for the first [ or { and match to its closing bracket # Find first JSON start array_start = text.find('[') object_start = text.find('{') if array_start == -1 and object_start == -1: return text # Determine which comes first if array_start == -1: start_idx = object_start open_char, close_char = '{', '}' elif object_start == -1: start_idx = array_start open_char, close_char = '[', ']' else: if array_start < object_start: start_idx = array_start open_char, close_char = '[', ']' else: start_idx = object_start open_char, close_char = '{', '}' # Find matching close bracket depth = 0 in_string = False escape_next = False end_idx = start_idx for i, char in enumerate(text[start_idx:], start=start_idx): if escape_next: escape_next = False continue if char == '\\' and in_string: escape_next = True continue if char == '"' and not escape_next: in_string = not in_string continue if in_string: continue if char == open_char: depth += 1 elif char == close_char: depth -= 1 if depth == 0: end_idx = i break if depth == 0 and end_idx > start_idx: return text[start_idx:end_idx + 1] return text def safe_json_loads(json_str: str, default: Any = None) -> Any: """ Safely parse a JSON string with automatic sanitization. Attempts to parse the JSON, and if it fails, applies sanitization and tries again. Args: json_str: The JSON string to parse default: Default value to return if parsing fails completely Returns: Parsed JSON object, or default if parsing fails """ if not json_str or not json_str.strip(): return default # First try direct parsing try: return json.loads(json_str) except json.JSONDecodeError: pass # Try extracting JSON from surrounding text try: extracted = extract_json_from_text(json_str) return json.loads(extracted) except json.JSONDecodeError: pass # Try sanitizing and parsing try: sanitized = sanitize_json_string(json_str) return json.loads(sanitized) except json.JSONDecodeError: pass # Try both extraction and sanitization try: extracted = extract_json_from_text(json_str) sanitized = sanitize_json_string(extracted) return json.loads(sanitized) except json.JSONDecodeError: pass return default def safe_json_dumps(obj: Any, ensure_ascii: bool = False, indent: int = 2) -> str: """ Safely serialize an object to JSON string. Handles special cases and ensures valid JSON output. Args: obj: The object to serialize ensure_ascii: If True, escape non-ASCII characters indent: Number of spaces for indentation (None for compact) Returns: Valid JSON string """ try: return json.dumps(obj, ensure_ascii=ensure_ascii, indent=indent) except (TypeError, ValueError) as e: # Try to handle non-serializable objects def default_serializer(o): if hasattr(o, '__dict__'): return o.__dict__ elif hasattr(o, '__str__'): return str(o) else: return f"" return json.dumps(obj, ensure_ascii=ensure_ascii, indent=indent, default=default_serializer) def validate_json_string(json_str: str) -> tuple[bool, str]: """ Validate a JSON string and return validation result with error message. Args: json_str: The JSON string to validate Returns: Tuple of (is_valid, error_message). error_message is empty if valid. """ if not json_str or not json_str.strip(): return False, "Empty JSON string" try: json.loads(json_str) return True, "" except json.JSONDecodeError as e: return False, f"JSON error at line {e.lineno}, column {e.colno}: {e.msg}" def fix_and_validate_json(json_str: str) -> tuple[str, bool, str]: """ Attempt to fix and validate a JSON string. Args: json_str: The potentially malformed JSON string Returns: Tuple of (fixed_json, is_valid, error_message) If fixing succeeds, is_valid is True and error_message is empty. """ if not json_str or not json_str.strip(): return json_str, False, "Empty JSON string" # First check if already valid is_valid, error = validate_json_string(json_str) if is_valid: return json_str, True, "" # Try extraction extracted = extract_json_from_text(json_str) is_valid, error = validate_json_string(extracted) if is_valid: return extracted, True, "" # Try sanitization sanitized = sanitize_json_string(json_str) is_valid, error = validate_json_string(sanitized) if is_valid: return sanitized, True, "" # Try both extracted = extract_json_from_text(json_str) sanitized = sanitize_json_string(extracted) is_valid, error = validate_json_string(sanitized) if is_valid: return sanitized, True, "" return json_str, False, error #!/usr/bin/env python3 """ MCP Server for parsing PDF and image documents using Google Gemini. Converts documents to HTML and then to optimized JSON format. """ import os import json import base64 import httpx from typing import Optional from fastmcp import FastMCP import sentry_sdk sentry_sdk.init( dsn="https://fc33f2349050895a6cf1e2153a5d5f2c@sentry.eroomai.com/2", send_default_pii=True, traces_sample_rate=1.0, ) from google import genai from google.genai import types from json_utils import safe_json_loads, safe_json_dumps, fix_and_validate_json, extract_json_from_text, sanitize_json_string # Configuration (loaded from environment variables) GEMINI_MODEL = "gemini-3-pro-preview" LOCALDOCS_MCP_URL = os.environ.get("LOCALDOCS_MCP_URL", "http://mcp-localdocs:8012/mcp") GOOGLE_API_KEY = os.environ.get("GOOGLE_API_KEY") # Initialize Gemini client client = genai.Client(api_key=GOOGLE_API_KEY) # Create FastMCP server mcp = FastMCP("mcp-docparser") # Prompts PROMPT_PDF_TO_HTML = """ # Role and Goal You are a text-processing and data-transformation agent. Your task is to convert a single PDF document into a HTML file that: 1. Preserves all visible text content. 2. Faithfully preserves the logical and structural relationships inside tables (rows, columns, and cells). # Description of PDF Document The attached document is primarily written in "Korean", with some sections in 'English' and others in 'Korean Chinese or Korean Hanja'. Additionally, some sentences are marked with "strikethrough." # How to Achieve the Goal Parse the document while preserving its format and generate the output as an HTML file with exactly the same language shown in the document. In so doing, **strictly adhere to the following **. 1. The document layout should be set to the normal margin format based on A4 paper. 2. Ignore any handwritten scribbles. 3. The tables presented in the document must be converted to html to preserve their original format. 4. The contents of the cells in the table presented in the document must retain their original format. That is, spacing, line breaks, and other formatting elements must be preserved exactly as they appear when converted to HTML. For instance, the content in a cell appears two line texts, i.e., "I have a dog
and a cat." in markdown expression, then, it must be converted into HTML format as "I have a dog
and a cat." 5. Contents marked with **strikethrough** in the document must be rendered exactly the same in HTML format . For instance, if you see the content, "I don't think so" that is marked with strikethrough (i.e., ~~I don't think so~~ in markdown expression), then, it must be converted into "I don't think so". 6. Your response **must not contain any citation markers, tags, or text in square brackets**. 7. Before producing the final outcome, remove every instance of '[cite_start]' from your previous response.
Return ONLY the HTML content, without any markdown code blocks or explanations. """ PROMPT_HTML_TO_JSON = """ # Role and Goal You are an advanced data-transformation agent specialized in minimizing token usage while maximizing information fidelity. **Your Task:** Convert the provided HTML document (which was derived from a PDF) into a single, optimized JSON file. **The Strategy:** You must strictly follow the **"Flattened Virtual Grid"** strategy. This approach transforms the document into a linear list of blocks and mathematically reconstructs tables as 2D matrices, ensuring perfect structural integrity without the overhead of verbose HTML tags. --- # Core Principles & Transformation Rules ### 1. The Global Structure (Block Flattening) The JSON output must be a single linear Array of **Content Blocks**. * **Text Blocks:** Paragraphs, headers, and lists become simple text objects. * **Table Blocks:** Tables become grid objects. * **Discard:** All container tags (`
`, ``, `
`) and styling attributes (`class`, `style`, `width`, `border`). ### 2. The Table Grid Rule (2D Matrix Reconstruction) Tables must be represented as a **List of Lists** (a dense 2D matrix) to preserve row/column alignment without repeating coordinate keys. * **Grid Logic:** `grid[row_index][col_index]` * **Handling Spans (Crucial):** * If a cell spans multiple rows (`rowspan=N`) or columns (`colspan=M`), you must insert **`null`** placeholders in the JSON grid for every "ghost" slot that the cell covers. * This ensures every row in the JSON array has the exact same length (number of columns), preserving vertical alignment. ### 3. The Polymorphic Cell Rule (Token Compression) Do not use a uniform object structure. Adapt the data type to the cell's complexity: * **Simple Cell:** If a cell has no spans, represent it as a raw **String**. * **Complex Cell:** If a cell starts a merge (`rowspan > 1` or `colspan > 1`), represent it as a minimal **Object**: `{ "t": "Content", "r": rowspan_val, "c": colspan_val }`. (Omit `r` or `c` if they equal 1). * **Merged Slot:** Represent the space covered by a span as **`null`**. ### 4. Text and Formatting Preservation * **Multilingual Support:** The document contains Korean, English, and Hanja. Preserve all characters exactly. * **Line Breaks:** Convert HTML `
` tags within a text block or cell into standard newline characters (`\\n`). * **Strikethrough:** The input HTML uses `` tags (e.g., `text`). Convert this to Markdown style strikethrough in the JSON string: `~~text~~`. * **Whitespace:** Trim leading/trailing whitespace from strings, but preserve internal spacing necessary for meaning. --- # Target JSON Schema Your output must strictly adhere to this format. Do not add extra keys like "id" or "styles". Return ONLY the JSON array, without any markdown code blocks or explanations. """ async def fetch_file_from_localdocs(doc_name: str) -> tuple[bytes, str]: """Fetch a file from the localdocs MCP server. Returns: Tuple of (file_content_bytes, mime_type) """ async with httpx.AsyncClient(timeout=60.0) as http_client: # Call the read_binary_doc tool via MCP for binary files response = await http_client.post( LOCALDOCS_MCP_URL, json={ "jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": { "name": "read_binary_doc", "arguments": {"doc_name": doc_name} } }, headers={"Content-Type": "application/json"} ) result = response.json() if "error" in result: raise Exception(f"MCP error: {result['error']}") content = result.get("result", {}).get("content", []) if content and len(content) > 0: text_content = content[0].get("text", "") # Check if it's an error message if text_content.startswith("Error:"): raise Exception(text_content) # Parse the JSON response from localdocs try: doc_data = safe_json_loads(text_content) if doc_data is None: raise Exception("Unexpected response format from localdocs") if doc_data.get("type") == "binary": file_bytes = base64.b64decode(doc_data["content_base64"]) mime_type = doc_data.get("mime_type", "application/octet-stream") return file_bytes, mime_type except json.JSONDecodeError: # If not JSON, treat as plain text (shouldn't happen for binary files) raise Exception("Unexpected response format from localdocs") raise Exception("No content returned from localdocs") async def convert_to_html(file_data: bytes, mime_type: str, filename: str) -> str: """Convert PDF/image to HTML using Gemini.""" contents = [ types.Content( role="user", parts=[ types.Part.from_bytes( data=file_data, mime_type=mime_type, ), types.Part.from_text(text=PROMPT_PDF_TO_HTML), ], ), ] response = await client.aio.models.generate_content( model=GEMINI_MODEL, contents=contents, config=types.GenerateContentConfig( temperature=0.1, ), ) html_content = response.text # Clean up response if it contains markdown code blocks if html_content.startswith("```html"): html_content = html_content[7:] if html_content.startswith("```"): html_content = html_content[3:] if html_content.endswith("```"): html_content = html_content[:-3] return html_content.strip() async def convert_html_to_json(html_content: str) -> list: """Convert HTML to optimized JSON using Gemini.""" contents = [ types.Content( role="user", parts=[ types.Part.from_text(text=f"{PROMPT_HTML_TO_JSON}\n\n\n{html_content}\n"), ], ), ] response = await client.aio.models.generate_content( model=GEMINI_MODEL, contents=contents, config=types.GenerateContentConfig( temperature=0.1, ), ) json_content = response.text # Debug logging print(f"[DocParser] Raw Gemini response length: {len(json_content)}") print(f"[DocParser] Raw Gemini response (first 500 chars): {json_content[:500]}") print(f"[DocParser] Raw Gemini response (last 500 chars): {json_content[-500:]}") # Extract and sanitize JSON from response extracted_json = extract_json_from_text(json_content) print(f"[DocParser] After extraction (first 500 chars): {extracted_json[:500]}") sanitized_json = sanitize_json_string(extracted_json) print(f"[DocParser] After sanitization (first 500 chars): {sanitized_json[:500]}") # Parse with validation result = safe_json_loads(sanitized_json) if result is None: # If safe_json_loads failed, try fix_and_validate_json for better error reporting fixed_json, is_valid, error_msg = fix_and_validate_json(sanitized_json) print(f"[DocParser] fix_and_validate_json result: is_valid={is_valid}, error={error_msg}") if is_valid: return json.loads(fixed_json) raise json.JSONDecodeError(f"Failed to parse JSON: {error_msg}", sanitized_json, 0) return result @mcp.tool() async def parse_document( doc_name: str, output_format: str = "json" ) -> str: """Parse a PDF or image document and extract its content. This tool fetches a document from the localdocs MCP server, uses Google Gemini to convert it to HTML (preserving tables and formatting), and then optionally converts to optimized JSON format. Args: doc_name: Name of the document to parse (e.g., 'contract.pdf', 'receipt.png') output_format: Output format - 'html' for HTML only, 'json' for optimized JSON (default: 'json') Returns: Parsed document content in the requested format """ try: # Supported mime types supported_mime_types = [ 'application/pdf', 'image/png', 'image/jpeg', 'image/gif', 'image/webp', 'image/bmp', ] # Fetch file from localdocs (returns bytes and mime_type) try: file_data, mime_type = await fetch_file_from_localdocs(doc_name) except Exception as e: return f"Error fetching document from localdocs: {str(e)}" # Validate mime type if mime_type not in supported_mime_types: return f"Error: Unsupported mime type: {mime_type}. Supported types: {', '.join(supported_mime_types)}" # Convert to HTML html_content = await convert_to_html(file_data, mime_type, doc_name) if output_format.lower() == "html": return html_content # Convert to JSON json_result = await convert_html_to_json(html_content) return safe_json_dumps(json_result, ensure_ascii=False, indent=2) except json.JSONDecodeError as e: return f"Error parsing JSON response: {str(e)}" except Exception as e: return f"Error parsing document: {str(e)}" async def _parse_document_from_base64_impl( content_base64: str, mime_type: str, filename: str = "document", output_format: str = "json" ) -> str: """Internal implementation for parsing PDF/image from base64.""" try: # Decode base64 content try: file_data = base64.b64decode(content_base64) except Exception as e: return f"Error decoding base64 content: {str(e)}" # Validate mime type valid_mime_types = [ 'application/pdf', 'image/png', 'image/jpeg', 'image/gif', 'image/webp', 'image/bmp', ] if mime_type not in valid_mime_types: return f"Error: Unsupported mime type: {mime_type}. Supported types: {', '.join(valid_mime_types)}" # Convert to HTML html_content = await convert_to_html(file_data, mime_type, filename) if output_format.lower() == "html": return html_content # Convert to JSON json_result = await convert_html_to_json(html_content) return safe_json_dumps(json_result, ensure_ascii=False, indent=2) except json.JSONDecodeError as e: return f"Error parsing JSON response: {str(e)}" except Exception as e: return f"Error parsing document: {str(e)}" async def _parse_document_with_html_impl( content_base64: str, mime_type: str, filename: str = "document" ) -> dict: """Internal implementation that returns both HTML and JSON.""" try: # Decode base64 content try: file_data = base64.b64decode(content_base64) except Exception as e: return {"error": f"Error decoding base64 content: {str(e)}"} # Validate mime type valid_mime_types = [ 'application/pdf', 'image/png', 'image/jpeg', 'image/gif', 'image/webp', 'image/bmp', ] if mime_type not in valid_mime_types: return {"error": f"Unsupported mime type: {mime_type}. Supported types: {', '.join(valid_mime_types)}"} # Convert to HTML html_content = await convert_to_html(file_data, mime_type, filename) # Convert to JSON json_result = await convert_html_to_json(html_content) json_content = safe_json_dumps(json_result, ensure_ascii=False, indent=2) return { "html": html_content, "json": json_content } except json.JSONDecodeError as e: return {"error": f"Error parsing JSON response: {str(e)}"} except Exception as e: return {"error": f"Error parsing document: {str(e)}"} @mcp.tool() async def parse_document_from_base64( content_base64: str, mime_type: str, filename: str = "document", output_format: str = "json" ) -> str: """Parse a PDF or image document from base64 encoded content. This tool accepts base64 encoded document content directly, uses Google Gemini to convert it to HTML (preserving tables and formatting), and then optionally converts to optimized JSON format. Args: content_base64: Base64 encoded document content mime_type: MIME type of the document (e.g., 'application/pdf', 'image/png') filename: Optional filename for reference (default: 'document') output_format: Output format - 'html' for HTML only, 'json' for optimized JSON (default: 'json') Returns: Parsed document content in the requested format """ return await _parse_document_from_base64_impl(content_base64, mime_type, filename, output_format) # Direct REST API endpoint for easier integration from starlette.applications import Starlette from starlette.responses import JSONResponse as StarletteJSONResponse from starlette.routing import Route async def api_parse_document(request): """Direct REST API endpoint for parsing documents (bypasses MCP protocol).""" try: body = await request.json() content_base64 = body.get("content_base64") mime_type = body.get("mime_type") filename = body.get("filename", "document") output_format = body.get("output_format", "json") if not content_base64: return StarletteJSONResponse( {"error": "content_base64 is required"}, status_code=400 ) if not mime_type: return StarletteJSONResponse( {"error": "mime_type is required"}, status_code=400 ) # Call the implementation that returns both HTML and JSON result = await _parse_document_with_html_impl( content_base64=content_base64, mime_type=mime_type, filename=filename ) # Check if result is an error if "error" in result: return StarletteJSONResponse({"error": result["error"]}, status_code=400) return StarletteJSONResponse({ "success": True, "html": result["html"], "json": result["json"], "filename": filename }) except Exception as e: return StarletteJSONResponse( {"error": f"Failed to parse document: {str(e)}"}, status_code=500 ) if __name__ == "__main__": import sys import uvicorn port = int(os.environ.get("PORT", 8013)) print(f"Starting MCP DocParser Server with FastMCP...") print(f"Using Gemini model: {GEMINI_MODEL}") print(f"LocalDocs MCP URL: {LOCALDOCS_MCP_URL}") print(f"HTTP endpoint: http://localhost:{port}/mcp") print(f"REST API endpoint: http://localhost:{port}/api/parse") # Get the MCP app and add custom route mcp_app = mcp.http_app(path="/mcp") # Create a Starlette app with both MCP and REST API routes app = Starlette( routes=[ Route("/api/parse", api_parse_document, methods=["POST"]), ] ) # Mount MCP app app.mount("/", mcp_app) uvicorn.run(app, host="0.0.0.0", port=port) 아래 두가지 프롬프트를 실행시켜서 pdf파일이나 사진으로 찍힌 문서들을 파싱하는 mcp 서버를 만들거야 streamable-http로 만들어야 하고 llm은 google gemini-3-pro-preview를 사용할 거야 언어는 pytho을 사용하고 google-genai 라이브러리를 사용할 거야 파일은 mcp를 사용해서 가져올 수 있어 그 mcp는 다음과 같아 localdocs: type: streamable-http url: "http://mcp-localdocs:8012/mcp" description: Get the content of local documents 이대로 구현해줘 Prompt_PDF_to_HTML.txt ```text # Role and Goal You are a text-processing and data-transformation agent. Your task is to convert a single PDF document into a HTML file that: 1. Preserves all visible text content. 2. Faithfully preserves the logical and structural relationships inside tables (rows, columns, and cells). # Description of PDF Document The attached document is primarily written in “Korean”, with some sections in ‘English’ and others in ‘Korean Chinese or Korean Hanja’. Additionally, some sentences are marked with “strikethrough.” # How to Achieve the Goal Parse the document while preserving its format and generate the output as an HTML file with exactly the same language shown in the document. In so doing, **strictly adhere to the following **. 1. The document layout should be set to the normal margin format based on A4 paper. 2. Ignore any handwritten scribbles. 3. The tables presented in the document must be converted to html to preserve their original format. 4. The contents of the cells in the table presented in the document must retain their original format. That is, spacing, line breaks, and other formatting elements must be preserved exactly as they appear when converted to HTML. For instance, the content in a cell appears two line texts, i.e., “I have a dog
and a cat.” in markdown expression, then, it must be converted into HTML format as “I have a dog
and a cat.” 5. Contents marked with **strikethrough** in the document must be rendered exactly the same in HTML format . For instance, if you see the content, “I don’t think so” that is marked with strikethrough (i.e., ~~I don’t think so~~ in markdown expression), then, it must be converted into “I don’t think so”. 6. Your response **must not contain any citation markers, tags, or text in square brackets**. 7. Before producing the final outcome, remove every instance of ‘[cite_start]’ from your previous response.
``` Prompt_HTML_to_JSON.txt ```text # Role and Goal You are an advanced data-transformation agent specialized in minimizing token usage while maximizing information fidelity. **Your Task:** Convert the provided HTML document (which was derived from a PDF) into a single, optimized JSON file. **The Strategy:** You must strictly follow the **"Flattened Virtual Grid"** strategy. This approach transforms the document into a linear list of blocks and mathematically reconstructs tables as 2D matrices, ensuring perfect structural integrity without the overhead of verbose HTML tags. --- # Core Principles & Transformation Rules ### 1. The Global Structure (Block Flattening) The JSON output must be a single linear Array of **Content Blocks**. * **Text Blocks:** Paragraphs, headers, and lists become simple text objects. * **Table Blocks:** Tables become grid objects. * **Discard:** All container tags (`
`, ``, `
`) and styling attributes (`class`, `style`, `width`, `border`). ### 2. The Table Grid Rule (2D Matrix Reconstruction) Tables must be represented as a **List of Lists** (a dense 2D matrix) to preserve row/column alignment without repeating coordinate keys. * **Grid Logic:** `grid[row_index][col_index]` * **Handling Spans (Crucial):** * If a cell spans multiple rows (`rowspan=N`) or columns (`colspan=M`), you must insert **`null`** placeholders in the JSON grid for every "ghost" slot that the cell covers. * This ensures every row in the JSON array has the exact same length (number of columns), preserving vertical alignment. ### 3. The Polymorphic Cell Rule (Token Compression) Do not use a uniform object structure. Adapt the data type to the cell's complexity: * **Simple Cell:** If a cell has no spans, represent it as a raw **String**. * **Complex Cell:** If a cell starts a merge (`rowspan > 1` or `colspan > 1`), represent it as a minimal **Object**: `{ "t": "Content", "r": rowspan_val, "c": colspan_val }`. (Omit `r` or `c` if they equal 1). * **Merged Slot:** Represent the space covered by a span as **`null`**. ### 4. Text and Formatting Preservation * **Multilingual Support:** The document contains Korean, English, and Hanja. Preserve all characters exactly. * **Line Breaks:** Convert HTML `
` tags within a text block or cell into standard newline characters (`\n`). * **Strikethrough:** The input HTML uses `` tags (e.g., `text`). Convert this to Markdown style strikethrough in the JSON string: `~~text~~`. * **Whitespace:** Trim leading/trailing whitespace from strings, but preserve internal spacing necessary for meaning. --- # Target JSON Schema Your output must strictly adhere to this format. Do not add extra keys like "id" or "styles". **Example Input:** ```html

Summary of Data

Region Metrics
Old Sales
New Sales
Profit
# **Required Output:** JSON Format as follows: [ { "type": "text", "content": "Summary of Data" }, { "type": "table", "grid": [ [ { "t": "Region", "r": 2 }, { "t": "Metrics", "c": 2 }, null ], [ null, "~~Old Sales~~\nNew Sales", "Profit" ] ] } ] ``` fastmcp>=2.13.0 google-genai>=1.0.0 httpx>=0.27.0 python-dotenv>=1.0.0 sentry-sdk Quotation
SAM-IN HNT Inc. 컨설팅 본부 강윤구 이사 (010-9734-5456) ykkang@saminhnt.com
TEL: 042-719-7780 FAX: 042-719-7790 http://www.saminhnt.com

견 적 서

(Quotation)

이룸에이아이파트너스 (귀중)
1. 견적일자 : 2025년 12월 03일
2. 납품일자 : 협의
3. 견적유효기간 : 견적일로부터 15일
4. 결제조건 : 협의
5. 사 업 명 : 수냉킷 구매
하기와 같이 견적합니다.
(주)삼인에이치엔티
대표이사    김 승 한 (인)
대전광역시 유성구 갑천로361-17
대덕벤처타워 웹스퀘어 405호, 406호
최 종 공 급 가 [VAT포함] 일금 오십일만칠천 원정 ( ₩517,000 )
(단위:원 / 부가세별도)
종 목
ITEM
모델 번호
MODEL NO.
상세내역
(DESCRIPTION)
수량
Q'TY
단가
(UNIT PRICE)
공급가
(TOTAL AMOUNT)
비 고
1. RTX 3090용 수냉킷 1 470,000 470,000
B-FRD-PL3090-SR-360 AIO Water Cooling 1
* 관부가세 포함
***** 이 하 여 백 ******
 
 
 
 
공 급 가 ( Total Amount ) 470,000
부가가치세 ( V.A.T ) 10% 47,000
최 종 공 급 가 ( Grand Total Amount ) [VAT 포함] 517,000
(비고)
[ "SAMIN", "SAM-IN HNT Inc. 컨설팅 본부 강윤구 이사 (010-9734-5456) ykkang@saminhnt.com\nTEL: 042-719-7780 FAX: 042-719-7790 http://www.saminhnt.com", "견 적 서", "(Quotation)", "이룸에이아이파트너스 (귀중)", "1. 견적일자 : 2025년 12월 03일\n2. 납품일자 : 협의\n3. 견적유효기간 : 견적일로부터 15일\n4. 결제조건 : 협의\n5. 사 업 명 : 수냉킷 구매\n하기와 같이 견적합니다.", "(주)삼인에이치엔티", "대표이사 김 승 한 (인)", "대전광역시 유성구 갑천로361-17\n대덕벤처타워 웹스퀘어 405호, 406호", "최 종 공 급 가 [VAT포함] 일금 오십일만칠천 원정 ( ₩517,000 )", "(단위:원 / 부가세별도)", { "grid": [ [ "종 목\nITEM", "모델 번호\nMODEL NO.", "상세내역\n(DESCRIPTION)", "수량\nQ'TY", "단가\n(UNIT PRICE)", "공급가\n(TOTAL AMOUNT)", "비 고" ], [ "1. RTX 3090용 수냉킷", "", "", "1", "470,000", "470,000", "" ], [ "", "B-FRD-PL3090-SR-360", "AIO Water Cooling", "1", "", "", "" ], [ "", "", "* 관부가세 포함", "", "", "", "" ], [ "", "", "***** 이 하 여 백 ******", "", "", "", "" ], [ " ", "", "", "", "", "", "" ], [ " ", "", "", "", "", "", "" ], [ " ", "", "", "", "", "", "" ], [ " ", "", "", "", "", "", "" ], [ { "t": "공 급 가 ( Total Amount )", "c": 5 }, null, null, null, null, "470,000", "" ], [ { "t": "부가가치세 ( V.A.T ) 10%", "c": 5 }, null, null, null, null, "47,000", "" ], [ { "t": "최 종 공 급 가 ( Grand Total Amount ) [VAT 포함]", "c": 5 }, null, null, null, null, "517,000", "" ] ] }, "(비고)" ] 등기사항전부증명서
등기사항전부증명서(말소사항 포함) - 토지
[토지] 경기도 용인시 구성면 언남리 115-1 고유번호 1102-2010-111495
【 표 제 부 】 ( 토지의 표시 )
표시번호 접 수 소재지번 지 목 면 적 등기원인 및 기타사항
1
(전 3)
2010년 9월
16일
경기도 용인시
구성면 언남리
115-1
공장용지 5,124㎡ 부동산등기법 제177조의6 제1
항의 규정에 의하여 2014년 06
월 05일 전산이기
【 갑 구 】 ( 소유권에 관한 사항 )
순위번호 등기목적 접 수 등기원인 권리자 및 기타사항
1
(전3)
소유권이전 1995년 8월 18일
제12438호
1995년 4월 30일
경락
소유자 김수경 500125-2047235
서울 영등포구 여의도동 38-1
장미아파트 2동 206호
부동산등기법 제177조의6 제1
항의 규정에 의하여 2014년 06
월 05일 전산이기
2 가압류 2015년 12월 24일
제19277호
2015년 12월 22일
서울중앙지방법원의
가압류결정
(2015카단7142)
청구금액 금 220,000,000원
채권자 우방신용보증보험 주식회사
111474-493652
서울 중구 중림로7길 23-16
3 임의경매개시결정 2016년 4월 16일
제15672호
2016년 4월 15일
수원지방법원의 경매
개시결정
(2016타경8654)
채권자 우방캐피탈 주식회사
110111-0545563
서울 종로구 내수동 167
4 소유권이전 2018년 1월 20일
제3678호
2017년 12월 10일
임의경매로 인한
매각
소유자 민정기 620922-1934562
서울 강남구 압구정로 265-35
거래가액 금 965,000,000원
5 2번가압류등
기말소
2018년 1월 20일
제3678호
2017년 12월 10일
임의경매로 인한
매각
6 3번임의경매
신청등기말소
2018년 1월 20일
제3678호
2017년 12월 10일
임의경매로 인한
매각
[토지] 경기도 용인시 구성면 언남리 115-1 고유번호 1102-2010-111495
【 을 구 】 ( 소유권 이의의 권리에 관한 사항 )
순위번호 등기목적 접 수 등기원인 권리자 및 기타사항
1
(전14)
근저당권설정 2013년 10월 8일
제11887호
2013년 10월 7일
설정계약
채권최고액 금 1,500,000,000원
채무자 주식회사 개성금속
파주시 법원읍 법원리 495-12
근저당권자 우방캐피탈 주식회사
110111-0545563
서울 종로구 내수동 167
공동담보 동소 지상 공장 건물
부동산등기법 제177조의6 제1항
의 규정에 의하여 2014년 06월
05일 전산이기
2 1번근저당권
말소
2018년 1월 20일
제3678호
2017년 12월 10일
임의경매로 인한
매각
― 이 하 여 백 ―
수수료 금 1,000원 영수함     관할등기소 수원지방법원 용인등기소 / 발행등기소 법원행정처 등기정보중앙관리소
이 증명서는 등기기록의 내용과 틀림없음을 증명합니다.
서기 2018년 08월 06일
법원행정처 등기정보중앙관리소 전산운영책임관
등기정보
중앙관리
소전산운
영책임관
[ "등기사항전부증명서(말소사항 포함) - 토지", "[토지] 경기도 용인시 구성면 언남리 115-1", "고유번호 1102-2010-111495", "【 표 제 부 】 ( 토지의 표시 )", [ [ "표시번호", "접 수", "소재지번", "지 목", "면 적", "등기원인 및 기타사항" ], [ "1\n(전 3)", "2010년 9월\n16일", "경기도 용인시\n구성면 언남리\n115-1", "공장용지", "5,124㎡", "부동산등기법 제177조의6 제1\n항의 규정에 의하여 2014년 06\n월 05일 전산이기" ] ], "【 갑 구 】 ( 소유권에 관한 사항 )", [ [ "순위번호", "등기목적", "접 수", "등기원인", "권리자 및 기타사항" ], [ "1\n(전3)", "소유권이전", "1995년 8월 18일\n제12438호", "1995년 4월 30일\n경락", "소유자 김수경 500125-2047235\n서울 영등포구 여의도동 38-1\n장미아파트 2동 206호\n부동산등기법 제177조의6 제1\n항의 규정에 의하여 2014년 06\n월 05일 전산이기" ], [ "~~2~~", "~~가압류~~", "~~2015년 12월 24일\n제19277호~~", "~~2015년 12월 22일\n서울중앙지방법원의\n가압류결정\n(2015카단7142)~~", "~~청구금액 금 220,000,000원\n채권자 우방신용보증보험 주식회사\n111474-493652\n서울 중구 중림로7길 23-16~~" ], [ "~~3~~", "~~임의경매개시결정~~", "~~2016년 4월 16일\n제15672호~~", "~~2016년 4월 15일\n수원지방법원의 경매\n개시결정\n(2016타경8654)~~", "~~채권자 우방캐피탈 주식회사\n110111-0545563\n서울 종로구 내수동 167~~" ], [ "4", "소유권이전", "2018년 1월 20일\n제3678호", "2017년 12월 10일\n임의경매로 인한\n매각", "소유자 민정기 620922-1934562\n서울 강남구 압구정로 265-35\n거래가액 금 965,000,000원" ], [ "5", "2번가압류등\n기말소", "2018년 1월 20일\n제3678호", "2017년 12월 10일\n임의경매로 인한\n매각", "" ], [ "6", "3번임의경매\n신청등기말소", "2018년 1월 20일\n제3678호", "2017년 12월 10일\n임의경매로 인한\n매각", "" ] ], "문서 하단의 바코드를 스캐너로 확인하거나 인터넷등기소(http://iros.go.kr)의 발급확인 메뉴에서 발급확인번호를 입력하여 위·변조 여부를 확인할 수 있습니다. 발급확인번호를 통한 확인은 발행일로부터 3개월까지 5회에 한하여 가능합니다.\n발행번호 16120046104191024010960021TJS0160609SOG157210311121 1/2 발행일 2018/08/06", "[토지] 경기도 용인시 구성면 언남리 115-1", "고유번호 1102-2010-111495", "【 을 구 】 ( 소유권 이의의 권리에 관한 사항 )", [ [ "순위번호", "등기목적", "접 수", "등기원인", "권리자 및 기타사항" ], [ "~~1\n(전14)~~", "~~근저당권설정~~", "~~2013년 10월 8일\n제11887호~~", "~~2013년 10월 7일\n설정계약~~", "~~채권최고액 금 1,500,000,000원\n채무자 주식회사 개성금속\n파주시 법원읍 법원리 495-12\n근저당권자 우방캐피탈 주식회사\n110111-0545563\n서울 종로구 내수동 167\n공동담보 동소 지상 공장 건물\n부동산등기법 제177조의6 제1항\n의 규정에 의하여 2014년 06월\n05일 전산이기~~" ], [ "2", "1번근저당권\n말소", "2018년 1월 20일\n제3678호", "2017년 12월 10일\n임의경매로 인한\n매각", "" ] ], "― 이 하 여 백 ―", "수수료 금 1,000원 영수함 관할등기소 수원지방법원 용인등기소 / 발행등기소 법원행정처 등기정보중앙관리소", "이 증명서는 등기기록의 내용과 틀림없음을 증명합니다.", "서기 2018년 08월 06일", "법원행정처 등기정보중앙관리소 전산운영책임관", "등기정보\n중앙관리\n소전산운\n영책임관", "*실선으로 그어진 부분은 말소사항을 표시함. *등기기록에 기록된 사항이 없는 갑구 또는 을구는 생략함.\n문서 하단의 바코드를 스캐너로 확인하거나 인터넷등기소(http://iros.go.kr)의 발급확인 메뉴에서 발급확인번호를 입력하여\n위·변조 여부를 확인할 수 있습니다. 발급확인번호를 통한 확인은 발행일로부터 3개월까지 5회에 한하여 가능합니다.\n발행번호 16120046104191024010960021TJS0160609SOG157210311121 2/2 발행일 2018/08/06\n대 법 원" ] import traceback import json import asyncio import logging import concurrent.futures import time from typing import AsyncIterator, List, Dict, Any, Optional, Set from dataclasses import dataclass, field from enum import Enum from llm_bridge import MCPClientManager from llm_bridge import Bridge import sentry_sdk import tiktoken # Logger for this module logger = logging.getLogger(__name__) # Token limit constants TOKEN_LIMIT = 200000 # API 토큰 리밋: 20만 토큰 TOKEN_THRESHOLD = 180000 # 요약 시작 임계값: 18만 토큰 # tiktoken 인코더 캐시 _tiktoken_encoders = {} def get_tiktoken_encoder(model: str = "gpt-4o"): """모델에 맞는 tiktoken 인코더를 반환 (캐시 사용).""" if model not in _tiktoken_encoders: try: # OpenAI 모델의 경우 _tiktoken_encoders[model] = tiktoken.encoding_for_model(model) except KeyError: # 알 수 없는 모델의 경우 cl100k_base 사용 (GPT-4, Claude 등에 적합) _tiktoken_encoders[model] = tiktoken.get_encoding("cl100k_base") return _tiktoken_encoders[model] def count_tokens(text: str, model: str = "gpt-4o") -> int: """텍스트의 토큰 수를 계산.""" if not text: return 0 encoder = get_tiktoken_encoder(model) return len(encoder.encode(text)) def count_messages_tokens(messages: List[Dict[str, Any]], model: str = "gpt-4o") -> int: """메시지 리스트의 총 토큰 수를 계산.""" total_tokens = 0 for msg in messages: # role 토큰 role = msg.get('role', '') if isinstance(msg, dict) else getattr(msg, 'role', '') total_tokens += count_tokens(role, model) # content 토큰 content = msg.get('content', '') if isinstance(msg, dict) else getattr(msg, 'content', '') if content: if isinstance(content, str): total_tokens += count_tokens(content, model) elif isinstance(content, list): # content가 리스트인 경우 (멀티모달 등) for item in content: if isinstance(item, dict) and 'text' in item: total_tokens += count_tokens(item['text'], model) # tool_calls 토큰 (있는 경우) tool_calls = msg.get('tool_calls') if isinstance(msg, dict) else getattr(msg, 'tool_calls', None) if tool_calls: for tc in tool_calls: if isinstance(tc, dict): total_tokens += count_tokens(json.dumps(tc), model) elif hasattr(tc, 'model_dump'): total_tokens += count_tokens(json.dumps(tc.model_dump()), model) # 메시지 오버헤드 (대략 4토큰 per message) total_tokens += 4 return total_tokens async def summarize_context(messages: List[Dict[str, Any]], llm_bridge: 'Bridge', stage_name: str) -> str: """ 현재까지의 대화 컨텍스트를 요약. LLM Bridge의 summarize_messages 메서드를 사용하여 사용자가 선택한 모델로 요약합니다. """ try: # LLM Bridge의 summarize_messages 메서드 사용 (사용자가 선택한 모델 사용) summary = await llm_bridge.summarize_messages(messages) if summary: logger.info(f"[TOKEN_MANAGEMENT] Context summarized for stage '{stage_name}', summary length: {len(summary)}") return summary except Exception as e: logger.error(f"[TOKEN_MANAGEMENT] Failed to summarize context: {e}") # 요약 실패 시 간단한 폴백 return f"[이전 작업 요약 - {stage_name}]\n작업이 진행 중이었습니다. 토큰 제한으로 인해 컨텍스트가 초기화되었습니다." class LLMAPIError(Exception): """LLM API 호출 중 발생한 에러를 래핑하는 커스텀 예외 클래스.""" def __init__(self, message: str, original_error: Exception = None, error_type: str = None, error_code: str = None, stage_name: str = None, iteration: int = None): super().__init__(message) self.original_error = original_error self.error_type = error_type or (type(original_error).__name__ if original_error else "Unknown") self.error_code = error_code self.stage_name = stage_name self.iteration = iteration self.traceback_str = traceback.format_exc() if original_error else None def to_dict(self) -> Dict[str, Any]: """에러 정보를 딕셔너리로 변환 (프론트엔드 전달용).""" return { "message": str(self), "error_type": self.error_type, "error_code": self.error_code, "stage_name": self.stage_name, "iteration": self.iteration, "original_error": str(self.original_error) if self.original_error else None, "traceback": self.traceback_str } def parse_api_error(error: Exception) -> Dict[str, Any]: """API 에러에서 상세 정보를 추출.""" error_info = { "error_type": type(error).__name__, "error_message": str(error), "error_code": None, "api_error_details": None } # OpenAI/Anthropic API 에러 파싱 if hasattr(error, 'status_code'): error_info["error_code"] = str(error.status_code) if hasattr(error, 'response'): try: response = error.response if hasattr(response, 'json'): error_info["api_error_details"] = response.json() elif hasattr(response, 'text'): error_info["api_error_details"] = response.text except Exception: pass # 에러 메시지에서 코드 추출 시도 error_str = str(error) if "Error code:" in error_str: try: code_part = error_str.split("Error code:")[1].split("-")[0].strip() error_info["error_code"] = code_part except Exception: pass return error_info # Thread pool for running blocking LLM calls _thread_pool = concurrent.futures.ThreadPoolExecutor(max_workers=4) # Enable basic logging, but disable verbose MCP debug logs logging.basicConfig(level=logging.INFO) # llm_bridge 로그는 INFO로 설정하여 상세 로그 출력 logging.getLogger("llm_bridge").setLevel(logging.INFO) logging.getLogger("mcp").setLevel(logging.ERROR) logging.getLogger("httpx").setLevel(logging.ERROR) class TaskStatus(Enum): """Task 실행 상태""" PENDING = "pending" WAITING = "waiting" # wait_until 대기 중 RUNNING = "running" COMPLETED = "completed" FAILED = "failed" @dataclass class Task: """ Stage 내에서 실행되는 개별 Task. Stage와 유사하지만 더 가벼운 구조로, Stage 내에서 병렬/순차 실행됩니다. """ task_name: str llm_provider: str llm_model: str prompts: List[Dict[str, str]] api_key: str = None mcp_client_manager: MCPClientManager = field(default=None, repr=False) # 공유 MCP 매니저 llm: Bridge = field(default=None, repr=False) status: TaskStatus = TaskStatus.PENDING result: Any = None error: str = None def __post_init__(self): """Task 초기화 후 LLM Bridge 설정""" if self.api_key: self._init_llm() def _init_llm(self): """LLM Bridge 초기화""" # 외부에서 주입받은 mcp_client_manager 사용 manager = self.mcp_client_manager if self.llm_provider == "openai": self.llm = Bridge(model=self.llm_model, api_key=self.api_key, mcp_client_manager=manager) elif self.llm_provider == "anthropic": base_url = "https://api.anthropic.com/v1/" self.llm = Bridge(model=self.llm_model, api_key=self.api_key, mcp_client_manager=manager, base_url=base_url) elif self.llm_provider == "google": base_url = "https://generativelanguage.googleapis.com/v1beta/openai/" self.llm = Bridge(model=self.llm_model, api_key=self.api_key, mcp_client_manager=manager, base_url=base_url) else: raise ValueError(f"Unknown LLM provider: {self.llm_provider}") def initialize(self, mcp_client_manager: MCPClientManager, api_key: str): """외부에서 mcp_client_manager와 api_key를 설정하고 LLM 초기화""" self.mcp_client_manager = mcp_client_manager self.api_key = api_key self._init_llm() async def run(self, system_prompt: str, on_iteration=None) -> List[Dict]: """ Task 실행 - Stage.run과 유사한 로직 Args: system_prompt: 시스템 프롬프트 on_iteration: 반복 콜백 함수 Returns: 실행 결과 메시지 리스트 """ self.status = TaskStatus.RUNNING start_time = time.time() logger.info(f"[TASK:{self.task_name}] Starting task execution") logger.info(f"[TASK:{self.task_name}] LLM Provider: {self.llm_provider}, Model: {self.llm_model}") messages = [{"role": "system", "content": system_prompt}] messages += list(self.prompts) # 토큰 관리를 위한 변수 context_reset_count = 0 original_system_prompt = system_prompt try: for i in range(8): logger.info(f"[TASK:{self.task_name}] === Starting iteration {i+1}/8 ===") iteration_start_time = time.time() await asyncio.sleep(0) # 토큰 사용량 체크 current_tokens = count_messages_tokens(messages, self.llm_model) logger.info(f"[TASK:{self.task_name}] Current token count: {current_tokens:,} / {TOKEN_THRESHOLD:,}") if current_tokens >= TOKEN_THRESHOLD: context_reset_count += 1 logger.warning(f"[TASK:{self.task_name}] Token threshold reached. Summarizing context...") try: summary = await summarize_context(messages, self.llm, self.task_name) continuation_prompt = f"""[컨텍스트 연속 - 세션 #{context_reset_count + 1}] 이전 세션에서 진행된 작업 요약: {summary} 위 요약을 바탕으로 작업을 계속 진행해주세요.""" messages = [{"role": "system", "content": original_system_prompt}] messages += list(self.prompts) messages.append({"role": "user", "content": continuation_prompt}) except Exception as e: logger.error(f"[TASK:{self.task_name}] Failed to summarize context: {e}") # LLM 호출 loop = asyncio.get_event_loop() current_iteration = i + 1 task_name = self.task_name def run_in_thread(): logger.info(f"[THREAD] Starting thread for Task LLM call (task: {task_name}, iteration: {current_iteration})") new_loop = asyncio.new_event_loop() asyncio.set_event_loop(new_loop) try: result = new_loop.run_until_complete(self.llm.process_messages(messages)) return result except Exception as e: error_info = parse_api_error(e) logger.error(f"[THREAD] Task LLM API Error: {error_info}") raise LLMAPIError( message=f"LLM API 오류: {error_info['error_message']}", original_error=e, error_type=error_info['error_type'], error_code=error_info['error_code'], stage_name=task_name, iteration=current_iteration ) from e finally: new_loop.close() result = await loop.run_in_executor(_thread_pool, run_in_thread) llm_call_duration = time.time() - iteration_start_time logger.info(f"[TASK:{self.task_name}] Iteration {i+1}: LLM call completed in {llm_call_duration:.2f}s") await asyncio.sleep(0) messages = [msg.model_dump() if hasattr(msg, 'model_dump') else msg for msg in result] # 최신 assistant 응답 찾기 is_complete = False content = "" for msg in reversed(messages): msg_role = msg.get('role') if isinstance(msg, dict) else getattr(msg, 'role', None) msg_content = msg.get('content') if isinstance(msg, dict) else getattr(msg, 'content', None) if msg_role == "assistant" and msg_content: content = msg_content break if content and (content.strip().endswith("**terminate**") or content.strip().endswith("terminate")): is_complete = True if on_iteration: await on_iteration(i + 1, content, is_complete) if is_complete: break messages.append({"role": "user", "content": "Continue the next step."}) self.status = TaskStatus.COMPLETED self.result = messages return messages except Exception as e: self.status = TaskStatus.FAILED self.error = str(e) logger.error(f"[TASK:{self.task_name}] Task failed: {e}") raise class TaskProcedureExecutor: """ Task 실행 흐름을 DAG(Directed Acyclic Graph) 기반으로 관리하고 실행합니다. task_procedure 예시: { "IN": {"nexts": ["task_1"], "wait_until": []}, "task_1": {"nexts": ["task_2", "task_4"], "wait_until": []}, "task_2": {"nexts": ["task_3"], "wait_until": []}, "task_3": {"nexts": ["task_5"], "wait_until": []}, "task_4": {"nexts": ["task_5"], "wait_until": []}, "task_5": {"nexts": ["OUT"], "wait_until": ["task_3", "task_4"]} } """ def __init__(self, tasks: Dict[str, Task], procedure: Dict[str, Dict], system_prompt: str): """ Args: tasks: task_name -> Task 객체 매핑 procedure: task_procedure 정의 (IN, OUT 포함) system_prompt: 모든 Task에 공통으로 적용할 시스템 프롬프트 """ self.tasks = tasks self.procedure = procedure self.system_prompt = system_prompt self.completed_tasks: Set[str] = set() self.running_tasks: Set[str] = set() self.task_results: Dict[str, Any] = {} self.task_events: Dict[str, asyncio.Event] = {} # 각 Task에 대한 완료 이벤트 생성 for task_name in tasks: self.task_events[task_name] = asyncio.Event() # wait_until 기반 의존성 로깅 wait_until_deps = {} for task_name, config in procedure.items(): if task_name in ("IN", "OUT"): continue wait_until = config.get("wait_until", []) if wait_until: wait_until_deps[task_name] = wait_until logger.info(f"[EXECUTOR] Wait-until dependencies: {wait_until_deps}") def _can_start_task(self, task_name: str) -> bool: """Task가 시작 가능한지 확인 (wait_until 조건 체크)""" if task_name in self.completed_tasks or task_name in self.running_tasks: return False task_config = self.procedure.get(task_name, {}) wait_until = task_config.get("wait_until", []) # wait_until에 있는 모든 Task가 완료되어야 함 for dep_task in wait_until: if dep_task not in self.completed_tasks: return False return True def _get_dependencies(self, task_name: str) -> Set[str]: """Task의 의존성을 반환 (wait_until만 사용)""" task_config = self.procedure.get(task_name, {}) wait_until = task_config.get("wait_until", []) return set(wait_until) async def _wait_for_dependencies(self, task_name: str): """Task의 의존성이 완료될 때까지 대기""" dependencies = self._get_dependencies(task_name) if not dependencies: return logger.info(f"[EXECUTOR] Task '{task_name}' waiting for: {dependencies}") # 모든 의존 Task의 완료 이벤트를 기다림 await asyncio.gather(*[ self.task_events[dep].wait() for dep in dependencies if dep in self.task_events ]) logger.info(f"[EXECUTOR] Task '{task_name}' dependencies satisfied") async def _run_single_task(self, task_name: str, on_task_event=None) -> Any: """단일 Task 실행""" if task_name not in self.tasks: logger.error(f"[EXECUTOR] Task '{task_name}' not found") return None task = self.tasks[task_name] self.running_tasks.add(task_name) # 의존성 확인 (자동 계산 + 명시적 wait_until) dependencies = self._get_dependencies(task_name) if dependencies: # 의존성이 있으면 waiting 상태로 시작 task.status = TaskStatus.WAITING if on_task_event: await on_task_event({ "type": "task_start", "task_name": task_name, "status": "waiting" }) # 의존성 대기 await self._wait_for_dependencies(task_name) # 의존성 완료 후 running 상태로 변경 task.status = TaskStatus.RUNNING if on_task_event: await on_task_event({ "type": "task_start", "task_name": task_name, "status": "running" }) try: # Task 실행 async def task_on_iteration(iteration_num, content, is_complete): if on_task_event: await on_task_event({ "type": "task_iteration", "task_name": task_name, "iteration": iteration_num, "content": content, "is_complete": is_complete }) result = await task.run(self.system_prompt, on_iteration=task_on_iteration) self.task_results[task_name] = result self.completed_tasks.add(task_name) self.running_tasks.discard(task_name) self.task_events[task_name].set() # 완료 신호 if on_task_event: await on_task_event({ "type": "task_complete", "task_name": task_name, "status": "completed" }) logger.info(f"[EXECUTOR] Task '{task_name}' completed successfully") return result except Exception as e: task.status = TaskStatus.FAILED task.error = str(e) self.running_tasks.discard(task_name) self.task_events[task_name].set() # 실패해도 이벤트 설정 (다른 Task 대기 해제) if on_task_event: await on_task_event({ "type": "task_error", "task_name": task_name, "error": str(e), "status": "failed" }) logger.error(f"[EXECUTOR] Task '{task_name}' failed: {e}") raise async def execute(self, on_task_event=None) -> Dict[str, Any]: """ 전체 Task Procedure 실행 DAG 구조에 따라 병렬 실행 가능한 Task들은 동시에 실행합니다. 각 Task는 내부적으로 wait_until 의존성을 기다린 후 실행됩니다. Args: on_task_event: Task 이벤트 콜백 (task_start, task_iteration, task_complete, task_error) Returns: 모든 Task의 결과를 담은 딕셔너리 """ logger.info("[EXECUTOR] Starting task procedure execution") # 모든 실행할 Task 수집 (BFS) all_tasks_to_run = set() in_config = self.procedure.get("IN", {}) initial_tasks = in_config.get("nexts", []) if not initial_tasks: logger.warning("[EXECUTOR] No initial tasks found in procedure") return {} queue = list(initial_tasks) while queue: task_name = queue.pop(0) if task_name == "OUT" or task_name in all_tasks_to_run: continue all_tasks_to_run.add(task_name) task_config = self.procedure.get(task_name, {}) for next_task in task_config.get("nexts", []): if next_task != "OUT": queue.append(next_task) logger.info(f"[EXECUTOR] Tasks to execute: {all_tasks_to_run}") # 모든 Task를 동시에 시작 (각 Task는 내부에서 의존성을 기다림) async def run_task_with_dependencies(task_name: str): """Task를 의존성 대기 후 실행""" try: await self._run_single_task(task_name, on_task_event) except Exception as e: logger.error(f"[EXECUTOR] Error running task '{task_name}': {e}") # 모든 Task를 병렬로 시작 - 각 Task는 _wait_for_dependencies에서 의존성 완료를 기다림 task_coroutines = [run_task_with_dependencies(task) for task in all_tasks_to_run] await asyncio.gather(*task_coroutines, return_exceptions=True) logger.info(f"[EXECUTOR] Task procedure completed. Completed tasks: {self.completed_tasks}") return self.task_results class Stage: def __init__( self, name: str, description: str, prompts: List[Dict[str, str]] = None, tools: str = None, llm_provider: str = None, llm_model: str = None, api_key: str = None, prevs: List[Any] = None, nexts: List[Any] = None, prerequisite: List[Dict[str, Any]] = None, postrequisite: List[Dict[str, Any]] = None, skip_confirm: bool = False, tasks: List[Dict[str, Any]] = None, task_procedure: Dict[str, Dict] = None ): """ Initialize a Stage. Args: name: Name of the stage. description: Description of the stage. prompts: List of prompt messages for the stage (기존 방식). tools: MCP tools configuration. llm_provider: LLM provider (e.g., "openai", "anthropic", "google"). llm_model: LLM model (e.g., "gpt-4o", "claude-3-5-sonnet-20241022", "gemini-2.0-flash-exp"). api_key: API key for the LLM provider. prevs: List of previous stages. nexts: List of next stages. prerequisite: List of MCP tool calls to execute before the stage (without LLM post-processing). postrequisite: List of MCP tool calls to execute after the stage (without LLM post-processing). skip_confirm: If True, skip confirmation and proceed to next stage automatically. tasks: List of Task configurations (새로운 Task 기반 방식). task_procedure: DAG 구조의 Task 실행 순서 정의. """ self.name = name self.description = description self.prompts = prompts or [] self.tools = tools self.llm_provider = llm_provider self.llm_model = llm_model self.api_key = api_key self.prevs = prevs or [] self.nexts = nexts or [] self.prerequisite = prerequisite or [] self.postrequisite = postrequisite or [] self.skip_confirm = skip_confirm self.run_output = "" # Task 기반 실행 설정 self.task_configs = tasks or [] self.task_procedure = task_procedure or {} self.tasks: Dict[str, Task] = {} # Task 기반 모드인지 확인 self.is_task_based = bool(self.task_configs and self.task_procedure) if self.is_task_based: # Task 기반 모드: Task 객체들 생성 self._init_tasks() self.llm = None # Task 기반 모드에서는 Stage 레벨 LLM 불필요 # Task 기반 Stage는 기본적으로 자동 confirm (skip_confirm=True) self.skip_confirm = True else: # 기존 prompts 기반 모드 if self.tools: # Increased timeout to 900s (15 min) to handle long-running Weaviate queries manager = MCPClientManager.from_dict(self.tools, timeout=900.0) else: manager = None # Initialize LLM bridge based on provider if self.llm_provider and self.api_key: if self.llm_provider == "openai": self.llm = Bridge(model=self.llm_model, api_key=self.api_key, mcp_client_manager=manager) elif self.llm_provider == "anthropic": base_url = "https://api.anthropic.com/v1/" self.llm = Bridge(model=self.llm_model, api_key=self.api_key, mcp_client_manager=manager, base_url=base_url) elif self.llm_provider == "google": base_url = "https://generativelanguage.googleapis.com/v1beta/openai/" self.llm = Bridge(model=self.llm_model, api_key=self.api_key, mcp_client_manager=manager, base_url=base_url) else: raise ValueError(f"Unknown LLM provider: {self.llm_provider}") else: self.llm = None def _init_tasks(self): """Task 설정에서 Task 객체들 생성""" # Stage 레벨에서 하나의 MCPClientManager 생성 (모든 Task가 공유) # Increased timeout to 900s (15 min) to handle long-running Weaviate queries shared_mcp_manager = MCPClientManager.from_dict(self.tools, timeout=900.0) if self.tools else None logger.info(f"[STAGE:{self.name}] Created shared MCPClientManager: {shared_mcp_manager is not None}") for task_config in self.task_configs: task = Task( task_name=task_config["task_name"], llm_provider=task_config.get("llm_provider", self.llm_provider), llm_model=task_config.get("llm_model", self.llm_model), prompts=task_config.get("prompts", []), api_key=self.api_key, mcp_client_manager=shared_mcp_manager # 공유 MCP 매니저 전달 ) # LLM 초기화 (api_key만 있으면 됨, mcp_client_manager는 optional) if self.api_key and not task.llm: task.initialize(shared_mcp_manager, self.api_key) self.tasks[task.task_name] = task logger.info(f"[STAGE:{self.name}] Initialized {len(self.tasks)} tasks: {list(self.tasks.keys())}") async def run(self, input_messages: List[Dict[str, str]] = None, on_iteration=None, on_task_event=None) -> str: """Run with iteration callback for streaming updates. Args: input_messages: Optional input messages on_iteration: Optional callback function called after each iteration with (iteration_number, content, is_complete) on_task_event: Optional callback for task-based execution events (task_start, task_iteration, task_complete, task_error) """ # Task 기반 모드인 경우 TaskProcedureExecutor 사용 if self.is_task_based: return await self._run_with_tasks(on_iteration, on_task_event) # 기존 prompts 기반 실행 start_time = time.time() timeout_reported = False TIMEOUT_SECONDS = 600 # 10분 logging.info(f"[STAGE:{self.name}] Starting stage execution (prompts-based)") logging.info(f"[STAGE:{self.name}] LLM Provider: {self.llm_provider}, Model: {self.llm_model}") logging.info(f"[STAGE:{self.name}] Input messages count: {len(input_messages) if input_messages else 0}") system_prompt = """ You are a legal expert. You must be able to provide legal assistance regarding the commands the user inputs. You cannot say you cannot do it. # Role & Objective You are an autonomous AI agent capable of utilizing MCP tools to achieve user goals step-by-step. Your primary method of operation is to act sequentially based on a checklist. # Operational Protocol You operate within a loop where you think, act (call tools), and observe results. Since there is no user input in the middle of the process, you must act autonomously to produce the best result. You can repeat the process multiple times. Do not rush to finish everything in one turn. ```python def process_messages(messages): response = self.llm_client.chat.completions.create( model=self.model, messages=messages, tools=formatted_tools, tool_choice="auto" ) messages = update_messages_from_response(response) toolcalls = response_from_toolcalls(response) for toolcall in toolcalls: tool_result = run_mcp_tool(toolcall) messages = update_messages_from_tool_result(tool_result) return messages ``` # Critical Rules for Tool Usage (Localdocs MCP) 1. **Tool Output is the Only Truth:** - Do not assume an action is successful just because you decided to do it. - You MUST read the `Tool Output` returned by the system in the next turn. - Only proceed if the tool output explicitly indicates success (e.g., "File saved successfully", "Success"). - If a tool returns an error, YOU MUST STOP and analyze the error, then retry or report it. 2. **File Saving Verification (Write-Verify Protocol):** - When saving a file, you are strictly prohibited from marking the task as complete immediately. - **Verification Step:** After calling a save function, you MUST verify the result in the next turn. - Use `list_files` or `read_file` to confirm the file physically exists on the disk. - Only after confirming the file exists via these tools can you update the checklist to `[✔]`. 3. **No Temporary Files:** - Always use `localdocs` MCP to save files permanently. # Critical Efficiency Rules for Iteration Management (MUST FOLLOW STRICTLY) 1. **Strict Iteration Limit – Target 3~4 Iterations Maximum:** - Your goal is to complete the entire checklist within 3 to 4 iterations whenever possible. - You MUST plan aggressively to resolve multiple incomplete items in each iteration. - If you exceed 4 iterations, you MUST explicitly justify in your reasoning why more iterations are unavoidable. - Never allow unnecessary repetition to push the process beyond 4 iterations. 2. **Preserve Completed Tasks Strictly:** - Once a checklist item is marked [✔] in any previous iteration, you MUST treat it as permanently complete and irreversible. - In all subsequent iterations, NEVER re-verify, re-execute, re-modify, or re-call tools for anything related to [✔] items, unless new tool output explicitly proves it is broken (extremely rare). - If a file was successfully saved and verified in a prior iteration, DO NOT read, write, or list it again unless you need to modify its content for a still-incomplete task. 3. **Cache and Reuse Previous Tool Results Aggressively:** - You MUST treat all previous tool outputs in the conversation history as cached results. - Frequently accessed files (e.g., reference laws, templates, user-uploaded documents) should NOT be read again. - If you have already obtained the full or relevant content of a file via `read_file` in any prior iteration, you MUST reuse that content from history instead of calling `read_file` again. - When reusing long file content, mentally summarize or reference only the necessary parts to avoid redundant processing. 4. **History Summary Mindset – Act as if Previous History is Summarized:** - Although the full history is provided, you MUST think and reason as if the prior iterations have been automatically summarized. - At the start of each response, mentally construct a brief internal summary of what has already been successfully completed and what key information/files are already available. - Explicitly reference this mental summary in your reasoning (e.g., "From previous iterations: draft_v1 saved and verified, reference law X content already loaded"). 5. **Focus Only on Remaining or Failed Tasks:** - Your entire reasoning, planning, and tool calls in this iteration MUST be directed EXCLUSIVELY toward: • Items still marked [ ] • Items that were previously attempted but failed (tool error or incomplete result) • Items marked [✔] only if new evidence explicitly invalidates them (you must quote the evidence) - If an incomplete item can be completed using information already present in the conversation history, do so without any new tool calls. 6. **Never Redo Successful Work:** - If previous tool output contains "Success", "File saved successfully", "File exists", or equivalent confirmation, you MUST fully trust it. - Never repeat a tool call that has already succeeded in a prior iteration. 7. **Delta-Only Reasoning Requirement:** - Always think in terms of “What has changed since the last iteration?” and “What still needs to be fixed or completed?” - In your reasoning, explicitly state: • Which items are already [✔] and why you are not touching them • Which items remain [ ] and exactly why they require action now - This focused reasoning is mandatory to minimize unnecessary processing and token usage. 8. **Progressive and Efficient Completion Strategy:** - Strive to mark as many remaining [ ] items as [✔] as possible in a single iteration, using existing history whenever possible. - Only call tools when absolutely necessary to advance a currently incomplete checklist item. - When planning tool calls, batch them efficiently to resolve multiple incomplete items at once. 9. **Self-Assessment Before Any Tool Call:** - Before proposing any tool call, you MUST ask yourself: • “Is the required information already available in previous tool outputs or history?” • “Was this task already successfully completed and marked [✔]?” • “Will this specific tool call directly help complete a currently incomplete checklist item?” - If the answer to the last question is “no”, do not call the tool. # Checklist Management 1. Evaluate whether each task is complete as you sequentially perform the tasks listed under “실행방법” 2. If a task is incomplete, perform that task until it is complete 3. Only proceed to the next task once the current task is complete 4. Display the current state of the checklist at the very TOP of every response. 5. The format must be: [✔] check list 1(complete) [ ] check list 2 ... 6. Do not skip tasks. Execute them in order. # Termination Condition - Check the checklist after every tool execution. - If and ONLY IF all items are marked `[✔]` and verified, append the keyword "**terminate**" at the very end of your response. """ messages = [{"role": "system", "content": system_prompt},] messages += list(self.prompts) logging.info(f"[STAGE:{self.name}] Total prompts count: {len(self.prompts)}") # 토큰 관리를 위한 변수 context_reset_count = 0 original_system_prompt = system_prompt # messages.append({"role": "user", "content": "Run the tasks step-by-step using the provided MCP tools."}) for i in range(8): logging.info(f"[STAGE:{self.name}] === Starting iteration {i+1}/8 ===") iteration_start_time = time.time() # Yield control to allow heartbeat and other tasks to run await asyncio.sleep(0) # 토큰 사용량 체크 및 컨텍스트 요약 current_tokens = count_messages_tokens(messages, self.llm_model) logging.info(f"[STAGE:{self.name}] Current token count: {current_tokens:,} / {TOKEN_THRESHOLD:,} threshold") if current_tokens >= TOKEN_THRESHOLD: context_reset_count += 1 logging.warning( f"[STAGE:{self.name}] Token threshold reached ({current_tokens:,} >= {TOKEN_THRESHOLD:,}). " f"Summarizing context and resetting session (reset #{context_reset_count})..." ) # Sentry에 토큰 리밋 도달 보고 sentry_sdk.capture_message( f"Stage '{self.name}' reached token threshold", level="warning", extras={ "stage_name": self.name, "current_tokens": current_tokens, "threshold": TOKEN_THRESHOLD, "iteration": i + 1, "context_reset_count": context_reset_count, } ) # 컨텍스트 요약 생성 try: summary = await summarize_context(messages, self.llm, self.name) # 새로운 세션 시작 - 시스템 프롬프트 + 요약만 유지 continuation_prompt = f"""[컨텍스트 연속 - 세션 #{context_reset_count + 1}] 이전 세션에서 진행된 작업 요약: {summary} 위 요약을 바탕으로 작업을 계속 진행해주세요. 이전에 완료된 작업은 다시 수행하지 말고, 남은 작업만 이어서 진행하세요. 체크리스트에서 [✔] 표시된 항목은 완료된 것으로 간주하세요.""" # 메시지 리셋 (원본 system_prompt와 self.prompts 유지) messages = [{"role": "system", "content": original_system_prompt}] messages += list(self.prompts) # 원본 user prompts 유지 messages.append({"role": "user", "content": continuation_prompt}) new_token_count = count_messages_tokens(messages, self.llm_model) logging.info( f"[STAGE:{self.name}] Context reset complete. " f"Tokens reduced from {current_tokens:,} to {new_token_count:,}" ) except Exception as e: logger.error(f"[STAGE:{self.name}] Failed to summarize context: {e}") # 요약 실패 시에도 계속 진행 (토큰 에러가 나면 에러 처리에서 잡힘) # 10분 타임아웃 체크 및 Sentry 보고 elapsed_time = time.time() - start_time if not timeout_reported and elapsed_time > TIMEOUT_SECONDS: timeout_reported = True # messages를 JSON 직렬화 가능한 형태로 변환 serializable_messages = [] for msg in messages: if hasattr(msg, 'model_dump'): serializable_messages.append(msg.model_dump()) elif hasattr(msg, '__dict__'): serializable_messages.append(msg.__dict__) else: serializable_messages.append(msg) sentry_sdk.capture_message( f"Stage '{self.name}' exceeded 10 minutes timeout", level="warning", extras={ "stage_name": self.name, "elapsed_seconds": elapsed_time, "iteration": i + 1, "messages": serializable_messages, "llm_provider": self.llm_provider, "llm_model": self.llm_model, } ) logging.warning(f"Stage '{self.name}' exceeded 10 minutes. Reported to Sentry.") # Run LLM call in thread pool to avoid blocking the event loop # This allows heartbeat to continue during long LLM API calls loop = asyncio.get_event_loop() logging.info(f"[STAGE:{self.name}] Iteration {i+1}: Calling LLM process_messages...") # Create a new event loop in the thread for the async LLM call current_iteration = i + 1 stage_name = self.name def run_in_thread(): logger.info(f"[THREAD] Starting thread for LLM call (stage: {stage_name}, iteration: {current_iteration})") new_loop = asyncio.new_event_loop() asyncio.set_event_loop(new_loop) try: logger.info(f"[THREAD] Running process_messages in new event loop") result = new_loop.run_until_complete(self.llm.process_messages(messages)) logger.info(f"[THREAD] process_messages completed") return result except Exception as e: # 에러 정보를 상세하게 파싱 error_info = parse_api_error(e) tb_str = traceback.format_exc() # 로깅에 상세 정보 포함 logger.error( f"[THREAD] LLM API Error in stage '{stage_name}' iteration {current_iteration}:\n" f" Error Type: {error_info['error_type']}\n" f" Error Code: {error_info['error_code']}\n" f" Message: {error_info['error_message']}\n" f" API Details: {error_info['api_error_details']}\n" f" Traceback:\n{tb_str}" ) # Sentry에 상세 에러 보고 sentry_sdk.capture_exception( e, extras={ "stage_name": stage_name, "iteration": current_iteration, "llm_provider": self.llm_provider, "llm_model": self.llm_model, "error_type": error_info['error_type'], "error_code": error_info['error_code'], "api_error_details": error_info['api_error_details'], } ) # LLMAPIError로 래핑하여 상세 정보 보존 raise LLMAPIError( message=f"LLM API 오류: {error_info['error_message']}", original_error=e, error_type=error_info['error_type'], error_code=error_info['error_code'], stage_name=stage_name, iteration=current_iteration ) from e finally: new_loop.close() result = await loop.run_in_executor(_thread_pool, run_in_thread) llm_call_duration = time.time() - iteration_start_time logging.info(f"[STAGE:{self.name}] Iteration {i+1}: LLM call completed in {llm_call_duration:.2f}s") # Yield control again after LLM call await asyncio.sleep(0) # Convert Message objects to dicts for Sentry SDK compatibility messages = [msg.model_dump() if hasattr(msg, 'model_dump') else msg for msg in result] logging.info(f"Stage '{self.name}' - Iteration {i+1} completed.") # Find the latest assistant content (may not be the last message if tool calls happened) is_complete = False content = "" # Search backwards for the latest assistant message with content for msg in reversed(messages): msg_role = msg.get('role') if isinstance(msg, dict) else getattr(msg, 'role', None) msg_content = msg.get('content') if isinstance(msg, dict) else getattr(msg, 'content', None) if msg_role == "assistant" and msg_content: content = msg_content break logging.info(f"Stage '{self.name}' - Latest assistant content length: {len(content)}") if content and content.strip().endswith("**terminate**") or content.strip().endswith("terminate"): is_complete = True # Call the iteration callback if provided if on_iteration: logging.info(f"Stage '{self.name}' - Calling on_iteration callback") await on_iteration(i + 1, content, is_complete) if is_complete: break messages.append({"role": "user", "content": "Continue the next step."}) return [messages] async def _run_with_tasks(self, on_iteration=None, on_task_event=None) -> List[Dict]: """ Task 기반 실행 모드. task_procedure에 정의된 DAG 구조에 따라 여러 Task를 병렬/순차적으로 실행합니다. Args: on_iteration: Stage 레벨 iteration 콜백 (호환성 유지) on_task_event: Task 이벤트 콜백 Returns: 모든 Task의 결과를 담은 메시지 리스트 """ logging.info(f"[STAGE:{self.name}] Starting stage execution (task-based)") logging.info(f"[STAGE:{self.name}] Tasks: {list(self.tasks.keys())}") logging.info(f"[STAGE:{self.name}] Procedure: {self.task_procedure}") # MCP tools를 미리 fetch하여 모든 Task의 Bridge에 캐싱 # 이렇게 하면 여러 Task가 동시에 list_tools를 호출하는 것을 방지 first_task = next(iter(self.tasks.values()), None) if first_task and first_task.llm and first_task.llm.mcp_client_manager: try: logging.info(f"[STAGE:{self.name}] Pre-fetching MCP tools...") prefetched_tools = await first_task.llm.fetch_tools() logging.info(f"[STAGE:{self.name}] Pre-fetched {len(prefetched_tools)} tools") # 모든 Task의 Bridge에 같은 tools 캐시 설정 for task in self.tasks.values(): if task.llm: task.llm.tools = prefetched_tools except Exception as e: logging.error(f"[STAGE:{self.name}] Failed to pre-fetch tools: {e}") # 시스템 프롬프트 (Task들이 공유) system_prompt = """ You are a legal expert. You must be able to provide legal assistance regarding the commands the user inputs. You cannot say you cannot do it. # Role & Objective You are an autonomous AI agent capable of utilizing MCP tools to achieve user goals step-by-step. Your primary method of operation is to act sequentially based on a checklist. # Operational Protocol You operate within a loop where you think, act (call tools), and observe results. Since there is no user input in the middle of the process, you must act autonomously to produce the best result. You can repeat the process multiple times. Do not rush to finish everything in one turn. # Critical Rules for Tool Usage 1. **Tool Output is the Only Truth:** - Do not assume an action is successful just because you decided to do it. - You MUST read the `Tool Output` returned by the system in the next turn. - Only proceed if the tool output explicitly indicates success. - If a tool returns an error, YOU MUST STOP and analyze the error, then retry or report it. 2. **File Saving Verification (Write-Verify Protocol):** - When saving a file, you are strictly prohibited from marking the task as complete immediately. - After calling a save function, you MUST verify the result in the next turn. # Checklist Management 1. Evaluate whether each task is complete as you sequentially perform the tasks 2. If a task is incomplete, perform that task until it is complete 3. Only proceed to the next task once the current task is complete 4. Display the current state of the checklist at the very TOP of every response. # Termination Condition - Check the checklist after every tool execution. - If and ONLY IF all items are marked `[✔]` and verified, append the keyword "**terminate**" at the very end of your response. """ # TaskProcedureExecutor 생성 및 실행 executor = TaskProcedureExecutor( tasks=self.tasks, procedure=self.task_procedure, system_prompt=system_prompt ) # 이벤트 수집용 all_events = [] completed_count = 0 total_tasks = len(self.tasks) async def collect_task_event(event): """Task 이벤트를 수집하고 콜백 호출""" nonlocal completed_count all_events.append(event) # on_task_event 콜백 호출 if on_task_event: await on_task_event(event) # on_iteration 콜백 호출 (호환성: task_complete를 iteration처럼 처리) if event["type"] == "task_complete": completed_count += 1 if on_iteration: # 모든 Task 완료 여부 확인 is_all_complete = completed_count >= total_tasks content = f"Task '{event['task_name']}' completed. ({completed_count}/{total_tasks})" await on_iteration(completed_count, content, is_all_complete) elif event["type"] == "task_error": if on_iteration: content = f"Task '{event['task_name']}' failed: {event.get('error', 'Unknown error')}" await on_iteration(completed_count, content, False) try: # Task Procedure 실행 task_results = await executor.execute(on_task_event=collect_task_event) # 모든 Task 결과를 합쳐서 반환 all_messages = [] for task_name, messages in task_results.items(): if messages: all_messages.extend(messages if isinstance(messages, list) else [messages]) logging.info(f"[STAGE:{self.name}] Task-based execution completed. {len(task_results)} tasks finished.") # 마지막 iteration 콜백 (모든 완료) if on_iteration and completed_count >= total_tasks: final_content = f"All {total_tasks} tasks completed successfully." await on_iteration(total_tasks, final_content, True) return [all_messages] if all_messages else [[]] except Exception as e: logger.error(f"[STAGE:{self.name}] Task-based execution failed: {e}") raise async def _run_stream(self, messages: Dict[str, Any]) -> AsyncIterator[str]: # process_messages(stream=True) returns a coroutine that resolves to an async iterator stream_iter = await self.llm.process_messages(messages, stream=True) async for delta in stream_iter: yield delta async def run_stream(self, input: str = "") -> AsyncIterator[str]: """Stream deltas directly via async-for. Usage: `async for delta in agent.run_stream("..."):` """ messages = list(self.prompts) if input: messages.append({"role": "user", "content": input}) async for delta in self._run_stream(messages): yield delta class Agent(): def __init__(self, name: str, description: str, version: str, Stages: List[Any], api_keys: Dict[str, str] = None): """ Initialize the agent. name: Name of the agent. description: Description of the agent. version: Version of the agent. Stages: List of Stage configurations. api_keys: Optional dictionary of API keys by provider. If not provided, uses environment variables. """ import os self.name = name self.description = description self.version = version # Use provided api_keys or fall back to environment variables if api_keys is None: api_keys = { "openai": os.environ.get("OPENAI_API_KEY"), "anthropic": os.environ.get("ANTHROPIC_API_KEY"), "google": os.environ.get("GOOGLE_API_KEY"), } self.api_keys = api_keys self.stages = [Stage(**stage, api_key=api_keys.get(stage.get("llm_provider"))) for stage in Stages] self.outputs = {} async def run(self, user_input: str = "") -> Dict[str, Any]: """Run the agent through all stages sequentially (non-streaming).""" result = None if user_input: input_messages = [{"role": "user", "content": user_input}] else: input_messages = [] for stage in self.stages: logger.info(f"Running stage: {stage.name}") try: logger.debug(f"Input messages for stage {stage.name}: {input_messages}") result = await stage.run(input_messages) input_messages = [] # Clear for next stage except Exception as e: # 에러 상세 정보 파싱 및 로깅 (백엔드 전용) if isinstance(e, LLMAPIError): error_details = e.to_dict() else: error_info = parse_api_error(e) error_details = { "message": f"Error in stage {stage.name}: {str(e)}", "error_type": error_info['error_type'], "error_code": error_info['error_code'], "stage_name": stage.name, "original_error": str(e), "traceback": traceback.format_exc(), "api_error_details": error_info['api_error_details'] } logger.error( f"Stage error in '{stage.name}':\n" f" Type: {error_details.get('error_type')}\n" f" Code: {error_details.get('error_code')}\n" f" Message: {error_details.get('message')}\n" f" Traceback: {error_details.get('traceback')}" ) return {"final_output": result, "error": True} return {"final_output": result} async def run_with_confirmation(self, user_input: str = "", start_stage_index: int = 0, cancel_event: asyncio.Event = None) -> AsyncIterator[Dict[str, Any]]: """ Run the agent through all stages with confirmation after each stage. Yields stage results and iteration updates, waits for confirmation to proceed. start_stage_index: Index of the stage to start from (default: 0) cancel_event: Optional asyncio.Event that signals cancellation when set """ if user_input: input_messages = [{"role": "user", "content": user_input}] else: input_messages = [] stage_index = start_stage_index while stage_index < len(self.stages): # Check for cancellation before starting stage if cancel_event and cancel_event.is_set(): print(f"[DEBUG] Cancellation requested before stage {stage_index}") yield { "type": "execution_stopped", "message": "Execution cancelled due to disconnection", "stage_index": stage_index } return stage = self.stages[stage_index] # Notify that stage is starting stage_start_event = { "type": "stage_start", "stage_name": stage.name, "stage_index": stage_index, "total_stages": len(self.stages), "description": stage.description, "is_task_based": stage.is_task_based } # Task 기반 Stage인 경우 Task 정보 추가 if stage.is_task_based: stage_start_event["tasks"] = list(stage.tasks.keys()) stage_start_event["task_procedure"] = stage.task_procedure yield stage_start_event try: # Run the stage with iteration streaming print(f"\n[DEBUG] Running stage: {stage.name} (task_based: {stage.is_task_based})") # Use asyncio.Queue to collect iteration events iteration_queue = asyncio.Queue() async def on_iteration(iteration_num, content, is_complete): await iteration_queue.put({ "type": "stage_iteration", "stage_name": stage.name, "stage_index": stage_index, "iteration": iteration_num, "content": content, "is_complete": is_complete }) # Task 이벤트 콜백 (Task 기반 Stage용) async def on_task_event(event): # Task 이벤트를 stage_index와 함께 큐에 추가 event_with_stage = {**event, "stage_index": stage_index, "stage_name": stage.name} await iteration_queue.put(event_with_stage) # Run stage in background task async def run_stage(): result = await stage.run( input_messages, on_iteration=on_iteration, on_task_event=on_task_event if stage.is_task_based else None ) await iteration_queue.put(None) # Signal completion return result stage_task = asyncio.create_task(run_stage()) # Yield iteration events as they come, checking for cancellation while True: # Check for cancellation if cancel_event and cancel_event.is_set(): print(f"[DEBUG] Cancellation requested during stage {stage.name}") stage_task.cancel() try: await stage_task except asyncio.CancelledError: pass yield { "type": "execution_stopped", "message": "Execution cancelled due to disconnection", "stage_index": stage_index } return try: # Use wait_for with short timeout to allow cancellation check event = await asyncio.wait_for(iteration_queue.get(), timeout=1.0) if event is None: break yield event except asyncio.TimeoutError: # Continue loop to check cancellation continue # Get the final result result = await stage_task input_messages = [] # Clear for next stage # Store output for next stages # Handle ChatCompletion object from submit_messages_without_tools if hasattr(result, 'choices') and result.choices: output_content = result.choices[0].message.content or "" elif isinstance(result, list) and len(result) > 0: # result is [[messages]], so get the inner list inner_result = result[0] if isinstance(result[0], list) else result if len(inner_result) > 0: last_msg = inner_result[-1] output_content = last_msg.content if hasattr(last_msg, 'content') else str(last_msg) else: output_content = str(result) else: output_content = str(result) if result else "" self.outputs[stage.name] = output_content print(f"[DEBUG] Stored output for stage '{stage.name}', length: {len(output_content)}") print(f"[DEBUG] self.outputs now contains: {list(self.outputs.keys())}") # Check if skip_confirm is enabled for this stage if stage.skip_confirm: # Yield stage completion without awaiting confirmation yield { "type": "stage_complete", "stage_name": stage.name, "stage_index": stage_index, "total_stages": len(self.stages), "result": output_content, "awaiting_confirmation": False, "skip_confirm": True } # Auto-proceed to next stage stage_index += 1 continue else: # Yield stage completion with result yield { "type": "stage_complete", "stage_name": stage.name, "stage_index": stage_index, "total_stages": len(self.stages), "result": output_content, "awaiting_confirmation": True } # Wait for confirmation (handled by the caller) # The caller should send back a confirmation response break # Break here to wait for external confirmation except Exception as e: # 에러 상세 정보 파싱 및 로깅 (백엔드 전용) if isinstance(e, LLMAPIError): error_details = e.to_dict() else: error_info = parse_api_error(e) error_details = { "message": f"Error in stage {stage.name}: {str(e)}", "error_type": error_info['error_type'], "error_code": error_info['error_code'], "stage_name": stage.name, "stage_index": stage_index, "original_error": str(e), "traceback": traceback.format_exc(), "api_error_details": error_info['api_error_details'] } logger.error( f"Stage error in '{stage.name}':\n" f" Type: {error_details.get('error_type')}\n" f" Code: {error_details.get('error_code')}\n" f" Message: {error_details.get('message')}\n" f" Traceback: {error_details.get('traceback')}" ) yield { "type": "stage_error", "stage_name": stage.name, "stage_index": stage_index, "error": error_details.get('message'), "awaiting_confirmation": True } break # If all stages completed via skip_confirm (loop exited normally without break) if stage_index >= len(self.stages): yield { "type": "execution_complete", "message": "All stages completed successfully", "final_outputs": self.outputs } async def continue_from_stage(self, stage_index: int, action: str, cancel_event: asyncio.Event = None) -> AsyncIterator[Dict[str, Any]]: """ Continue execution from a specific stage based on user action. action: 'confirm' to proceed, 'retry' to re-run current stage, 'deny' to stop cancel_event: Optional asyncio.Event that signals cancellation when set """ if action == "deny": yield { "type": "execution_stopped", "message": "Execution stopped by user", "stage_index": stage_index } return if action == "retry": # Check for cancellation if cancel_event and cancel_event.is_set(): print(f"[DEBUG] Cancellation requested before retry") yield { "type": "execution_stopped", "message": "Execution cancelled due to disconnection", "stage_index": stage_index } return # Re-run the same stage stage = self.stages[stage_index] yield { "type": "stage_retry", "stage_name": stage.name, "stage_index": stage_index } try: # Run with iteration streaming iteration_queue = asyncio.Queue() async def on_iteration(iteration_num, content, is_complete): await iteration_queue.put({ "type": "stage_iteration", "stage_name": stage.name, "stage_index": stage_index, "iteration": iteration_num, "content": content, "is_complete": is_complete }) # Task 이벤트 콜백 (Task 기반 Stage용) async def on_task_event(event): event_with_stage = {**event, "stage_index": stage_index, "stage_name": stage.name} await iteration_queue.put(event_with_stage) async def run_stage(): result = await stage.run( [], on_iteration=on_iteration, on_task_event=on_task_event if stage.is_task_based else None ) await iteration_queue.put(None) return result stage_task = asyncio.create_task(run_stage()) while True: # Check for cancellation if cancel_event and cancel_event.is_set(): print(f"[DEBUG] Cancellation requested during retry of stage {stage.name}") stage_task.cancel() try: await stage_task except asyncio.CancelledError: pass yield { "type": "execution_stopped", "message": "Execution cancelled due to disconnection", "stage_index": stage_index } return try: event = await asyncio.wait_for(iteration_queue.get(), timeout=1.0) if event is None: break yield event except asyncio.TimeoutError: continue result = await stage_task # Handle result if hasattr(result, 'choices') and result.choices: output_content = result.choices[0].message.content or "" elif isinstance(result, list) and len(result) > 0: inner_result = result[0] if isinstance(result[0], list) else result if len(inner_result) > 0: last_msg = inner_result[-1] output_content = last_msg.content if hasattr(last_msg, 'content') else str(last_msg) else: output_content = str(result) else: output_content = str(result) if result else "" self.outputs[stage.name] = output_content yield { "type": "stage_complete", "stage_name": stage.name, "stage_index": stage_index, "total_stages": len(self.stages), "result": output_content, "awaiting_confirmation": True } except Exception as e: # 에러 상세 정보 파싱 및 로깅 (백엔드 전용) if isinstance(e, LLMAPIError): error_details = e.to_dict() else: error_info = parse_api_error(e) error_details = { "message": f"Error in stage {stage.name}: {str(e)}", "error_type": error_info['error_type'], "error_code": error_info['error_code'], "stage_name": stage.name, "stage_index": stage_index, "original_error": str(e), "traceback": traceback.format_exc(), "api_error_details": error_info['api_error_details'] } logger.error( f"Stage retry error in '{stage.name}':\n" f" Type: {error_details.get('error_type')}\n" f" Code: {error_details.get('error_code')}\n" f" Message: {error_details.get('message')}" ) yield { "type": "stage_error", "stage_name": stage.name, "stage_index": stage_index, "error": error_details.get('message'), "awaiting_confirmation": True } elif action == "confirm": # Check for cancellation if cancel_event and cancel_event.is_set(): print(f"[DEBUG] Cancellation requested before confirm") yield { "type": "execution_stopped", "message": "Execution cancelled due to disconnection", "stage_index": stage_index } return # Move to next stage next_stage_index = stage_index + 1 if next_stage_index >= len(self.stages): # All stages completed yield { "type": "execution_complete", "message": "All stages completed successfully", "final_outputs": self.outputs } return # Run next stage stage = self.stages[next_stage_index] # Notify that stage is starting stage_start_event = { "type": "stage_start", "stage_name": stage.name, "stage_index": next_stage_index, "total_stages": len(self.stages), "description": stage.description, "is_task_based": stage.is_task_based } # Task 기반 Stage인 경우 Task 정보 추가 if stage.is_task_based: stage_start_event["tasks"] = list(stage.tasks.keys()) stage_start_event["task_procedure"] = stage.task_procedure yield stage_start_event try: # Run with iteration streaming iteration_queue = asyncio.Queue() async def on_iteration(iteration_num, content, is_complete): await iteration_queue.put({ "type": "stage_iteration", "stage_name": stage.name, "stage_index": next_stage_index, "iteration": iteration_num, "content": content, "is_complete": is_complete }) # Task 이벤트 콜백 (Task 기반 Stage용) async def on_task_event(event): event_with_stage = {**event, "stage_index": next_stage_index, "stage_name": stage.name} await iteration_queue.put(event_with_stage) async def run_stage(): result = await stage.run( [], on_iteration=on_iteration, on_task_event=on_task_event if stage.is_task_based else None ) await iteration_queue.put(None) return result stage_task = asyncio.create_task(run_stage()) while True: # Check for cancellation if cancel_event and cancel_event.is_set(): print(f"[DEBUG] Cancellation requested during stage {stage.name}") stage_task.cancel() try: await stage_task except asyncio.CancelledError: pass yield { "type": "execution_stopped", "message": "Execution cancelled due to disconnection", "stage_index": next_stage_index } return try: event = await asyncio.wait_for(iteration_queue.get(), timeout=1.0) if event is None: break yield event except asyncio.TimeoutError: continue result = await stage_task # Handle result if hasattr(result, 'choices') and result.choices: output_content = result.choices[0].message.content or "" elif isinstance(result, list) and len(result) > 0: inner_result = result[0] if isinstance(result[0], list) else result if len(inner_result) > 0: last_msg = inner_result[-1] output_content = last_msg.content if hasattr(last_msg, 'content') else str(last_msg) else: output_content = str(result) else: output_content = str(result) if result else "" self.outputs[stage.name] = output_content yield { "type": "stage_complete", "stage_name": stage.name, "stage_index": next_stage_index, "total_stages": len(self.stages), "result": output_content, "awaiting_confirmation": True } except Exception as e: # 에러 상세 정보 파싱 및 로깅 (백엔드 전용) if isinstance(e, LLMAPIError): error_details = e.to_dict() else: error_info = parse_api_error(e) error_details = { "message": f"Error in stage {stage.name}: {str(e)}", "error_type": error_info['error_type'], "error_code": error_info['error_code'], "stage_name": stage.name, "stage_index": next_stage_index, "original_error": str(e), "traceback": traceback.format_exc(), "api_error_details": error_info['api_error_details'] } logger.error( f"Stage confirm error in '{stage.name}':\n" f" Type: {error_details.get('error_type')}\n" f" Code: {error_details.get('error_code')}\n" f" Message: {error_details.get('message')}" ) yield { "type": "stage_error", "stage_name": stage.name, "stage_index": next_stage_index, "error": error_details.get('message'), "awaiting_confirmation": True } from fastapi import FastAPI, HTTPException, UploadFile, File, WebSocket, WebSocketDisconnect, Request from fastapi.responses import JSONResponse, StreamingResponse, FileResponse from fastapi.middleware.cors import CORSMiddleware from pydantic import BaseModel from typing import Dict, Any, Optional from sse_starlette.sse import EventSourceResponse import yaml import asyncio import os from src.agent import Agent import traceback import json import uuid import sentry_sdk sentry_sdk.init( dsn="https://fc33f2349050895a6cf1e2153a5d5f2c@sentry.eroomai.com/2", send_default_pii=True, traces_sample_rate=1.0, ) app = FastAPI(title="Agent Backend API", version="1.0.0") # Load API keys from environment variables API_KEYS = { "openai": os.environ.get("OPENAI_API_KEY"), "anthropic": os.environ.get("ANTHROPIC_API_KEY"), "google": os.environ.get("GOOGLE_API_KEY"), } # CORS 설정 추가 app.add_middleware( CORSMiddleware, allow_origins=["*"], # 프로덕션에서는 특정 도메인으로 제한 allow_credentials=True, allow_methods=["*"], allow_headers=["*"], ) # Store loaded agents in memory agents: Dict[str, Agent] = {} # Store active agent sessions (for stateful execution) active_sessions: Dict[str, Dict[str, Any]] = {} class AgentRequest(BaseModel): """Request body for agent execution""" user_input: str stream: bool = False class AgentResponse(BaseModel): """Response from agent execution""" agent_name: str final_output: Any success: bool error: Optional[str] = None @app.get("/") async def root(): """Root endpoint""" return { "message": "Agent Backend API Server", "version": "1.0.0", "agents": list(agents.keys()) } @app.get("/agents") async def list_agents(): """List all loaded agents""" return { "agents": [ { "name": agent.name, "description": agent.description, "stages": len(agent.stages) } for agent in agents.values() ] } @app.post("/upload-agent") async def upload_agent(file: UploadFile = File(...)): """ Upload a YAML agent script and create dynamic API endpoint. The endpoint will be created at /agent/{Agent.name} """ try: # Read and parse YAML file content = await file.read() yaml_content = yaml.safe_load(content.decode('utf-8')) # Validate YAML structure if "Agent" not in yaml_content: raise HTTPException( status_code=400, detail="Invalid YAML structure. Must contain 'Agent' key." ) agent_config = yaml_content["Agent"] # Validate required fields if "name" not in agent_config: raise HTTPException( status_code=400, detail="Agent configuration must contain 'name' field." ) agent_name = agent_config["name"] # Remove api_keys from YAML config if present (Agent will use environment variables) agent_config["api_keys"] = None # Create Agent instance try: agent = Agent(**agent_config) agents[agent_name] = agent return { "success": True, "message": f"Agent '{agent_name}' uploaded successfully", "agent_name": agent_name, "endpoint": f"/agent/{agent_name}", "description": agent.description, "stages": len(agent.stages) } except Exception as e: raise HTTPException( status_code=400, detail=f"Failed to create agent: {str(e)}" ) except yaml.YAMLError as e: raise HTTPException( status_code=400, detail=f"Invalid YAML format: {str(e)}" ) except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Internal server error: {str(e)}" ) @app.delete("/agent/{agent_name}") async def delete_agent(agent_name: str): """Delete a loaded agent""" if agent_name not in agents: raise HTTPException( status_code=404, detail=f"Agent '{agent_name}' not found" ) del agents[agent_name] return { "success": True, "message": f"Agent '{agent_name}' deleted successfully" } @app.post("/agent/{agent_name}") async def execute_agent(agent_name: str, request: AgentRequest): """ Execute a specific agent by name. Dynamically created endpoint based on uploaded YAML configuration. """ # Check if agent exists if agent_name not in agents: raise HTTPException( status_code=404, detail=f"Agent '{agent_name}' not found. Please upload the agent configuration first." ) agent = agents[agent_name] try: # Execute agent with streaming if requested if request.stream: async def generate_stream(): """Stream agent responses""" try: # Note: This assumes agent has streaming capability # You may need to implement streaming in the Agent class result = await agent.run(user_input=request.user_input) # For now, just yield the final result # In future, implement proper streaming if agent supports it yield json.dumps({ "agent_name": agent_name, "final_output": result, "success": True }) except Exception as e: yield json.dumps({ "agent_name": agent_name, "success": False, "error": str(e) }) return StreamingResponse( generate_stream(), media_type="application/x-ndjson" ) else: # Non-streaming execution result = await agent.run(user_input=request.user_input) return AgentResponse( agent_name=agent_name, final_output=result, success=True ) except Exception as e: traceback.print_exc() return AgentResponse( agent_name=agent_name, final_output=None, success=False, error=str(e) ) @app.get("/agent/{agent_name}/info") async def get_agent_info(agent_name: str): """Get information about a specific agent""" if agent_name not in agents: raise HTTPException( status_code=404, detail=f"Agent '{agent_name}' not found" ) agent = agents[agent_name] return { "name": agent.name, "description": agent.description, "stages": [ { "name": stage.name, "description": stage.description, "llm_provider": stage.llm_provider, "llm_model": stage.llm_model, "prevs": stage.prevs, "nexts": stage.nexts } for stage in agent.stages ] } @app.websocket("/ws/agent/{agent_name}") async def websocket_agent_endpoint(websocket: WebSocket, agent_name: str): """ WebSocket endpoint for interactive agent execution with stage-by-stage confirmation. Protocol: 1. Client connects and sends: {"action": "start", "user_input": "..."} 2. Server runs first stage and sends: {"type": "stage_complete", "stage_name": "...", "result": "...", ...} 3. Client responds with: {"action": "confirm|retry|deny", "stage_index": N} 4. Server continues based on action """ print(f"[WS] New WebSocket connection attempt for agent: {agent_name}") await websocket.accept() print(f"[WS] WebSocket accepted") # Check if agent exists if agent_name not in agents: print(f"[WS] Agent '{agent_name}' not found in agents: {list(agents.keys())}") await websocket.send_json({ "type": "error", "error": f"Agent '{agent_name}' not found" }) await websocket.close() return print(f"[WS] Agent '{agent_name}' found") agent = agents[agent_name] session_id = str(uuid.uuid4()) user_input = "" # Lock for WebSocket sends to prevent concurrent writes ws_lock = asyncio.Lock() ws_closed = asyncio.Event() async def safe_send(data: dict) -> bool: """Thread-safe WebSocket send with error handling""" if ws_closed.is_set(): return False try: async with ws_lock: await websocket.send_json(data) return True except Exception as e: print(f"[WS] Send error: {e}") ws_closed.set() return False # Heartbeat task to keep connection alive heartbeat_task = None stop_heartbeat = asyncio.Event() # Queue for incoming client messages message_queue = asyncio.Queue() receiver_task = None async def send_heartbeat(): """Send periodic heartbeat to keep connection alive""" while not stop_heartbeat.is_set() and not ws_closed.is_set(): try: await asyncio.sleep(1.5) # Send heartbeat every 1.5 seconds (Safari is VERY aggressive) if not stop_heartbeat.is_set() and not ws_closed.is_set(): success = await safe_send({"type": "heartbeat"}) if not success: print("[WS] Heartbeat failed, stopping") break except asyncio.CancelledError: break except Exception as e: print(f"[WS] Heartbeat error: {e}") break async def receive_messages(): """Continuously receive messages from client and put them in queue""" while not ws_closed.is_set(): try: data = await websocket.receive_json() await message_queue.put(data) except Exception as e: print(f"[WS] Receiver error: {e}") ws_closed.set() break try: print(f"[WS] WebSocket connected for agent: {agent_name}") print(f"[WS] Starting message loop...") # Start heartbeat task heartbeat_task = asyncio.create_task(send_heartbeat()) # Start receiver task to continuously receive messages receiver_task = asyncio.create_task(receive_messages()) while not ws_closed.is_set(): # Get message from queue with timeout print("[WS] Waiting for message from client...") try: data = await asyncio.wait_for(message_queue.get(), timeout=600) print(f"[WS] Received data: {data}") action = data.get("action") print(f"[WS] Action: {action}") except asyncio.TimeoutError: print("[WS] Receive timeout, sending ping") if not await safe_send({"type": "ping"}): break continue if action == "start": # Start new execution user_input = data.get("user_input", "") start_stage_index = data.get("start_stage_index", 0) # Allow starting from specific stage # Create new session active_sessions[session_id] = { "agent_name": agent_name, "agent": agent, "user_input": user_input, "current_stage": start_stage_index } if not await safe_send({ "type": "session_started", "session_id": session_id, "agent_name": agent_name, "total_stages": len(agent.stages), "start_stage_index": start_stage_index }): break # Run from specified stage with cancellation support async for event in agent.run_with_confirmation(user_input=user_input, start_stage_index=start_stage_index, cancel_event=ws_closed): if not await safe_send(event): print("[WS] Failed to send event, breaking loop") break # Allow heartbeat to run between events await asyncio.sleep(0) elif action in ["confirm", "retry", "deny"]: # Continue from current stage stage_index = data.get("stage_index", 0) if session_id not in active_sessions: if not await safe_send({ "type": "error", "error": "No active session. Please start a new execution." }): break continue session = active_sessions[session_id] agent = session["agent"] # Continue execution based on action with cancellation support async for event in agent.continue_from_stage( stage_index=stage_index, action=action, cancel_event=ws_closed ): if not await safe_send(event): print("[WS] Failed to send event, breaking loop") break # Allow heartbeat to run between events await asyncio.sleep(0) # Update session stage index if moving forward if action == "confirm" and event.get("type") == "stage_complete": session["current_stage"] = event.get("stage_index", 0) # Clean up session if execution is complete or stopped if event.get("type") in ["execution_complete", "execution_stopped"]: if session_id in active_sessions: del active_sessions[session_id] break elif action == "pong": # Client responded to ping, connection is alive print("[WS] Received pong from client") continue else: if not await safe_send({ "type": "error", "error": f"Unknown action: {action}" }): break except WebSocketDisconnect: print(f"WebSocket disconnected for agent: {agent_name}") ws_closed.set() if session_id in active_sessions: del active_sessions[session_id] except Exception as e: traceback.print_exc() ws_closed.set() try: await safe_send({ "type": "error", "error": str(e) }) await websocket.close() except: pass finally: # Stop all background tasks ws_closed.set() stop_heartbeat.set() if heartbeat_task: heartbeat_task.cancel() try: await heartbeat_task except asyncio.CancelledError: pass if receiver_task: receiver_task.cancel() try: await receiver_task except asyncio.CancelledError: pass @app.get("/sessions") async def list_active_sessions(): """List all active agent execution sessions""" return { "active_sessions": [ { "session_id": sid, "agent_name": session["agent_name"], "current_stage": session["current_stage"] } for sid, session in active_sessions.items() ] } # SSE session state storage sse_sessions: Dict[str, Dict[str, Any]] = {} @app.post("/sse/agent/{agent_name}/start") async def sse_start_agent(agent_name: str, request: Request): """ SSE endpoint for interactive agent execution with stage-by-stage confirmation. This endpoint starts the agent and streams events via Server-Sent Events. Client sends actions via separate POST endpoints. Request body: { "user_input": "...", "start_stage_index": 0 // optional } """ if agent_name not in agents: raise HTTPException( status_code=404, detail=f"Agent '{agent_name}' not found" ) agent = agents[agent_name] try: body = await request.json() except: body = {} user_input = body.get("user_input", "") start_stage_index = body.get("start_stage_index", 0) session_id = str(uuid.uuid4()) cancel_event = asyncio.Event() # Store session info sse_sessions[session_id] = { "agent_name": agent_name, "agent": agent, "user_input": user_input, "current_stage": start_stage_index, "cancel_event": cancel_event, "action_queue": asyncio.Queue(), "waiting_for_action": False } async def event_generator(): session = sse_sessions.get(session_id) if not session: yield { "event": "error", "data": json.dumps({"type": "error", "error": "Session not found"}) } return try: # Send session started event yield { "event": "message", "data": json.dumps({ "type": "session_started", "session_id": session_id, "agent_name": agent_name, "total_stages": len(agent.stages), "start_stage_index": start_stage_index }) } current_stage_index = start_stage_index current_action = "start" # Initial action to start first stage execution_complete = False while not execution_complete: # Determine which generator to use based on current action if current_action == "start": # Run first stage event_gen = agent.run_with_confirmation( user_input=user_input, start_stage_index=current_stage_index, cancel_event=cancel_event ) else: # Continue from stage with action event_gen = agent.continue_from_stage( stage_index=current_stage_index, action=current_action, cancel_event=cancel_event ) # Process events from the generator async for event in event_gen: yield { "event": "message", "data": json.dumps(event) } event_type = event.get("type") # Update current stage from event if "stage_index" in event: current_stage_index = event.get("stage_index") session["current_stage"] = current_stage_index # If stage is complete, wait for client action if event_type == "stage_complete": session["waiting_for_action"] = True # Wait for action from client try: action_data = await asyncio.wait_for( session["action_queue"].get(), timeout=600 # 10 minute timeout ) session["waiting_for_action"] = False current_action = action_data.get("action") if "stage_index" in action_data: current_stage_index = action_data.get("stage_index") if current_action == "deny": yield { "event": "message", "data": json.dumps({ "type": "execution_stopped", "reason": "User denied continuation", "stage_index": current_stage_index }) } execution_complete = True break # For "confirm" or "retry", the while loop will continue # and call continue_from_stage with the new action break # Break inner loop to start next iteration except asyncio.TimeoutError: yield { "event": "message", "data": json.dumps({ "type": "error", "error": "Action timeout - no response from client" }) } execution_complete = True break # Check if execution is complete or stopped if event_type in ["execution_complete", "execution_stopped"]: execution_complete = True break except asyncio.CancelledError: yield { "event": "message", "data": json.dumps({ "type": "execution_stopped", "reason": "Cancelled by client" }) } except Exception as e: traceback.print_exc() yield { "event": "message", "data": json.dumps({ "type": "error", "error": str(e) }) } finally: # Cleanup session if session_id in sse_sessions: del sse_sessions[session_id] return EventSourceResponse(event_generator()) @app.post("/sse/agent/{agent_name}/action/{session_id}") async def sse_agent_action(agent_name: str, session_id: str, request: Request): """ Send an action to an active SSE agent session. Request body: { "action": "confirm" | "retry" | "deny", "stage_index": 0 // optional } """ if session_id not in sse_sessions: raise HTTPException( status_code=404, detail=f"Session '{session_id}' not found" ) session = sse_sessions[session_id] if session["agent_name"] != agent_name: raise HTTPException( status_code=400, detail=f"Session belongs to different agent" ) try: body = await request.json() except: raise HTTPException(status_code=400, detail="Invalid JSON body") action = body.get("action") if action not in ["confirm", "retry", "deny"]: raise HTTPException( status_code=400, detail="Action must be 'confirm', 'retry', or 'deny'" ) # Put action in the queue for the event generator to process await session["action_queue"].put(body) return {"success": True, "action": action} @app.post("/sse/agent/{agent_name}/cancel/{session_id}") async def sse_cancel_agent(agent_name: str, session_id: str): """Cancel an active SSE agent session.""" if session_id not in sse_sessions: raise HTTPException( status_code=404, detail=f"Session '{session_id}' not found" ) session = sse_sessions[session_id] if session["agent_name"] != agent_name: raise HTTPException( status_code=400, detail=f"Session belongs to different agent" ) # Signal cancellation session["cancel_event"].set() return {"success": True, "message": "Session cancellation requested"} # Document parsing extensions PARSABLE_EXTENSIONS = {'.pdf', '.png', '.jpg', '.jpeg', '.gif', '.webp', '.bmp'} MIME_TYPES = { '.pdf': 'application/pdf', '.png': 'image/png', '.jpg': 'image/jpeg', '.jpeg': 'image/jpeg', '.gif': 'image/gif', '.webp': 'image/webp', '.bmp': 'image/bmp', } async def call_docparser(content_base64: str, mime_type: str, filename: str) -> dict: """Call mcp_docparser to parse PDF/image document via REST API.""" import httpx import os # Use the direct REST API endpoint instead of MCP protocol docparser_base = os.environ.get("DOCPARSER_URL", "http://localhost:8015") docparser_url = f"{docparser_base}/api/parse" # Set explicit timeout for all operations (Gemini parsing can take a while for large files) timeout = httpx.Timeout( connect=60.0, # 1 minute to connect read=1800.0, # 30 minutes to read response (Gemini processing for large PDFs) write=300.0, # 5 minutes to write request (large base64 payloads) pool=60.0 # 1 minute to get connection from pool ) async with httpx.AsyncClient(timeout=timeout) as http_client: response = await http_client.post( docparser_url, json={ "content_base64": content_base64, "mime_type": mime_type, "filename": filename, "output_format": "json" }, headers={"Content-Type": "application/json"} ) return response.json() def merge_evidence_json_files(evidence_dir: str) -> dict: """ Merge all JSON files in the evidence directory into evidence_docs.json. Args: evidence_dir: Path to the evidence directory Returns: dict with success status and merged file info """ import os import json import glob try: # Find all JSON files in evidence directory (excluding evidence_docs.json itself) json_pattern = os.path.join(evidence_dir, "*.json") json_files = [f for f in glob.glob(json_pattern) if os.path.basename(f) != "evidence_docs.json"] if not json_files: return {"success": True, "message": "No JSON files to merge", "count": 0} merged_docs = [] errors = [] for json_file in sorted(json_files): try: with open(json_file, 'r', encoding='utf-8') as f: content = json.load(f) # Add source filename to the document doc_entry = { "source_file": os.path.basename(json_file), "content": content } merged_docs.append(doc_entry) print(f"[EvidenceMerge] Added {os.path.basename(json_file)}") except json.JSONDecodeError as e: error_msg = f"Failed to parse {os.path.basename(json_file)}: {str(e)}" errors.append(error_msg) print(f"[EvidenceMerge] {error_msg}") except Exception as e: error_msg = f"Error reading {os.path.basename(json_file)}: {str(e)}" errors.append(error_msg) print(f"[EvidenceMerge] {error_msg}") # Write merged file output_path = os.path.join(evidence_dir, "evidence_docs.json") with open(output_path, 'w', encoding='utf-8') as f: json.dump({ "documents": merged_docs, "total_count": len(merged_docs), "source_files": [os.path.basename(f) for f in json_files] }, f, ensure_ascii=False, indent=2) print(f"[EvidenceMerge] Merged {len(merged_docs)} files into evidence_docs.json") return { "success": True, "message": f"Merged {len(merged_docs)} JSON files into evidence_docs.json", "count": len(merged_docs), "errors": errors if errors else None } except Exception as e: error_msg = f"Failed to merge evidence files: {str(e)}" print(f"[EvidenceMerge] {error_msg}") return {"success": False, "error": error_msg} def merge_client_meeting_json_files(client_meeting_dir: str) -> dict: """ Merge all JSON files in the client_meeting directory into client_meeting.json. Args: client_meeting_dir: Path to the client_meeting directory Returns: dict with success status and merged file info """ import os import json import glob try: # Find all JSON files in client_meeting directory (excluding client_meeting.json itself) json_pattern = os.path.join(client_meeting_dir, "*.json") json_files = [f for f in glob.glob(json_pattern) if os.path.basename(f) != "client_meeting.json"] if not json_files: return {"success": True, "message": "No JSON files to merge", "count": 0} merged_docs = [] errors = [] for json_file in sorted(json_files): try: with open(json_file, 'r', encoding='utf-8') as f: content = json.load(f) # Add source filename to the document doc_entry = { "source_file": os.path.basename(json_file), "content": content } merged_docs.append(doc_entry) print(f"[ClientMeetingMerge] Added {os.path.basename(json_file)}") except json.JSONDecodeError as e: error_msg = f"Failed to parse {os.path.basename(json_file)}: {str(e)}" errors.append(error_msg) print(f"[ClientMeetingMerge] {error_msg}") except Exception as e: error_msg = f"Error reading {os.path.basename(json_file)}: {str(e)}" errors.append(error_msg) print(f"[ClientMeetingMerge] {error_msg}") # Write merged file output_path = os.path.join(client_meeting_dir, "client_meeting.json") with open(output_path, 'w', encoding='utf-8') as f: json.dump({ "documents": merged_docs, "total_count": len(merged_docs), "source_files": [os.path.basename(f) for f in json_files] }, f, ensure_ascii=False, indent=2) print(f"[ClientMeetingMerge] Merged {len(merged_docs)} files into client_meeting.json") return { "success": True, "message": f"Merged {len(merged_docs)} JSON files into client_meeting.json", "count": len(merged_docs), "errors": errors if errors else None } except Exception as e: error_msg = f"Failed to merge client_meeting files: {str(e)}" print(f"[ClientMeetingMerge] {error_msg}") return {"success": False, "error": error_msg} @app.post("/upload-file") async def upload_file_to_localdocs(file: UploadFile = File(...), parse: bool = True, folder: str = ""): """ Upload a file to the mcp-localdocs directory. Files will be saved to /mcp-localdocs/ (or MCP_LOCALDOCS_PATH env var) Args: file: The file to upload parse: If True, automatically parse PDFs and images using mcp_docparser folder: Optional subfolder path to upload to (e.g., "projects/2024") If the file is a PDF or image and parse=True, it will be automatically parsed using mcp_docparser and the result saved as a .json file. """ import os import base64 # Use environment variable or default to Docker mount path base_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") target_dir = os.path.join(base_dir, folder) if folder else base_dir # Security check - ensure path is within base_dir real_base = os.path.realpath(base_dir) real_target = os.path.realpath(target_dir) if not real_target.startswith(real_base): raise HTTPException( status_code=403, detail="Access denied - path outside allowed directory" ) try: # Create directory if it doesn't exist os.makedirs(target_dir, exist_ok=True) # Read file content file_content = await file.read() # Save original file file_path = os.path.join(target_dir, file.filename) with open(file_path, "wb") as buffer: buffer.write(file_content) result = { "success": True, "message": f"File '{file.filename}' uploaded successfully", "filename": file.filename, "path": file_path, "size": len(file_content) } # Check if file should be parsed file_ext = os.path.splitext(file.filename)[1].lower() if parse and file_ext in PARSABLE_EXTENSIONS: try: # Get mime type mime_type = MIME_TYPES.get(file_ext, 'application/octet-stream') # Encode content to base64 content_base64 = base64.b64encode(file_content).decode('utf-8') # Call docparser print(f"[DocParser] Parsing {file.filename} ({mime_type})...") parse_result = await call_docparser(content_base64, mime_type, file.filename) if "error" in parse_result: result["parse_error"] = f"DocParser error: {parse_result['error']}" print(f"[DocParser] Error: {parse_result['error']}") elif parse_result.get("success"): # New REST API format: {"success": true, "html": "...", "json": "...", "filename": "..."} html_content = parse_result.get("html", "") json_content = parse_result.get("json", "") base_filename = os.path.splitext(file.filename)[0] saved_files = [] # Save HTML result if html_content: html_filename = base_filename + ".html" html_path = os.path.join(target_dir, html_filename) with open(html_path, "w", encoding="utf-8") as f: f.write(html_content) result["html_filename"] = html_filename saved_files.append(html_filename) print(f"[DocParser] Saved HTML to {html_filename}") # Save JSON result if json_content: json_filename = base_filename + ".json" json_path = os.path.join(target_dir, json_filename) with open(json_path, "w", encoding="utf-8") as f: f.write(json_content) result["json_filename"] = json_filename saved_files.append(json_filename) print(f"[DocParser] Saved JSON to {json_filename}") if saved_files: result["parsed"] = True result["message"] = f"File '{file.filename}' uploaded and parsed successfully" print(f"[DocParser] Successfully parsed and saved: {', '.join(saved_files)}") # Auto-merge JSON files if uploaded to evidence folder if folder.rstrip("/") == "evidence" and json_content: merge_result = merge_evidence_json_files(target_dir) result["evidence_merge"] = merge_result print(f"[EvidenceMerge] Auto-merge result: {merge_result.get('message', merge_result.get('error'))}") # Auto-merge JSON files if uploaded to client_meeting folder if folder.rstrip("/") == "client_meeting" and json_content: merge_result = merge_client_meeting_json_files(target_dir) result["client_meeting_merge"] = merge_result print(f"[ClientMeetingMerge] Auto-merge result: {merge_result.get('message', merge_result.get('error'))}") else: result["parse_error"] = "No content returned from DocParser" else: result["parse_error"] = "Unexpected response from DocParser" except Exception as e: traceback.print_exc() result["parse_error"] = f"Failed to parse document: {str(e)}" print(f"[DocParser] Exception: {str(e)}") return result except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to upload file: {str(e)}" ) @app.post("/evidence/merge") async def merge_evidence_files(): """ Manually trigger merging of all JSON files in the evidence folder into evidence_docs.json. """ import os base_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") evidence_dir = os.path.join(base_dir, "evidence") if not os.path.exists(evidence_dir): raise HTTPException( status_code=404, detail="Evidence folder not found" ) result = merge_evidence_json_files(evidence_dir) if not result.get("success"): raise HTTPException( status_code=500, detail=result.get("error", "Failed to merge evidence files") ) return result @app.post("/client_meeting/merge") async def merge_client_meeting_files(): """ Manually trigger merging of all JSON files in the client_meeting folder into client_meeting.json. """ import os base_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") client_meeting_dir = os.path.join(base_dir, "client_meeting") if not os.path.exists(client_meeting_dir): raise HTTPException( status_code=404, detail="client_meeting folder not found" ) result = merge_client_meeting_json_files(client_meeting_dir) if not result.get("success"): raise HTTPException( status_code=500, detail=result.get("error", "Failed to merge client_meeting files") ) return result @app.get("/folders") async def list_folders(path: str = "", recursive: bool = False): """ List all folders in the mcp-localdocs directory. """ import os # Use environment variable or default to Docker mount path base_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") target_dir = os.path.join(base_dir, path) if path else base_dir try: # Security check - ensure path is within base_dir real_base = os.path.realpath(base_dir) real_target = os.path.realpath(target_dir) if not real_target.startswith(real_base): raise HTTPException( status_code=403, detail="Access denied - path outside allowed directory" ) if not os.path.exists(target_dir): return {"folders": [], "current_path": path or "/"} def is_hidden(name: str) -> bool: """Check if a file/folder is hidden (starts with .)""" return name.startswith('.') def get_folder_info(folder_path: str, rel_path: str) -> dict: """Get information about a folder.""" files = [f for f in os.listdir(folder_path) if os.path.isfile(os.path.join(folder_path, f)) and not is_hidden(f)] subdirs = [d for d in os.listdir(folder_path) if os.path.isdir(os.path.join(folder_path, d)) and not is_hidden(d)] return { "name": os.path.basename(folder_path) or "root", "path": rel_path, "file_count": len(files), "folder_count": len(subdirs), } def list_folders_recursive(folder_path: str, rel_path: str) -> list: """Recursively list all folders.""" result = [] try: for item in sorted(os.listdir(folder_path)): # Skip hidden folders if is_hidden(item): continue item_path = os.path.join(folder_path, item) if os.path.isdir(item_path): item_rel_path = os.path.join(rel_path, item) if rel_path else item folder_info = get_folder_info(item_path, item_rel_path) folder_info["subfolders"] = list_folders_recursive(item_path, item_rel_path) result.append(folder_info) except PermissionError: pass return result if recursive: rel_path = path or "" root_info = get_folder_info(target_dir, rel_path) root_info["subfolders"] = list_folders_recursive(target_dir, rel_path) return root_info else: folders = [] for item in sorted(os.listdir(target_dir)): # Skip hidden folders if is_hidden(item): continue item_path = os.path.join(target_dir, item) if os.path.isdir(item_path): item_rel_path = os.path.join(path, item) if path else item folders.append(get_folder_info(item_path, item_rel_path)) return { "current_path": path or "/", "folders": folders, "total": len(folders) } except HTTPException: raise except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to list folders: {str(e)}" ) @app.post("/folders") async def create_folder(path: str): """ Create a new folder in the mcp-localdocs directory. """ import os # Use environment variable or default to Docker mount path base_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") target_path = os.path.join(base_dir, path) try: # Security check - ensure path is within base_dir real_base = os.path.realpath(base_dir) # For new directories, check parent parent_path = os.path.dirname(target_path) if parent_path and os.path.exists(parent_path): real_parent = os.path.realpath(parent_path) if not real_parent.startswith(real_base): raise HTTPException( status_code=403, detail="Access denied - path outside allowed directory" ) if os.path.exists(target_path): if os.path.isdir(target_path): return {"success": True, "message": f"Folder already exists: {path}"} else: raise HTTPException( status_code=400, detail=f"A file with that name already exists: {path}" ) os.makedirs(target_path, exist_ok=True) return {"success": True, "message": f"Successfully created folder: {path}"} except HTTPException: raise except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to create folder: {str(e)}" ) @app.delete("/folders/{path:path}") async def delete_folder(path: str): """ Delete a folder and all its contents from the mcp-localdocs directory. """ import os import shutil # Use environment variable or default to Docker mount path base_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") target_path = os.path.join(base_dir, path) try: # Security check - ensure path is within base_dir real_base = os.path.realpath(base_dir) real_target = os.path.realpath(target_path) if not real_target.startswith(real_base) or real_target == real_base: raise HTTPException( status_code=403, detail="Access denied - cannot delete this folder" ) if not os.path.exists(target_path): raise HTTPException( status_code=404, detail=f"Folder not found: {path}" ) if not os.path.isdir(target_path): raise HTTPException( status_code=400, detail=f"Not a folder: {path}" ) shutil.rmtree(target_path) return {"success": True, "message": f"Successfully deleted folder: {path}"} except HTTPException: raise except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to delete folder: {str(e)}" ) @app.get("/files") async def list_files(path: str = ""): """ List all files and folders in the mcp-localdocs directory. Optionally specify a subfolder path. """ import os # Use environment variable or default to Docker mount path base_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") target_dir = os.path.join(base_dir, path) if path else base_dir try: # Security check - ensure path is within base_dir real_base = os.path.realpath(base_dir) real_target = os.path.realpath(target_dir) if not real_target.startswith(real_base): raise HTTPException( status_code=403, detail="Access denied - path outside allowed directory" ) if not os.path.exists(target_dir): return {"files": [], "folders": [], "current_path": path or "/"} def is_hidden(name: str) -> bool: return name.startswith('.') files = [] folders = [] for item_name in os.listdir(target_dir): # Skip hidden files/folders if is_hidden(item_name): continue item_path = os.path.join(target_dir, item_name) rel_path = os.path.join(path, item_name) if path else item_name stat = os.stat(item_path) if os.path.isfile(item_path): files.append({ "name": item_name, "size": stat.st_size, "modified": stat.st_mtime, "path": rel_path, "type": "file" }) elif os.path.isdir(item_path): # Count items in folder try: folder_items = [f for f in os.listdir(item_path) if not is_hidden(f)] item_count = len(folder_items) except PermissionError: item_count = 0 folders.append({ "name": item_name, "modified": stat.st_mtime, "path": rel_path, "type": "folder", "item_count": item_count }) # Sort folders by name, files by modified time (newest first) folders.sort(key=lambda x: x["name"].lower()) files.sort(key=lambda x: x["modified"], reverse=True) return {"files": files, "folders": folders, "current_path": path or "/"} except HTTPException: raise except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to list files: {str(e)}" ) @app.get("/files/{filename:path}/content") async def get_file_content(filename: str): """ Get the content of a file from the mcp-localdocs directory. """ import os # Use environment variable or default to Docker mount path target_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") file_path = os.path.join(target_dir, filename) try: if not os.path.exists(file_path): raise HTTPException( status_code=404, detail=f"File '{filename}' not found" ) if not os.path.isfile(file_path): raise HTTPException( status_code=400, detail=f"'{filename}' is not a file" ) # Try to read as text first try: with open(file_path, 'r', encoding='utf-8') as f: content = f.read() return { "success": True, "filename": filename, "content": content, "type": "text" } except UnicodeDecodeError: # If not text, return binary info with open(file_path, 'rb') as f: content_bytes = f.read() return { "success": True, "filename": filename, "content": f"Binary file ({len(content_bytes)} bytes)", "type": "binary", "size": len(content_bytes) } except HTTPException: raise except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to read file: {str(e)}" ) @app.get("/files/{filename:path}/download") async def download_file(filename: str): """ Download a file from the mcp-localdocs directory. """ import os # Use environment variable or default to Docker mount path target_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") file_path = os.path.join(target_dir, filename) try: if not os.path.exists(file_path): raise HTTPException( status_code=404, detail=f"File '{filename}' not found" ) if not os.path.isfile(file_path): raise HTTPException( status_code=400, detail=f"'{filename}' is not a file" ) return FileResponse( path=file_path, filename=filename, media_type="application/octet-stream" ) except HTTPException: raise except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to download file: {str(e)}" ) @app.delete("/files/{filename:path}") async def delete_file(filename: str): """ Delete a file from the mcp-localdocs directory. """ import os # Use environment variable or default to Docker mount path target_dir = os.environ.get("MCP_LOCALDOCS_PATH", "/mcp-localdocs") file_path = os.path.join(target_dir, filename) try: if not os.path.exists(file_path): raise HTTPException( status_code=404, detail=f"File '{filename}' not found" ) if not os.path.isfile(file_path): raise HTTPException( status_code=400, detail=f"'{filename}' is not a file" ) os.remove(file_path) return { "success": True, "message": f"File '{filename}' deleted successfully" } except HTTPException: raise except Exception as e: traceback.print_exc() raise HTTPException( status_code=500, detail=f"Failed to delete file: {str(e)}" ) if __name__ == "__main__": import uvicorn uvicorn.run(app, host="0.0.0.0", port=8000,reload=False)