618 lines
32 KiB
Markdown
618 lines
32 KiB
Markdown
### 질문
|
|
|
|
> 내가 개발하는 인공지능 에이전트는 작업 1 단계에서 GPT 모델을 API를 통해 호출하여 사용한다. 이 때, 작업 2에서 동일한 GPT 모델을 api를 통해 호출하여 작업 1에서 얻은 결과를 input file로 사용한다면, 작업 1의 결과물 파일을 kv-caching 처럼 저장한 후, 작업 2에서 추가적으로 토큰을 사용하지 않고도 GPT 모델이 caching 자료를 hit하여 사용할 수 있는가? 즉, 그렇게 설계를 할 수 있는가?
|
|
|
|
### 답변
|
|
|
|
결론부터 말하면, **사용자 측 설계로 작업 1의 결과 파일을 “KV cache 자체”처럼 저장해 두었다가, 작업 2에서 그 내부 캐시를 직접 hit시키며 추가 입력 토큰 없이 재사용하는 방식은 현재 OpenAI API에서 일반적으로 가능한 설계가 아닙니다.** OpenAI의 공식 문서상 Prompt Caching은 **서버 측에서 자동으로 동작하는 기능**이며, **정확히 동일한 프롬프트 prefix**가 반복될 때 비용과 지연을 줄여 주는 메커니즘입니다. 또한 OpenAI는 Prompt Caching에서 **모델의 key/value tensor가 보존될 수 있다**고 설명하지만, 이는 플랫폼 내부 동작이지 사용자가 외부 파일로 저장·재주입하는 인터페이스가 아닙니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
따라서 질문을 엄밀하게 나누면 답은 다음과 같습니다.
|
|
**(1) “KV cache를 내가 저장해서 다음 API 호출 때 넣을 수 있는가?” → 아니오.**
|
|
**(2) “작업 1 결과를 작업 2에서 거의 같은 prefix로 다시 보내면 서버 측 Prompt Cache hit를 기대할 수 있는가?” → 예, 일부는 가능.**
|
|
다만 이 경우에도 **입력 토큰이 완전히 0이 되는 것은 아닙니다.** OpenAI는 Prompt Caching이 입력 비용을 크게 낮출 수 있다고 설명하지만, 이는 **할인된 cached input 요금**이지 무과금이 아닙니다. 공식 가격표도 일반 input과 **cached input을 별도 가격 항목**으로 표시하고 있습니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
더 구체적으로 말하면, OpenAI의 Prompt Caching은 **1024 tokens 이상인 프롬프트**에서 자동으로 활성화되며, **정확한 prefix 일치**가 있어야 cache hit가 가능합니다. 문서에는 정적 내용은 앞부분에 두고, 동적 내용은 뒤에 두라고 명시되어 있습니다. 즉, 작업 1의 결과물을 파일이든 텍스트든 작업 2에서 다시 쓰려면, 그 내용이 **작업 2 요청의 앞부분(prefix)** 에 거의 동일하게 포함되어야 합니다. 단순히 “같은 파일을 참조한다”거나 “같은 의미의 내용을 요약해서 넣는다”는 정도로는 KV cache hit를 보장할 수 없습니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
또한 OpenAI는 Prompt Caching에 대해, **캐시된 것은 프롬프트 자체에 대한 계산 결과일 뿐이며, 최종 출력은 매번 새로 생성된다**고 분명히 적고 있습니다. 즉, 작업 1의 출력 파일을 다시 사용하더라도, 작업 2의 응답 생성은 새로 수행됩니다. 다시 말해 이것은 **“중간 은닉상태를 사용자에게 공개한 후 재주입하는 persistent KV cache”** 가 아니라, **반복되는 입력 prefix에 대한 서버 측 prefill 재사용**에 가깝습니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
질문에서 말한 “작업 1 결과물 파일을 input file로 사용”하는 부분도 중요합니다. OpenAI의 File Inputs 문서는, 파일을 입력으로 넣을 수는 있지만 **파일 내용은 결국 모델 컨텍스트로 파싱되어 들어가며 token usage가 발생할 수 있다**고 명시합니다. 특히 PDF는 추출 텍스트와 페이지 이미지가 함께 컨텍스트에 포함되어 토큰 사용량이 커질 수 있습니다. 즉, 파일이라고 해서 토큰을 우회하는 것이 아닙니다. 같은 파일을 다시 넣는 경우에도, 그 파일이 모델 입력으로 해석되는 이상 **토큰 문맥에 포함**되며, 운이 좋으면 서버 측 prompt cache가 일부 적중할 수 있을 뿐입니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/file-inputs/?utm_source=chatgpt.com))
|
|
|
|
한편, OpenAI는 작업 간 문맥 연결을 위해 **`previous_response_id`** 나 **Conversation state** 를 제공하지만, 이것도 질문하신 의미의 “토큰 없는 KV 재사용”과는 다릅니다. 공식 문서에는 **`previous_response_id`를 쓰더라도 체인에 포함된 이전 입력 토큰들은 모두 input tokens로 과금된다**고 명시되어 있습니다. 따라서 이 기능은 상태 관리에는 유용하지만, “작업 1의 긴 출력을 무료로 작업 2에 넘긴다”는 구조는 아닙니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/conversation-state/))
|
|
|
|
따라서 설계 관점에서 가장 정확한 정리는 다음과 같습니다.
|
|
|
|
**가능한 설계**
|
|
|
|
1. **서버 측 Prompt Caching을 노린 반복 prefix 설계**
|
|
작업 1 결과 중 반복해서 참조될 큰 정적 블록을 작업 2 프롬프트의 앞부분에 동일하게 배치하고, `prompt_cache_key`와 필요시 `prompt_cache_retention=24h`를 사용해 hit 확률을 높일 수 있습니다. 이 경우 비용과 지연은 줄어들 수 있습니다. ([OpenAI Platform](https://platform.openai.com/docs/api-reference/responses/list?ref=test-ippon.ghost.io))
|
|
2. **Conversation / `previous_response_id` 기반 상태 연결**
|
|
작업 1과 작업 2를 하나의 연속 세션처럼 다룰 수 있습니다. 구현은 단순해지지만, 이전 문맥이 토큰 계산에서 사라지는 것은 아닙니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/conversation-state/?utm_source=chatgpt.com))
|
|
3. **작업 1 결과를 압축·구조화해서 작업 2에 전달**
|
|
예를 들어 작업 1 출력 전체를 다시 넣지 말고, JSON key facts, table, extracted variables, entity graph, summary state 등으로 압축해 넘기면 실제 토큰 비용을 줄일 수 있습니다. 이것은 공식 문서의 직접 진술이라기보다, 토큰 과금 구조상 합리적인 엔지니어링 귀결입니다. **이 부분은 제 추론이지만 근거는 충분합니다.** ([OpenAI Developers](https://developers.openai.com/api/docs/guides/conversation-state/))
|
|
|
|
**불가능하거나 오해하기 쉬운 설계**
|
|
|
|
1. **사용자가 KV cache blob를 파일로 저장했다가 다음 REST 호출에 재주입하는 방식**
|
|
현재 공식 문서 기준으로 그런 공개 API 인터페이스는 보이지 않습니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
2. **작업 1 결과 파일을 다시 넣기만 하면 토큰이 0이 되는 방식**
|
|
아닙니다. 최대한 잘 되어도 cached input 요금이며, `previous_response_id`를 써도 이전 입력 토큰은 청구됩니다. ([OpenAI](https://openai.com/api/pricing/?utm_source=chatgpt.com))
|
|
|
|
실무적으로는 다음 구조가 가장 타당합니다.
|
|
|
|
- **반복적으로 거의 같은 긴 컨텍스트**를 쓰는 경우: Prompt Caching 최적화
|
|
- **멀티턴 상태 연결**이 중요한 경우: Responses API의 `previous_response_id` 또는 Conversations
|
|
- **큰 산출물을 후속 단계에서 부분적으로만 재사용**하는 경우: 작업 1에서 바로 압축된 machine-readable state를 함께 생성
|
|
- **장기 보존/검색**이 필요한 경우: 벡터 DB나 외부 저장소에 보관 후 retrieval, 다만 retrieval로 뽑은 내용도 최종적으로 모델에 넣는 순간 토큰은 사용됨
|
|
|
|
따라서 당신의 원문 질문에 가장 정확히 답하면 다음과 같습니다.
|
|
|
|
**“작업 1의 결과물 파일을 KV-caching처럼 저장해 두고, 작업 2에서 추가 토큰 없이 동일 GPT 모델이 그것을 hit하여 쓰게 설계할 수 있는가?”**
|
|
→ **엄밀한 의미에서는 불가능합니다.**
|
|
다만 **동일한 prefix를 반복 입력하는 방식으로 OpenAI의 서버 측 Prompt Caching을 유도하여 비용과 지연을 크게 줄이는 설계는 가능합니다.** 그러나 그것은 **0-token 재사용도 아니고, 사용자가 직접 KV cache를 소유·주입하는 설계도 아닙니다.** ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
원하시면 다음 답변에서 제가 바로
|
|
**“작업 1 → 작업 2 파이프라인을 Prompt Caching 친화적으로 설계하는 구체적 API 아키텍처”** 를 Python 예제로 작성해 드리겠습니다.
|
|
|
|
---
|
|
|
|
### 질문
|
|
|
|
> “작업 1 → 작업 2 파이프라인을 Prompt Caching 친화적으로 설계하는 구체적 API 아키텍처” 를 Python 예제로 작성하라.
|
|
|
|
### 답변
|
|
|
|
아래 설계가 현재 OpenAI API의 특성을 가장 정교하게 반영합니다. 핵심은 **“작업 1의 동적 결과물 자체를 KV cache처럼 재사용하려고 하지 말고, 작업 1과 작업 2가 공유하는 거대한 정적 prefix를 매 호출마다 동일하게 반복하여 서버 측 Prompt Cache를 최대한 적중시키는 것”** 입니다. Prompt Caching은 자동으로 동작하며, 1024 tokens 이상의 프롬프트에서 활성화되고, `prompt_cache_key`와 `prompt_cache_retention`을 통해 적중률을 높일 수 있습니다. 또한 OpenAI는 정적·반복 콘텐츠를 앞부분에, 동적 콘텐츠를 뒷부분에 두라고 권고하고 있고, 응답의 `usage.prompt_tokens_details.cached_tokens`로 실제 cache hit를 관찰할 수 있습니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
동시에, **`previous_response_id`는 상태 연결에는 유용하지만 비용 절감 장치로 이해하면 안 됩니다.** OpenAI는 `previous_response_id`를 사용하더라도 체인에 포함된 이전 입력 토큰들이 계속 input tokens로 과금된다고 명시합니다. 따라서 “Prompt Caching 친화적 설계”와 “세션 연결 설계”는 분리해서 생각하는 편이 정확합니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/conversation-state/))
|
|
|
|
또한 **파일 입력은 토큰을 우회하지 않습니다.** Responses API는 `input_file`로 파일을 받을 수 있고, 파일은 `file_id`, `file_url`, 혹은 인라인 데이터로 전달할 수 있습니다. 파일 업로드 시 `purpose="user_data"`를 쓰는 것이 권장됩니다. 그러나 파일은 결국 모델 문맥으로 해석되며, 특히 PDF는 추출 텍스트와 페이지 이미지가 함께 반영되어 토큰 사용량이 커질 수 있습니다. 따라서 “작업 1의 긴 보고서 파일”을 그대로 작업 2에 다시 넣는 구조는 가능하되, **그것을 Prompt Cache의 주된 최적화 수단으로 삼는 것은 권장되지 않습니다.** ([OpenAI Developers](https://developers.openai.com/api/docs/guides/file-inputs/))
|
|
|
|
------
|
|
|
|
## 권장 아키텍처
|
|
|
|
가장 바람직한 구조는 다음과 같습니다.
|
|
|
|
1. **거대한 공통 prefix를 별도 상수로 관리한다.**
|
|
여기에는 에이전트의 헌장, 도메인 온톨로지, 용어 정의, 평가 루브릭, JSON schema 설명, 출력 규율, 금지사항 등 **반복적으로 항상 들어가는 긴 정적 텍스트**를 넣습니다. 이 부분이 작업 1과 작업 2에서 **byte-for-byte로 최대한 동일**해야 합니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
2. **작업 1은 “긴 자연어 보고서”보다 “작은 구조화 상태(state JSON)”를 만든다.**
|
|
작업 2가 필요한 것은 대개 작업 1의 모든 수사적 문장보다, 핵심 사실·추출값·판단근거·불확실성·에러 상태입니다. 따라서 작업 1의 산출물은 후속 단계용 `stage1_state.json`으로 압축해 두는 편이 유리합니다. 이것은 문서의 직접 문구라기보다는 토큰 과금과 캐시 메커니즘에 따른 엔지니어링 권고입니다.
|
|
3. **작업 2는 동일한 공통 prefix를 다시 보내고, 그 뒤에 작은 `stage1_state.json`과 새로운 지시를 붙인다.**
|
|
이렇게 하면 캐시는 **공통 prefix 부분**에서 최대한 적중하고, 작업 1의 동적 결과는 작기 때문에 추가 토큰 부담이 제한됩니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
4. **`prompt_cache_key`는 동일 workflow family에 대해 일관되게 유지한다.**
|
|
OpenAI는 동일 prefix를 공유하는 요청들에서 `prompt_cache_key`를 일관되게 쓰라고 권고하며, 각 unique prefix-`prompt_cache_key` 조합당 분당 15회 미만으로 유지하라고 안내합니다. `prompt_cache_retention="24h"`를 사용하면 캐시 유지 시간을 연장할 수 있습니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
5. **실측한다.**
|
|
`responses.input_tokens.count`로 사전 토큰 수를 확인하고, 실제 응답의 `cached_tokens`를 로그에 남겨 캐시 적중률을 검증해야 합니다. Prompt Caching은 “되겠지”라고 추정할 일이 아니라, 계측해야 할 대상입니다. ([OpenAI Developers](https://developers.openai.com/api/reference/python/resources/responses/subresources/input_tokens/methods/count/?utm_source=chatgpt.com))
|
|
|
|
------
|
|
|
|
## Python 예제
|
|
|
|
아래 코드는 다음을 구현합니다.
|
|
|
|
- 작업 1과 작업 2가 공유하는 **긴 공통 prefix**
|
|
- 작업 1에서 **구조화 상태 JSON** 생성
|
|
- 작업 2에서 그 JSON을 사용해 후속 판단 수행
|
|
- `prompt_cache_key`, `prompt_cache_retention="24h"` 사용
|
|
- `input_tokens.count`로 사전 토큰 계측
|
|
- 실제 `cached_tokens` 로깅
|
|
- 비교를 위한 **비권장 file-input variant** 포함
|
|
|
|
```python
|
|
import os
|
|
import json
|
|
import hashlib
|
|
from pathlib import Path
|
|
from typing import Any, Dict
|
|
|
|
from openai import OpenAI
|
|
|
|
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
|
|
|
|
MODEL = "gpt-5.1" # 예시. 동일 모델을 작업 1/2에 일관되게 사용.
|
|
CACHE_RETENTION = "24h"
|
|
ARTIFACT_DIR = Path("artifacts")
|
|
ARTIFACT_DIR.mkdir(exist_ok=True)
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 1) 작업 1과 작업 2가 공유하는 "긴 정적 prefix"
|
|
# - 실제 운영에서는 이 부분을 충분히 길고(보통 1024 tokens 이상),
|
|
# byte-for-byte 안정적으로 유지해야 Prompt Cache에 유리합니다.
|
|
# ---------------------------------------------------------------------
|
|
SHARED_PREFIX = """
|
|
[Agent Charter v3.2]
|
|
You are a production-grade research and analysis agent.
|
|
You must:
|
|
- preserve evidentiary discipline,
|
|
- separate observed facts from inferences,
|
|
- return machine-readable outputs when requested,
|
|
- surface uncertainty explicitly,
|
|
- avoid rhetorical filler,
|
|
- keep category labels stable across runs,
|
|
- use the ontology and definitions below exactly as written.
|
|
|
|
[Ontology]
|
|
- entity
|
|
- claim
|
|
- evidence
|
|
- contradiction
|
|
- uncertainty
|
|
- decision_rule
|
|
- risk_flag
|
|
- missing_information
|
|
|
|
[Evaluation Rubric]
|
|
1. Extract facts exactly.
|
|
2. Normalize dates and quantities.
|
|
3. Preserve provenance references.
|
|
4. Separate direct evidence from model inference.
|
|
5. Emit stable keys for downstream use.
|
|
6. Do not invent fields outside schema.
|
|
7. Mark unknowns explicitly as null or empty array.
|
|
8. Keep labels deterministic.
|
|
|
|
[Style Constraints]
|
|
- no markdown unless explicitly requested
|
|
- no prose outside schema
|
|
- no omitted required keys
|
|
- no additional properties
|
|
|
|
[Important]
|
|
The downstream pipeline depends on stable JSON keys.
|
|
Treat the JSON schema as a contract.
|
|
"""
|
|
|
|
STAGE1_SUFFIX = """
|
|
[Stage 1 Objective]
|
|
Read the source material and extract a compact intermediate state that can be used
|
|
by downstream stages. Do not produce a long narrative report unless specifically requested.
|
|
"""
|
|
|
|
STAGE2_SUFFIX = """
|
|
[Stage 2 Objective]
|
|
Consume the stage1_state JSON as authoritative intermediate state.
|
|
Then perform the downstream decision or transformation task.
|
|
Do not restate the entire input unless necessary.
|
|
"""
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 2) Structured Outputs schema
|
|
# ---------------------------------------------------------------------
|
|
STAGE1_SCHEMA: Dict[str, Any] = {
|
|
"type": "object",
|
|
"additionalProperties": False,
|
|
"properties": {
|
|
"document_id": {"type": "string"},
|
|
"summary": {"type": "string"},
|
|
"entities": {
|
|
"type": "array",
|
|
"items": {"type": "string"}
|
|
},
|
|
"facts": {
|
|
"type": "array",
|
|
"items": {
|
|
"type": "object",
|
|
"additionalProperties": False,
|
|
"properties": {
|
|
"fact": {"type": "string"},
|
|
"evidence": {"type": "string"},
|
|
"confidence": {"type": "number"}
|
|
},
|
|
"required": ["fact", "evidence", "confidence"]
|
|
}
|
|
},
|
|
"uncertainties": {
|
|
"type": "array",
|
|
"items": {"type": "string"}
|
|
},
|
|
"risk_flags": {
|
|
"type": "array",
|
|
"items": {"type": "string"}
|
|
},
|
|
"recommended_next_inputs": {
|
|
"type": "array",
|
|
"items": {"type": "string"}
|
|
}
|
|
},
|
|
"required": [
|
|
"document_id",
|
|
"summary",
|
|
"entities",
|
|
"facts",
|
|
"uncertainties",
|
|
"risk_flags",
|
|
"recommended_next_inputs"
|
|
]
|
|
}
|
|
|
|
STAGE2_SCHEMA: Dict[str, Any] = {
|
|
"type": "object",
|
|
"additionalProperties": False,
|
|
"properties": {
|
|
"decision": {"type": "string"},
|
|
"rationale": {"type": "string"},
|
|
"used_facts": {
|
|
"type": "array",
|
|
"items": {"type": "string"}
|
|
},
|
|
"remaining_uncertainties": {
|
|
"type": "array",
|
|
"items": {"type": "string"}
|
|
},
|
|
"action_items": {
|
|
"type": "array",
|
|
"items": {"type": "string"}
|
|
}
|
|
},
|
|
"required": [
|
|
"decision",
|
|
"rationale",
|
|
"used_facts",
|
|
"remaining_uncertainties",
|
|
"action_items"
|
|
]
|
|
}
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 3) Cache key
|
|
# - 동일한 workflow family / 동일 prefix version 에 대해 일관되게 유지
|
|
# - PII를 직접 넣지 말고 hash 사용 권장
|
|
# ---------------------------------------------------------------------
|
|
def make_prompt_cache_key(tenant: str, workflow_family: str, prefix_version: str = "v3.2") -> str:
|
|
raw = f"{tenant}|{workflow_family}|{prefix_version}"
|
|
return hashlib.sha256(raw.encode("utf-8")).hexdigest()[:48]
|
|
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 4) 유틸리티
|
|
# ---------------------------------------------------------------------
|
|
def usage_summary(resp) -> Dict[str, Any]:
|
|
usage = getattr(resp, "usage", None)
|
|
if usage is None:
|
|
return {}
|
|
|
|
prompt_tokens = getattr(usage, "prompt_tokens", None)
|
|
completion_tokens = getattr(usage, "completion_tokens", None)
|
|
total_tokens = getattr(usage, "total_tokens", None)
|
|
|
|
prompt_details = getattr(usage, "prompt_tokens_details", None)
|
|
cached_tokens = getattr(prompt_details, "cached_tokens", None) if prompt_details else None
|
|
|
|
return {
|
|
"prompt_tokens": prompt_tokens,
|
|
"completion_tokens": completion_tokens,
|
|
"total_tokens": total_tokens,
|
|
"cached_tokens": cached_tokens,
|
|
}
|
|
|
|
|
|
def upload_user_file(path: str) -> str:
|
|
with open(path, "rb") as f:
|
|
uploaded = client.files.create(
|
|
file=f,
|
|
purpose="user_data",
|
|
)
|
|
return uploaded.id
|
|
|
|
|
|
def count_tokens_for_request(model: str, instructions: str, input_payload: Any, text_format: Dict[str, Any] | None = None):
|
|
kwargs = {
|
|
"model": model,
|
|
"instructions": instructions,
|
|
"input": input_payload,
|
|
}
|
|
if text_format is not None:
|
|
kwargs["text"] = {"format": text_format}
|
|
|
|
resp = client.responses.input_tokens.count(**kwargs)
|
|
return resp.input_tokens
|
|
|
|
|
|
def render_stage1_report(state: Dict[str, Any]) -> str:
|
|
"""사람이 읽는 audit report는 로컬에서 렌더링. 후속 단계에는 이 긴 보고서를 되도록 쓰지 않는다."""
|
|
lines = []
|
|
lines.append(f"# Stage 1 Report")
|
|
lines.append(f"- document_id: {state['document_id']}")
|
|
lines.append("")
|
|
lines.append("## Summary")
|
|
lines.append(state["summary"])
|
|
lines.append("")
|
|
lines.append("## Entities")
|
|
for e in state["entities"]:
|
|
lines.append(f"- {e}")
|
|
lines.append("")
|
|
lines.append("## Facts")
|
|
for i, fact in enumerate(state["facts"], 1):
|
|
lines.append(f"{i}. Fact: {fact['fact']}")
|
|
lines.append(f" - Evidence: {fact['evidence']}")
|
|
lines.append(f" - Confidence: {fact['confidence']}")
|
|
lines.append("")
|
|
lines.append("## Uncertainties")
|
|
for u in state["uncertainties"]:
|
|
lines.append(f"- {u}")
|
|
lines.append("")
|
|
lines.append("## Risk Flags")
|
|
for r in state["risk_flags"]:
|
|
lines.append(f"- {r}")
|
|
lines.append("")
|
|
lines.append("## Recommended Next Inputs")
|
|
for r in state["recommended_next_inputs"]:
|
|
lines.append(f"- {r}")
|
|
return "\n".join(lines)
|
|
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 5) 작업 1
|
|
# ---------------------------------------------------------------------
|
|
def run_stage1(
|
|
source_file_id: str,
|
|
stage1_request: str,
|
|
tenant: str = "tenant_alpha",
|
|
workflow_family: str = "research_pipeline_v1",
|
|
) -> Dict[str, Any]:
|
|
instructions = SHARED_PREFIX + "\n\n" + STAGE1_SUFFIX
|
|
cache_key = make_prompt_cache_key(tenant, workflow_family, prefix_version="v3.2")
|
|
|
|
input_payload = [
|
|
{
|
|
"role": "user",
|
|
"content": [
|
|
{
|
|
"type": "input_text",
|
|
"text": stage1_request
|
|
},
|
|
{
|
|
"type": "input_file",
|
|
"file_id": source_file_id
|
|
}
|
|
]
|
|
}
|
|
]
|
|
|
|
stage1_format = {
|
|
"type": "json_schema",
|
|
"name": "stage1_state",
|
|
"strict": True,
|
|
"schema": STAGE1_SCHEMA
|
|
}
|
|
|
|
estimated_input_tokens = count_tokens_for_request(
|
|
model=MODEL,
|
|
instructions=instructions,
|
|
input_payload=input_payload,
|
|
text_format=stage1_format,
|
|
)
|
|
|
|
response = client.responses.create(
|
|
model=MODEL,
|
|
instructions=instructions,
|
|
input=input_payload,
|
|
text={"format": stage1_format},
|
|
prompt_cache_key=cache_key,
|
|
prompt_cache_retention=CACHE_RETENTION,
|
|
)
|
|
|
|
state = json.loads(response.output_text)
|
|
|
|
state_path = ARTIFACT_DIR / "stage1_state.json"
|
|
state_path.write_text(json.dumps(state, ensure_ascii=False, indent=2), encoding="utf-8")
|
|
|
|
report_md = render_stage1_report(state)
|
|
report_path = ARTIFACT_DIR / "stage1_report.md"
|
|
report_path.write_text(report_md, encoding="utf-8")
|
|
|
|
return {
|
|
"response_id": response.id,
|
|
"state": state,
|
|
"state_path": str(state_path),
|
|
"report_path": str(report_path),
|
|
"estimated_input_tokens": estimated_input_tokens,
|
|
"usage": usage_summary(response),
|
|
}
|
|
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 6) 작업 2 (권장)
|
|
# - 작업 1의 긴 자연어 보고서 대신 compact state JSON 사용
|
|
# - 동일한 SHARED_PREFIX를 다시 보내 cache hit를 노림
|
|
# ---------------------------------------------------------------------
|
|
def run_stage2_from_state(
|
|
stage1_state: Dict[str, Any],
|
|
stage2_request: str,
|
|
tenant: str = "tenant_alpha",
|
|
workflow_family: str = "research_pipeline_v1",
|
|
) -> Dict[str, Any]:
|
|
instructions = SHARED_PREFIX + "\n\n" + STAGE2_SUFFIX
|
|
cache_key = make_prompt_cache_key(tenant, workflow_family, prefix_version="v3.2")
|
|
|
|
input_payload = [
|
|
{
|
|
"role": "user",
|
|
"content": [
|
|
{
|
|
"type": "input_text",
|
|
"text": (
|
|
"Use the following stage1_state JSON as the authoritative intermediate state.\n\n"
|
|
f"{json.dumps(stage1_state, ensure_ascii=False)}\n\n"
|
|
f"Stage 2 task:\n{stage2_request}"
|
|
)
|
|
}
|
|
]
|
|
}
|
|
]
|
|
|
|
stage2_format = {
|
|
"type": "json_schema",
|
|
"name": "stage2_result",
|
|
"strict": True,
|
|
"schema": STAGE2_SCHEMA
|
|
}
|
|
|
|
estimated_input_tokens = count_tokens_for_request(
|
|
model=MODEL,
|
|
instructions=instructions,
|
|
input_payload=input_payload,
|
|
text_format=stage2_format,
|
|
)
|
|
|
|
response = client.responses.create(
|
|
model=MODEL,
|
|
instructions=instructions,
|
|
input=input_payload,
|
|
text={"format": stage2_format},
|
|
prompt_cache_key=cache_key,
|
|
prompt_cache_retention=CACHE_RETENTION,
|
|
)
|
|
|
|
result = json.loads(response.output_text)
|
|
|
|
result_path = ARTIFACT_DIR / "stage2_result.json"
|
|
result_path.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
|
|
|
return {
|
|
"response_id": response.id,
|
|
"result": result,
|
|
"result_path": str(result_path),
|
|
"estimated_input_tokens": estimated_input_tokens,
|
|
"usage": usage_summary(response),
|
|
}
|
|
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 7) 작업 2 (비권장 변형)
|
|
# - 작업 1 보고서 파일을 그대로 input_file로 다시 주입
|
|
# - 가능은 하지만, 캐시 관점에서는 stage1_state 방식보다 일반적으로 덜 효율적
|
|
# ---------------------------------------------------------------------
|
|
def run_stage2_from_stage1_report_file(
|
|
stage1_report_path: str,
|
|
stage2_request: str,
|
|
tenant: str = "tenant_alpha",
|
|
workflow_family: str = "research_pipeline_v1",
|
|
) -> Dict[str, Any]:
|
|
instructions = SHARED_PREFIX + "\n\n" + STAGE2_SUFFIX
|
|
cache_key = make_prompt_cache_key(tenant, workflow_family, prefix_version="v3.2")
|
|
|
|
stage1_report_file_id = upload_user_file(stage1_report_path)
|
|
|
|
input_payload = [
|
|
{
|
|
"role": "user",
|
|
"content": [
|
|
{"type": "input_text", "text": stage2_request},
|
|
{"type": "input_file", "file_id": stage1_report_file_id},
|
|
]
|
|
}
|
|
]
|
|
|
|
stage2_format = {
|
|
"type": "json_schema",
|
|
"name": "stage2_result",
|
|
"strict": True,
|
|
"schema": STAGE2_SCHEMA
|
|
}
|
|
|
|
estimated_input_tokens = count_tokens_for_request(
|
|
model=MODEL,
|
|
instructions=instructions,
|
|
input_payload=input_payload,
|
|
text_format=stage2_format,
|
|
)
|
|
|
|
response = client.responses.create(
|
|
model=MODEL,
|
|
instructions=instructions,
|
|
input=input_payload,
|
|
text={"format": stage2_format},
|
|
prompt_cache_key=cache_key,
|
|
prompt_cache_retention=CACHE_RETENTION,
|
|
)
|
|
|
|
result = json.loads(response.output_text)
|
|
|
|
return {
|
|
"response_id": response.id,
|
|
"result": result,
|
|
"estimated_input_tokens": estimated_input_tokens,
|
|
"usage": usage_summary(response),
|
|
}
|
|
|
|
|
|
# ---------------------------------------------------------------------
|
|
# 8) 예시 실행
|
|
# ---------------------------------------------------------------------
|
|
if __name__ == "__main__":
|
|
# 예: 원문 PDF/CSV/TXT 등을 먼저 업로드
|
|
source_file_id = upload_user_file("source_document.pdf")
|
|
|
|
stage1 = run_stage1(
|
|
source_file_id=source_file_id,
|
|
stage1_request=(
|
|
"Extract the core entities, factual claims, uncertainties, and risk flags from the source file. "
|
|
"Return only the compact intermediate state."
|
|
),
|
|
)
|
|
|
|
print("Stage 1 usage:", json.dumps(stage1["usage"], indent=2))
|
|
print("Stage 1 estimated input tokens:", stage1["estimated_input_tokens"])
|
|
print("Stage 1 state file:", stage1["state_path"])
|
|
print("Stage 1 report file:", stage1["report_path"])
|
|
|
|
stage2 = run_stage2_from_state(
|
|
stage1_state=stage1["state"],
|
|
stage2_request=(
|
|
"Based on the intermediate state, produce a downstream decision with rationale, "
|
|
"remaining uncertainties, and action items."
|
|
),
|
|
)
|
|
|
|
print("Stage 2 usage:", json.dumps(stage2["usage"], indent=2))
|
|
print("Stage 2 estimated input tokens:", stage2["estimated_input_tokens"])
|
|
print("Stage 2 result file:", stage2["result_path"])
|
|
```
|
|
|
|
------
|
|
|
|
## 이 코드가 Prompt Caching 친화적인 이유
|
|
|
|
첫째, **작업 1과 작업 2가 동일한 `SHARED_PREFIX`를 프롬프트의 맨 앞에 반복**합니다. OpenAI는 정적·반복 내용은 앞에, 동적 내용은 뒤에 두라고 권고하고 있고, 이것이 가장 기본적인 cache-hit 최적화입니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
둘째, **동일한 `prompt_cache_key`를 재사용**합니다. OpenAI는 공통 prefix를 공유하는 요청들에 대해 `prompt_cache_key`를 일관되게 사용할 것을 권고하고, `prompt_cache_retention="24h"`로 캐시 유지 시간을 연장할 수 있다고 설명합니다. ([OpenAI Developers](https://developers.openai.com/api/reference/resources/responses/methods/create/))
|
|
|
|
셋째, **작업 1의 결과를 긴 prose report가 아니라 compact JSON state로 저장**합니다. Prompt Caching은 공통 prefix를 재사용할 때 강력하며, 작업 1의 동적 결과 자체는 본질적으로 “새로운 텍스트”입니다. 그러므로 후속 단계 입력은 가능한 한 작고 구조화되어야 합니다. 이 부분은 문서의 직접 문구라기보다, 토큰 과금 구조와 cache 메커니즘에서 도출되는 설계 원칙입니다.
|
|
|
|
넷째, **실제 cache hit를 `cached_tokens`로 측정**합니다. Prompt Caching은 1024 tokens 이상에서 의미가 생기며, 응답 usage에 `cached_tokens`가 기록됩니다. 따라서 운영에서는 이 값을 반드시 로그에 남겨야 합니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
------
|
|
|
|
## `previous_response_id`를 여기에 왜 넣지 않았는가
|
|
|
|
넣을 수는 있습니다. 그러나 이 질문의 목적이 **Prompt Caching 친화적 비용 구조**라면, `previous_response_id`는 주 해법이 아닙니다. OpenAI는 `previous_response_id`를 사용해도 이전 입력 토큰이 과금된다고 명시합니다. 그러므로 상태 연결은 편의 기능이고, 비용 절감의 주축은 **공통 prefix 캐싱 + 동적 상태 압축**이어야 합니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/conversation-state/))
|
|
|
|
------
|
|
|
|
## 파일 기반 파이프라인을 꼭 유지해야 한다면
|
|
|
|
가능한 최소한의 형태는 다음 원칙을 따르는 편이 낫습니다.
|
|
|
|
- 작업 1의 결과를 **두 개**로 나눕니다.
|
|
`stage1_report.md`는 인간 검토용, `stage1_state.json`은 기계 후속 단계용입니다.
|
|
- 작업 2는 원칙적으로 `stage1_state.json`만 사용합니다.
|
|
정말 필요할 때만 `stage1_report.md`를 `input_file`로 추가합니다.
|
|
- PDF를 후속 단계 입력으로 반복해서 넣는 경우는 특히 주의해야 합니다.
|
|
OpenAI는 PDF 입력이 추출 텍스트와 페이지 이미지를 함께 문맥에 포함시켜 토큰 사용량이 커질 수 있다고 명시합니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/file-inputs/))
|
|
|
|
------
|
|
|
|
## 운영 체크리스트
|
|
|
|
실제 운영에서는 아래 네 가지를 반드시 보십시오.
|
|
|
|
1. **공통 prefix가 정말 byte-for-byte 동일한가**
|
|
2. **공통 prefix가 충분히 길어 1024 tokens 이상인가**
|
|
3. **`cached_tokens`가 실제로 증가하는가**
|
|
4. **동일 prefix-`prompt_cache_key` 조합을 과도하게 분산시키지 않았는가**
|
|
OpenAI는 해당 조합당 분당 15회 미만을 권고합니다. ([OpenAI Developers](https://developers.openai.com/api/docs/guides/prompt-caching/))
|
|
|
|
------
|
|
|