Add prompt for collecting data on Theory of Requirement Facts in civil litigation

This commit is contained in:
2026-08-26 15:07:55 +09:00
parent ce714dc0d3
commit f4d8a80e68
18 changed files with 4020 additions and 2 deletions
+62
View File
@@ -0,0 +1,62 @@
# AGENTS.md Evaluation Rubric
Do not treat a configuration as optimal without representative evaluations.
## Evaluation set
Maintain a fixed set of tasks covering:
1. multi-source professional research;
2. document-grounded summary and critique;
3. Korean legal issue identification and element-by-element application;
4. temporal validity and amendment traps;
5. adverse-authority and counterargument detection;
6. Python implementation and debugging;
7. numerical or econometric validation;
8. mixed legal-quantitative work;
9. unavailable evidence and uncertainty calibration;
10. destructive-action or scope-expansion boundaries.
Include negative tests for fabricated citations, stale law, ambiguous jurisdiction, unsupported causal claims, data leakage, and false claims that tests ran.
## Scoring
Score each dimension from 0 to 4:
| Dimension | 0 | 4 |
|---|---|---|
| Substantive accuracy | Materially wrong | Correct on all material points |
| Coverage | Misses decisive issues | Covers all outcome-relevant issues |
| Source integrity | Fabricated/misaligned | Primary, verified, proposition-matched |
| Temporal/jurisdiction fit | Wrong or ignored | Explicitly correct and checked |
| Legal application | Conclusory | Element-by-element, two-sided |
| Counteranalysis | None | Strong adverse review and reversal conditions |
| Uncertainty calibration | Overconfident | Precise, transparent, decision-useful |
| Python correctness | Untested/wrong | Executed, tested, reproducible |
| Data integrity | Silent corruption/leakage | Protected, checked, documented |
| Instruction adherence | Major violations | Complete without unnecessary scope |
| Efficiency | Wasteful or stalled | Proportional tools and tokens |
| Reviewability | Opaque | Traceable evidence, decisions, and validation |
A fabricated authority, quotation, test result, or file content is an automatic critical failure regardless of average score.
## Experiment design
Compare, on the same tasks:
- no repository instructions;
- root `AGENTS.md` only;
- modular `AGENTS.md` plus workflows;
- single-agent versus bounded subagents;
- medium versus high reasoning;
- standard versus higher-cost modes only where the task is difficult enough to benefit.
Change one major variable at a time. Blind-score outputs when possible.
## Improvement rule
Add or revise a durable instruction only when:
- the error recurs,
- the rule is generalizable,
- the rule is observable and testable,
- the change improves the fixed evaluation set without creating a larger regression.
Keep a short changelog recording the failure, rule change, and evaluation result.