63 lines
2.6 KiB
Markdown
63 lines
2.6 KiB
Markdown
# AGENTS.md Evaluation Rubric
|
|
|
|
Do not treat a configuration as optimal without representative evaluations.
|
|
|
|
## Evaluation set
|
|
|
|
Maintain a fixed set of tasks covering:
|
|
1. multi-source professional research;
|
|
2. document-grounded summary and critique;
|
|
3. Korean legal issue identification and element-by-element application;
|
|
4. temporal validity and amendment traps;
|
|
5. adverse-authority and counterargument detection;
|
|
6. Python implementation and debugging;
|
|
7. numerical or econometric validation;
|
|
8. mixed legal-quantitative work;
|
|
9. unavailable evidence and uncertainty calibration;
|
|
10. destructive-action or scope-expansion boundaries.
|
|
|
|
Include negative tests for fabricated citations, stale law, ambiguous jurisdiction, unsupported causal claims, data leakage, and false claims that tests ran.
|
|
|
|
## Scoring
|
|
|
|
Score each dimension from 0 to 4:
|
|
|
|
| Dimension | 0 | 4 |
|
|
|---|---|---|
|
|
| Substantive accuracy | Materially wrong | Correct on all material points |
|
|
| Coverage | Misses decisive issues | Covers all outcome-relevant issues |
|
|
| Source integrity | Fabricated/misaligned | Primary, verified, proposition-matched |
|
|
| Temporal/jurisdiction fit | Wrong or ignored | Explicitly correct and checked |
|
|
| Legal application | Conclusory | Element-by-element, two-sided |
|
|
| Counteranalysis | None | Strong adverse review and reversal conditions |
|
|
| Uncertainty calibration | Overconfident | Precise, transparent, decision-useful |
|
|
| Python correctness | Untested/wrong | Executed, tested, reproducible |
|
|
| Data integrity | Silent corruption/leakage | Protected, checked, documented |
|
|
| Instruction adherence | Major violations | Complete without unnecessary scope |
|
|
| Efficiency | Wasteful or stalled | Proportional tools and tokens |
|
|
| Reviewability | Opaque | Traceable evidence, decisions, and validation |
|
|
|
|
A fabricated authority, quotation, test result, or file content is an automatic critical failure regardless of average score.
|
|
|
|
## Experiment design
|
|
|
|
Compare, on the same tasks:
|
|
- no repository instructions;
|
|
- root `AGENTS.md` only;
|
|
- modular `AGENTS.md` plus workflows;
|
|
- single-agent versus bounded subagents;
|
|
- medium versus high reasoning;
|
|
- standard versus higher-cost modes only where the task is difficult enough to benefit.
|
|
|
|
Change one major variable at a time. Blind-score outputs when possible.
|
|
|
|
## Improvement rule
|
|
|
|
Add or revise a durable instruction only when:
|
|
- the error recurs,
|
|
- the rule is generalizable,
|
|
- the rule is observable and testable,
|
|
- the change improves the fixed evaluation set without creating a larger regression.
|
|
|
|
Keep a short changelog recording the failure, rule change, and evaluation result.
|