|
|
@@ -0,0 +1,29 @@
|
|
|
+# Todo 52.6 Benchmark Report Auditability Plan
|
|
|
+
|
|
|
+**Status:** in_progress
|
|
|
+
|
|
|
+## Goal
|
|
|
+
|
|
|
+Make benchmark reports complete and self-explanatory even under concurrent writers or orchestration-level failures.
|
|
|
+
|
|
|
+## Scope
|
|
|
+
|
|
|
+- Aggregate successful model-call elapsed samples and tool handler latency samples with p50/p95 and low-sample warnings.
|
|
|
+- Reserve each output basename using an exclusive lock/reservation file before choosing paths, so concurrent processes cannot publish mixed JSON/Markdown pairs.
|
|
|
+- Add redacted runner/system errors to the report model; count a catastrophic runner failure in overall attempted/failed values even when no case result exists.
|
|
|
+- Emit low-confidence warnings for zero as well as 1–19 successful samples.
|
|
|
+- Redact first, then truncate errors to at most 2048 UTF-8 bytes without producing invalid text.
|
|
|
+
|
|
|
+## Files
|
|
|
+
|
|
|
+- Modify: `src/agent_lab/infrastructure/benchmark_reporting.py`
|
|
|
+- Modify: `src/agent_lab/presentation/benchmark_cli.py`
|
|
|
+- Modify: `tests/test_benchmark_reporting.py`
|
|
|
+- Modify: `tests/test_benchmark_cli.py`
|
|
|
+
|
|
|
+## Verification
|
|
|
+
|
|
|
+- `uv run pytest tests/test_benchmark_reporting.py tests/test_benchmark_cli.py -q`
|
|
|
+- concurrent writer stress test
|
|
|
+- `uv run pytest`
|
|
|
+- `git diff --check`
|