|
|
@@ -0,0 +1,39 @@
|
|
|
+# Todo 52.12 Benchmark Metric Integrity and Parallel Proof Plan
|
|
|
+
|
|
|
+**Status:** in_progress
|
|
|
+
|
|
|
+## Goal
|
|
|
+
|
|
|
+Ensure benchmark latency, token-cost, and multi-event concurrency results measure the claimed behavior rather than consumer overhead, persistence defaults, or event cardinality alone.
|
|
|
+
|
|
|
+## Observation
|
|
|
+
|
|
|
+- `TimingChatClient.elapsed_ms` spans the outer generator's `yield`, so downstream Runtime/audit/queue work and consumer pauses inflate the value described as pure model-call elapsed.
|
|
|
+- Successful calls without a provider usage chunk are persisted as zero-token rows and currently become a false `0 tokens` benchmark result.
|
|
|
+- The parallel case verifies two successful events but does not prove their handlers overlap; a serial executor can still pass the benchmark.
|
|
|
+
|
|
|
+## Scope
|
|
|
+
|
|
|
+- Measure model elapsed as the accumulated time awaiting the wrapped provider iterator, excluding time while the caller processes yielded items.
|
|
|
+- Preserve provider TTFT and first-visible timing semantics independently from provider-await elapsed.
|
|
|
+- Require complete usage coverage across all model/fallback calls before reporting token totals. Missing usage produces nullable token metrics plus an explicit semantic failure instead of zero or a partial sum.
|
|
|
+- Add a benchmark-owned execution probe for the parallel volume/schedule case that deterministically requires both handlers to start before either completes.
|
|
|
+- Record the parallel-overlap result in `BenchmarkRunResult`, include it in report output, and fail the parallel case if overlap is absent.
|
|
|
+- Keep mock and live runs on the same probe/registry path and keep non-parallel case behavior unchanged.
|
|
|
+
|
|
|
+## Files
|
|
|
+
|
|
|
+- Modify: `src/agent_lab/application/benchmark.py`
|
|
|
+- Modify: `src/agent_lab/infrastructure/benchmark_reporting.py`
|
|
|
+- Modify: `tests/test_benchmark_runner.py`
|
|
|
+- Modify: `tests/test_benchmark_reporting.py`
|
|
|
+- Modify: `README.md`
|
|
|
+
|
|
|
+## Verification
|
|
|
+
|
|
|
+- `uv run pytest tests/test_benchmark_runner.py tests/test_benchmark_reporting.py -q`
|
|
|
+- `uv run pytest tests/test_benchmark_config.py tests/test_benchmark_runner.py tests/test_benchmark_reporting.py tests/test_benchmark_cli.py tests/test_tool_invocation_comparison.py -q`
|
|
|
+- zero-network mock CLI still reports 12 attempted runs and validates parallel overlap
|
|
|
+- `uv run pytest`
|
|
|
+- `git diff --check`
|
|
|
+
|