For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Build a packaged benchmark CLI that runs deterministic mock or live OpenAI-compatible comparisons for both invocation modes and writes secret-safe JSON/Markdown reports.
Architecture: Add a benchmark application module for config, case catalog, run models, timing and aggregation; a reporting infrastructure module for redaction/rendering/atomic writes; and a thin presentation CLI. Each run uses DebugRuntime.start_session() with a temporary SQLite store and timing chat-client wrapper.
Tech Stack: Python 3.11+, asyncio, Pydantic, argparse, httpx-based existing OpenAI-compatible client, SQLite, pytest, uv.
Files:
src/agent_lab/application/benchmark.pytests/test_benchmark_config.pyModify: tests/test_tool_invocation_comparison.py
[ ] Write failing tests for strict schema version, URL validation, known case/mode IDs, positive runs, and rejection of secret fields.
[ ] Verify RED with uv run pytest tests/test_benchmark_config.py -q.
[ ] Implement strict Pydantic models BenchmarkConfig, BenchmarkTarget, BenchmarkCase, BenchmarkExpectation, and immutable built-in case catalog.
[ ] Implement build_benchmark_request(case, mode, model) using the existing DebugRunRequest, identical params, and built-in enabled tools.
[ ] Expose deterministic mock rounds from the production case catalog and refactor the existing comparison test to reuse the catalog instead of duplicating prompts/expectations.
[ ] Run config and comparison tests, then commit with feat: add benchmark case catalog.
Files:
src/agent_lab/application/benchmark.pyCreate: tests/test_benchmark_runner.py
[ ] Write failing tests for provider TTFT versus visible TTFT, per-call elapsed/usage, persisted wall time, detected events/sources, tool/fallback counts, semantic failures, partial-failure continuation, and temporary-database isolation.
[ ] Verify RED with uv run pytest tests/test_benchmark_runner.py -q.
[ ] Implement TimingChatClient around the existing stream protocol, recording first item, first message delta, usage, elapsed, and fallback classification.
[ ] Implement deterministic MockBenchmarkChatClient that never constructs HTTP transport and emits mode-correct streams for the six cases.
[ ] Implement BenchmarkRunner.run() over case → mode → iteration, creating/closing one runtime, chat client, and temporary SQLite store per run.
[ ] Implement semantic evaluation from output, audit, and usage ledgers and return ordered BenchmarkRunResult records without stopping after an individual failure.
[ ] Run runner/config/comparison tests, then commit with feat: run isolated agent benchmarks.
Files:
src/agent_lab/infrastructure/benchmark_reporting.pysrc/agent_lab/presentation/benchmark_cli.pytests/test_benchmark_reporting.pytests/test_benchmark_cli.pyModify: pyproject.toml
[ ] Write failing tests for linear-interpolated p50/p95, failed-sample exclusion, low-sample warnings, redaction, JSON/Markdown consistency, atomic writes, CLI overrides, mock/live API-key requirements, exit codes 0/1/2, and continued report writing after run failures.
[ ] Verify RED with uv run pytest tests/test_benchmark_reporting.py tests/test_benchmark_cli.py -q.
[ ] Implement aggregation into one typed BenchmarkReport model shared by JSON and Markdown.
[ ] Implement centralized redaction and 2 KiB error truncation before any rendered output.
[ ] Implement atomic matching .json/.md writes under the configured output directory.
[ ] Implement argparse with --config, --base-url, --model, repeatable --case/--mode, --runs, --output-dir, and --mock; apply CLI > JSON and environment-only API key.
[ ] Register agent-lab-benchmark = agent_lab.presentation.benchmark_cli:main and verify --help.
[ ] Run reporting/CLI/runner tests, then commit with feat: add benchmark reports and cli.
Files:
benchmarks/single-target.example.json.gitignoreREADME.mddocs/plans/todo-52-agent-benchmark-runner.mdModify: docs/plans/todos.md
[ ] Add a secret-free example covering both modes and all six cases; ignore outputs/benchmarks/.
[ ] Document .env, config copy/edit, mock command, live command, output fields, exit codes, fairness controls, and the limits of scripted/mock adapters.
[ ] Run agent-lab-benchmark --config benchmarks/single-target.example.json --mock --output-dir /tmp/agent-lab-benchmark and inspect both reports.
[ ] Run focused benchmark tests, uv run pytest, CLI help, app import, node --check, and git diff --check.
[ ] Perform specification review then code-quality review; add Todo 52.x rows for post-execution fixes.
[ ] Record evidence and Workflow Evolution findings, mark Todo 52 done, and commit with docs: close benchmark runner todo.