# Todo 51 Comparison Matrix Closeout Plan **Status:** in_progress ## Goal Close the service experiment with one repeatable behavior matrix, current operational documentation, explicit comparison controls, and bounded conclusions. ## Scenario Matrix Run the same six business scenarios in both `dual_agent` and `chat_agent_tools` modes (12 parameterized runs): 1. ordinary chat with all built-ins enabled and no event; 2. absolute device-volume adjustment with silent success; 3. fixed RFC3339 schedule creation with one deterministic confirmation; 4. web search with first answer, tool result, then exactly one second answer; 5. session termination with exactly one farewell and terminal completion; 6. concurrent volume plus schedule, bounded at two workers, preserving input result order and one schedule template. The matrix compares shared user-visible semantics and normalized business results. It deliberately allows provider transcript and event-selection protocol differences. ## Fair Comparison Controls - Same model parameters, enabled event set, user inputs, output budget, registry, plugin schemas, handlers, event budget, concurrency, and timeouts. - Direct-mode provider tool protocol tokens remain counted. - Token totals include only real model calls; tool rows have zero tokens. - `turn_wall_time_ms` is the only turn/session wall-clock total; overlapping model/tool spans are not summed. - Deduplicated/replayed event clones count as one logical tool execution. - Web search requires two ChatAgent answers in both modes. - Provider TTFT and visible-text TTFT remain distinct. ## Documentation Scope - Rewrite `README.md` as the operational entry for architecture, built-ins, modes, configuration, UI/WebSocket/API use, metrics, verification, and current adapter limitations. - Calibrate `docs/plans/2026-07-12-chat-event-agent-production-events-design.md` to implemented namespaced event names and link this closeout; keep it as the architecture fact source. - Record matrix evidence, conclusions, and remaining risks in this file. - Update `docs/plans/todos.md` only when final verification passes. ## Files - Add: `tests/test_tool_invocation_comparison.py` - Modify: `README.md` - Modify: `docs/plans/2026-07-12-chat-event-agent-production-events-design.md` - Modify: this file and `docs/plans/todos.md` during final Update. ## Requirement Execution Loop ### Plan 1. Add the 12-run RED comparison matrix with small local scripted clients and injected production plugin ports. 2. Fix only genuine shared-contract gaps exposed by the matrix; use Todo 51.x for any post-execution repair. 3. Update README and architecture design names/status from the verified implementation. 4. Run matrix, focused, full, JS syntax, app import, and diff checks. 5. Perform specification and quality review, then record evidence and bounded experiment conclusions. ## Verification - `uv run pytest tests/test_tool_invocation_comparison.py -q` - `uv run pytest tests/test_debug_runtime.py tests/test_builtin_event_plugins.py tests/test_sqlite_store.py tests/test_websocket_api.py -q` - `uv run pytest` - `node --check src/agent_lab/presentation/static/app.js` - `uv run python -c "from agent_lab.main import app; print(app.title)"` - `git diff --check`