# Agent Lab MVP Todos ## Workflow - Main session owns this file and evaluates progress after each todo. - Work on one todo at a time: write its plan, execute it, verify it, then update this file. - Do not plan all implementation details upfront. Later todos may change after each evaluation. - Detailed plans live beside this file as `docs/plans/todo-N-*.md`. - Status values: `pending`, `in_progress`, `done`, `blocked`. ## Current State - Branch: `codex/agent-lab-mvp`. - MVP committed at `6babafa` with FastAPI, WebSocket, static UI, OpenAI-compatible streaming, EventAgent handoff, provider-compatibility guards, and tests. - Post-MVP work now evolves toward the original queue/event-driven debugging-console requirements. ## Coverage Review - OpenAI Chat Completions compatible streaming and provider payload guards are covered by todos 2, 5, and 7. - Queue/event-driven ChatAgent and EventAgent handoff is covered by todo 9. - EventAgent tool loading and UI management is covered by todo 10. - Prompt/pre-message editing, parameter controls, and reusable prompt sets are covered by todos 3 and 11. - Chat-window token, cached-token, TTFT, elapsed-time, and per-round stats are covered by todo 12. - Review follow-up gaps in real queue boundaries, API validation, packaging, and tool error isolation are covered by todos 13-16. ## Todos | ID | Status | Plan | Goal | Verification | | --- | --- | --- | --- | --- | | 1 | done | `docs/plans/todo-1-project-foundation.md` | Establish uv/FastAPI project skeleton, DDD package layout, and core failing tests for parser, event agent, and runtime routing. | `uv run pytest` fails only because implementation is missing, not because tests or imports are malformed. | | 2 | done | `docs/plans/todo-2-domain-runtime.md` | Implement domain models, OpenAI Chat Completions compatible stream parsing, EventAgent tools, and the direct async runtime loop. | Core tests pass with mocked ChatAgent client. | | 3 | done | `docs/plans/todo-3-websocket-api.md` | Add FastAPI config, HTTP routes, WebSocket protocol, static debug page, and assistant transcript persistence for multi-turn LLM calls. | WebSocket contract tests pass; app imports successfully. | | 4 | done | `docs/plans/todo-4-real-llm-smoke.md` | Document env-based real LLM run flow, add a lightweight manual smoke helper, and make sure settings names are clear. | App starts with env-based config; docs show exact `uv` commands and required env vars. | | 5 | done | `docs/plans/todo-5-tool-call-history.md` | Harden Chat Completions history so assistant tool calls and tool replies are sent in provider-compatible order. | Tests prove the next model call includes assistant `tool_calls` before `tool` replies. | | 6 | done | `docs/plans/todo-6-cleanup-closeout.md` | Clean temporary/generated files, update contributor docs, and run final verification. | Git status contains only intended source/docs files; `uv run pytest` passes. | | 7 | done | `docs/plans/todo-7-provider-compatibility.md` | Fix review findings for real Chat Completions compatible providers: default model fallback, role-aware message serialization, and configurable usage streaming. | Tests prove provider payload shape and full suite passes. | | 8 | done | `docs/plans/todo-8-tool-pre-message-guard.md` | Prevent manually configured pre-messages from producing invalid provider `tool` payloads without `tool_call_id`. | Tests prove invalid tool pre-messages are rejected and UI no longer exposes `tool` role. | | 9 | done | `docs/plans/todo-9-explicit-runtime-queues.md` | Refactor runtime internals to explicit input, output, and event queues while preserving the WebSocket stream contract. | Tests prove ChatAgent consumes input queue, EventAgent consumes event queue, and output queue drives streamed responses. | | 10 | done | `docs/plans/todo-10-tool-registry-management.md` | Replace the single hardcoded EventAgent tool checkbox with a backend tool registry and UI-managed loaded tools. | Tests prove available tools are listed by API and selected tools control EventAgent behavior. | | 11 | done | `docs/plans/todo-11-prompt-workspace.md` | Improve prompt/pre-message editing with browser-side persistence and reusable prompt sets. | UI tests or focused JS tests prove prompts persist and can be restored. | | 12 | done | `docs/plans/todo-12-round-stats.md` | Move round statistics to structured backend events for token counts, cached tokens, TTFT, elapsed time, and per-turn summaries. | `uv run pytest tests/test_debug_runtime.py tests/test_websocket_api.py`; `uv run pytest`. | | 13 | done | `docs/plans/todo-13-real-queue-websocket-loop.md` | Make runtime queues real producer/consumer boundaries and let WebSocket run separate upstream/downstream tasks. | `uv run pytest tests/test_debug_runtime.py tests/test_websocket_api.py`; `uv run pytest`. | | 14 | done | `docs/plans/todo-14-message-role-validation.md` | Reject provider-invalid API message roles and malformed tool replies before network calls. | `uv run pytest tests/test_websocket_api.py tests/test_event_agent.py`; `uv run pytest`. | | 15 | done | `docs/plans/todo-15-static-package-data.md` | Include static UI assets in installed package builds. | `uv run pytest tests/test_packaging.py`; `uv run pytest`; `uv build`; wheel inspection. | | 16 | done | `docs/plans/todo-16-tool-error-isolation.md` | Convert EventAgent tool handler exceptions into structured tool-result errors instead of aborting the session. | `uv run pytest tests/test_event_agent.py tests/test_debug_runtime.py`; `uv run pytest`. | | 17 | done | `docs/plans/todo-17-agent-config-audit-tools.md` | Align ChatAgent/EventAgent config, add extra_body defaults, default EventAgent to one event round, add mock tools, audit events, and backend logging. | `uv run pytest` passes (`42 passed`, one existing Starlette deprecation warning). | | 18 | done | `docs/plans/todo-18-prompt-modal-tool-ui.md` | Move prompt workspace into a modal and improve visible tool-management UI. | `uv run pytest` passes (`46 passed`), plus local browser check confirms modal and tool UI render. | | 19 | done | `docs/plans/todo-19-latency-oriented-polish.md` | Do one optimization pass focused on reducing perceived reply wait time and cleaning frontend/backend rough edges. | `uv run pytest` passes (`49 passed`), plus local browser check confirms tools are visible in the first viewport. | | 20 | done | `docs/plans/todo-20-agent-config-modals.md` | Split ChatAgent and EventAgent settings into separate configuration modals so the sidebar is no longer crowded. | `uv run pytest` passes. | | 21 | done | `docs/plans/todo-21-event-agent-context-summary.md` | Make EventAgent config part of parameter generation context and enqueue an aggregated tool-result summary for the next ChatAgent round. | `uv run pytest` passes (`50 passed`, one existing Starlette deprecation warning). | | 22 | done | `docs/plans/todo-22-text-event-protocol.md` | Replace ChatAgent tool calls with streamed text event parsing, inject generated event descriptions as a system message, and harden Agent config buttons. | `uv run pytest` passes (`50 passed`, one existing Starlette deprecation warning), plus local browser click check. | | 23 | done | `docs/plans/todo-23-agent-prompt-and-event-llm.md` | Move Prompt Config into ChatAgent config, show available-event prompt by default, and make EventAgent use LLM tool calls to generate tool arguments. | `uv run pytest` passes (`51 passed`, one existing Starlette deprecation warning); changed Python modules compile; local server returned page/static assets/tools API. | | 24 | done | `docs/plans/todo-24-chat-agent-prompt-list.md` | Put all ChatAgent prompts, including generated EventAgent event rules, into one ordered prompt list with add/delete/drag controls. | `uv run pytest` passes (`51 passed`, one existing Starlette deprecation warning), plus local browser prompt-list check. | | 25 | done | `docs/plans/todo-25-prompt-list-message-types.md` | Make the ChatAgent prompt list visually compact and show the message type for each prompt item. | `uv run pytest` passes (`51 passed`, one existing Starlette deprecation warning); local service returned `/health` and versioned JS asset. | | 26 | done | `docs/plans/todo-26-workspace-snapshot-save.md` | Move save/load/delete out of ChatAgent config and save one workspace snapshot covering prompts, Agent config, and selected tools. | `uv run pytest tests/test_websocket_api.py -q` passes (`27 passed`, one existing Starlette deprecation warning). | | 27 | done | `docs/plans/todo-27-event-agent-tool-arguments.md` | Make EventAgent tool argument generation robust for compatible providers that do not emit the expected tool-call finish reason or ignore tool calls. | `uv run pytest` passes (`53 passed`, one existing Starlette deprecation warning). | | 28 | done | `docs/plans/todo-28-session-turns.md` | Split WebSocket session lifetime from user turns so one conversation can run multiple user messages while each turn has its own event budget. | `uv run pytest` passes (`56 passed`, one existing Starlette deprecation warning). | | 29 | done | `docs/plans/todo-29-rich-audit-log.md` | Add audit entries for ChatAgent/EventAgent model parameters, prompts/messages, results, and usage without logging secrets. | `uv run pytest` passes (`57 passed`, one existing Starlette deprecation warning). | | 30 | done | `docs/plans/todo-30-sqlite-session-store.md` | Add a SQLite-backed store and REST APIs for sessions, chat messages, audit replay data, and usage summaries. | `uv run pytest tests/test_sqlite_store.py tests/test_websocket_api.py -q` passes (`36 passed`, one existing Starlette deprecation warning); `uv run pytest` passes (`66 passed`, one existing warning). | | 31 | done | `docs/plans/todo-31-runtime-persistence.md` | Persist runtime Agent config snapshots, user/assistant messages, audit entries, and per-call usage during WebSocket turns. | `uv run pytest tests/test_sqlite_store.py tests/test_debug_runtime.py tests/test_websocket_api.py -q` passes (`55 passed`, one existing warning); `uv run pytest` passes (`66 passed`, one existing warning). | | 32 | done | `docs/plans/todo-32-session-ui.md` | Add new-session/session-list UI, separate audit replay panel, and single-call/turn/session usage displays. | `node --check src/agent_lab/presentation/static/app.js`; `uv run pytest tests/test_websocket_api.py -q` passes (`34 passed`, one existing warning); `uv run pytest` passes (`66 passed`, one existing warning). | | 33 | done | `docs/plans/todo-33-audit-replay-stream-ui.md` | Redesign Audit Replay as a clickable event stream, keep audit events out of chat, and fix fixed/scrolling layout boundaries. | `node --check src/agent_lab/presentation/static/app.js`; `uv run pytest tests/test_websocket_api.py -q` passes (`36 passed`, one existing warning); `uv run pytest` passes (`68 passed`, one existing warning); `git diff --check` passes. | | 34 | done | `docs/plans/todo-34-event-agent-history-boundary.md` | Restrict EventAgent context to conversational user/assistant messages instead of the full ChatAgent context list. | `uv run pytest tests/test_debug_runtime.py -q` passes (`21 passed`); `node --check src/agent_lab/presentation/static/app.js`; `uv run pytest` passes (`70 passed`, one existing warning); `git diff --check` passes. | | 35 | done | `docs/plans/todo-35-audit-modal-raw-llm-events.md` | Make Audit Replay a modal, fix round event metadata after EventAgent follow-up rounds, and expose raw LLM response chunks. | `node --check src/agent_lab/presentation/static/app.js`; `uv run pytest tests/test_debug_runtime.py tests/test_event_agent.py tests/test_websocket_api.py -q` passes (`63 passed`, one existing warning); `uv run pytest` passes (`71 passed`, one existing warning); `git diff --check` passes. | | 36 | done | `docs/plans/todo-36-audit-detail-compact-ui.md` | Make Audit Detail modal metadata and event lists compact, and version CSS to avoid stale oversized styling. | `node --check src/agent_lab/presentation/static/app.js`; `uv run pytest tests/test_websocket_api.py -q` passes (`36 passed`, one existing warning); `uv run pytest` passes (`71 passed`, one existing warning); `git diff --check` passes. | | 37 | done | `docs/plans/todo-37-audit-detail-flow-layout.md` | Fully replace Audit Detail grid/card layout with compact flow layout so metadata cannot stretch vertically. | `node --check src/agent_lab/presentation/static/app.js`; `uv run pytest tests/test_websocket_api.py -q` passes (`36 passed`, one existing warning); `uv run pytest` passes (`71 passed`, one existing warning); `git diff --check` passes. | | 38 | done | `docs/plans/todo-38-audit-round-event-semantics.md` | Make `events_enabled` mean current-round enabled events and move full configured event list to `configured_events`. | `node --check src/agent_lab/presentation/static/app.js`; `uv run pytest tests/test_debug_runtime.py tests/test_websocket_api.py -q` passes (`58 passed`, one existing warning); `uv run pytest` passes (`71 passed`, one existing warning); `git diff --check` passes. | | 39 | done | `docs/plans/todo-39-audit-turn-relative-time.md` | Add per-audit-event elapsed time relative to the user message that started the current turn. | `uv run pytest` passes (`72 passed`, one existing Starlette deprecation warning); `node --check src/agent_lab/presentation/static/app.js`; `git diff --check` passes. | | 40 | done | `docs/plans/todo-40-chat-message-stream-audit.md` | Add audit events before ChatAgent visible message streaming starts and after it ends. | `uv run pytest` passes (`73 passed`, one existing Starlette deprecation warning); `node --check src/agent_lab/presentation/static/app.js`; `git diff --check` passes. | | 41 | done | `docs/plans/todo-41-session-button-feedback.md` | Add explicit Sessions-panel feedback for New, Refresh, and Load clicks. | `uv run pytest` passes (`73 passed`, one existing Starlette deprecation warning); `node --check src/agent_lab/presentation/static/app.js`; `git diff --check` passes. | | 42 | done | `docs/plans/todo-42-session-switch-idle-socket.md` | Allow switching sessions after a turn completes while the reusable WebSocket is still open. | `uv run pytest tests/test_websocket_api.py -q` passes (`37 passed`, one existing Starlette deprecation warning); `uv run pytest` passes (`74 passed`, one existing warning); `node --check src/agent_lab/presentation/static/app.js`; `git diff --check` passes. | | 43 | done | `docs/plans/todo-43-typed-provider-tool-calls.md` | Preserve native provider tool calls as a distinct stream event without changing the existing text-event path. | `uv run pytest tests/test_openai_stream_parser.py -q`; focused typed-event regressions; full suite passes (`82 passed`, one existing warning). | | 43.1 | done | `docs/plans/todo-43.1-typed-event-consumer-regression.md` | Repair Runtime and EventAgent consumers after typed stream kinds exposed legacy `event` assumptions. | Runtime consumes `text_event`, EventAgent consumes `provider_tool_call`; focused and full tests pass. | | 43.2 | done | `docs/plans/todo-43.2-remove-legacy-event-consumption.md` | Remove legacy `event` consumption from production paths and migrate remaining fixtures to typed stream kinds. | Wrong-source negative tests and real client `[DONE]` integration pass; full suite passes (`82 passed`). | | 44 | in_progress | `docs/plans/todo-44-provider-tool-transcript.md` | Add provider-compatible assistant tool-call and tool-result message serialization plus explicit tool-choice behavior. | Client/domain tests prove legal transcripts serialize and malformed tool history is rejected. | | 45 | pending | `docs/plans/todo-45-configurable-tool-invocation-mode.md` | Add `dual_agent` and `chat_agent_tools` runtime modes that share one registry and user-visible behavior. | Runtime tests prove mode-specific requests, direct execution, valid provider history, and dual-mode regression safety. | | 46 | pending | `docs/plans/todo-46-flat-event-policies-hybrid-resolution.md` | Extend flat event definitions with deterministic-first argument resolution, LLM fallback, and generic result policies. | Tests prove deterministic events skip EventAgent LLM and ambiguous cases use one bounded fallback. | | 47 | pending | `docs/plans/todo-47-production-event-definitions.md` | Add experimental terminate-session, device-volume, schedule, and web-search event definitions with normalized contracts. | Focused tests cover valid, ambiguous, invalid, idempotent, and failure results for all four events. | | 48 | pending | `docs/plans/todo-48-parallel-event-batches.md` | Execute multiple flat events with bounded parallelism, failure isolation, aggregation, and a terminal barrier. | Runtime tests prove independent concurrency, stable result order, isolated failures, and termination after sibling completion or timeout. | | 49 | pending | `docs/plans/todo-49-tool-mode-ui-session-restore.md` | Add UI mode selection, workspace/session restoration, and clear EventAgent applicability feedback. | Static/WebSocket tests prove requests, snapshots, loaded sessions, and audit context preserve the selected mode. | | 50 | pending | `docs/plans/todo-50-comparable-model-tool-metrics.md` | Record mode-aware model/tool spans and include EventAgent usage so both architectures can be compared fairly. | Persistence/API tests report agent, call kind, tokens, TTFT, tool latency, fallback usage, and total turn wall time without double counting. | | 51 | pending | `docs/plans/todo-51-comparison-matrix-closeout.md` | Add repeatable comparison fixtures, update project documentation, and close the experimental implementation. | Both modes pass the same scenario matrix; full tests, JS syntax, app import, and documentation checks pass. |