ソースを参照

docs: plan configurable tool invocation modes

Problem: the runtime has no backend switch for comparing the existing dual-agent text-event flow with direct ChatAgent provider tools on the same registry.

Risk: direct mode requires a valid assistant/tool transcript, enabled-name enforcement, and event-budget handling without changing the default dual-agent behavior.
zhenyu.hu 2 週間 前
コミット
a3ba43b36f
1 ファイル変更83 行追加0 行削除
  1. 83 0
      docs/plans/todo-46-configurable-tool-invocation-mode.md

+ 83 - 0
docs/plans/todo-46-configurable-tool-invocation-mode.md

@@ -0,0 +1,83 @@
+# Todo 46 Configurable Tool Invocation Mode Plan
+
+**Status:** in_progress
+
+## Goal
+
+Add a backend-selectable comparison between the existing fast-reply `dual_agent` path and a `chat_agent_tools` path where ChatAgent receives provider tools directly, while both modes use the same registry, schemas, handlers, event budget, and externally visible tool results.
+
+## Decisions
+
+- Add `tool_invocation_mode` to `DebugRunRequest`, defaulting to `dual_agent` for backward compatibility.
+- Keep mode at the orchestration/request level, not in model parameters or process settings.
+- `dual_agent` keeps the existing compact text-event prompt, EventAgent queue, deterministic-first resolution, optional LLM fallback, and summary follow-up.
+- `chat_agent_tools` sends the enabled registry schemas to ChatAgent, consumes only `provider_tool_call`, executes those calls through the shared `EventKernel` with `PROVIDER_RESOLVED`, and continues with a valid assistant/tool provider transcript.
+- Provider-resolved arguments are not rewritten by EventAgent fallback in this todo.
+- UI selection and reconnect/session transcript restoration remain Todo 49.
+
+## Scope
+
+- Validate the two supported mode values and reject unknown values.
+- Mode-select ChatAgent prompt injection, tools, and accepted stream event type.
+- Execute one or multiple direct provider calls through the shared registry/kernel with enabled-name enforcement.
+- Preserve provider call ID, name, arguments, raw arguments, assistant visible content, tool-call order, and tool-reply order.
+- Append direct assistant `tool_calls` followed immediately by matching tool messages before the next ChatAgent request.
+- Reuse `max_event_loops`; after budget exhaustion, send no tools and ignore further provider tool calls.
+- Include mode in audit/session config snapshots without changing the current UI payload.
+- Cover both one-shot `run()` and reusable session/turn runtime paths.
+
+## Boundaries
+
+- Do not modify parser, provider HTTP client, or message model unless a failing contract proves a gap; typed calls and transcript serialization already exist.
+- Do not route direct mode through the EventAgent request/summary queue.
+- Do not let direct transcripts contaminate `dual_agent`; dual mode must continue to ignore provider tool calls.
+- Do not implement built-in production event plugins, bounded batch scheduling, UI controls, restore migrations, or comparison metrics in this todo.
+- Unknown or disabled provider tool names must produce a structured tool result and never execute a handler.
+
+## Files
+
+- Modify: `src/agent_lab/application/contracts.py`
+- Modify: `src/agent_lab/application/runtime.py`
+- Modify: `tests/test_debug_runtime.py`
+- Modify: `tests/test_websocket_api.py`
+- Modify: `src/agent_lab/application/tools.py` only if an enabled-name provider execution facade is needed to avoid Runtime reaching into registry internals.
+
+## Requirement Execution Loop
+
+### Plan
+
+1. Add RED contract tests for default, valid modes, and invalid mode rejection.
+2. Add RED runtime tests proving mode-specific prompt/tool payloads and typed stream consumption.
+3. Add RED direct-mode transcript tests for mixed visible content plus one/multiple provider tool calls.
+4. Add RED execution tests for enabled, disabled, unknown, handler failure, and exhausted event budget.
+5. Implement the smallest mode helpers shared by one-shot and session turn loops; preserve existing audit/output contracts.
+6. Run focused/full verification and independent specification/quality reviews.
+
+### Act
+
+- In `dual_agent`, assert ChatAgent receives `tools=[]`, receives compact event instructions, ignores `provider_tool_call`, and continues through EventAgent summary exactly as before.
+- In `chat_agent_tools`, assert ChatAgent receives exactly the enabled schemas and no compact text-event system block.
+- Collect only provider calls in direct mode; emit existing `event` output/audit records with full provider arguments.
+- Create one assistant message containing both visible content and ordered `tool_calls`.
+- Execute calls using `enabled_names`; create ordered `tool` replies with matching IDs, emit `tool_result`, append them to active history, and continue ChatAgent.
+- Assert the second provider request transcript is `assistant(tool_calls)` followed by all matching tool replies.
+- Assert direct mode ignores text events and dual mode ignores provider calls.
+- Assert the next round sends `tools=[]` after `max_event_loops` is reached.
+- Include `tool_invocation_mode` in session config and relevant round/request audit details.
+
+### Observe
+
+Compare both paths against the same fake registry and scenarios. Confirm the architectural difference is limited to model/tool orchestration, not schemas, handlers, results, or event budget.
+
+### Update
+
+- Record verification evidence and review findings.
+- Add `46.x` rows for any post-commit correction.
+- Mark Todo 46 done only when dual-agent regressions and direct provider transcripts both pass.
+
+## Verification
+
+- `UV_CACHE_DIR=.uv-cache uv run pytest tests/test_debug_runtime.py tests/test_websocket_api.py -q`
+- `UV_CACHE_DIR=.uv-cache uv run pytest`
+- `UV_CACHE_DIR=.uv-cache uv run python -c "from agent_lab.main import app; print(app.title)"`
+- `git diff --check`