|
|
il y a 2 semaines | |
|---|---|---|
| .vscode | il y a 3 semaines | |
| benchmarks | il y a 2 semaines | |
| docs | il y a 2 semaines | |
| scripts | il y a 3 semaines | |
| src | il y a 2 semaines | |
| tests | il y a 2 semaines | |
| .env.example | il y a 3 semaines | |
| .gitignore | il y a 2 semaines | |
| AGENTS.md | il y a 3 semaines | |
| LICENSE | il y a 3 semaines | |
| README.md | il y a 2 semaines | |
| pyproject.toml | il y a 2 semaines | |
| uv.lock | il y a 2 semaines |
Agent Lab is a FastAPI debug console for comparing ChatAgent capability-invocation protocols on one shared event runtime. The implementation uses a generic event kernel plus a flat registry of built-in plugins; invocation mode changes how the model selects capabilities, not how accepted events are normalized, validated, scheduled, executed, persisted, or reported.
The architecture source of truth is the production events design. The repeatable experiment and closeout scope is tracked by Todo 51.
The generic kernel owns event lookup, deterministic resolution, validation, bounded fallback, deduplication, conflict-aware concurrency, timeouts, terminal barriers, normalized results, result-policy aggregation, audit events, and metrics. Runtime and the kernel do not branch on a built-in event name.
The four production-shaped plugins are flat, namespaced definitions:
| Event | Result policy | Current default behavior |
|---|---|---|
session.terminate |
terminate |
Records an in-memory terminated result and ends the debug session on success. |
device.volume.adjust |
silent_success |
Validates and records an in-memory volume action; success adds no second answer. |
calendar.schedule.create |
template_follow_up |
Records an in-memory schedule and emits a deterministic confirmation. |
knowledge.web.search |
llm_follow_up |
Returns deterministic mock sources, then asks ChatAgent for one grounded follow-up. |
The default registry also retains the lab-only handoff_note, mock_search, and
mock_ticket tools. GET /api/tools is the authoritative runtime catalog.
Set tool_invocation_mode per request:
| Mode | Selection protocol | Argument source | Mode-specific cost or dependency |
|---|---|---|---|
dual_agent |
ChatAgent receives a compact catalog and emits hidden text events after visible content. | Shared deterministic resolvers run first; EventAgent may make a bounded fallback model call for incomplete supported events. | Depends on marker adherence and adds fallback calls when deterministic resolution is incomplete. |
chat_agent_tools |
ChatAgent receives enabled provider tool schemas and emits native tool calls. | ChatAgent proposes arguments; the shared kernel still normalizes, validates, authorizes, deduplicates, schedules, and executes them. | Full schemas and provider tool transcript remain in ChatAgent context and usage; EventAgent model/prompt settings are inactive for provider-resolved calls. |
Both modes share the same enabled definitions, JSON schemas, handlers/adapters, event budget, batch concurrency, timeout, conflict and terminal policies, normalized results, and user-visible success/failure contract. Provider-native assistant/tool transcript details are intentionally allowed to differ.
For knowledge.web.search, both modes implement the same visible sequence:
The first answer must not wait for the search adapter, and the second answer must use the compact returned evidence. This is a shared behavior contract, not a provider transcript normalization trick.
The default application wiring is safe for local experiments, but it is not production integration:
example.invalid mock sources and does not access the internet.Inject implementations of SessionTerminationPort, DeviceVolumePort,
CalendarSchedulePort, and WebSearchPort when testing service-owned adapters.
Real device, calendar, session, and internet integration belongs outside the
default Agent Lab wiring.
Install dependencies and create a local environment file:
uv sync
cp .env.example .env
Supported settings use the AGENT_LAB_ prefix:
AGENT_LAB_OPENAI_API_KEY: API key for the compatible provider;AGENT_LAB_OPENAI_BASE_URL: provider /v1 root;AGENT_LAB_OPENAI_DEFAULT_MODEL: model used when a request leaves model blank;AGENT_LAB_OPENAI_INCLUDE_USAGE: disable if the provider rejects streamed usage options;AGENT_LAB_REQUEST_TIMEOUT_SECONDS: provider HTTP timeout;AGENT_LAB_DATABASE_PATH: SQLite session/audit/usage database path (defaults to .agent_lab.sqlite3).Do not commit secrets. Start the app with:
uv run uvicorn agent_lab.main:app --reload
Open http://127.0.0.1:8000. The UI can select either invocation mode, enable
tools/events, configure ChatAgent and applicable EventAgent settings, set event
loop/concurrency/batch limits, create or restore sessions, replay messages and
audit data, and inspect per-call, per-turn, and per-session usage.
agent-lab-benchmark runs the same built-in scenarios through dual_agent and
chat_agent_tools, then writes matching machine-readable and review-friendly
reports. It runs in process against the production runtime and event kernel; it
does not start FastAPI or add WebSocket/network timing to the measured turn.
The API key is read only from .env or the environment through
AGENT_LAB_OPENAI_API_KEY. Do not put a key in the benchmark JSON or commit it.
The JSON config supplies the OpenAI-compatible base_url, model, run count,
modes, and cases:
cp benchmarks/single-target.example.json benchmarks/single-target.local.json
Edit the copied file before a live run. The checked-in example deliberately uses
https://provider.example/v1 and replace-with-model-id placeholders and
contains no secret.
Start with the deterministic mock benchmark. This command does not require an API key and makes no network request:
uv run agent-lab-benchmark \
--config benchmarks/single-target.example.json \
--mock
For a live OpenAI-compatible provider, set the key locally, configure a real
base_url and model in the copied JSON, and omit --mock:
export AGENT_LAB_OPENAI_API_KEY='replace-locally'
uv run agent-lab-benchmark \
--config benchmarks/single-target.local.json
CLI values override JSON values. --base-url, --model, and --runs replace
their scalar settings; repeat --case or --mode to replace the corresponding
JSON list. For example, this still performs a zero-network mock run:
uv run agent-lab-benchmark \
--config benchmarks/single-target.example.json \
--mock \
--runs 3 \
--case web_search_two_answers \
--mode dual_agent \
--mode chat_agent_tools
The six built-in cases are:
| Case | Expected behavior |
|---|---|
ordinary_chat |
One visible answer and no event. |
device_volume_silent |
One volume event with silent success and no model follow-up. |
calendar_schedule_template |
One schedule event followed by the deterministic confirmation template. |
web_search_two_answers |
A fast first answer, one search event, and exactly one grounded second answer. |
session_terminate |
One farewell, one terminate event, and a terminal session. |
parallel_volume_schedule |
Volume and schedule events both succeed and their async handlers are proven to overlap. |
Each run records provider TTFT (first provider stream item), visible TTFT (first non-whitespace user-visible delta), turn wall time, per-call model latency, tool handler latency, prompt/completion/total/cached tokens, model/fallback/tool call counts, and event names, sources, statuses, and semantic correctness. Provider TTFT and visible TTFT are intentionally separate; model and tool spans may overlap and are not added to manufacture a total latency.
Per-call model elapsed time accumulates only time awaiting the provider stream
iterator. It excludes downstream runtime work and consumer backpressure after a
yield, while provider TTFT and visible TTFT remain wall-clock arrival metrics.
Token totals are reported only when every model and argument-fallback call
returns usage; otherwise all token metrics are null and the run fails with an
explicit missing-usage semantic error. The parallel volume/schedule case uses a
benchmark-owned barrier probe and passes only when both event handlers are
active at the same time, not merely when two successful event results exist.
By default, reports are published as one complete bundle:
outputs/benchmarks/<timestamp>/report.json
outputs/benchmarks/<timestamp>/report.md
The timestamp path is a single relative symlink published only after both files are fully written and synced; collisions receive a numeric suffix. JSON is the machine-readable source of truth and Markdown is rendered from the same report model.
Exit codes are stable:
0: all runs passed and reports were written;1: at least one request or semantic check failed, but reports were written;2: CLI, config, environment, runner-cardinality, or report/output failure;130: interruption.For a fair mode comparison, keep the provider/model, generation parameters, input cases, run count, enabled events, limits, timeouts, adapters, and run conditions unchanged. Use fresh paired runs and inspect failed samples instead of comparing only averages. p95 is reported with a low-confidence warning when fewer than 20 successful samples exist in a case/mode group.
Interpret results within these limits:
Connect to ws://127.0.0.1:8000/ws/debug. The smallest valid initial payload is:
{
"user_message": "Hello",
"chat_agent": {}
}
This uses the default dual_agent mode and default EventAgent configuration. A
minimal built-in event request can make the comparison controls explicit:
{
"user_message": "Set the device volume to 30",
"chat_agent": {},
"event_agent": {
"enabled_tools": ["device.volume.adjust"]
},
"tool_invocation_mode": "dual_agent"
}
Change only tool_invocation_mode to chat_agent_tools for the other protocol
while keeping model parameters, enabled events, event limits, input, and adapter
wiring equal. After session_started, the same socket accepts later turns as
{"type":"user_message","content":"..."} until a terminal done or error.
ws_smoke.py is not a dual-mode benchmarkWith the server running, scripts/ws_smoke.py sends one convenience request and
prints streamed JSON:
uv run python scripts/ws_smoke.py --message "Debug this agent handoff path."
It defaults to dual_agent, exercises the handoff_note path, has no paired
mode runner or fairness controls, and does not collect the Todo 51 scenario
matrix. Use it only for WebSocket connectivity and payload smoke testing.
| Method and path | Purpose |
|---|---|
GET /health |
Process health. |
GET / |
Static debug UI. |
GET /api/tools |
Current flat registry catalog and schemas. |
POST /api/sessions |
Create a persisted session with optional title and config. |
GET /api/sessions |
List persisted sessions. |
GET /api/sessions/{session_id} |
Read session metadata and immutable comparison config. |
GET /api/sessions/{session_id}/messages |
Replay persisted visible/provider messages. |
GET /api/sessions/{session_id}/audit |
Read ordered runtime audit events. |
GET /api/sessions/{session_id}/usage |
Read comparable call, turn, and session metrics. |
/usage Fair-Comparison ContractGET /api/sessions/{session_id}/usage returns calls, turns, and session.
Interpret it with these controls:
argument_fallback rows and
contribute their model tokens;turn_wall_time_ms is the turn/session wall-clock measure. Do not add model
elapsed_ms or overlapping tool_latency_ms spans to estimate total latency;tool_count or fallback count;ttft_ms measures the first provider stream item. Visible-text TTFT
is a separate runtime/audit observation and must not be conflated with it.Run the closeout matrix and focused runtime checks:
uv run pytest tests/test_tool_invocation_comparison.py -q
uv run pytest tests/test_debug_runtime.py tests/test_builtin_event_plugins.py tests/test_sqlite_store.py tests/test_websocket_api.py -q
Run the repository gates:
uv run pytest
node --check src/agent_lab/presentation/static/app.js
uv run python -c "from agent_lab.main import app; print(app.title)"
git diff --check
These tests establish repeatable application semantics with scripted model clients and injected/default mock ports. They do not prove live-provider content plus-tool streaming behavior, real network search quality, real device/calendar side effects, production credentials, or production latency. Run provider and service integration smokes separately before drawing deployment conclusions.