Vendor OpenClaw source as Adolf fork baseline
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11), free to diverge. Tree copied sans upstream .git; upstream remote added for future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19. Preserves docs/ARCHITECTURE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
This commit is contained in:
143
qa/scenarios/agents/instruction-followthrough-repo-contract.yaml
Normal file
143
qa/scenarios/agents/instruction-followthrough-repo-contract.yaml
Normal file
@@ -0,0 +1,143 @@
|
||||
title: Instruction followthrough repo contract
|
||||
|
||||
scenario:
|
||||
id: instruction-followthrough-repo-contract
|
||||
surface: repo-contract
|
||||
coverage:
|
||||
primary:
|
||||
- agents.instructions
|
||||
secondary:
|
||||
- runtime.first-action
|
||||
objective: Verify the agent reads repo instruction files first, follows the required tool order, and completes the first feasible action instead of stopping at a plan.
|
||||
successCriteria:
|
||||
- Agent reads the seeded instruction files before writing the requested artifact.
|
||||
- Agent writes the requested artifact in the same run instead of returning only a plan.
|
||||
- Agent does not ask for permission before the first feasible action.
|
||||
- Final reply makes the completed read/write sequence explicit.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- src/agents/system-prompt.ts
|
||||
- src/agents/embedded-agent-runner/run/incomplete-turn.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent reads repo instructions first, then completes the first bounded followthrough task without stalling.
|
||||
config:
|
||||
requiredChannelDriver: qa-channel
|
||||
workspaceFiles:
|
||||
AGENT.md: |-
|
||||
# Repo contract
|
||||
|
||||
Step order:
|
||||
1. Read AGENT.md.
|
||||
2. Read SOUL.md.
|
||||
3. Read FOLLOWTHROUGH_INPUT.md.
|
||||
4. Write ./repo-contract-summary.txt.
|
||||
5. Reply with three labeled lines exactly once: Read, Wrote, Status.
|
||||
|
||||
Do not stop after planning.
|
||||
Do not ask for permission before the first feasible action.
|
||||
SOUL.md: |-
|
||||
# Execution style
|
||||
|
||||
Stay brief, honest, and action-first.
|
||||
If the next tool action is feasible, do it before replying.
|
||||
FOLLOWTHROUGH_INPUT.md: |-
|
||||
Mission: prove you followed the repo contract.
|
||||
Evidence path: AGENT.md -> SOUL.md -> FOLLOWTHROUGH_INPUT.md -> repo-contract-summary.txt
|
||||
prompt: |-
|
||||
Repo contract followthrough check. Read AGENT.md, SOUL.md, and FOLLOWTHROUGH_INPUT.md first.
|
||||
Then follow the repo contract exactly, write ./repo-contract-summary.txt, and reply with
|
||||
three labeled lines: Read, Wrote, Status.
|
||||
Do not stop after planning and do not ask for permission before the first feasible action.
|
||||
expectedReplyAll:
|
||||
- "read:"
|
||||
- "wrote:"
|
||||
- "status:"
|
||||
expectedArtifactAll:
|
||||
- "repo contract"
|
||||
expectedArtifactAny:
|
||||
- "evidence path"
|
||||
- "agent.md"
|
||||
- "followthrough"
|
||||
forbiddenNeedles:
|
||||
- need permission
|
||||
- need your approval
|
||||
- can you approve
|
||||
- i would
|
||||
- i can
|
||||
- next i would
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: follows repo instructions instead of stopping at a plan
|
||||
actions:
|
||||
- call: reset
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Object.entries(config.workspaceFiles ?? {})"
|
||||
item: workspaceFile
|
||||
actions:
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, String(workspaceFile[0]))"
|
||||
- expr: "`${String(workspaceFile[1] ?? '').trimEnd()}\\n`"
|
||||
- utf8
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'repo-contract-summary.txt')"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:repo-contract
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 40000)
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(() => { const normalize = (value) => normalizeLowercaseStringOrEmpty(value); const matches = (value) => { const normalized = normalize(value); return normalized && config.expectedArtifactAll.every((needle) => normalized.includes(normalize(needle))) && config.expectedArtifactAny.some((needle) => normalized.includes(normalize(needle))); }; return fs.readFile(artifactPath, 'utf8').then((value) => matches(value) ? value : undefined).catch(() => undefined); })()"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- set: normalizedArtifact
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(artifact)"
|
||||
- assert:
|
||||
expr: "config.expectedArtifactAll.every((needle) => normalizedArtifact.includes(normalizeLowercaseStringOrEmpty(needle))) && config.expectedArtifactAny.some((needle) => normalizedArtifact.includes(normalizeLowercaseStringOrEmpty(needle)))"
|
||||
message:
|
||||
expr: "`repo contract artifact missing expected followthrough signals: ${artifact}`"
|
||||
- set: expectedReplyAll
|
||||
value:
|
||||
expr: config.expectedReplyAll.map(normalizeLowercaseStringOrEmpty)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && expectedReplyAll.every((needle) => normalizeLowercaseStringOrEmpty(candidate.text).includes(needle))).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!config.forbiddenNeedles.some((needle) => normalizeLowercaseStringOrEmpty(outbound.text).includes(needle))"
|
||||
message:
|
||||
expr: "`repo contract followthrough bounced for permission or stalled: ${outbound.text}`"
|
||||
- set: followthroughDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].filter((request) => /repo contract followthrough check/i.test(String(request.allInputText ?? ''))) : []"
|
||||
- assert:
|
||||
expr: "!env.mock || followthroughDebugRequests.filter((request) => request.plannedToolName === 'read').length >= 3"
|
||||
message:
|
||||
expr: "`expected three read tool calls before write, saw plannedToolNames=${JSON.stringify(followthroughDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
- assert:
|
||||
expr: "!env.mock || followthroughDebugRequests.some((request) => request.plannedToolName === 'write')"
|
||||
message:
|
||||
expr: "`expected write tool call during repo contract followthrough, saw plannedToolNames=${JSON.stringify(followthroughDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
- assert:
|
||||
expr: "!env.mock || (() => { const readIndices = followthroughDebugRequests.map((r, i) => r.plannedToolName === 'read' ? i : -1).filter(i => i >= 0); const firstWrite = followthroughDebugRequests.findIndex((r) => r.plannedToolName === 'write'); return readIndices.length >= 3 && firstWrite >= 0 && readIndices[2] < firstWrite; })()"
|
||||
message:
|
||||
expr: "`expected all 3 reads before any write during repo contract followthrough, saw plannedToolNames=${JSON.stringify(followthroughDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
detailsExpr: outbound.text
|
||||
108
qa/scenarios/agents/subagent-completion-direct-fallback.yaml
Normal file
108
qa/scenarios/agents/subagent-completion-direct-fallback.yaml
Normal file
@@ -0,0 +1,108 @@
|
||||
title: Subagent completion direct fallback
|
||||
|
||||
scenario:
|
||||
id: subagent-completion-direct-fallback
|
||||
surface: subagents
|
||||
coverage:
|
||||
primary:
|
||||
- agents.subagents
|
||||
secondary:
|
||||
- runtime.delivery
|
||||
- channels.qa-channel
|
||||
objective: Verify a yielded parent still receives a successful subagent result through direct fallback delivery when the dormant announce turn produces no visible reply.
|
||||
successCriteria:
|
||||
- Parent launches a native subagent.
|
||||
- Parent yields instead of waiting in-turn.
|
||||
- Subagent completion result is delivered to the original QA DM without a thread id.
|
||||
- Durable task delivery is marked delivered, not failed.
|
||||
docsRefs:
|
||||
- docs/tools/subagents.md
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- src/agents/subagent-announce-delivery.ts
|
||||
- src/agents/subagent-registry-lifecycle.ts
|
||||
- src/agents/tools/sessions-yield-tool.ts
|
||||
- extensions/qa-lab/src/providers/mock-openai/server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Reproduce yielded-parent subagent completion delivery and require frozen-result fallback to the QA DM.
|
||||
config:
|
||||
prompt: "Subagent direct fallback QA check: spawn one native subagent worker. The worker must finish with exactly QA-SUBAGENT-DIRECT-FALLBACK-OK. After spawning it, call sessions_yield and wait for the completion event. Do not use ACP."
|
||||
expectedMarker: QA-SUBAGENT-DIRECT-FALLBACK-OK
|
||||
expectedLabel: qa-direct-fallback-worker
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: yielded parent receives child completion through direct fallback
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 120000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 120000
|
||||
- call: reset
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:subagent-direct-fallback:${randomUUID().slice(0, 8)}`"
|
||||
- try:
|
||||
actions:
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'outbound' && String(message.text ?? '').includes(config.expectedMarker)).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 180000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "String(outbound.text ?? '').trim().includes(config.expectedMarker)"
|
||||
message:
|
||||
expr: "`fallback completion marker missing from outbound QA DM: ${recentOutboundSummary(state)}`"
|
||||
catchAs: fallbackError
|
||||
catch:
|
||||
- set: fallbackDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].slice(-20).map((request) => ({ plannedToolName: request.plannedToolName ?? null, plannedToolArgs: request.plannedToolArgs ?? null, prompt: String(request.prompt ?? '').slice(0, 280), allInputText: String(request.allInputText ?? '').slice(0, 280), toolOutput: request.toolOutput ? String(request.toolOutput).slice(0, 280) : null })) : []"
|
||||
- set: fallbackTasks
|
||||
value:
|
||||
expr: "(await runQaCli(env, ['tasks', 'list', '--json', '--runtime', 'subagent'], { timeoutMs: liveTurnTimeoutMs(env, 60000), json: true }).catch((error) => ({ error: String(error?.message ?? error) })))"
|
||||
- throw:
|
||||
expr: "`subagent fallback marker missing: ${fallbackError?.message ?? fallbackError}; outbound=${recentOutboundSummary(state, 8)} tasks=${JSON.stringify(fallbackTasks)} requests=${JSON.stringify(fallbackDebugRequests)}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: fallbackDebugRequests
|
||||
value:
|
||||
expr: "[...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))]"
|
||||
- assert:
|
||||
expr: "fallbackDebugRequests.some((request) => !request.toolOutput && /subagent direct fallback qa check/i.test(String(request.allInputText ?? '')) && request.plannedToolName === 'sessions_spawn' && request.plannedToolArgs?.label === config.expectedLabel)"
|
||||
message:
|
||||
expr: "`expected sessions_spawn for yielded fallback scenario, saw ${JSON.stringify(fallbackDebugRequests.map((request) => ({ plannedToolName: request.plannedToolName ?? null, plannedToolArgs: request.plannedToolArgs ?? null })))}`"
|
||||
- assert:
|
||||
expr: "fallbackDebugRequests.some((request) => /subagent direct fallback qa check/i.test(String(request.allInputText ?? '')) && request.plannedToolName === 'sessions_yield')"
|
||||
message:
|
||||
expr: "`expected sessions_yield for yielded fallback scenario, saw ${JSON.stringify(fallbackDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
- call: waitForCondition
|
||||
saveAs: deliveredTask
|
||||
args:
|
||||
- lambda:
|
||||
expr: "(async () => { const payload = await runQaCli(env, ['tasks', 'list', '--json', '--runtime', 'subagent'], { timeoutMs: liveTurnTimeoutMs(env, 60000), json: true }); return (payload.tasks ?? []).find((task) => task.label === config.expectedLabel && task.deliveryStatus === 'delivered' && task.status === 'succeeded') ?? null; })()"
|
||||
- expr: liveTurnTimeoutMs(env, 60000)
|
||||
- 250
|
||||
- assert:
|
||||
expr: "deliveredTask.deliveryStatus === 'delivered'"
|
||||
message:
|
||||
expr: "`expected delivered task status for ${config.expectedLabel}, got ${JSON.stringify(deliveredTask)}`"
|
||||
detailsExpr: "outbound.text"
|
||||
261
qa/scenarios/agents/subagent-fanout-synthesis.yaml
Normal file
261
qa/scenarios/agents/subagent-fanout-synthesis.yaml
Normal file
@@ -0,0 +1,261 @@
|
||||
title: Subagent fanout synthesis
|
||||
|
||||
scenario:
|
||||
id: subagent-fanout-synthesis
|
||||
surface: subagents
|
||||
coverage:
|
||||
primary:
|
||||
- agents.subagents
|
||||
secondary:
|
||||
- agents.synthesis
|
||||
objective: Verify the agent can delegate multiple bounded subagent tasks and fold both results back into one parent reply.
|
||||
successCriteria:
|
||||
- Parent flow launches at least two bounded subagent tasks.
|
||||
- Both delegated results are acknowledged in the main flow.
|
||||
- Final answer synthesizes both worker outputs in one reply.
|
||||
docsRefs:
|
||||
- docs/tools/subagents.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/agents/subagent-spawn.ts
|
||||
- src/agents/system-prompt.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can delegate multiple bounded subagent tasks and fold both results back into one parent reply.
|
||||
config:
|
||||
prompt: |-
|
||||
Subagent fanout synthesis check: delegate exactly two bounded subagents sequentially using sessions_spawn, not ACP.
|
||||
First spawn exactly one child with label qa-fanout-alpha and task: verify that `HEARTBEAT.md` exists and reply exactly `ok` if it does. Wait for that child to finish.
|
||||
Then spawn exactly one child with label qa-fanout-beta and task: verify that `repo/qa/scenarios/agents/subagent-fanout-synthesis.yaml` exists and reply exactly `ok` if it does. Wait for that child to finish.
|
||||
Do not spawn any more children after qa-fanout-beta finishes.
|
||||
Then reply with exactly these two lines and nothing else:
|
||||
subagent-1: ok
|
||||
subagent-2: ok
|
||||
expectedReplyAny:
|
||||
- "subagent-1: ok"
|
||||
- "subagent-2: ok"
|
||||
expectedReplyGroups:
|
||||
- - alpha-ok
|
||||
- subagent_one_ok
|
||||
- subagent one ok
|
||||
- "subagent-1: ok"
|
||||
- - beta-ok
|
||||
- subagent_two_ok
|
||||
- subagent two ok
|
||||
- "subagent-2: ok"
|
||||
expectedChildLabels:
|
||||
- qa-fanout-alpha
|
||||
- qa-fanout-beta
|
||||
expectedChildCompletionMarkers:
|
||||
- ALPHA-OK
|
||||
- BETA-OK
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: spawns sequential workers and folds both results back into the parent reply
|
||||
actions:
|
||||
- set: attempts
|
||||
value:
|
||||
expr: "env.providerMode === 'mock-openai' ? 1 : 2"
|
||||
- set: lastError
|
||||
value: null
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Array.from({ length: attempts }, (_, index) => index + 1)"
|
||||
item: attempt
|
||||
actions:
|
||||
- if:
|
||||
expr: "lastError === '__done__'"
|
||||
then:
|
||||
- set: skippedAttempt
|
||||
value:
|
||||
expr: attempt
|
||||
else:
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 120000
|
||||
- call: reset
|
||||
- set: alphaLabel
|
||||
value:
|
||||
expr: "env.providerMode === 'mock-openai' ? config.expectedChildLabels[0] : `${config.expectedChildLabels[0]}-${attempt}`"
|
||||
- set: betaLabel
|
||||
value:
|
||||
expr: "env.providerMode === 'mock-openai' ? config.expectedChildLabels[1] : `${config.expectedChildLabels[1]}-${attempt}`"
|
||||
- set: prompt
|
||||
value:
|
||||
expr: "`Subagent fanout synthesis check: delegate exactly two bounded subagents sequentially using sessions_spawn, not ACP.\nFirst spawn exactly one child with label ${alphaLabel} and task: verify that \\`HEARTBEAT.md\\` exists and reply exactly \\`ok\\` if it does. Wait for that child to finish.\nThen spawn exactly one child with label ${betaLabel} and task: verify that \\`repo/qa/scenarios/agents/subagent-fanout-synthesis.yaml\\` exists and reply exactly \\`ok\\` if it does. Wait for that child to finish.\nDo not spawn any more children after ${betaLabel} finishes.\nThen reply with exactly these two lines and nothing else:\nsubagent-1: ok\nsubagent-2: ok`"
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:fanout:${attempt}:${randomUUID().slice(0, 8)}`"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
ref: prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- call: waitForAgentHistoryReply
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- lambda:
|
||||
params: [text]
|
||||
expr: "config.expectedReplyGroups.every((group) => group.some((needle) => normalizeLowercaseStringOrEmpty(text).includes(needle)))"
|
||||
- expr: "30000"
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- call: readRawQaSessionStore
|
||||
saveAs: store
|
||||
args:
|
||||
- ref: env
|
||||
- set: childRows
|
||||
value:
|
||||
expr: "Object.values(store).filter((entry) => entry.spawnedBy === sessionKey)"
|
||||
- set: sawAlpha
|
||||
value:
|
||||
expr: "childRows.some((entry) => entry.label === alphaLabel)"
|
||||
- set: sawBeta
|
||||
value:
|
||||
expr: "childRows.some((entry) => entry.label === betaLabel)"
|
||||
- assert:
|
||||
expr: "sawAlpha && sawBeta"
|
||||
message:
|
||||
expr: "`fanout child sessions missing (alpha=${String(sawAlpha)} beta=${String(sawBeta)})`"
|
||||
# Tool-call assertion (criterion 2 of the
|
||||
# parity completion gate in #64227): the
|
||||
# scenario must have actually invoked
|
||||
# `sessions_spawn` at least twice with
|
||||
# distinct labels, not just ended up with
|
||||
# two rows in the session store through
|
||||
# prose trickery. The session store alone
|
||||
# can be populated by other flows or by a
|
||||
# model that fabricates "delegation"
|
||||
# narration. `plannedToolName` on the
|
||||
# mock's `/debug/requests` log is the
|
||||
# tool-call ground truth: two recorded
|
||||
# sessions_spawn requests with distinct
|
||||
# labels means the model really dispatched
|
||||
# both subagents.
|
||||
- set: fanoutSpawnRequests
|
||||
value:
|
||||
expr: "[...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].filter((request) => request.plannedToolName === 'sessions_spawn' && /subagent fanout synthesis check/i.test(String(request.allInputText ?? '')))"
|
||||
- assert:
|
||||
expr: "fanoutSpawnRequests.length >= 2"
|
||||
message:
|
||||
expr: "`expected at least two sessions_spawn tool calls during subagent fanout scenario, saw ${fanoutSpawnRequests.length}`"
|
||||
- set: details
|
||||
value:
|
||||
expr: "outbound.text"
|
||||
- set: lastError
|
||||
value: __done__
|
||||
catchAs: attemptError
|
||||
catch:
|
||||
- if:
|
||||
expr: "/timed out after/i.test(formatErrorMessage(attemptError))"
|
||||
then:
|
||||
- call: readRawQaSessionStore
|
||||
saveAs: timeoutStore
|
||||
args:
|
||||
- ref: env
|
||||
- set: timeoutChildEntries
|
||||
value:
|
||||
expr: "Object.entries(timeoutStore).map(([key, entry]) => ({ ...entry, key })).filter((entry) => entry.spawnedBy === sessionKey)"
|
||||
- set: timeoutChildRows
|
||||
value:
|
||||
expr: "timeoutChildEntries"
|
||||
- set: timeoutAlphaSessionKey
|
||||
value:
|
||||
expr: "timeoutChildEntries.find((entry) => entry.label === alphaLabel)?.key ?? ''"
|
||||
- set: timeoutBetaSessionKey
|
||||
value:
|
||||
expr: "timeoutChildEntries.find((entry) => entry.label === betaLabel)?.key ?? ''"
|
||||
- set: timeoutSawAlpha
|
||||
value:
|
||||
expr: "timeoutChildRows.some((entry) => entry.label === alphaLabel)"
|
||||
- set: timeoutSawBeta
|
||||
value:
|
||||
expr: "timeoutChildRows.some((entry) => entry.label === betaLabel)"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: timeoutSpawnRequests
|
||||
value:
|
||||
expr: "[...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].filter((request) => request.plannedToolName === 'sessions_spawn' && /subagent fanout synthesis check/i.test(String(request.allInputText ?? '')))"
|
||||
- if:
|
||||
expr: "timeoutSawAlpha && timeoutSawBeta && timeoutSpawnRequests.length >= 2"
|
||||
then:
|
||||
- set: details
|
||||
value: "subagent-1: ok\nsubagent-2: ok"
|
||||
- set: lastError
|
||||
value: __done__
|
||||
else:
|
||||
- set: timeoutAlphaTranscript
|
||||
value:
|
||||
expr: "timeoutAlphaSessionKey ? await readSessionTranscriptSummary(env, timeoutAlphaSessionKey) : null"
|
||||
- set: timeoutBetaTranscript
|
||||
value:
|
||||
expr: "timeoutBetaSessionKey ? await readSessionTranscriptSummary(env, timeoutBetaSessionKey) : null"
|
||||
- set: timeoutAlphaOk
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(timeoutAlphaTranscript?.finalText) === 'ok'"
|
||||
- set: timeoutBetaOk
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(timeoutBetaTranscript?.finalText) === 'ok'"
|
||||
- if:
|
||||
expr: "timeoutSawAlpha && timeoutSawBeta && timeoutAlphaOk && timeoutBetaOk"
|
||||
then:
|
||||
- set: details
|
||||
value: "subagent-1: ok\nsubagent-2: ok"
|
||||
- set: lastError
|
||||
value: __done__
|
||||
- if:
|
||||
expr: "lastError !== '__done__'"
|
||||
then:
|
||||
- set: lastError
|
||||
value:
|
||||
ref: attemptError
|
||||
- if:
|
||||
expr: "lastError !== '__done__' && attempt < attempts"
|
||||
then:
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 120000
|
||||
catch:
|
||||
- set: ignoredRetryWait
|
||||
value: true
|
||||
- assert:
|
||||
expr: "lastError === '__done__'"
|
||||
message:
|
||||
expr: "lastError instanceof Error ? formatErrorMessage(lastError) : String(lastError ?? 'fanout retry exhausted')"
|
||||
- if:
|
||||
# Codex completes child sessions through its app-server path but
|
||||
# does not relay the child marker back onto the parent QA channel.
|
||||
# The shared assertions above already prove both child tool calls
|
||||
# and child session rows; keep this transport-only proof OpenClaw-specific.
|
||||
expr: "Boolean(env.mock) && env.gateway.runtimeEnv.OPENCLAW_QA_FORCE_RUNTIME !== 'codex'"
|
||||
then:
|
||||
- forEach:
|
||||
items:
|
||||
expr: "config.expectedChildCompletionMarkers"
|
||||
item: childCompletionMarker
|
||||
actions:
|
||||
- call: waitForOutboundMessage
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "String(candidate.text ?? '').trim() === childCompletionMarker"
|
||||
- 30000
|
||||
detailsExpr: "details"
|
||||
60
qa/scenarios/agents/subagent-forked-context.yaml
Normal file
60
qa/scenarios/agents/subagent-forked-context.yaml
Normal file
@@ -0,0 +1,60 @@
|
||||
title: Subagent forked context
|
||||
|
||||
scenario:
|
||||
id: subagent-forked-context
|
||||
surface: subagents
|
||||
coverage:
|
||||
primary:
|
||||
- agents.subagents
|
||||
objective: Verify the agent can choose forked subagent context when the child needs the current transcript.
|
||||
successCriteria:
|
||||
- Agent launches a native subagent with context=fork.
|
||||
- Subagent uses the forked requester transcript to recover the visible code.
|
||||
- Subagent request remains bounded and does not switch to ACP.
|
||||
- User-visible output includes the delegated result and the visible code.
|
||||
docsRefs:
|
||||
- docs/tools/subagents.md
|
||||
- docs/concepts/session-tool.md
|
||||
codeRefs:
|
||||
- src/agents/tools/sessions-spawn-tool.ts
|
||||
- src/agents/subagent-spawn.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Ask the agent to delegate work that depends on the current transcript and assert sessions_spawn carries context=fork.
|
||||
config:
|
||||
requiredChannelDriver: qa-channel
|
||||
contextNeedle: FORKED-CONTEXT-ALPHA
|
||||
prompt: "Forked subagent context QA check. The visible code in this current conversation is FORKED-CONTEXT-ALPHA. Delegate to a native subagent to report the visible code from the requester transcript. Do not include the visible code in the child task text; the child must recover it from forked transcript context. Use forked context if the child needs the current transcript; otherwise it will not know the code. A spawn-accepted result is not the answer. Wait for the child completion, then make sure user-visible output includes the visible code."
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: forks current transcript context for the child
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:forked-context
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && String(candidate.text ?? '').includes(config.contextNeedle) && !normalizeLowercaseStringOrEmpty(candidate.text).includes('waiting')).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "env.mock || String(outbound.text ?? '').includes(config.contextNeedle)"
|
||||
message:
|
||||
expr: "`expected live final answer to include fork-only context code ${config.contextNeedle}, got: ${outbound.text}`"
|
||||
- set: forkDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))] : []"
|
||||
- assert:
|
||||
expr: "!env.mock || forkDebugRequests.some((request) => !request.toolOutput && /forked subagent context qa check/i.test(String(request.allInputText ?? '')) && request.plannedToolName === 'sessions_spawn' && (request.plannedToolArgs?.context === 'fork' || /context\\s*=\\s*fork/i.test(String(request.allInputText ?? ''))))"
|
||||
message:
|
||||
expr: "`expected sessions_spawn context=fork during forked context scenario, saw ${JSON.stringify(forkDebugRequests.map((request) => ({ plannedToolName: request.plannedToolName ?? null, plannedToolArgs: request.plannedToolArgs ?? null })))} `"
|
||||
detailsExpr: outbound.text
|
||||
74
qa/scenarios/agents/subagent-handoff.yaml
Normal file
74
qa/scenarios/agents/subagent-handoff.yaml
Normal file
@@ -0,0 +1,74 @@
|
||||
title: Subagent handoff
|
||||
|
||||
scenario:
|
||||
id: subagent-handoff
|
||||
surface: subagents
|
||||
coverage:
|
||||
primary:
|
||||
- agents.subagents
|
||||
objective: Verify the agent can delegate a bounded task to a subagent and fold the result back into the main thread.
|
||||
successCriteria:
|
||||
- Agent launches a bounded subagent task.
|
||||
- Subagent result is acknowledged in the main flow.
|
||||
- Final answer attributes delegated work clearly.
|
||||
docsRefs:
|
||||
- docs/tools/subagents.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/agents/system-prompt.ts
|
||||
- extensions/qa-lab/src/report.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can delegate a bounded task to a subagent and fold the result back into the main thread.
|
||||
config:
|
||||
requiredChannelDriver: qa-channel
|
||||
prompt: "Delegate one bounded QA task to a subagent. Wait for the subagent to finish. Then reply with three labeled sections exactly once: Delegated task, Result, Evidence. Include the child result itself, not 'waiting'."
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: delegates a bounded task and reports the result
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:subagent
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- call: waitForAgentHistoryReply
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: env
|
||||
- agent:qa:subagent
|
||||
- lambda:
|
||||
params: [text]
|
||||
expr: "(() => { const lower = normalizeLowercaseStringOrEmpty(text); return lower.includes('delegated task') && lower.includes('result') && lower.includes('evidence') && !lower.includes('waiting'); })()"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!['failed to delegate','could not delegate','subagent unavailable'].some((needle) => normalizeLowercaseStringOrEmpty(outbound.text).includes(needle))"
|
||||
message:
|
||||
expr: "`subagent handoff reported failure: ${outbound.text}`"
|
||||
# Parity gate criterion 2 (no fake progress / fake tool completion):
|
||||
# require an actual sessions_spawn tool call. Without this, a model
|
||||
# could produce the three labeled sections ("Delegated task", "Result",
|
||||
# "Evidence") as free-form prose without ever delegating to a real
|
||||
# subagent. The assertion is pinned to THIS scenario by matching the
|
||||
# scenario-unique prompt substring "Delegate one bounded QA task"
|
||||
# (not a broad /delegate|subagent/ regex) so the earlier
|
||||
# subagent-fanout-synthesis scenario — which also contains "delegate"
|
||||
# and produces its own pre-tool sessions_spawn request — cannot
|
||||
# satisfy the assertion here. The match is also constrained to
|
||||
# pre-tool requests (no toolOutput) because the mock only plans
|
||||
# sessions_spawn on requests with no toolOutput; the follow-up
|
||||
# request after the tool runs has plannedToolName unset.
|
||||
- set: subagentDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))] : []"
|
||||
- assert:
|
||||
expr: "!env.mock || subagentDebugRequests.some((request) => !request.toolOutput && /delegate one bounded qa task/i.test(String(request.allInputText ?? '')) && request.plannedToolName === 'sessions_spawn')"
|
||||
message:
|
||||
expr: "`expected sessions_spawn tool call during subagent handoff scenario, saw plannedToolNames=${JSON.stringify(subagentDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
detailsExpr: outbound.text
|
||||
174
qa/scenarios/agents/subagent-stale-child-links.yaml
Normal file
174
qa/scenarios/agents/subagent-stale-child-links.yaml
Normal file
@@ -0,0 +1,174 @@
|
||||
title: Subagent stale child links
|
||||
|
||||
scenario:
|
||||
id: subagent-stale-child-links
|
||||
surface: subagents
|
||||
coverage:
|
||||
primary:
|
||||
- agents.subagents
|
||||
secondary:
|
||||
- gateway.sessions-list
|
||||
objective: Verify restarted gateways hide stale persisted subagent child links without hiding live or fresh children.
|
||||
successCriteria:
|
||||
- Old ended subagent run records are not exposed as current children.
|
||||
- Old store-only spawnedBy and parentSessionKey rows are not exposed as current children.
|
||||
- Child-side ACP store rows from sibling agents are not exposed as current children.
|
||||
- Live subagent runs and fresh dashboard children remain visible.
|
||||
docsRefs:
|
||||
- docs/tools/subagents.md
|
||||
- docs/concepts/qa-e2e-automation.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/gateway/session-utils.ts
|
||||
- src/agents/subagent-run-liveness.ts
|
||||
- extensions/qa-lab/src/gateway-child.ts
|
||||
execution:
|
||||
kind: flow
|
||||
suiteIsolation: isolated
|
||||
isolationReason: Seeds persisted gateway session/subagent state and restarts the gateway.
|
||||
summary: Seed stale subagent session state on disk, restart the real gateway, then assert sessions.list filters only the stale child links.
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: restarted gateway filters stale subagent child links
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- set: mainKey
|
||||
value: "agent:qa:main"
|
||||
- set: staleRunKey
|
||||
value: "agent:qa:subagent:qa-stale-ended"
|
||||
- set: staleOrphanKey
|
||||
value: "agent:qa:subagent:qa-orphan"
|
||||
- set: staleAcpKey
|
||||
value: "agent:claude:acp:qa-stale-acp"
|
||||
- set: freshDashboardKey
|
||||
value: "agent:qa:dashboard:qa-fresh-child"
|
||||
- set: liveRunKey
|
||||
value: "agent:qa:subagent:qa-live-child"
|
||||
- call: env.gateway.restartAfterStateMutation
|
||||
args:
|
||||
- lambda:
|
||||
params:
|
||||
- ctx
|
||||
async: true
|
||||
expr: |-
|
||||
await (async () => {
|
||||
const now = Date.now();
|
||||
const old = now - 2 * 60 * 60 * 1000;
|
||||
const recent = now - 5000;
|
||||
const qaSessionsDir = path.join(ctx.stateDir, "agents", "qa", "sessions");
|
||||
const claudeSessionsDir = path.join(ctx.stateDir, "agents", "claude", "sessions");
|
||||
const subagentDir = path.join(ctx.stateDir, "subagents");
|
||||
await fs.mkdir(qaSessionsDir, { recursive: true });
|
||||
await fs.mkdir(claudeSessionsDir, { recursive: true });
|
||||
await fs.mkdir(subagentDir, { recursive: true });
|
||||
await fs.writeFile(path.join(subagentDir, "runs.json"), `${JSON.stringify({
|
||||
version: 2,
|
||||
runs: {
|
||||
"run-stale-ended": {
|
||||
runId: "run-stale-ended",
|
||||
childSessionKey: staleRunKey,
|
||||
controllerSessionKey: mainKey,
|
||||
requesterSessionKey: mainKey,
|
||||
requesterDisplayKey: "main",
|
||||
task: "old ended ghost",
|
||||
cleanup: "keep",
|
||||
createdAt: old - 60000,
|
||||
startedAt: old - 50000,
|
||||
endedAt: old,
|
||||
outcome: { status: "ok" },
|
||||
},
|
||||
"run-live-visible": {
|
||||
runId: "run-live-visible",
|
||||
childSessionKey: liveRunKey,
|
||||
controllerSessionKey: mainKey,
|
||||
requesterSessionKey: mainKey,
|
||||
requesterDisplayKey: "main",
|
||||
task: "live child remains visible",
|
||||
cleanup: "keep",
|
||||
createdAt: recent,
|
||||
startedAt: recent,
|
||||
},
|
||||
},
|
||||
}, null, 2)}\n`, "utf8");
|
||||
await fs.writeFile(path.join(qaSessionsDir, "sessions.json"), `${JSON.stringify({
|
||||
[mainKey]: {
|
||||
sessionId: "sess-main",
|
||||
updatedAt: now,
|
||||
},
|
||||
[staleRunKey]: {
|
||||
sessionId: "sess-stale-run",
|
||||
updatedAt: old,
|
||||
spawnedBy: mainKey,
|
||||
status: "done",
|
||||
endedAt: old,
|
||||
},
|
||||
[staleOrphanKey]: {
|
||||
sessionId: "sess-orphan",
|
||||
updatedAt: old,
|
||||
parentSessionKey: mainKey,
|
||||
},
|
||||
[freshDashboardKey]: {
|
||||
sessionId: "sess-fresh-dashboard",
|
||||
updatedAt: now,
|
||||
parentSessionKey: mainKey,
|
||||
},
|
||||
[liveRunKey]: {
|
||||
sessionId: "sess-live-child",
|
||||
updatedAt: recent,
|
||||
spawnedBy: mainKey,
|
||||
},
|
||||
}, null, 2)}\n`, "utf8");
|
||||
await fs.writeFile(path.join(claudeSessionsDir, "sessions.json"), `${JSON.stringify({
|
||||
[staleAcpKey]: {
|
||||
sessionId: "sess-acp-stale",
|
||||
updatedAt: old,
|
||||
spawnedBy: mainKey,
|
||||
status: "done",
|
||||
endedAt: old,
|
||||
},
|
||||
}, null, 2)}\n`, "utf8");
|
||||
})()
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: env.gateway.call
|
||||
saveAs: listed
|
||||
args:
|
||||
- "sessions.list"
|
||||
- {}
|
||||
- timeoutMs: 60000
|
||||
- call: env.gateway.call
|
||||
saveAs: filtered
|
||||
args:
|
||||
- "sessions.list"
|
||||
- spawnedBy:
|
||||
ref: mainKey
|
||||
- timeoutMs: 60000
|
||||
- set: mainChildren
|
||||
value:
|
||||
expr: "(listed.sessions.find((session) => session.key === mainKey)?.childSessions ?? [])"
|
||||
- set: filteredKeys
|
||||
value:
|
||||
expr: "filtered.sessions.map((session) => session.key)"
|
||||
- assert:
|
||||
expr: "mainChildren.includes(freshDashboardKey)"
|
||||
message:
|
||||
expr: "`fresh dashboard child missing from main children: ${JSON.stringify(mainChildren)}`"
|
||||
- assert:
|
||||
expr: "mainChildren.includes(liveRunKey)"
|
||||
message:
|
||||
expr: "`live subagent child missing from main children: ${JSON.stringify(mainChildren)}`"
|
||||
- assert:
|
||||
expr: "filteredKeys.includes(freshDashboardKey) && filteredKeys.includes(liveRunKey)"
|
||||
message:
|
||||
expr: "`spawnedBy filter dropped live/fresh children: ${JSON.stringify(filteredKeys)}`"
|
||||
- assert:
|
||||
expr: "![staleRunKey, staleOrphanKey, staleAcpKey].some((key) => mainChildren.includes(key) || filteredKeys.includes(key))"
|
||||
message:
|
||||
expr: "`stale child leaked through sessions.list (main=${JSON.stringify(mainChildren)} filtered=${JSON.stringify(filteredKeys)})`"
|
||||
detailsExpr: "({ mainChildren, filteredKeys })"
|
||||
Reference in New Issue
Block a user