Vendor OpenClaw source as Adolf fork baseline
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11), free to diverge. Tree copied sans upstream .git; upstream remote added for future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19. Preserves docs/ARCHITECTURE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
This commit is contained in:
64
qa/scenarios/workspace/lobster-invaders-build.yaml
Normal file
64
qa/scenarios/workspace/lobster-invaders-build.yaml
Normal file
@@ -0,0 +1,64 @@
|
||||
title: Build Lobster Invaders
|
||||
|
||||
scenario:
|
||||
id: lobster-invaders-build
|
||||
surface: workspace
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.artifacts
|
||||
secondary:
|
||||
- workspace.builds
|
||||
objective: Verify the agent can read the repo, create a tiny playable artifact, and report what changed.
|
||||
successCriteria:
|
||||
- Agent inspects source before coding.
|
||||
- Agent builds a tiny playable Lobster Invaders artifact.
|
||||
- Agent explains how to run or view the artifact.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/web/dashboard.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/report.ts
|
||||
- extensions/qa-lab/web/src/app.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can read the repo, create a tiny playable artifact, and report what changed.
|
||||
config:
|
||||
prompt: Read the QA kickoff context first, then build a tiny Lobster Invaders HTML game at ./lobster-invaders.html in this workspace and tell me where it is.
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: creates the artifact after reading context
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:lobster-invaders
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForOutboundMessage
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'lobster-invaders.html')"
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => null))?.includes('Lobster Invaders') ? await fs.readFile(artifactPath, 'utf8').catch(() => null) : undefined)"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
- 250
|
||||
- assert:
|
||||
expr: "artifact.includes('Lobster Invaders')"
|
||||
message: missing Lobster Invaders artifact
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).some((request) => (request.toolOutput ?? '').includes('QA mission'))"
|
||||
message: expected pre-write read evidence
|
||||
detailsExpr: "'lobster-invaders.html'"
|
||||
251
qa/scenarios/workspace/long-running-release-audit.yaml
Normal file
251
qa/scenarios/workspace/long-running-release-audit.yaml
Normal file
@@ -0,0 +1,251 @@
|
||||
title: Long-running release audit
|
||||
|
||||
scenario:
|
||||
id: long-running-release-audit
|
||||
surface: workspace
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.long-running-task
|
||||
secondary:
|
||||
- workspace.repo-discovery
|
||||
- workspace.artifacts
|
||||
objective: Verify a model can complete a sustained repo audit that requires reading several scattered files, ignoring stale decoys, writing artifacts, and checking its own output.
|
||||
successCriteria:
|
||||
- Agent starts from the seeded project README instead of guessing.
|
||||
- Agent inspects docs and source files across multiple directories.
|
||||
- Agent writes a JSON audit and a Markdown handoff with all required findings.
|
||||
- Each JSON finding carries a boolean verified field; missing current source evidence is verified false.
|
||||
- Agent verifies the generated JSON before finalizing.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- qa/scenarios/index.yaml
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/scenario-flow-runner.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify sustained repo-style auditing with scattered evidence and artifact verification.
|
||||
config:
|
||||
reportFile: audit-fixture/release-audit.json
|
||||
handoffFile: audit-fixture/release-handoff.md
|
||||
expectedFindings:
|
||||
- REL-GATEWAY-417
|
||||
- REL-CHANNEL-238
|
||||
- REL-CRON-904
|
||||
- REL-MEMORY-552
|
||||
- REL-PLUGIN-319
|
||||
- REL-INSTALL-846
|
||||
- REL-DOCS-611
|
||||
workspaceFiles:
|
||||
audit-fixture/README.md: |-
|
||||
# Finch Relay
|
||||
|
||||
Release readiness task:
|
||||
|
||||
The release checklist was renamed during docs cleanup.
|
||||
Find the current checklist under `docs/`, then inspect the referenced
|
||||
source and notes. Produce:
|
||||
|
||||
- `release-audit.json`
|
||||
- `release-handoff.md`
|
||||
|
||||
Do not use `docs/archive/old-release-notes.md` as current evidence.
|
||||
Historical path that may appear in stale notes: `docs/release-checklist.md`.
|
||||
audit-fixture/docs/current-readiness-checklist.md: |-
|
||||
# Release Checklist
|
||||
|
||||
Current release readiness requires checking eight areas:
|
||||
|
||||
1. gateway reconnect handling
|
||||
2. channel delivery ordering
|
||||
3. cron duplicate prevention
|
||||
4. memory recall fallback
|
||||
5. plugin runtime loading
|
||||
6. control UI status copy
|
||||
7. installer update path
|
||||
8. docs update status
|
||||
|
||||
Useful current sources:
|
||||
|
||||
- `src/gateway/reconnect.ts`
|
||||
- `src/channels/delivery.ts`
|
||||
- `src/scheduling/cron.ts`
|
||||
- `src/memory/recall.ts`
|
||||
- `src/plugins/runtime.ts`
|
||||
- `ui/control-panel.ts`
|
||||
- `install/update.ts`
|
||||
- `docs/operator-notes.md`
|
||||
|
||||
The archive folder contains stale notes and should not be treated as
|
||||
current release evidence.
|
||||
audit-fixture/docs/operator-notes.md: |-
|
||||
# Operator Notes
|
||||
|
||||
Current docs update status:
|
||||
|
||||
Finding id: REL-DOCS-611
|
||||
Status: docs mention reconnect, cron, memory, plugin, and installer checks,
|
||||
but the channel ordering and UI notes still need maintainer handoff.
|
||||
audit-fixture/docs/archive/old-release-notes.md: |-
|
||||
# Old Release Notes
|
||||
|
||||
Stale finding id: REL-STALE-000
|
||||
This file is from a previous release and should not appear in the new
|
||||
release audit.
|
||||
audit-fixture/src/gateway/reconnect.ts: |-
|
||||
export const gatewayReconnectReleaseFinding = {
|
||||
id: "REL-GATEWAY-417",
|
||||
area: "gateway reconnect handling",
|
||||
status: "retry jitter verified, resume token fallback still needs manual spot check",
|
||||
};
|
||||
audit-fixture/src/channels/delivery.ts: |-
|
||||
export const channelDeliveryReleaseFinding = {
|
||||
id: "REL-CHANNEL-238",
|
||||
area: "channel delivery ordering",
|
||||
status: "thread replies preserve ordering, root-channel fallback needs handoff note",
|
||||
};
|
||||
audit-fixture/src/scheduling/cron.ts: |-
|
||||
export const cronDuplicateReleaseFinding = {
|
||||
id: "REL-CRON-904",
|
||||
area: "cron duplicate prevention",
|
||||
status: "single-run lock verified for restart wakeups",
|
||||
};
|
||||
audit-fixture/src/memory/recall.ts: |-
|
||||
export const memoryRecallReleaseFinding = {
|
||||
id: "REL-MEMORY-552",
|
||||
area: "memory recall fallback",
|
||||
status: "fallback summary survives empty memory search, but ranking sample needs second reviewer",
|
||||
};
|
||||
audit-fixture/src/plugins/runtime.ts: |-
|
||||
export const pluginRuntimeReleaseFinding = {
|
||||
id: "REL-PLUGIN-319",
|
||||
area: "plugin runtime loading",
|
||||
status: "bundled runtime manifest loads cleanly after restart",
|
||||
};
|
||||
audit-fixture/install/update.ts: |-
|
||||
export const installerUpdateReleaseFinding = {
|
||||
id: "REL-INSTALL-846",
|
||||
area: "installer update path",
|
||||
status: "update smoke passed from previous stable tag",
|
||||
};
|
||||
prompt: |-
|
||||
Do a release readiness audit for the small project under `audit-fixture/`.
|
||||
Start from `audit-fixture/README.md`, find the current checklist, inspect the referenced docs/source, then create `audit-fixture/release-audit.json` and `audit-fixture/release-handoff.md`.
|
||||
|
||||
The JSON should include current finding ids, source files, statuses, and a boolean `verified` on each finding.
|
||||
If a current checklist source is referenced but missing from the fixture, include that source as a finding with `verified: false`.
|
||||
The Markdown handoff should summarize what is ready and what needs follow-up.
|
||||
Check your generated JSON before finalizing.
|
||||
Final reply exactly: RELEASE-AUDIT-COMPLETE
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: completes the sustained release audit with verified artifacts
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Object.entries(config.workspaceFiles ?? {})"
|
||||
item: workspaceFile
|
||||
actions:
|
||||
- set: seededPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, String(workspaceFile[0]))"
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- expr: "path.dirname(seededPath)"
|
||||
- recursive: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: seededPath
|
||||
- expr: "`${String(workspaceFile[1] ?? '').trimEnd()}\\n`"
|
||||
- utf8
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:release-audit:${randomUUID().slice(0, 8)}`"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 120000)
|
||||
- set: reportPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, config.reportFile)"
|
||||
- set: handoffPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, config.handoffFile)"
|
||||
- call: waitForCondition
|
||||
saveAs: reportText
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "fs.readFile(reportPath, 'utf8').then((value) => config.expectedFindings.every((finding) => value.includes(finding)) ? value : undefined).catch(() => undefined)"
|
||||
- expr: liveTurnTimeoutMs(env, 60000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- call: waitForCondition
|
||||
saveAs: handoffText
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "fs.readFile(handoffPath, 'utf8').then((value) => config.expectedFindings.every((finding) => value.includes(finding)) && !value.includes('REL-STALE-000') ? value : undefined).catch(() => undefined)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- set: report
|
||||
value:
|
||||
expr: "JSON.parse(reportText)"
|
||||
- assert:
|
||||
expr: "['src/gateway/reconnect.ts', 'src/channels/delivery.ts', 'src/scheduling/cron.ts', 'src/memory/recall.ts', 'src/plugins/runtime.ts', 'install/update.ts', 'docs/operator-notes.md'].every((file) => JSON.stringify(report).includes(file))"
|
||||
message:
|
||||
expr: "`report missing expected source refs: ${reportText}`"
|
||||
- assert:
|
||||
expr: "config.expectedFindings.every((finding) => JSON.stringify(report).includes(finding))"
|
||||
message:
|
||||
expr: "`report missing expected finding ids: ${reportText}`"
|
||||
- assert:
|
||||
expr: "!JSON.stringify(Array.isArray(report.findings) ? report.findings : report).includes('REL-STALE-000') && !handoffText.includes('REL-STALE-000')"
|
||||
message:
|
||||
expr: "`stale archive finding leaked into audit: report=${reportText}\\nhandoff=${handoffText}`"
|
||||
- set: reportFindings
|
||||
value:
|
||||
expr: "Array.isArray(report) ? report : (Array.isArray(report.findings) ? report.findings : [])"
|
||||
- set: uiFinding
|
||||
value:
|
||||
expr: "reportFindings.find((finding) => JSON.stringify(finding).includes('ui/control-panel.ts'))"
|
||||
- assert:
|
||||
expr: "uiFinding && uiFinding.verified === false"
|
||||
message:
|
||||
expr: "`missing UI evidence was not marked unverified: report=${reportText}\\nhandoff=${handoffText}`"
|
||||
- assert:
|
||||
expr: "reportFindings.length > 0 && reportFindings.every((finding) => typeof finding?.verified === 'boolean')"
|
||||
message:
|
||||
expr: "`each report finding must include a boolean verified field: ${reportText}`"
|
||||
- call: waitForAgentHistoryReply
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- lambda:
|
||||
params: [text]
|
||||
expr: "text.trim() === 'RELEASE-AUDIT-COMPLETE'"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- call: readRawQaSessionStore
|
||||
saveAs: store
|
||||
args:
|
||||
- ref: env
|
||||
- set: sessionEntry
|
||||
value:
|
||||
expr: "store[sessionKey]"
|
||||
- assert:
|
||||
expr: "Boolean(sessionEntry)"
|
||||
message:
|
||||
expr: "`missing QA session entry for ${sessionKey}`"
|
||||
detailsExpr: "`${outbound.text}\\n${reportText}\\n\\n${handoffText}`"
|
||||
156
qa/scenarios/workspace/medium-game-plan-codex-harness.yaml
Normal file
156
qa/scenarios/workspace/medium-game-plan-codex-harness.yaml
Normal file
@@ -0,0 +1,156 @@
|
||||
title: Medium game plan Codex harness
|
||||
|
||||
scenario:
|
||||
id: medium-game-plan-codex-harness
|
||||
surface: workspace
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.planning
|
||||
secondary:
|
||||
- models.codex-cli
|
||||
objective: Verify the Codex app-server harness can plan and build a medium-complex self-contained browser game.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary model is openai/gpt-5.5 with the Codex harness forced.
|
||||
- The scenario forces the Codex embedded harness.
|
||||
- The prompt explicitly asks the agent to enter plan mode before editing.
|
||||
- The agent writes a self-contained HTML game with a canvas loop, controls, scoring, waves, pause, and restart.
|
||||
docsRefs:
|
||||
- docs/plugins/sdk-agent-harness.md
|
||||
- docs/gateway/configuration-reference.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/codex/harness.ts
|
||||
- src/agents/harness/selection.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model openai/gpt-5.5 --alt-model openai/gpt-5.5 --fast --thinking medium --scenario medium-game-plan-codex-harness`.
|
||||
config:
|
||||
requiredProvider: codex
|
||||
requiredModel: gpt-5.5
|
||||
harnessRuntime: codex
|
||||
artifactFile: star-garden-defenders-codex.html
|
||||
gameTitle: Star Garden Defenders
|
||||
minBytes: 5000
|
||||
buildPrompt: |-
|
||||
Enter plan mode first and write a short implementation plan before editing.
|
||||
|
||||
Then build a medium-complex, self-contained browser game at ./star-garden-defenders-codex.html.
|
||||
|
||||
Game: Star Garden Defenders.
|
||||
Requirements:
|
||||
- one HTML file only; no external assets, fonts, scripts, or network calls
|
||||
- canvas-based arcade loop with requestAnimationFrame
|
||||
- keyboard controls and mouse or pointer support
|
||||
- player movement, enemy waves, collectibles or power-ups, collision handling
|
||||
- score, lives or health, wave number, pause, restart, and game-over state
|
||||
- polished inline CSS and clear on-screen controls
|
||||
- after writing the file, reply with the filename and the main systems implemented
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: confirms GPT-5.5 Codex harness target
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
|
||||
message:
|
||||
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: patchConfig
|
||||
saveAs: patchResult
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
agents:
|
||||
defaults:
|
||||
models:
|
||||
expr: "({ [env.primaryModel]: { agentRuntime: { id: config.harnessRuntime } } })"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: readConfigSnapshot
|
||||
saveAs: snapshot
|
||||
args:
|
||||
- ref: env
|
||||
- assert:
|
||||
expr: "snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id === config.harnessRuntime"
|
||||
message:
|
||||
expr: "`expected ${env.primaryModel} agentRuntime.id=${config.harnessRuntime}, got ${JSON.stringify(snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime)}`"
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} runtime=${snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id}` : `mock mode: parsed ${scenario.id}`"
|
||||
- name: builds the medium game artifact
|
||||
actions:
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:medium-game-codex
|
||||
message:
|
||||
expr: config.buildPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 420000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.artifactFile)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, config.artifactFile)"
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => '')).includes(config.gameTitle) ? await fs.readFile(artifactPath, 'utf8').catch(() => '') : undefined)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- 500
|
||||
- set: artifactLower
|
||||
value:
|
||||
expr: normalizeLowercaseStringOrEmpty(artifact)
|
||||
- assert:
|
||||
expr: "artifact.length >= config.minBytes"
|
||||
message:
|
||||
expr: "`expected medium game artifact >= ${config.minBytes} bytes, got ${artifact.length}`"
|
||||
- assert:
|
||||
expr: "artifactLower.includes('star garden defenders') && artifactLower.includes('<canvas') && artifactLower.includes('requestanimationframe')"
|
||||
message: missing title, canvas, or animation loop
|
||||
- assert:
|
||||
expr: "artifactLower.includes('keydown') || artifactLower.includes('keyup')"
|
||||
message: missing keyboard controls
|
||||
- assert:
|
||||
expr: "artifactLower.includes('score') && artifactLower.includes('wave') && artifactLower.includes('pause') && artifactLower.includes('restart')"
|
||||
message: missing score, wave, pause, or restart systems
|
||||
- assert:
|
||||
expr: "outbound.text.includes(config.artifactFile)"
|
||||
message:
|
||||
expr: "`final reply did not mention ${config.artifactFile}: ${outbound.text}`"
|
||||
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live medium-game build' : `${config.artifactFile} bytes=${artifact.length}`"
|
||||
156
qa/scenarios/workspace/medium-game-plan-openclaw-harness.yaml
Normal file
156
qa/scenarios/workspace/medium-game-plan-openclaw-harness.yaml
Normal file
@@ -0,0 +1,156 @@
|
||||
title: Medium game plan OpenClaw harness
|
||||
|
||||
scenario:
|
||||
id: medium-game-plan-openclaw-harness
|
||||
surface: workspace
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.planning
|
||||
secondary:
|
||||
- agents.openclaw-harness
|
||||
objective: Verify GPT-5.5 can use the OpenClaw harness to plan and build a medium-complex self-contained browser game.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary model is openai/gpt-5.5.
|
||||
- The scenario forces the embedded OpenClaw harness before the build turn.
|
||||
- The prompt explicitly asks the agent to enter plan mode before editing.
|
||||
- The agent writes a self-contained HTML game with a canvas loop, controls, scoring, waves, pause, and restart.
|
||||
docsRefs:
|
||||
- docs/plugins/sdk-agent-harness.md
|
||||
- docs/gateway/configuration-reference.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/agents/harness/selection.ts
|
||||
- src/agents/harness/builtin-openclaw.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model openai/gpt-5.5 --alt-model openai/gpt-5.5 --fast --thinking medium --scenario medium-game-plan-openclaw-harness`.
|
||||
config:
|
||||
requiredProvider: openai
|
||||
requiredModel: gpt-5.5
|
||||
harnessRuntime: openclaw
|
||||
artifactFile: star-garden-defenders-openclaw.html
|
||||
gameTitle: Star Garden Defenders
|
||||
minBytes: 5000
|
||||
buildPrompt: |-
|
||||
Enter plan mode first and write a short implementation plan before editing.
|
||||
|
||||
Then build a medium-complex, self-contained browser game at ./star-garden-defenders-openclaw.html.
|
||||
|
||||
Game: Star Garden Defenders.
|
||||
Requirements:
|
||||
- one HTML file only; no external assets, fonts, scripts, or network calls
|
||||
- canvas-based arcade loop with requestAnimationFrame
|
||||
- keyboard controls and mouse or pointer support
|
||||
- player movement, enemy waves, collectibles or power-ups, collision handling
|
||||
- score, lives or health, wave number, pause, restart, and game-over state
|
||||
- polished inline CSS and clear on-screen controls
|
||||
- after writing the file, reply with the filename and the main systems implemented
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: confirms GPT-5.5 OpenClaw harness target
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
|
||||
message:
|
||||
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: patchConfig
|
||||
saveAs: patchResult
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
agents:
|
||||
defaults:
|
||||
models:
|
||||
expr: "({ [env.primaryModel]: { agentRuntime: { id: config.harnessRuntime } } })"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: readConfigSnapshot
|
||||
saveAs: snapshot
|
||||
args:
|
||||
- ref: env
|
||||
- assert:
|
||||
expr: "snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id === config.harnessRuntime"
|
||||
message:
|
||||
expr: "`expected ${env.primaryModel} agentRuntime.id=${config.harnessRuntime}, got ${JSON.stringify(snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime)}`"
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} runtime=${snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id}` : `mock mode: parsed ${scenario.id}`"
|
||||
- name: builds the medium game artifact
|
||||
actions:
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:medium-game-openclaw
|
||||
message:
|
||||
expr: config.buildPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 420000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.artifactFile)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, config.artifactFile)"
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => '')).includes(config.gameTitle) ? await fs.readFile(artifactPath, 'utf8').catch(() => '') : undefined)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- 500
|
||||
- set: artifactLower
|
||||
value:
|
||||
expr: normalizeLowercaseStringOrEmpty(artifact)
|
||||
- assert:
|
||||
expr: "artifact.length >= config.minBytes"
|
||||
message:
|
||||
expr: "`expected medium game artifact >= ${config.minBytes} bytes, got ${artifact.length}`"
|
||||
- assert:
|
||||
expr: "artifactLower.includes('star garden defenders') && artifactLower.includes('<canvas') && artifactLower.includes('requestanimationframe')"
|
||||
message: missing title, canvas, or animation loop
|
||||
- assert:
|
||||
expr: "artifactLower.includes('keydown') || artifactLower.includes('keyup')"
|
||||
message: missing keyboard controls
|
||||
- assert:
|
||||
expr: "artifactLower.includes('score') && artifactLower.includes('wave') && artifactLower.includes('pause') && artifactLower.includes('restart')"
|
||||
message: missing score, wave, pause, or restart systems
|
||||
- assert:
|
||||
expr: "outbound.text.includes(config.artifactFile)"
|
||||
message:
|
||||
expr: "`final reply did not mention ${config.artifactFile}: ${outbound.text}`"
|
||||
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live medium-game build' : `${config.artifactFile} bytes=${artifact.length}`"
|
||||
77
qa/scenarios/workspace/source-docs-discovery-report.yaml
Normal file
77
qa/scenarios/workspace/source-docs-discovery-report.yaml
Normal file
@@ -0,0 +1,77 @@
|
||||
title: Source and docs discovery report
|
||||
|
||||
scenario:
|
||||
id: source-docs-discovery-report
|
||||
surface: discovery
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.repo-discovery
|
||||
secondary:
|
||||
- docs.discovery
|
||||
objective: Verify the agent can read repo docs and source, expand the QA plan, and publish a worked or did-not-work report.
|
||||
successCriteria:
|
||||
- Agent reads docs and source before proposing more tests.
|
||||
- Agent identifies extra candidate scenarios beyond the seed list.
|
||||
- Agent ends with a worked or failed QA report.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/web/dashboard.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/report.ts
|
||||
- extensions/qa-lab/src/self-check.ts
|
||||
- src/agents/system-prompt.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can read repo docs and source, expand the QA plan, and publish a worked or did-not-work report.
|
||||
config:
|
||||
requiredFiles:
|
||||
- repo/qa/scenarios/index.yaml
|
||||
- repo/extensions/qa-lab/src/suite.ts
|
||||
- repo/docs/help/testing.md
|
||||
prompt: Read the seeded docs and source plan. The full repo is mounted under ./repo/. Explicitly inspect repo/qa/scenarios/index.yaml, repo/extensions/qa-lab/src/suite.ts, and repo/docs/help/testing.md, then report grouped into Worked, Failed, Blocked, and Follow-up. Mention at least two extra QA scenarios beyond the seed list.
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: reads seeded material and emits a protocol report
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:discovery
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && hasDiscoveryLabels(candidate.text)).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!reportsMissingDiscoveryFiles(outbound.text)"
|
||||
message:
|
||||
expr: "`discovery report still missed repo files: ${outbound.text}`"
|
||||
- assert:
|
||||
expr: "!reportsDiscoveryScopeLeak(outbound.text)"
|
||||
message:
|
||||
expr: "`discovery report drifted beyond scope: ${outbound.text}`"
|
||||
# Parity gate criterion 2 (no fake progress / fake tool completion):
|
||||
# require an actual read tool call before the prose report. Without this,
|
||||
# a model could fabricate a plausible Worked/Failed/Blocked/Follow-up
|
||||
# report without ever touching the repo files the prompt names. The
|
||||
# debug request log is fetched once and reused for both the assertion
|
||||
# and its failure-message diagnostic. Each request's allInputText is
|
||||
# lowercased inline at match time (the real prompt writes it as
|
||||
# "Worked, Failed, Blocked") so the contains check is case-insensitive.
|
||||
- set: discoveryDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))] : []"
|
||||
- assert:
|
||||
expr: "!env.mock || discoveryDebugRequests.some((request) => String(request.allInputText ?? '').toLowerCase().includes('worked, failed, blocked') && request.plannedToolName === 'read')"
|
||||
message:
|
||||
expr: "`expected at least one read tool call during discovery report scenario, saw plannedToolNames=${JSON.stringify(discoveryDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
detailsExpr: outbound.text
|
||||
Reference in New Issue
Block a user