Vendor OpenClaw source as Adolf fork baseline
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled

Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11),
free to diverge. Tree copied sans upstream .git; upstream remote added for
future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19.
Preserves docs/ARCHITECTURE.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
This commit is contained in:
2026-07-05 09:36:54 +00:00
parent 3216769225
commit bedb527145
21108 changed files with 6010766 additions and 0 deletions

View File

@@ -0,0 +1,64 @@
title: Build Lobster Invaders
scenario:
id: lobster-invaders-build
surface: workspace
coverage:
primary:
- workspace.artifacts
secondary:
- workspace.builds
objective: Verify the agent can read the repo, create a tiny playable artifact, and report what changed.
successCriteria:
- Agent inspects source before coding.
- Agent builds a tiny playable Lobster Invaders artifact.
- Agent explains how to run or view the artifact.
docsRefs:
- docs/help/testing.md
- docs/web/dashboard.md
codeRefs:
- extensions/qa-lab/src/report.ts
- extensions/qa-lab/web/src/app.ts
execution:
kind: flow
summary: Verify the agent can read the repo, create a tiny playable artifact, and report what changed.
config:
prompt: Read the QA kickoff context first, then build a tiny Lobster Invaders HTML game at ./lobster-invaders.html in this workspace and tell me where it is.
flow:
steps:
- name: creates the artifact after reading context
actions:
- call: reset
- call: runAgentPrompt
args:
- ref: env
- sessionKey: agent:qa:lobster-invaders
message:
expr: config.prompt
timeoutMs:
expr: liveTurnTimeoutMs(env, 30000)
- call: waitForOutboundMessage
args:
- ref: state
- lambda:
params: [candidate]
expr: "candidate.conversation.id === 'qa-operator'"
- set: artifactPath
value:
expr: "path.join(env.gateway.workspaceDir, 'lobster-invaders.html')"
- call: waitForCondition
saveAs: artifact
args:
- lambda:
async: true
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => null))?.includes('Lobster Invaders') ? await fs.readFile(artifactPath, 'utf8').catch(() => null) : undefined)"
- expr: liveTurnTimeoutMs(env, 20000)
- 250
- assert:
expr: "artifact.includes('Lobster Invaders')"
message: missing Lobster Invaders artifact
- assert:
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).some((request) => (request.toolOutput ?? '').includes('QA mission'))"
message: expected pre-write read evidence
detailsExpr: "'lobster-invaders.html'"

View File

@@ -0,0 +1,251 @@
title: Long-running release audit
scenario:
id: long-running-release-audit
surface: workspace
coverage:
primary:
- workspace.long-running-task
secondary:
- workspace.repo-discovery
- workspace.artifacts
objective: Verify a model can complete a sustained repo audit that requires reading several scattered files, ignoring stale decoys, writing artifacts, and checking its own output.
successCriteria:
- Agent starts from the seeded project README instead of guessing.
- Agent inspects docs and source files across multiple directories.
- Agent writes a JSON audit and a Markdown handoff with all required findings.
- Each JSON finding carries a boolean verified field; missing current source evidence is verified false.
- Agent verifies the generated JSON before finalizing.
docsRefs:
- docs/help/testing.md
- qa/scenarios/index.yaml
codeRefs:
- extensions/qa-lab/src/suite.ts
- extensions/qa-lab/src/scenario-flow-runner.ts
execution:
kind: flow
summary: Verify sustained repo-style auditing with scattered evidence and artifact verification.
config:
reportFile: audit-fixture/release-audit.json
handoffFile: audit-fixture/release-handoff.md
expectedFindings:
- REL-GATEWAY-417
- REL-CHANNEL-238
- REL-CRON-904
- REL-MEMORY-552
- REL-PLUGIN-319
- REL-INSTALL-846
- REL-DOCS-611
workspaceFiles:
audit-fixture/README.md: |-
# Finch Relay
Release readiness task:
The release checklist was renamed during docs cleanup.
Find the current checklist under `docs/`, then inspect the referenced
source and notes. Produce:
- `release-audit.json`
- `release-handoff.md`
Do not use `docs/archive/old-release-notes.md` as current evidence.
Historical path that may appear in stale notes: `docs/release-checklist.md`.
audit-fixture/docs/current-readiness-checklist.md: |-
# Release Checklist
Current release readiness requires checking eight areas:
1. gateway reconnect handling
2. channel delivery ordering
3. cron duplicate prevention
4. memory recall fallback
5. plugin runtime loading
6. control UI status copy
7. installer update path
8. docs update status
Useful current sources:
- `src/gateway/reconnect.ts`
- `src/channels/delivery.ts`
- `src/scheduling/cron.ts`
- `src/memory/recall.ts`
- `src/plugins/runtime.ts`
- `ui/control-panel.ts`
- `install/update.ts`
- `docs/operator-notes.md`
The archive folder contains stale notes and should not be treated as
current release evidence.
audit-fixture/docs/operator-notes.md: |-
# Operator Notes
Current docs update status:
Finding id: REL-DOCS-611
Status: docs mention reconnect, cron, memory, plugin, and installer checks,
but the channel ordering and UI notes still need maintainer handoff.
audit-fixture/docs/archive/old-release-notes.md: |-
# Old Release Notes
Stale finding id: REL-STALE-000
This file is from a previous release and should not appear in the new
release audit.
audit-fixture/src/gateway/reconnect.ts: |-
export const gatewayReconnectReleaseFinding = {
id: "REL-GATEWAY-417",
area: "gateway reconnect handling",
status: "retry jitter verified, resume token fallback still needs manual spot check",
};
audit-fixture/src/channels/delivery.ts: |-
export const channelDeliveryReleaseFinding = {
id: "REL-CHANNEL-238",
area: "channel delivery ordering",
status: "thread replies preserve ordering, root-channel fallback needs handoff note",
};
audit-fixture/src/scheduling/cron.ts: |-
export const cronDuplicateReleaseFinding = {
id: "REL-CRON-904",
area: "cron duplicate prevention",
status: "single-run lock verified for restart wakeups",
};
audit-fixture/src/memory/recall.ts: |-
export const memoryRecallReleaseFinding = {
id: "REL-MEMORY-552",
area: "memory recall fallback",
status: "fallback summary survives empty memory search, but ranking sample needs second reviewer",
};
audit-fixture/src/plugins/runtime.ts: |-
export const pluginRuntimeReleaseFinding = {
id: "REL-PLUGIN-319",
area: "plugin runtime loading",
status: "bundled runtime manifest loads cleanly after restart",
};
audit-fixture/install/update.ts: |-
export const installerUpdateReleaseFinding = {
id: "REL-INSTALL-846",
area: "installer update path",
status: "update smoke passed from previous stable tag",
};
prompt: |-
Do a release readiness audit for the small project under `audit-fixture/`.
Start from `audit-fixture/README.md`, find the current checklist, inspect the referenced docs/source, then create `audit-fixture/release-audit.json` and `audit-fixture/release-handoff.md`.
The JSON should include current finding ids, source files, statuses, and a boolean `verified` on each finding.
If a current checklist source is referenced but missing from the fixture, include that source as a finding with `verified: false`.
The Markdown handoff should summarize what is ready and what needs follow-up.
Check your generated JSON before finalizing.
Final reply exactly: RELEASE-AUDIT-COMPLETE
flow:
steps:
- name: completes the sustained release audit with verified artifacts
actions:
- call: waitForGatewayHealthy
args:
- ref: env
- 60000
- call: reset
- forEach:
items:
expr: "Object.entries(config.workspaceFiles ?? {})"
item: workspaceFile
actions:
- set: seededPath
value:
expr: "path.join(env.gateway.workspaceDir, String(workspaceFile[0]))"
- call: fs.mkdir
args:
- expr: "path.dirname(seededPath)"
- recursive: true
- call: fs.writeFile
args:
- ref: seededPath
- expr: "`${String(workspaceFile[1] ?? '').trimEnd()}\\n`"
- utf8
- set: sessionKey
value:
expr: "`agent:qa:release-audit:${randomUUID().slice(0, 8)}`"
- call: runAgentPrompt
args:
- ref: env
- sessionKey:
ref: sessionKey
message:
expr: config.prompt
timeoutMs:
expr: liveTurnTimeoutMs(env, 120000)
- set: reportPath
value:
expr: "path.join(env.gateway.workspaceDir, config.reportFile)"
- set: handoffPath
value:
expr: "path.join(env.gateway.workspaceDir, config.handoffFile)"
- call: waitForCondition
saveAs: reportText
args:
- lambda:
async: true
expr: "fs.readFile(reportPath, 'utf8').then((value) => config.expectedFindings.every((finding) => value.includes(finding)) ? value : undefined).catch(() => undefined)"
- expr: liveTurnTimeoutMs(env, 60000)
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
- call: waitForCondition
saveAs: handoffText
args:
- lambda:
async: true
expr: "fs.readFile(handoffPath, 'utf8').then((value) => config.expectedFindings.every((finding) => value.includes(finding)) && !value.includes('REL-STALE-000') ? value : undefined).catch(() => undefined)"
- expr: liveTurnTimeoutMs(env, 30000)
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
- set: report
value:
expr: "JSON.parse(reportText)"
- assert:
expr: "['src/gateway/reconnect.ts', 'src/channels/delivery.ts', 'src/scheduling/cron.ts', 'src/memory/recall.ts', 'src/plugins/runtime.ts', 'install/update.ts', 'docs/operator-notes.md'].every((file) => JSON.stringify(report).includes(file))"
message:
expr: "`report missing expected source refs: ${reportText}`"
- assert:
expr: "config.expectedFindings.every((finding) => JSON.stringify(report).includes(finding))"
message:
expr: "`report missing expected finding ids: ${reportText}`"
- assert:
expr: "!JSON.stringify(Array.isArray(report.findings) ? report.findings : report).includes('REL-STALE-000') && !handoffText.includes('REL-STALE-000')"
message:
expr: "`stale archive finding leaked into audit: report=${reportText}\\nhandoff=${handoffText}`"
- set: reportFindings
value:
expr: "Array.isArray(report) ? report : (Array.isArray(report.findings) ? report.findings : [])"
- set: uiFinding
value:
expr: "reportFindings.find((finding) => JSON.stringify(finding).includes('ui/control-panel.ts'))"
- assert:
expr: "uiFinding && uiFinding.verified === false"
message:
expr: "`missing UI evidence was not marked unverified: report=${reportText}\\nhandoff=${handoffText}`"
- assert:
expr: "reportFindings.length > 0 && reportFindings.every((finding) => typeof finding?.verified === 'boolean')"
message:
expr: "`each report finding must include a boolean verified field: ${reportText}`"
- call: waitForAgentHistoryReply
saveAs: outbound
args:
- ref: env
- ref: sessionKey
- lambda:
params: [text]
expr: "text.trim() === 'RELEASE-AUDIT-COMPLETE'"
- expr: liveTurnTimeoutMs(env, 45000)
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
- call: readRawQaSessionStore
saveAs: store
args:
- ref: env
- set: sessionEntry
value:
expr: "store[sessionKey]"
- assert:
expr: "Boolean(sessionEntry)"
message:
expr: "`missing QA session entry for ${sessionKey}`"
detailsExpr: "`${outbound.text}\\n${reportText}\\n\\n${handoffText}`"

View File

@@ -0,0 +1,156 @@
title: Medium game plan Codex harness
scenario:
id: medium-game-plan-codex-harness
surface: workspace
coverage:
primary:
- workspace.planning
secondary:
- models.codex-cli
objective: Verify the Codex app-server harness can plan and build a medium-complex self-contained browser game.
successCriteria:
- A live-frontier run fails fast unless the selected primary model is openai/gpt-5.5 with the Codex harness forced.
- The scenario forces the Codex embedded harness.
- The prompt explicitly asks the agent to enter plan mode before editing.
- The agent writes a self-contained HTML game with a canvas loop, controls, scoring, waves, pause, and restart.
docsRefs:
- docs/plugins/sdk-agent-harness.md
- docs/gateway/configuration-reference.md
- docs/help/testing.md
codeRefs:
- extensions/codex/harness.ts
- src/agents/harness/selection.ts
- extensions/qa-lab/src/suite.ts
execution:
kind: flow
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model openai/gpt-5.5 --alt-model openai/gpt-5.5 --fast --thinking medium --scenario medium-game-plan-codex-harness`.
config:
requiredProvider: codex
requiredModel: gpt-5.5
harnessRuntime: codex
artifactFile: star-garden-defenders-codex.html
gameTitle: Star Garden Defenders
minBytes: 5000
buildPrompt: |-
Enter plan mode first and write a short implementation plan before editing.
Then build a medium-complex, self-contained browser game at ./star-garden-defenders-codex.html.
Game: Star Garden Defenders.
Requirements:
- one HTML file only; no external assets, fonts, scripts, or network calls
- canvas-based arcade loop with requestAnimationFrame
- keyboard controls and mouse or pointer support
- player movement, enemy waves, collectibles or power-ups, collision handling
- score, lives or health, wave number, pause, restart, and game-over state
- polished inline CSS and clear on-screen controls
- after writing the file, reply with the filename and the main systems implemented
flow:
steps:
- name: confirms GPT-5.5 Codex harness target
actions:
- set: selected
value:
expr: splitModelRef(env.primaryModel)
- assert:
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
message:
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
- assert:
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
message:
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
- if:
expr: "env.providerMode !== 'live-frontier'"
then:
- assert: "true"
else:
- call: patchConfig
saveAs: patchResult
args:
- env:
ref: env
patch:
agents:
defaults:
models:
expr: "({ [env.primaryModel]: { agentRuntime: { id: config.harnessRuntime } } })"
- call: waitForGatewayHealthy
args:
- ref: env
- 60000
- call: waitForQaChannelReady
args:
- ref: env
- 60000
- call: readConfigSnapshot
saveAs: snapshot
args:
- ref: env
- assert:
expr: "snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id === config.harnessRuntime"
message:
expr: "`expected ${env.primaryModel} agentRuntime.id=${config.harnessRuntime}, got ${JSON.stringify(snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime)}`"
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} runtime=${snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id}` : `mock mode: parsed ${scenario.id}`"
- name: builds the medium game artifact
actions:
- if:
expr: "env.providerMode !== 'live-frontier'"
then:
- assert: "true"
else:
- call: reset
- call: runAgentPrompt
args:
- ref: env
- sessionKey: agent:qa:medium-game-codex
message:
expr: config.buildPrompt
provider:
expr: selected?.provider
model:
expr: selected?.model
timeoutMs:
expr: resolveQaLiveTurnTimeoutMs(env, 420000, env.primaryModel)
- call: waitForOutboundMessage
saveAs: outbound
args:
- ref: state
- lambda:
params: [candidate]
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.artifactFile)"
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
- set: artifactPath
value:
expr: "path.join(env.gateway.workspaceDir, config.artifactFile)"
- call: waitForCondition
saveAs: artifact
args:
- lambda:
async: true
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => '')).includes(config.gameTitle) ? await fs.readFile(artifactPath, 'utf8').catch(() => '') : undefined)"
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
- 500
- set: artifactLower
value:
expr: normalizeLowercaseStringOrEmpty(artifact)
- assert:
expr: "artifact.length >= config.minBytes"
message:
expr: "`expected medium game artifact >= ${config.minBytes} bytes, got ${artifact.length}`"
- assert:
expr: "artifactLower.includes('star garden defenders') && artifactLower.includes('<canvas') && artifactLower.includes('requestanimationframe')"
message: missing title, canvas, or animation loop
- assert:
expr: "artifactLower.includes('keydown') || artifactLower.includes('keyup')"
message: missing keyboard controls
- assert:
expr: "artifactLower.includes('score') && artifactLower.includes('wave') && artifactLower.includes('pause') && artifactLower.includes('restart')"
message: missing score, wave, pause, or restart systems
- assert:
expr: "outbound.text.includes(config.artifactFile)"
message:
expr: "`final reply did not mention ${config.artifactFile}: ${outbound.text}`"
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live medium-game build' : `${config.artifactFile} bytes=${artifact.length}`"

View File

@@ -0,0 +1,156 @@
title: Medium game plan OpenClaw harness
scenario:
id: medium-game-plan-openclaw-harness
surface: workspace
coverage:
primary:
- workspace.planning
secondary:
- agents.openclaw-harness
objective: Verify GPT-5.5 can use the OpenClaw harness to plan and build a medium-complex self-contained browser game.
successCriteria:
- A live-frontier run fails fast unless the selected primary model is openai/gpt-5.5.
- The scenario forces the embedded OpenClaw harness before the build turn.
- The prompt explicitly asks the agent to enter plan mode before editing.
- The agent writes a self-contained HTML game with a canvas loop, controls, scoring, waves, pause, and restart.
docsRefs:
- docs/plugins/sdk-agent-harness.md
- docs/gateway/configuration-reference.md
- docs/help/testing.md
codeRefs:
- src/agents/harness/selection.ts
- src/agents/harness/builtin-openclaw.ts
- extensions/qa-lab/src/suite.ts
execution:
kind: flow
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model openai/gpt-5.5 --alt-model openai/gpt-5.5 --fast --thinking medium --scenario medium-game-plan-openclaw-harness`.
config:
requiredProvider: openai
requiredModel: gpt-5.5
harnessRuntime: openclaw
artifactFile: star-garden-defenders-openclaw.html
gameTitle: Star Garden Defenders
minBytes: 5000
buildPrompt: |-
Enter plan mode first and write a short implementation plan before editing.
Then build a medium-complex, self-contained browser game at ./star-garden-defenders-openclaw.html.
Game: Star Garden Defenders.
Requirements:
- one HTML file only; no external assets, fonts, scripts, or network calls
- canvas-based arcade loop with requestAnimationFrame
- keyboard controls and mouse or pointer support
- player movement, enemy waves, collectibles or power-ups, collision handling
- score, lives or health, wave number, pause, restart, and game-over state
- polished inline CSS and clear on-screen controls
- after writing the file, reply with the filename and the main systems implemented
flow:
steps:
- name: confirms GPT-5.5 OpenClaw harness target
actions:
- set: selected
value:
expr: splitModelRef(env.primaryModel)
- assert:
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
message:
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
- assert:
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
message:
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
- if:
expr: "env.providerMode !== 'live-frontier'"
then:
- assert: "true"
else:
- call: patchConfig
saveAs: patchResult
args:
- env:
ref: env
patch:
agents:
defaults:
models:
expr: "({ [env.primaryModel]: { agentRuntime: { id: config.harnessRuntime } } })"
- call: waitForGatewayHealthy
args:
- ref: env
- 60000
- call: waitForQaChannelReady
args:
- ref: env
- 60000
- call: readConfigSnapshot
saveAs: snapshot
args:
- ref: env
- assert:
expr: "snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id === config.harnessRuntime"
message:
expr: "`expected ${env.primaryModel} agentRuntime.id=${config.harnessRuntime}, got ${JSON.stringify(snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime)}`"
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} runtime=${snapshot.config.agents?.defaults?.models?.[env.primaryModel]?.agentRuntime?.id}` : `mock mode: parsed ${scenario.id}`"
- name: builds the medium game artifact
actions:
- if:
expr: "env.providerMode !== 'live-frontier'"
then:
- assert: "true"
else:
- call: reset
- call: runAgentPrompt
args:
- ref: env
- sessionKey: agent:qa:medium-game-openclaw
message:
expr: config.buildPrompt
provider:
expr: selected?.provider
model:
expr: selected?.model
timeoutMs:
expr: resolveQaLiveTurnTimeoutMs(env, 420000, env.primaryModel)
- call: waitForOutboundMessage
saveAs: outbound
args:
- ref: state
- lambda:
params: [candidate]
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.artifactFile)"
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
- set: artifactPath
value:
expr: "path.join(env.gateway.workspaceDir, config.artifactFile)"
- call: waitForCondition
saveAs: artifact
args:
- lambda:
async: true
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => '')).includes(config.gameTitle) ? await fs.readFile(artifactPath, 'utf8').catch(() => '') : undefined)"
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
- 500
- set: artifactLower
value:
expr: normalizeLowercaseStringOrEmpty(artifact)
- assert:
expr: "artifact.length >= config.minBytes"
message:
expr: "`expected medium game artifact >= ${config.minBytes} bytes, got ${artifact.length}`"
- assert:
expr: "artifactLower.includes('star garden defenders') && artifactLower.includes('<canvas') && artifactLower.includes('requestanimationframe')"
message: missing title, canvas, or animation loop
- assert:
expr: "artifactLower.includes('keydown') || artifactLower.includes('keyup')"
message: missing keyboard controls
- assert:
expr: "artifactLower.includes('score') && artifactLower.includes('wave') && artifactLower.includes('pause') && artifactLower.includes('restart')"
message: missing score, wave, pause, or restart systems
- assert:
expr: "outbound.text.includes(config.artifactFile)"
message:
expr: "`final reply did not mention ${config.artifactFile}: ${outbound.text}`"
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live medium-game build' : `${config.artifactFile} bytes=${artifact.length}`"

View File

@@ -0,0 +1,77 @@
title: Source and docs discovery report
scenario:
id: source-docs-discovery-report
surface: discovery
coverage:
primary:
- workspace.repo-discovery
secondary:
- docs.discovery
objective: Verify the agent can read repo docs and source, expand the QA plan, and publish a worked or did-not-work report.
successCriteria:
- Agent reads docs and source before proposing more tests.
- Agent identifies extra candidate scenarios beyond the seed list.
- Agent ends with a worked or failed QA report.
docsRefs:
- docs/help/testing.md
- docs/web/dashboard.md
- docs/channels/qa-channel.md
codeRefs:
- extensions/qa-lab/src/report.ts
- extensions/qa-lab/src/self-check.ts
- src/agents/system-prompt.ts
execution:
kind: flow
summary: Verify the agent can read repo docs and source, expand the QA plan, and publish a worked or did-not-work report.
config:
requiredFiles:
- repo/qa/scenarios/index.yaml
- repo/extensions/qa-lab/src/suite.ts
- repo/docs/help/testing.md
prompt: Read the seeded docs and source plan. The full repo is mounted under ./repo/. Explicitly inspect repo/qa/scenarios/index.yaml, repo/extensions/qa-lab/src/suite.ts, and repo/docs/help/testing.md, then report grouped into Worked, Failed, Blocked, and Follow-up. Mention at least two extra QA scenarios beyond the seed list.
flow:
steps:
- name: reads seeded material and emits a protocol report
actions:
- call: reset
- call: runAgentPrompt
args:
- ref: env
- sessionKey: agent:qa:discovery
message:
expr: config.prompt
timeoutMs:
expr: liveTurnTimeoutMs(env, 30000)
- call: waitForCondition
saveAs: outbound
args:
- lambda:
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && hasDiscoveryLabels(candidate.text)).at(-1)"
- expr: liveTurnTimeoutMs(env, 20000)
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
- assert:
expr: "!reportsMissingDiscoveryFiles(outbound.text)"
message:
expr: "`discovery report still missed repo files: ${outbound.text}`"
- assert:
expr: "!reportsDiscoveryScopeLeak(outbound.text)"
message:
expr: "`discovery report drifted beyond scope: ${outbound.text}`"
# Parity gate criterion 2 (no fake progress / fake tool completion):
# require an actual read tool call before the prose report. Without this,
# a model could fabricate a plausible Worked/Failed/Blocked/Follow-up
# report without ever touching the repo files the prompt names. The
# debug request log is fetched once and reused for both the assertion
# and its failure-message diagnostic. Each request's allInputText is
# lowercased inline at match time (the real prompt writes it as
# "Worked, Failed, Blocked") so the contains check is case-insensitive.
- set: discoveryDebugRequests
value:
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))] : []"
- assert:
expr: "!env.mock || discoveryDebugRequests.some((request) => String(request.allInputText ?? '').toLowerCase().includes('worked, failed, blocked') && request.plannedToolName === 'read')"
message:
expr: "`expected at least one read tool call during discovery report scenario, saw plannedToolNames=${JSON.stringify(discoveryDebugRequests.map((request) => request.plannedToolName ?? null))}`"
detailsExpr: outbound.text