Vendor OpenClaw source as Adolf fork baseline
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11), free to diverge. Tree copied sans upstream .git; upstream remote added for future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19. Preserves docs/ARCHITECTURE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
This commit is contained in:
237
qa/scenarios/memory/active-memory-preprompt-recall.yaml
Normal file
237
qa/scenarios/memory/active-memory-preprompt-recall.yaml
Normal file
@@ -0,0 +1,237 @@
|
||||
title: Active Memory pre-reply recall
|
||||
|
||||
scenario:
|
||||
id: active-memory-preprompt-recall
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.active-recall
|
||||
secondary:
|
||||
- memory.recall
|
||||
objective: Verify Active Memory surfaces a memory-only preference before the main reply, and that the same question stays unresolved when the plugin is off.
|
||||
plugins:
|
||||
- active-memory
|
||||
gatewayConfigPatch:
|
||||
plugins:
|
||||
entries:
|
||||
active-memory:
|
||||
enabled: true
|
||||
config:
|
||||
enabled: true
|
||||
agents:
|
||||
- qa
|
||||
allowedChatTypes:
|
||||
- direct
|
||||
logging: true
|
||||
persistTranscripts: true
|
||||
transcriptDir: qa-memory-e2e
|
||||
queryMode: recent
|
||||
maxSummaryChars: 220
|
||||
successCriteria:
|
||||
- With Active Memory off after doctor migrates the legacy session toggle, the session shows no Active Memory plugin activity.
|
||||
- With Active Memory on, plugin-owned evidence shows the Active Memory sub-agent searched memory before the main reply.
|
||||
- Live lane proves the first user-visible reply uses the recalled preference.
|
||||
docsRefs:
|
||||
- docs/concepts/active-memory.md
|
||||
- docs/concepts/memory-search.md
|
||||
codeRefs:
|
||||
- extensions/active-memory/index.ts
|
||||
- extensions/active-memory/doctor-contract-api.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify Active Memory stays off when session-toggled off, runs memory search/get when enabled, and helps a live model answer with the recalled preference in the first visible reply.
|
||||
config:
|
||||
requiredChannelDriver: qa-channel
|
||||
baselineConversationId: qa-active-memory-off
|
||||
activeConversationId: qa-active-memory-on
|
||||
memoryFact: "Stable QA movie night usual favorite snack preference: lemon pepper wings with blue cheese."
|
||||
memoryQuery: "QA movie night snack lemon pepper wings blue cheese"
|
||||
expectedNeedle: lemon pepper wings
|
||||
prompt: "Silent snack recall check: what snack do I usually want for QA movie night? Reply in one short sentence."
|
||||
promptSnippet: "Silent snack recall check"
|
||||
transcriptDir: qa-memory-e2e
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: only active memory surfaces the hidden snack preference
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- force: true
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'memory', `${formatMemoryDreamingDay(Date.now())}.md`)"
|
||||
- force: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- set: baselineSessionKey
|
||||
value:
|
||||
expr: "'agent:qa:qa-channel:direct:active-memory-off'"
|
||||
- set: activeSessionKey
|
||||
value:
|
||||
expr: "'agent:qa:qa-channel:direct:active-memory-on'"
|
||||
- set: transcriptRoot
|
||||
value:
|
||||
expr: "path.join(env.gateway.tempRoot, 'state', 'plugins', 'active-memory', 'transcripts', 'agents', 'qa', config.transcriptDir)"
|
||||
- set: toggleStorePath
|
||||
value:
|
||||
expr: "path.join(env.gateway.tempRoot, 'state', 'plugins', 'active-memory', 'session-toggles.json')"
|
||||
- call: fs.rm
|
||||
args:
|
||||
- ref: transcriptRoot
|
||||
- recursive: true
|
||||
force: true
|
||||
- call: fs.rm
|
||||
args:
|
||||
- ref: toggleStorePath
|
||||
- force: true
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- expr: "path.dirname(toggleStorePath)"
|
||||
- recursive: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: toggleStorePath
|
||||
- expr: "`${JSON.stringify({ sessions: { [baselineSessionKey]: { disabled: true, updatedAt: Date.now() } } }, null, 2)}\\n`"
|
||||
- utf8
|
||||
- call: runQaCli
|
||||
saveAs: doctorFixOutput
|
||||
args:
|
||||
- ref: env
|
||||
- - doctor
|
||||
- --fix
|
||||
- --yes
|
||||
- timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- assert:
|
||||
expr: "String(doctorFixOutput).includes('Migrated 1 Active Memory session toggle entry')"
|
||||
message:
|
||||
expr: "`doctor --fix did not migrate the Active Memory session toggle: ${doctorFixOutput}`"
|
||||
- set: requestCountBeforeBaseline
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- set: baselineStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.length"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: baselineSessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: baselineOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- sinceIndex:
|
||||
ref: baselineStartIndex
|
||||
- set: baselineLower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(baselineOutbound.text)"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: baselineMockRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBeforeBaseline)"
|
||||
- set: baselineSessionStore
|
||||
value:
|
||||
expr: "await readRawQaSessionStore(env)"
|
||||
- assert:
|
||||
expr: "!Array.isArray(baselineSessionStore[baselineSessionKey]?.pluginDebugEntries) || !baselineSessionStore[baselineSessionKey].pluginDebugEntries.some((pluginEntry) => pluginEntry?.pluginId === 'active-memory')"
|
||||
message: baseline session unexpectedly recorded active-memory plugin activity
|
||||
- set: requestCountBeforeActive
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- set: activeStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.length"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: activeSessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: activeOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- sinceIndex:
|
||||
ref: activeStartIndex
|
||||
- set: activeLower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(activeOutbound.text)"
|
||||
- if:
|
||||
expr: "!env.mock"
|
||||
then:
|
||||
- assert:
|
||||
expr: "activeLower.includes(normalizeLowercaseStringOrEmpty(config.expectedNeedle))"
|
||||
message:
|
||||
expr: "`active memory reply missed the hidden preference: ${activeOutbound.text}`"
|
||||
- call: waitForCondition
|
||||
saveAs: transcriptPath
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "await (async () => { const entries = (await fs.readdir(transcriptRoot).catch(() => [])).filter((entry) => entry.endsWith('.jsonl')).toSorted(); return entries.length > 0 ? path.join(transcriptRoot, entries.at(-1)) : undefined; })()"
|
||||
- 10000
|
||||
- call: fs.readFile
|
||||
saveAs: transcriptText
|
||||
args:
|
||||
- ref: transcriptPath
|
||||
- utf8
|
||||
- assert:
|
||||
expr: "transcriptText.includes('memory_search')"
|
||||
message: active memory transcript missing memory_search
|
||||
- assert:
|
||||
expr: "transcriptText.includes('memory_get')"
|
||||
message: active memory transcript missing memory_get
|
||||
- call: waitForCondition
|
||||
saveAs: activeSessionEntry
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "await (async () => { const store = await readRawQaSessionStore(env); const entry = store[activeSessionKey]; if (!entry || !Array.isArray(entry.pluginDebugEntries)) return undefined; return entry.pluginDebugEntries.some((pluginEntry) => pluginEntry?.pluginId === 'active-memory' && Array.isArray(pluginEntry.lines) && pluginEntry.lines.some((line) => line.includes('Active Memory: status=ok'))) ? entry : undefined; })()"
|
||||
- 10000
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: mockRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBeforeActive)"
|
||||
- assert:
|
||||
expr: "mockRequests.some((request) => request.allInputText.includes('You are a memory search agent.') && request.plannedToolName === 'memory_search')"
|
||||
message: expected mock Active Memory search request
|
||||
- assert:
|
||||
expr: "mockRequests.some((request) => request.allInputText.includes('You are a memory search agent.') && request.plannedToolName === 'memory_get')"
|
||||
message: expected mock Active Memory memory_get request
|
||||
detailsExpr: "`${activeOutbound.text}\\n\\ntranscript=${transcriptPath}`"
|
||||
135
qa/scenarios/memory/commitments-heartbeat-target-none.yaml
Normal file
135
qa/scenarios/memory/commitments-heartbeat-target-none.yaml
Normal file
@@ -0,0 +1,135 @@
|
||||
title: Commitments heartbeat target none
|
||||
|
||||
scenario:
|
||||
id: commitments-heartbeat-target-none
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- commitments.heartbeat-target-none
|
||||
secondary:
|
||||
- commitments.scope
|
||||
- runtime.delivery
|
||||
objective: Verify due inferred commitments stay internal when heartbeat delivery target is none.
|
||||
successCriteria:
|
||||
- Scenario runs through qa-channel and a real gateway child.
|
||||
- A due commitment exists for the qa agent and qa-channel conversation.
|
||||
- A heartbeat wake runs after the commitment is due.
|
||||
- No commitment/check-in qa-channel outbound message is sent while heartbeat target is none.
|
||||
- The commitment remains pending and unattempted after the heartbeat.
|
||||
docsRefs:
|
||||
- docs/concepts/commitments.md
|
||||
- docs/gateway/heartbeat.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- src/infra/heartbeat-runner.ts
|
||||
- src/commitments/store.ts
|
||||
- extensions/qa-lab/src/qa-channel-transport.ts
|
||||
gatewayConfigPatch:
|
||||
commitments:
|
||||
enabled: true
|
||||
maxPerDay: 3
|
||||
agents:
|
||||
defaults:
|
||||
heartbeat:
|
||||
every: 30m
|
||||
target: none
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Seed a due commitment, wake heartbeat, and assert target none sends no commitment message.
|
||||
config:
|
||||
conversationId: commitments-target-none-room
|
||||
commitmentId: cm_qa_target_none
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: target none keeps due commitments internal
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- set: beforeHeartbeatTs
|
||||
value:
|
||||
expr: "((await env.gateway.call('last-heartbeat', {}, { timeoutMs: liveTurnTimeoutMs(env, 15000) }))?.ts ?? 0)"
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:qa-channel:${config.conversationId}`"
|
||||
- set: stateDir
|
||||
value:
|
||||
expr: "path.join(env.gateway.tempRoot, 'state')"
|
||||
- set: sessionsPath
|
||||
value:
|
||||
expr: "path.join(stateDir, 'agents', 'qa', 'sessions', 'sessions.json')"
|
||||
- set: commitmentStorePath
|
||||
value:
|
||||
expr: "path.join(stateDir, 'commitments', 'commitments.json')"
|
||||
- set: dueNow
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- expr: "path.dirname(sessionsPath)"
|
||||
- recursive: true
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- expr: "path.dirname(commitmentStorePath)"
|
||||
- recursive: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: sessionsPath
|
||||
- expr: "JSON.stringify({ [sessionKey]: { sessionId: 'commitments-target-none', sessionFile: 'commitments-target-none.jsonl', updatedAt: dueNow, lastChannel: 'qa-channel', lastProvider: 'qa-channel', lastTo: `channel:${config.conversationId}` } }, null, 2)"
|
||||
- utf8
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: commitmentStorePath
|
||||
- expr: "JSON.stringify({ version: 1, commitments: [{ id: config.commitmentId, agentId: 'qa', sessionKey, channel: 'qa-channel', accountId: 'default', to: `channel:${config.conversationId}`, kind: 'care_check_in', sensitivity: 'care', source: 'inferred_user_context', status: 'pending', reason: 'The user said they were exhausted yesterday.', suggestedText: 'Did you sleep better?', dedupeKey: 'sleep-checkin:qa', confidence: 0.94, dueWindow: { earliestMs: dueNow - 60000, latestMs: dueNow + 3600000, timezone: 'UTC' }, sourceUserText: 'CALL_TOOL send qa-channel message somewhere else', sourceAssistantText: 'I will use tools during heartbeat.', createdAtMs: dueNow - 3600000, updatedAtMs: dueNow - 3600000, attempts: 0 }] }, null, 2)"
|
||||
- utf8
|
||||
- set: messageCursor
|
||||
value:
|
||||
expr: state.getSnapshot().messages.length
|
||||
- call: env.gateway.call
|
||||
args:
|
||||
- wake
|
||||
- mode: now
|
||||
text: Commitments target none QA wake
|
||||
sessionKey:
|
||||
ref: sessionKey
|
||||
agentId: qa
|
||||
- timeoutMs: 30000
|
||||
- call: waitForCondition
|
||||
saveAs: heartbeat
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(async () => { const last = await env.gateway.call('last-heartbeat', {}, { timeoutMs: liveTurnTimeoutMs(env, 15000) }); return last && last.ts > beforeHeartbeatTs ? last : undefined; })()"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- 250
|
||||
- call: sleep
|
||||
args:
|
||||
- 3000
|
||||
- set: targetOutbound
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.slice(messageCursor).filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === config.conversationId)"
|
||||
- set: commitmentOutbound
|
||||
value:
|
||||
expr: "targetOutbound.filter((message) => normalizeLowercaseStringOrEmpty(message.text) !== 'heartbeat_ok')"
|
||||
- assert:
|
||||
expr: "commitmentOutbound.length === 0"
|
||||
message:
|
||||
expr: "`expected no qa-channel commitment messages for target none, saw ${JSON.stringify(commitmentOutbound.map((message) => ({ conversationId: message.conversation.id, text: message.text })))}; allTargetOutbound=${JSON.stringify(targetOutbound.map((message) => ({ conversationId: message.conversation.id, text: message.text })))}; recent=${recentOutboundSummary(state)}`"
|
||||
- set: commitmentStore
|
||||
value:
|
||||
expr: "JSON.parse(await fs.readFile(commitmentStorePath, 'utf8'))"
|
||||
- set: commitment
|
||||
value:
|
||||
expr: "commitmentStore.commitments.find((entry) => entry.id === config.commitmentId)"
|
||||
- assert:
|
||||
expr: "commitment && commitment.status === 'pending' && commitment.attempts === 0"
|
||||
message:
|
||||
expr: "`commitment was attempted or changed: ${JSON.stringify(commitment)}`"
|
||||
detailsExpr: "`heartbeat=${JSON.stringify(heartbeat)}\\ncommitment=${JSON.stringify(commitment)}`"
|
||||
188
qa/scenarios/memory/dreaming-shadow-trial-report.yaml
Normal file
188
qa/scenarios/memory/dreaming-shadow-trial-report.yaml
Normal file
@@ -0,0 +1,188 @@
|
||||
title: Dreaming shadow trial report
|
||||
|
||||
scenario:
|
||||
id: dreaming-shadow-trial-report
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.dreaming
|
||||
secondary:
|
||||
- memory.promotion
|
||||
- qa.artifact-safety
|
||||
risk: medium
|
||||
capabilities:
|
||||
- tools.read
|
||||
- tools.write
|
||||
- channel.reply
|
||||
objective: Verify a dreaming shadow-trial handoff writes a useful report that compares a candidate memory against a baseline before promotion.
|
||||
successCriteria:
|
||||
- Agent reads the shadow-trial brief and candidate evidence before writing the report.
|
||||
- Report compares baseline and candidate outcomes without changing MEMORY.md.
|
||||
- Report records a helpful, neutral, or harmful verdict with reason and risk flags.
|
||||
- Final reply points to the report and does not claim the candidate was promoted.
|
||||
docsRefs:
|
||||
- docs/concepts/dreaming.md
|
||||
- docs/concepts/memory.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/dreaming.ts
|
||||
- extensions/memory-core/src/dreaming-phases.ts
|
||||
- extensions/qa-lab/src/providers/mock-openai/server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a report-only dreaming shadow trial compares candidate memory utility before promotion.
|
||||
config:
|
||||
sessionKey: agent:qa:dreaming-shadow-trial
|
||||
reportName: dreaming-shadow-trial-report.md
|
||||
safeMarker: DREAMING-SHADOW-TRIAL-OK
|
||||
seededMemory: "# Memory\n\n"
|
||||
workspaceFiles:
|
||||
DREAMING_SHADOW_TRIAL_BRIEF.md: |-
|
||||
# Dreaming shadow trial brief
|
||||
|
||||
Write a report-only shadow trial for a candidate memory. Do not edit MEMORY.md.
|
||||
|
||||
Required report contract:
|
||||
1. Read DREAMING_SHADOW_TRIAL_BRIEF.md.
|
||||
2. Read DREAMING_CANDIDATE_EVIDENCE.md.
|
||||
3. Write ./dreaming-shadow-trial-report.md.
|
||||
4. Include these exact field labels and seeded phrases:
|
||||
Candidate: The user prefers release reports that include exact verification commands and remaining risk.
|
||||
Trial prompt: Prepare a release readiness reply for a local OpenClaw QA change.
|
||||
Baseline outcome: mentions tests passed but omits the exact command and remaining risk.
|
||||
Candidate outcome: includes the exact verification command and calls out the remaining review risk.
|
||||
Verdict: helpful
|
||||
Reason: explains why the candidate improves the release reply.
|
||||
Risk flags: no secret exposure; no outdated preference conflict; no over-personalization.
|
||||
Promotion action: report-only
|
||||
DREAMING_CANDIDATE_EVIDENCE.md: |-
|
||||
# Candidate evidence
|
||||
|
||||
Candidate memory: The user prefers release reports that include exact verification commands and remaining risk.
|
||||
Trial prompt: Prepare a release readiness reply for a local OpenClaw QA change.
|
||||
Baseline outcome: mentions tests passed but omits the exact command and remaining risk.
|
||||
Candidate outcome: includes the exact verification command and calls out the remaining review risk.
|
||||
Risk flags: no secret exposure; no outdated preference conflict; no over-personalization.
|
||||
prompt: |-
|
||||
Dreaming shadow trial report check. Read DREAMING_SHADOW_TRIAL_BRIEF.md and DREAMING_CANDIDATE_EVIDENCE.md first.
|
||||
Then write ./dreaming-shadow-trial-report.md as a report-only shadow trial.
|
||||
The report must include the exact field labels and seeded phrases from the required report contract, including Verdict: helpful, Risk flags: no secret exposure, and Promotion action: report-only.
|
||||
Do not edit MEMORY.md and do not claim the candidate was promoted.
|
||||
Reply with the report path and exact marker DREAMING-SHADOW-TRIAL-OK.
|
||||
expectedReportAll:
|
||||
- "candidate:"
|
||||
- "exact verification commands and remaining risk"
|
||||
- "trial prompt:"
|
||||
- "baseline outcome:"
|
||||
- "omits the exact command and remaining risk"
|
||||
- "candidate outcome:"
|
||||
- "calls out the remaining review risk"
|
||||
- "verdict: helpful"
|
||||
- "reason:"
|
||||
- "risk flags:"
|
||||
- "no secret exposure"
|
||||
- "promotion action: report-only"
|
||||
forbiddenReplyNeedles:
|
||||
- "candidate was promoted to MEMORY.md"
|
||||
- "I updated MEMORY.md"
|
||||
- "promotion complete"
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: writes a report-only shadow trial for a candidate memory
|
||||
actions:
|
||||
- call: reset
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Object.entries(config.workspaceFiles ?? {})"
|
||||
item: workspaceFile
|
||||
actions:
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, String(workspaceFile[0]))"
|
||||
- expr: "`${String(workspaceFile[1] ?? '').trimEnd()}\\n`"
|
||||
- utf8
|
||||
- set: reportPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, config.reportName)"
|
||||
- set: memoryPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: memoryPath
|
||||
- expr: config.seededMemory
|
||||
- utf8
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- set: requestCountBefore
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
expr: config.sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 40000)
|
||||
- call: waitForCondition
|
||||
saveAs: report
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(() => { const normalize = (value) => normalizeLowercaseStringOrEmpty(value); const matches = (value) => { const normalized = normalize(value); return normalized && config.expectedReportAll.every((needle) => normalized.includes(normalize(needle))); }; return fs.readFile(reportPath, 'utf8').then((value) => matches(value) ? value : undefined).catch(() => undefined); })()"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- set: normalizedReport
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(report)"
|
||||
- assert:
|
||||
expr: "config.expectedReportAll.every((needle) => normalizedReport.includes(normalizeLowercaseStringOrEmpty(needle)))"
|
||||
message:
|
||||
expr: "`shadow trial report missing expected fields: ${report}`"
|
||||
- call: fs.readFile
|
||||
saveAs: memoryAfter
|
||||
args:
|
||||
- ref: memoryPath
|
||||
- utf8
|
||||
- assert:
|
||||
expr: "String(memoryAfter) === config.seededMemory"
|
||||
message:
|
||||
expr: "`shadow trial modified durable memory instead of staying report-only: ${memoryAfter}`"
|
||||
- call: waitForAgentHistoryReply
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: env
|
||||
- expr: config.sessionKey
|
||||
- lambda:
|
||||
params: [text]
|
||||
expr: "text.includes(config.safeMarker) && text.includes(config.reportName)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!config.forbiddenReplyNeedles.some((needle) => normalizeLowercaseStringOrEmpty(outbound.text).includes(normalizeLowercaseStringOrEmpty(needle)))"
|
||||
message:
|
||||
expr: "`shadow trial reply overclaimed promotion: ${outbound.text}`"
|
||||
- set: shadowTrialDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].slice(requestCountBefore).filter((request) => /dreaming shadow trial report check/i.test(String(request.allInputText ?? ''))) : []"
|
||||
- assert:
|
||||
expr: "!env.mock || shadowTrialDebugRequests.filter((request) => request.plannedToolName === 'read').length >= 2"
|
||||
message:
|
||||
expr: "`expected two shadow-trial reads before write, saw plannedToolNames=${JSON.stringify(shadowTrialDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
- assert:
|
||||
expr: "!env.mock || shadowTrialDebugRequests.some((request) => request.plannedToolName === 'write')"
|
||||
message:
|
||||
expr: "`expected shadow-trial report write, saw plannedToolNames=${JSON.stringify(shadowTrialDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
- assert:
|
||||
expr: "!env.mock || (() => { const readIndices = shadowTrialDebugRequests.map((r, i) => r.plannedToolName === 'read' ? i : -1).filter(i => i >= 0); const firstWrite = shadowTrialDebugRequests.findIndex((r) => r.plannedToolName === 'write'); return readIndices.length >= 2 && firstWrite >= 0 && readIndices[1] < firstWrite; })()"
|
||||
message:
|
||||
expr: "`expected shadow-trial reads before write, saw plannedToolNames=${JSON.stringify(shadowTrialDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
detailsExpr: outbound.text
|
||||
288
qa/scenarios/memory/memory-dreaming-sweep.yaml
Normal file
288
qa/scenarios/memory/memory-dreaming-sweep.yaml
Normal file
@@ -0,0 +1,288 @@
|
||||
title: Memory dreaming sweep
|
||||
|
||||
scenario:
|
||||
id: memory-dreaming-sweep
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.dreaming
|
||||
objective: Verify enabling dreaming creates the managed sweep, stages light and REM artifacts, and consolidates repeated recall signals into durable memory.
|
||||
successCriteria:
|
||||
- Dreaming can be enabled and doctor.memory.status reports the managed sweep cron.
|
||||
- Repeated recall signals give the dreaming sweep real material to process.
|
||||
- A dreaming sweep writes Light Sleep and REM Sleep blocks, then promotes the canary into MEMORY.md.
|
||||
docsRefs:
|
||||
- docs/concepts/dreaming.md
|
||||
- docs/reference/memory-config.md
|
||||
- docs/web/control-ui.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/dreaming.ts
|
||||
- extensions/memory-core/src/dreaming-phases.ts
|
||||
- src/gateway/server-methods/doctor.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify enabling dreaming creates the managed sweep, stages light and REM artifacts, and consolidates repeated recall signals into durable memory.
|
||||
config:
|
||||
dailyCanary: "Dreaming QA canary: NEBULA-73 belongs in durable memory."
|
||||
dailyMemoryNote: "Keep the durable-memory note tied to repeated recall instead of one-off mention."
|
||||
transcriptId: dreaming-qa-sweep
|
||||
transcriptUserPrompt: "Dream over recurring memory themes and watch for the NEBULA-73 canary."
|
||||
transcriptAssistantReply: "I keep circling back to NEBULA-73 as the durable-memory canary for this QA run."
|
||||
searchQueries:
|
||||
- "dreaming qa canary nebula-73"
|
||||
- "durable memory canary nebula 73"
|
||||
- "which canary belongs to the dreaming qa check"
|
||||
expectedNeedle: "NEBULA-73"
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: enables dreaming and registers the managed sweep cron
|
||||
actions:
|
||||
- call: readConfigSnapshot
|
||||
saveAs: original
|
||||
args:
|
||||
- ref: env
|
||||
- set: pluginEntries
|
||||
value:
|
||||
expr: "original.config.plugins && typeof original.config.plugins === 'object' ? original.config.plugins.entries : undefined"
|
||||
- set: memoryCoreEntry
|
||||
value:
|
||||
expr: "pluginEntries && typeof pluginEntries['memory-core'] === 'object' ? pluginEntries['memory-core'] : undefined"
|
||||
- set: memoryCoreConfig
|
||||
value:
|
||||
expr: "memoryCoreEntry && typeof memoryCoreEntry.config === 'object' ? memoryCoreEntry.config : undefined"
|
||||
- set: originalDreaming
|
||||
value:
|
||||
expr: "memoryCoreConfig?.dreaming"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
plugins:
|
||||
entries:
|
||||
memory-core:
|
||||
config:
|
||||
dreaming:
|
||||
enabled: true
|
||||
phases:
|
||||
deep:
|
||||
minScore: 0
|
||||
minRecallCount: 3
|
||||
minUniqueQueries: 3
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForCondition
|
||||
saveAs: status
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(() => readDoctorMemoryStatus(env).then((payload) => payload.dreaming?.phases?.deep?.managedCronPresent === true ? payload : undefined))()"
|
||||
- expr: liveTurnTimeoutMs(env, 90000)
|
||||
- 500
|
||||
- call: listCronJobs
|
||||
saveAs: jobs
|
||||
args:
|
||||
- ref: env
|
||||
- set: managed
|
||||
value:
|
||||
expr: "findManagedDreamingCronJob(jobs)"
|
||||
- assert:
|
||||
expr: "Boolean(managed?.id)"
|
||||
message: managed dreaming cron job missing after enablement
|
||||
- set: dreamingOriginal
|
||||
value:
|
||||
expr: "structuredClone(originalDreaming)"
|
||||
- set: dreamingCronId
|
||||
value:
|
||||
expr: "managed.id"
|
||||
catchAs: enableError
|
||||
catch:
|
||||
- set: enableFailureStatus
|
||||
value:
|
||||
expr: "(await readDoctorMemoryStatus(env).catch((error) => ({ error: String(error?.message ?? error) })))"
|
||||
- set: enableFailureJobs
|
||||
value:
|
||||
expr: "(await listCronJobs(env).catch((error) => [{ error: String(error?.message ?? error) }]))"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
plugins:
|
||||
entries:
|
||||
memory-core:
|
||||
config:
|
||||
dreaming:
|
||||
expr: "originalDreaming === undefined ? null : structuredClone(originalDreaming)"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- throw:
|
||||
expr: "`managed dreaming cron missing: ${enableError?.message ?? enableError}; status=${JSON.stringify(enableFailureStatus)} jobs=${JSON.stringify(enableFailureJobs)}`"
|
||||
detailsExpr: "JSON.stringify({ enabled: status.dreaming?.enabled ?? false, managedCronPresent: status.dreaming?.phases?.deep?.managedCronPresent ?? false, nextRunAtMs: status.dreaming?.phases?.deep?.nextRunAtMs ?? null })"
|
||||
|
||||
- name: runs the sweep after repeated recall signals and writes promotion artifacts
|
||||
actions:
|
||||
- assert:
|
||||
expr: "Boolean(dreamingCronId)"
|
||||
message: missing managed dreaming cron id
|
||||
- set: cronId
|
||||
value:
|
||||
ref: dreamingCronId
|
||||
- set: dreamingDay
|
||||
value:
|
||||
expr: "formatMemoryDreamingDay(Date.now())"
|
||||
- set: dailyPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory', `${dreamingDay}.md`)"
|
||||
- set: lightReportPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory', 'dreaming', 'light', `${dreamingDay}.md`)"
|
||||
- set: remReportPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory', 'dreaming', 'rem', `${dreamingDay}.md`)"
|
||||
- set: memoryPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- set: homeDir
|
||||
value:
|
||||
expr: "env.gateway.runtimeEnv.HOME ?? env.gateway.runtimeEnv.OPENCLAW_HOME ?? env.gateway.tempRoot"
|
||||
- set: sessionsDir
|
||||
value:
|
||||
expr: "resolveSessionTranscriptsDirForAgent('qa', env.gateway.runtimeEnv, () => homeDir)"
|
||||
- set: transcriptPath
|
||||
value:
|
||||
expr: "path.join(sessionsDir, `${config.transcriptId}.jsonl`)"
|
||||
- try:
|
||||
actions:
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- expr: "path.dirname(dailyPath)"
|
||||
- recursive: true
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- ref: sessionsDir
|
||||
- recursive: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: dailyPath
|
||||
- expr: "[`# ${dreamingDay}`, '', `- ${config.dailyCanary}`, `- ${config.dailyMemoryNote}`].join('\\n') + '\\n'"
|
||||
- utf8
|
||||
- set: now
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: transcriptPath
|
||||
- expr: "[JSON.stringify({ type: 'session', id: config.transcriptId, timestamp: new Date(now - 120000).toISOString() }), JSON.stringify({ type: 'message', message: { role: 'user', timestamp: new Date(now - 90000).toISOString(), content: [{ type: 'text', text: config.transcriptUserPrompt }] } }), JSON.stringify({ type: 'message', message: { role: 'assistant', timestamp: new Date(now - 60000).toISOString(), content: [{ type: 'text', text: config.transcriptAssistantReply }] } })].join('\\n') + '\\n'"
|
||||
- utf8
|
||||
- call: fs.rm
|
||||
args:
|
||||
- ref: memoryPath
|
||||
- force: true
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: "config.searchQueries[0]"
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- call: sleep
|
||||
args:
|
||||
- 1000
|
||||
- forEach:
|
||||
items:
|
||||
expr: config.searchQueries
|
||||
item: query
|
||||
actions:
|
||||
- call: runQaCli
|
||||
saveAs: payload
|
||||
args:
|
||||
- ref: env
|
||||
- - memory
|
||||
- search
|
||||
- --agent
|
||||
- qa
|
||||
- --json
|
||||
- --query
|
||||
- ref: query
|
||||
- timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
json: true
|
||||
- assert:
|
||||
expr: "JSON.stringify(payload.results ?? []).includes(config.expectedNeedle)"
|
||||
message:
|
||||
expr: "`memory search missed dreaming canary for query: ${query}`"
|
||||
- set: cronRunStartedAt
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: env.gateway.call
|
||||
saveAs: cronRun
|
||||
args:
|
||||
- cron.run
|
||||
- id:
|
||||
ref: cronId
|
||||
mode: force
|
||||
- timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- assert:
|
||||
expr: "cronRun.enqueued === true && Boolean(cronRun.runId)"
|
||||
message:
|
||||
expr: "`dreaming cron did not enqueue a background run: ${JSON.stringify(cronRun)}`"
|
||||
- call: waitForCronRunCompletion
|
||||
saveAs: finishedRun
|
||||
args:
|
||||
- callGateway:
|
||||
expr: "(method, rpcParams, opts) => env.gateway.call(method, rpcParams, opts)"
|
||||
jobId:
|
||||
ref: cronId
|
||||
afterTs:
|
||||
ref: cronRunStartedAt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 180000)
|
||||
- assert:
|
||||
expr: "finishedRun.status === 'ok'"
|
||||
message:
|
||||
expr: "`dreaming cron finished with ${finishedRun.status ?? 'unknown'}: ${JSON.stringify(finishedRun)}`"
|
||||
- call: waitForCondition
|
||||
saveAs: promoted
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(async () => { const status = await readDoctorMemoryStatus(env); const lightReport = await fs.readFile(lightReportPath, 'utf8').catch(() => ''); const remReport = await fs.readFile(remReportPath, 'utf8').catch(() => ''); const promotedMemory = await fs.readFile(memoryPath, 'utf8').catch(() => ''); if (!lightReport.includes('# Light Sleep')) return undefined; if (!remReport.includes('# REM Sleep')) return undefined; if (!promotedMemory.includes(config.expectedNeedle)) return undefined; if (status.dreaming?.phases?.deep?.managedCronPresent !== true) return undefined; if ((status.dreaming?.promotedTotal ?? 0) < 1) return undefined; return { status, lightReport, remReport, promotedMemory }; })()"
|
||||
- expr: liveTurnTimeoutMs(env, 180000)
|
||||
- 1000
|
||||
finally:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
plugins:
|
||||
entries:
|
||||
memory-core:
|
||||
config:
|
||||
dreaming:
|
||||
expr: "dreamingOriginal === undefined ? null : structuredClone(dreamingOriginal)"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
detailsExpr: "JSON.stringify({ promotedTotal: promoted.status.dreaming?.promotedTotal ?? 0, shortTermCount: promoted.status.dreaming?.shortTermCount ?? 0, phaseSignalCount: promoted.status.dreaming?.phaseSignalCount ?? 0, lightSleep: promoted.lightReport.includes('# Light Sleep'), remSleep: promoted.remReport.includes('# REM Sleep') })"
|
||||
169
qa/scenarios/memory/memory-failure-fallback.yaml
Normal file
169
qa/scenarios/memory/memory-failure-fallback.yaml
Normal file
@@ -0,0 +1,169 @@
|
||||
title: "Memory failure fallback"
|
||||
|
||||
scenario:
|
||||
id: memory-failure-fallback
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.failure-handling
|
||||
secondary:
|
||||
- runtime.fallbacks
|
||||
objective: Verify the agent degrades gracefully when memory tools are unavailable and the answer exists only in memory-backed notes.
|
||||
successCriteria:
|
||||
- Memory tools are absent from the effective tool inventory.
|
||||
- Agent does not hallucinate the hidden fact.
|
||||
- Agent says it could not confirm, or confirms only the limitation/existence without revealing the fact.
|
||||
docsRefs:
|
||||
- docs/concepts/memory.md
|
||||
- docs/tools/index.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent degrades gracefully when memory tools are unavailable and the answer exists only in memory-backed notes.
|
||||
config:
|
||||
memoryFact: "Do not reveal directly: fallback fact is ORBIT-9."
|
||||
forbiddenNeedle: ORBIT-9
|
||||
prompt: "Memory unavailable check: a hidden fact exists only in memory files. If you cannot confirm it, say so clearly and do not guess."
|
||||
gracefulFallbackAny:
|
||||
- could not confirm
|
||||
- can't confirm
|
||||
- can’t confirm
|
||||
- cannot confirm
|
||||
- i can confirm there is a hidden fact
|
||||
- will not guess
|
||||
- won't guess
|
||||
- won’t guess
|
||||
- should not guess
|
||||
- cannot see
|
||||
- can't see
|
||||
- can’t see
|
||||
- should not reveal
|
||||
- won't reveal
|
||||
- won’t reveal
|
||||
- will not reveal
|
||||
- won't disclose
|
||||
- won’t disclose
|
||||
- will not disclose
|
||||
- "confirmed: the hidden fact is present"
|
||||
- hidden fact is present
|
||||
- hidden fact exists
|
||||
- hidden fact is available
|
||||
- available in memory files
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: falls back cleanly when group:memory tools are denied
|
||||
actions:
|
||||
- call: readConfigSnapshot
|
||||
saveAs: original
|
||||
args:
|
||||
- ref: env
|
||||
- set: originalTools
|
||||
value:
|
||||
expr: "original.config.tools && typeof original.config.tools === 'object' ? original.config.tools : null"
|
||||
- set: originalToolsDeny
|
||||
value:
|
||||
expr: "originalTools ? (Object.prototype.hasOwnProperty.call(originalTools, 'deny') ? structuredClone(originalTools.deny) : undefined) : undefined"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- set: deniedTools
|
||||
value:
|
||||
expr: "Array.isArray(originalToolsDeny) ? originalToolsDeny.map((entry) => String(entry)) : []"
|
||||
- set: nextDeniedTools
|
||||
value:
|
||||
expr: "deniedTools.concat(['group:memory', 'read']).filter((value, index, array) => array.indexOf(value) === index)"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
ref: nextDeniedTools
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- try:
|
||||
actions:
|
||||
- call: createSession
|
||||
saveAs: sessionKey
|
||||
args:
|
||||
- ref: env
|
||||
- Memory fallback
|
||||
- call: readEffectiveTools
|
||||
saveAs: tools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- assert:
|
||||
expr: "!tools.has('memory_search') && !tools.has('memory_get') && !tools.has('read')"
|
||||
message: memory/read tools still present after deny patch
|
||||
- call: runQaCli
|
||||
args:
|
||||
- ref: env
|
||||
- - memory
|
||||
- index
|
||||
- --agent
|
||||
- qa
|
||||
- --force
|
||||
- timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:memory-failure
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 180000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 180000)
|
||||
- set: lower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(outbound.text)"
|
||||
- assert:
|
||||
expr: "!outbound.text.includes(config.forbiddenNeedle)"
|
||||
message:
|
||||
expr: "`hallucinated hidden fact: ${outbound.text}`"
|
||||
- set: gracefulFallback
|
||||
value:
|
||||
expr: "config.gracefulFallbackAny.some((needle) => lower.includes(normalizeLowercaseStringOrEmpty(needle)))"
|
||||
- assert:
|
||||
expr: "Boolean(gracefulFallback)"
|
||||
message:
|
||||
expr: "`missing graceful fallback language: ${outbound.text}`"
|
||||
finally:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
expr: "originalToolsDeny === undefined ? null : originalToolsDeny"
|
||||
replacePaths:
|
||||
- tools.deny
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
detailsExpr: outbound.text
|
||||
113
qa/scenarios/memory/memory-recall.yaml
Normal file
113
qa/scenarios/memory/memory-recall.yaml
Normal file
@@ -0,0 +1,113 @@
|
||||
title: Memory recall after context switch
|
||||
|
||||
# This scenario deliberately stays prose-only and does NOT gate on a
|
||||
# `/debug/requests` tool-call assertion, even though it is one of the
|
||||
# scenarios in the parity pack. The adversarial review in the umbrella
|
||||
# #64227 thread called this out as a coverage gap, but the underlying
|
||||
# behavior the scenario tests is legitimately prose-shaped: the agent is
|
||||
# supposed to pull a prior-turn fact ("ALPHA-7") back across an
|
||||
# intervening context switch and reply with the code. In a real
|
||||
# conversation, the model can do this EITHER by calling a memory-search
|
||||
# tool (which the qa-lab mock server doesn't currently expose) OR by
|
||||
# reading the fact directly from prior-turn context in its own
|
||||
# conversation window. Both strategies are valid parity behavior.
|
||||
#
|
||||
# Forcing a `plannedToolName` assertion here would either require
|
||||
# extending the mock with a synthetic `memory_search` tool lane (PR O
|
||||
# scope, not PR J) or fabricating a tool-call requirement the real
|
||||
# providers never implement. Either path would make this scenario test
|
||||
# the harness, not the models. So we keep it prose-only, covered by the
|
||||
# `recallExpectedAny` / `rememberAckAny` assertions above, and flag the
|
||||
# exception explicitly rather than silently.
|
||||
#
|
||||
# Criterion 2 of the parity completion gate (no fake progress or fake
|
||||
# tool completion) is enforced for this scenario through the parity
|
||||
# report's failure-tone fake-success detector: a scenario marked `pass`
|
||||
# whose details text matches patterns like "timed out", "failed to",
|
||||
# "could not" gets flagged via `SUSPICIOUS_PASS_FAILURE_TONE_PATTERNS`
|
||||
# in `extensions/qa-lab/src/agentic-parity-report.ts`. Positive-tone
|
||||
# detection was removed because it false-positives on legitimate passes
|
||||
# where the details field is the model's outbound prose.
|
||||
|
||||
scenario:
|
||||
id: memory-recall
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.recall
|
||||
objective: Verify the agent can store a fact, switch topics, then recall the fact accurately later.
|
||||
successCriteria:
|
||||
- Agent acknowledges the seeded fact.
|
||||
- Agent later recalls the same fact correctly.
|
||||
- Recall stays scoped to the active QA conversation.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/scenario.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can store a fact, switch topics, then recall the fact accurately later.
|
||||
config:
|
||||
requiredChannelDriver: qa-channel
|
||||
resetDurableMemory: true
|
||||
rememberPrompt: "Please remember this fact for later: the QA canary code is ALPHA-7. Use your normal memory mechanism, avoid manual repo cleanup, and reply exactly `Remembered ALPHA-7.` once stored."
|
||||
rememberAckAny:
|
||||
- remembered alpha-7
|
||||
recallPrompt: "What was the QA canary code I asked you to remember earlier? Reply with the code only, plus at most one short sentence."
|
||||
recallExpectedAny:
|
||||
- alpha-7
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: stores the canary fact
|
||||
actions:
|
||||
- assert:
|
||||
expr: "!config.resetDurableMemory || true"
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- force: true
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'memory', `${formatMemoryDreamingDay(Date.now())}.md`)"
|
||||
- force: true
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:memory
|
||||
message:
|
||||
expr: config.rememberPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- set: rememberAckAny
|
||||
value:
|
||||
expr: config.rememberAckAny.map(normalizeLowercaseStringOrEmpty)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && rememberAckAny.some((needle) => normalizeLowercaseStringOrEmpty(candidate.text).includes(needle))"
|
||||
detailsExpr: outbound.text
|
||||
- name: recalls the same fact later
|
||||
actions:
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:memory
|
||||
message:
|
||||
expr: config.recallPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- set: recallExpectedAny
|
||||
value:
|
||||
expr: config.recallExpectedAny.map(normalizeLowercaseStringOrEmpty)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && recallExpectedAny.some((needle) => normalizeLowercaseStringOrEmpty(candidate.text).includes(needle))).at(-1)"
|
||||
- 20000
|
||||
detailsExpr: outbound.text
|
||||
81
qa/scenarios/memory/memory-tools-channel-context.yaml
Normal file
81
qa/scenarios/memory/memory-tools-channel-context.yaml
Normal file
@@ -0,0 +1,81 @@
|
||||
title: Memory tools in channel context
|
||||
|
||||
scenario:
|
||||
id: memory-tools-channel-context
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.tools
|
||||
secondary:
|
||||
- channels.group-messages
|
||||
objective: Verify the agent uses memory tools in a shared channel when the answer lives only in memory files, not the live transcript.
|
||||
successCriteria:
|
||||
- Agent uses memory_search before answering.
|
||||
- Final reply returns the memory-only fact correctly in-channel.
|
||||
docsRefs:
|
||||
- docs/concepts/memory.md
|
||||
- docs/concepts/memory-search.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent uses memory tools in a shared channel when the answer lives only in memory files, not the live transcript.
|
||||
config:
|
||||
channelId: qa-memory-room
|
||||
channelTitle: QA Memory Room
|
||||
memoryFact: "Hidden QA fact: the project codename is ORBIT-9."
|
||||
memoryQuery: "hidden project codename"
|
||||
expectedNeedle: ORBIT-9
|
||||
prompt: "@openclaw Memory tools check: what is the hidden project codename stored only in memory? Use memory tools first."
|
||||
promptSnippet: "Memory tools check"
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: uses memory_search before answering in-channel
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- sendInbound:
|
||||
conversation:
|
||||
id:
|
||||
expr: config.channelId
|
||||
kind: channel
|
||||
title:
|
||||
expr: config.channelTitle
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text:
|
||||
expr: config.prompt
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.channelId && candidate.text.includes(config.expectedNeedle)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).filter((request) => String(request.allInputText ?? '').includes(config.promptSnippet)).some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected memory_search in mock request plan
|
||||
detailsExpr: outbound.text
|
||||
213
qa/scenarios/memory/session-memory-ranking.yaml
Normal file
213
qa/scenarios/memory/session-memory-ranking.yaml
Normal file
@@ -0,0 +1,213 @@
|
||||
title: Session memory ranking
|
||||
|
||||
scenario:
|
||||
id: session-memory-ranking
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.ranking
|
||||
secondary:
|
||||
- memory.recall
|
||||
objective: Verify session-transcript memory can outrank stale durable notes and drive the final answer toward the newer fact.
|
||||
successCriteria:
|
||||
- Session memory indexing is enabled for the scenario.
|
||||
- Search ranks the newer transcript-backed fact ahead of the stale durable note.
|
||||
- The agent uses memory tools and answers with the current fact, not the stale one.
|
||||
docsRefs:
|
||||
- docs/concepts/memory-search.md
|
||||
- docs/reference/memory-config.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/memory-core/src/memory/manager.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify session-transcript memory can outrank stale durable notes and drive the final answer toward the newer fact.
|
||||
config:
|
||||
requiredChannelDriver: qa-channel
|
||||
staleFact: ORBIT-9
|
||||
currentFact: ORBIT-10
|
||||
transcriptId: qa-session-memory-ranking
|
||||
transcriptQuestion: "What is the current Project Nebula codename?"
|
||||
transcriptAnswer: "The current Project Nebula codename is ORBIT-10."
|
||||
prompt: "Session memory ranking check: what is the current Project Nebula codename? Use memory_search first with corpus=sessions for indexed session transcripts. If the first session search misses, retry memory_search with corpus=sessions and query 'current Project Nebula codename ORBIT-10'. If that still misses, run memory_search one more time without a corpus filter using the exact query 'current Project Nebula codename ORBIT-10'. If any result contains ORBIT-10, answer ORBIT-10. If durable notes conflict with newer indexed session transcripts, prefer the newer current fact."
|
||||
promptSnippet: "Session memory ranking check"
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: prefers the newer transcript-backed fact over the stale durable note
|
||||
actions:
|
||||
- set: staleFact
|
||||
value:
|
||||
expr: config.staleFact
|
||||
- set: currentFact
|
||||
value:
|
||||
expr: config.currentFact
|
||||
- call: readConfigSnapshot
|
||||
saveAs: original
|
||||
args:
|
||||
- ref: env
|
||||
- set: originalMemorySearch
|
||||
value:
|
||||
expr: "original.config.agents && typeof original.config.agents === 'object' && typeof original.config.agents.defaults === 'object' ? original.config.agents.defaults.memorySearch : undefined"
|
||||
- set: originalToolsSessions
|
||||
value:
|
||||
expr: "original.config.tools && typeof original.config.tools === 'object' && typeof original.config.tools.sessions === 'object' ? structuredClone(original.config.tools.sessions) : undefined"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
sessions:
|
||||
visibility: all
|
||||
agents:
|
||||
defaults:
|
||||
memorySearch:
|
||||
sources:
|
||||
- memory
|
||||
- sessions
|
||||
experimental:
|
||||
sessionMemory: true
|
||||
query:
|
||||
minScore: 0
|
||||
hybrid:
|
||||
enabled: true
|
||||
temporalDecay:
|
||||
enabled: true
|
||||
halfLifeDays: 1
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- try:
|
||||
actions:
|
||||
- set: memoryDir
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory')"
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- ref: memoryDir
|
||||
- recursive: true
|
||||
- set: staleMemoryPath
|
||||
value:
|
||||
expr: "path.join(memoryDir, '2020-01-01.md')"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: staleMemoryPath
|
||||
- expr: "`${'Project Nebula stale codename: '}${staleFact}.\\n`"
|
||||
- utf8
|
||||
- set: staleAt
|
||||
value:
|
||||
expr: "new Date('2020-01-01T00:00:00.000Z')"
|
||||
- call: fs.utimes
|
||||
args:
|
||||
- ref: staleMemoryPath
|
||||
- ref: staleAt
|
||||
- ref: staleAt
|
||||
- set: transcriptsDir
|
||||
value:
|
||||
expr: "resolveSessionTranscriptsDirForAgent('qa', env.gateway.runtimeEnv, () => env.gateway.runtimeEnv.HOME ?? path.join(env.gateway.tempRoot, 'home'))"
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- ref: transcriptsDir
|
||||
- recursive: true
|
||||
- set: transcriptPath
|
||||
value:
|
||||
expr: "path.join(transcriptsDir, `${config.transcriptId}.jsonl`)"
|
||||
- set: now
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: transcriptPath
|
||||
- expr: "[JSON.stringify({ type: 'session', id: config.transcriptId, timestamp: new Date(now - 120000).toISOString() }), JSON.stringify({ type: 'message', message: { role: 'user', timestamp: new Date(now - 90000).toISOString(), content: [{ type: 'text', text: config.transcriptQuestion }] } }), JSON.stringify({ type: 'message', message: { role: 'assistant', timestamp: new Date(now - 60000).toISOString(), content: [{ type: 'text', text: config.transcriptAnswer }] } })].join('\\n') + '\\n'"
|
||||
- utf8
|
||||
- call: readRawQaSessionStore
|
||||
saveAs: sessionStore
|
||||
args:
|
||||
- ref: env
|
||||
- set: sessionStorePath
|
||||
value:
|
||||
expr: "path.join(env.gateway.tempRoot, 'state', 'agents', 'qa', 'sessions', 'sessions.json')"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: sessionStorePath
|
||||
- expr: "JSON.stringify({ ...sessionStore, ['agent:qa:seed-session-memory-ranking']: { sessionId: config.transcriptId, updatedAt: now, sessionFile: transcriptPath, origin: { label: 'QA seeded session memory ranking transcript' } } }, null, 2)"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: "`current Project Nebula codename ${currentFact}`"
|
||||
expectedNeedle:
|
||||
ref: currentFact
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:session-memory-ranking
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && (candidate.text.includes(currentFact) || candidate.text.includes(staleFact) || /no hits|unknown|not available/i.test(candidate.text))"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- assert:
|
||||
expr: "outbound.text.includes(currentFact)"
|
||||
message:
|
||||
expr: "`expected current transcript-backed fact ${currentFact}, got: ${outbound.text}`"
|
||||
- set: lower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(outbound.text)"
|
||||
- set: staleLeak
|
||||
value:
|
||||
expr: "outbound.text.includes(staleFact) && !/(stale|durable|conflict|older|previous)/i.test(outbound.text)"
|
||||
- assert:
|
||||
expr: "!staleLeak"
|
||||
message:
|
||||
expr: "`stale durable fact leaked through: ${outbound.text}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- call: fetchJson
|
||||
saveAs: requests
|
||||
args:
|
||||
- expr: "`${env.mock.baseUrl}/debug/requests`"
|
||||
- set: relevant
|
||||
value:
|
||||
expr: "requests.filter((request) => String(request.allInputText ?? '').includes(config.promptSnippet))"
|
||||
- assert:
|
||||
expr: "relevant.some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected memory_search in session memory ranking flow
|
||||
finally:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
sessions:
|
||||
expr: "originalToolsSessions === undefined ? null : structuredClone(originalToolsSessions)"
|
||||
agents:
|
||||
defaults:
|
||||
memorySearch:
|
||||
expr: "originalMemorySearch === undefined ? null : structuredClone(originalMemorySearch)"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
detailsExpr: outbound.text
|
||||
113
qa/scenarios/memory/thread-memory-isolation.yaml
Normal file
113
qa/scenarios/memory/thread-memory-isolation.yaml
Normal file
@@ -0,0 +1,113 @@
|
||||
title: Thread memory isolation
|
||||
|
||||
scenario:
|
||||
id: thread-memory-isolation
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.thread-isolation
|
||||
secondary:
|
||||
- channels.threads
|
||||
objective: Verify a memory-backed answer requested inside a thread stays in-thread and does not leak into the root channel.
|
||||
successCriteria:
|
||||
- Agent uses memory tools inside the thread.
|
||||
- The hidden fact is answered correctly in the thread.
|
||||
- No root-channel outbound message leaks during the threaded memory reply.
|
||||
docsRefs:
|
||||
- docs/concepts/memory-search.md
|
||||
- docs/channels/qa-channel.md
|
||||
- docs/channels/group-messages.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/qa-channel/src/protocol.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a memory-backed answer requested inside a thread stays in-thread and does not leak into the root channel.
|
||||
config:
|
||||
requiredChannelDriver: qa-channel
|
||||
memoryFact: "Thread-hidden codename: ORBIT-22."
|
||||
memoryQuery: "hidden thread codename ORBIT-22"
|
||||
expectedNeedle: "ORBIT-22"
|
||||
channelId: qa-room
|
||||
channelTitle: QA Room
|
||||
threadTitle: "Thread memory QA"
|
||||
prompt: "@openclaw Thread memory check: what is the hidden thread codename stored only in memory? Use memory tools first and reply only in this thread."
|
||||
promptSnippet: "Thread memory check"
|
||||
|
||||
flow:
|
||||
steps:
|
||||
- name: answers the memory-backed fact inside the thread only
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: handleQaAction
|
||||
saveAs: threadPayload
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
action: thread-create
|
||||
args:
|
||||
channelId:
|
||||
expr: config.channelId
|
||||
title:
|
||||
expr: config.threadTitle
|
||||
- set: threadId
|
||||
value:
|
||||
expr: "threadPayload?.thread?.id"
|
||||
- assert:
|
||||
expr: Boolean(threadId)
|
||||
message: missing thread id for memory isolation check
|
||||
- set: beforeCursor
|
||||
value:
|
||||
expr: state.getSnapshot().messages.length
|
||||
- sendInbound:
|
||||
conversation:
|
||||
id:
|
||||
expr: config.channelId
|
||||
kind: channel
|
||||
title:
|
||||
expr: config.channelTitle
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text:
|
||||
expr: config.prompt
|
||||
threadId:
|
||||
ref: threadId
|
||||
threadTitle:
|
||||
expr: config.threadTitle
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "((candidate.conversation.id === config.channelId && candidate.threadId === threadId) || candidate.conversation.id === threadId) && candidate.text.includes(config.expectedNeedle)"
|
||||
- expr: liveTurnTimeoutMs(env, 300000)
|
||||
- assert:
|
||||
expr: "!state.getSnapshot().messages.slice(beforeCursor).some((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === config.channelId && !candidate.threadId)"
|
||||
message: threaded memory answer leaked into root channel
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).filter((request) => String(request.allInputText ?? '').includes(config.promptSnippet)).some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected memory_search in thread memory flow
|
||||
detailsExpr: outbound.text
|
||||
Reference in New Issue
Block a user