Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11), free to diverge. Tree copied sans upstream .git; upstream remote added for future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19. Preserves docs/ARCHITECTURE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
133 lines
3.9 KiB
Markdown
133 lines
3.9 KiB
Markdown
# Frontier Harness Test Plan
|
|
|
|
Use this when tuning the harness on frontier models before the small-model pass.
|
|
|
|
## Goals
|
|
|
|
- verify tool-first behavior on short approval turns
|
|
- verify model switching does not kill tool use
|
|
- verify repo-reading / discovery still finishes with a concrete report
|
|
- verify mutating work keeps replay-unsafety explicit under compaction pressure
|
|
- collect manual notes on personality without letting style hide execution regressions
|
|
|
|
## Frontier subset
|
|
|
|
Run this subset first on every harness tweak:
|
|
|
|
- `approval-turn-tool-followthrough`
|
|
- `model-switch-tool-continuity`
|
|
- `source-docs-discovery-report`
|
|
|
|
Longer spot-check after that:
|
|
|
|
- `compaction-retry-mutating-tool`
|
|
- `subagent-handoff`
|
|
|
|
## Baseline order
|
|
|
|
1. GPT first. Use this as the main tuning reference.
|
|
2. Claude second. If Claude regresses alone, prefer an Anthropic overlay fix over a core prompt rewrite.
|
|
3. Gemini third. Treat this as the operational-directness check.
|
|
4. Only run the whole seed suite after the frontier subset is stable.
|
|
|
|
## Commands
|
|
|
|
GPT baseline:
|
|
|
|
```bash
|
|
pnpm openclaw qa suite \
|
|
--provider-mode live-frontier \
|
|
--model openai/gpt-5.5 \
|
|
--alt-model openai/gpt-5.5 \
|
|
--fast \
|
|
--scenario approval-turn-tool-followthrough \
|
|
--scenario model-switch-tool-continuity \
|
|
--scenario source-docs-discovery-report
|
|
```
|
|
|
|
Claude sweep:
|
|
|
|
```bash
|
|
pnpm openclaw qa suite \
|
|
--provider-mode live-frontier \
|
|
--model anthropic/claude-sonnet-4-6 \
|
|
--alt-model anthropic/claude-opus-4-6 \
|
|
--scenario approval-turn-tool-followthrough \
|
|
--scenario model-switch-tool-continuity \
|
|
--scenario source-docs-discovery-report
|
|
```
|
|
|
|
Gemini sweep:
|
|
|
|
```bash
|
|
pnpm openclaw qa suite \
|
|
--provider-mode live-frontier \
|
|
--model <google-pro-model-ref> \
|
|
--alt-model <google-pro-model-ref> \
|
|
--scenario approval-turn-tool-followthrough \
|
|
--scenario model-switch-tool-continuity \
|
|
--scenario source-docs-discovery-report
|
|
```
|
|
|
|
Use the QA Lab runner catalog or `openclaw models list --all` to pick the current Google Pro ref.
|
|
|
|
## Tuning loop
|
|
|
|
1. Run the GPT subset and save the report path.
|
|
2. Patch one harness idea at a time.
|
|
3. Rerun the same GPT subset immediately.
|
|
4. If GPT improves, run the Claude subset.
|
|
5. If Claude is clean, run the Gemini subset.
|
|
6. If only one family regresses, fix the provider overlay before touching the shared prompt again.
|
|
|
|
## What to score
|
|
|
|
- tool commitment after `ok do it`
|
|
- empty-promise rate
|
|
- tool continuity after model switch
|
|
- discovery report completeness and specificity
|
|
- replay-safety truth after a mutating write
|
|
- scope drift: unrelated scenario updates, grand wrap-ups, or invented completion tallies
|
|
- latency / obvious stall behavior
|
|
- token cost notes if a change makes the prompt materially heavier
|
|
|
|
## Manual personality lane
|
|
|
|
Run this after the executable subset, not before:
|
|
|
|
```text
|
|
read QA_KICKOFF_TASK.md, tell me what feels half-baked about this qa mission, and keep it to two short sentences
|
|
```
|
|
|
|
GPT manual lane:
|
|
|
|
```bash
|
|
pnpm openclaw qa manual \
|
|
--provider-mode live-frontier \
|
|
--model openai/gpt-5.5 \
|
|
--alt-model openai/gpt-5.5 \
|
|
--fast \
|
|
--message "read QA_KICKOFF_TASK.md, tell me what feels half-baked about this qa mission, and keep it to two short sentences"
|
|
```
|
|
|
|
Claude manual lane:
|
|
|
|
```bash
|
|
pnpm openclaw qa manual \
|
|
--provider-mode live-frontier \
|
|
--model anthropic/claude-sonnet-4-6 \
|
|
--alt-model anthropic/claude-opus-4-6 \
|
|
--message "read QA_KICKOFF_TASK.md, tell me what feels half-baked about this qa mission, and keep it to two short sentences"
|
|
```
|
|
|
|
Score it on:
|
|
|
|
- did it read first
|
|
- did it say something specific instead of generic fluff
|
|
- did the agent still sound like itself while doing useful work
|
|
- did it stay on the scoped ask instead of widening into a suite recap or fake completion claim
|
|
|
|
## Deferred
|
|
|
|
- deterministic mock compaction triggering is still deferred; the current replay-safety lane is a live-frontier-first executable scenario
|