Vendor OpenClaw source as Adolf fork baseline
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11), free to diverge. Tree copied sans upstream .git; upstream remote added for future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19. Preserves docs/ARCHITECTURE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
This commit is contained in:
252
.agents/skills/autoreview/SKILL.md
Normal file
252
.agents/skills/autoreview/SKILL.md
Normal file
@@ -0,0 +1,252 @@
|
|||||||
|
---
|
||||||
|
name: autoreview
|
||||||
|
description: "Auto Review closeout. Codex review is the default when no engine is set and is the recommended reviewer."
|
||||||
|
---
|
||||||
|
|
||||||
|
# Auto Review
|
||||||
|
|
||||||
|
Run the bundled structured review helper as a closeout check. This is code review, not Guardian `auto_review` approval routing.
|
||||||
|
|
||||||
|
Codex review is the default when no engine is set. It usually delivers the best review results and should remain the normal final closeout engine.
|
||||||
|
|
||||||
|
Use when:
|
||||||
|
|
||||||
|
- user asks for Codex review / Claude review / autoreview / second-model review
|
||||||
|
- after non-trivial code edits, before final/commit/ship
|
||||||
|
- reviewing a local branch or PR branch after fixes
|
||||||
|
|
||||||
|
## Contract
|
||||||
|
|
||||||
|
- Treat review output as advisory. Never blindly apply it.
|
||||||
|
- Verify every finding by reading the real code path and adjacent files.
|
||||||
|
- Read dependency docs/source/types when the finding depends on external behavior.
|
||||||
|
- Reject unrealistic edge cases, speculative risks, broad rewrites, and fixes that over-complicate the codebase.
|
||||||
|
- Prefer small fixes at the right ownership boundary; no refactor unless it clearly improves the bug class.
|
||||||
|
- When an accepted finding shows a bug class or repeated pattern, inspect the current PR scope for sibling instances before fixing.
|
||||||
|
- Fix the scoped bug class at once when practical; stop at touched surfaces, owner boundaries, and clear follow-up territory.
|
||||||
|
- Keep going until structured review returns no accepted/actionable findings only while the work remains inside the original task scope.
|
||||||
|
- If a review-triggered fix changes code, rerun focused tests and rerun the structured review helper.
|
||||||
|
- For security-audit suppression changes, verify accepted findings remain auditable: suppressed findings stay in structured output, active output keeps an unsuppressible suppression notice, and aggregate findings cannot hide unrelated active risk.
|
||||||
|
- Never switch or override the requested review engine/model. If the review hits model capacity, retry the same command a few times with the same engine/model.
|
||||||
|
- Be patient with large bundles. Structured review can take up to 30 minutes while the model call is active, especially with Codex tools or web search.
|
||||||
|
- Treat heartbeat lines like `review still running: ... elapsed=... pid=...` as healthy progress, not a hang. Let the helper continue while heartbeats are advancing. Pass `--stream-engine-output` when live engine text is useful; Codex, Claude, and cursor-agent filter tool/file chatter, other engines pass raw output through.
|
||||||
|
- Do not kill a review just because it has been quiet for 2-5 minutes, or because it is still running under the 30-minute window. Inspect the process only after missing multiple expected heartbeats, after 30 minutes, or after an obviously failed subprocess; prefer letting the same helper command finish.
|
||||||
|
- Tools are useful in review mode. The helper allows read-only inspection tools and web search by default so reviewers can check dependency contracts, upstream docs, and current behavior.
|
||||||
|
- Security perspective is always included, but it should not cripple legitimate functionality. Report security findings only when the change creates a concrete, actionable risk or removes an important safety check.
|
||||||
|
- For regression provenance, if no blamed PR is traceable, use the blamed commit as the provenance: commit SHA, date, and author username. Do not guess a merger or frame missing PR metadata as a separate finding.
|
||||||
|
- Do not invoke built-in `codex review`, nested reviewers, or reviewer panels from inside the review. The helper builds one bundle, calls one selected engine, validates one structured result, and stops.
|
||||||
|
- Stop as soon as the helper exits 0 with no accepted/actionable findings. Do not run an extra review just to get a nicer "clean" line, a second opinion, or clearer closeout wording.
|
||||||
|
- Treat the helper's successful exit plus absence of actionable findings as the clean review result, even if the underlying Codex CLI output is terse.
|
||||||
|
- Multi-reviewer panels are opt-in only. Use them when explicitly requested or when risk justifies the extra spend; the main agent still verifies every accepted finding before fixing.
|
||||||
|
- If rejecting a finding as intentional/not worth fixing, add a brief inline code comment only when it explains a real invariant or ownership decision that future reviewers should know.
|
||||||
|
- If `gh`/Gitcrawl reports `database disk image is malformed`, run `gitcrawl doctor --json` once to let the portable cache repair before retrying review; do not bypass the shim unless repair fails and freshness requires live GitHub.
|
||||||
|
- If Gitcrawl reports a portable manifest mismatch, source/runtime DB health error, or stale portable-store checkout, run `gitcrawl doctor --json` and inspect `source_db_health`, `runtime_db_health`, and `portable_store_status` before falling back to live GitHub.
|
||||||
|
- Do not push just to review. Push only when the user requested push/ship/PR update.
|
||||||
|
|
||||||
|
## Scope Governor
|
||||||
|
|
||||||
|
Autoreview is a closeout gate, not permission to rewrite the task.
|
||||||
|
|
||||||
|
Before the first review, freeze a scope baseline: original request or issue, target branch, intended behavior, owner boundary, changed files, and non-test LOC. For inherited or already-bloated branches, use the intended PR diff as the baseline rather than accepting all existing branch drift.
|
||||||
|
|
||||||
|
Before patching a finding, classify it:
|
||||||
|
|
||||||
|
- **In-scope blocker**: the finding is introduced by the current diff, affects the same owner boundary, and can be fixed without changing the task's contract.
|
||||||
|
- **Follow-up**: the finding is real but belongs to an adjacent bug class, sibling surface, cleanup, or broader hardening track.
|
||||||
|
- **Stop-and-escalate**: the finding requires a new protocol/config/storage/public API contract, a different owner boundary, a release-process change, or a design choice outside the original request.
|
||||||
|
|
||||||
|
Stop patching and report the scope break instead of continuing when:
|
||||||
|
|
||||||
|
- a narrow PR turns into an architecture change, protocol change, migration, or release-process change;
|
||||||
|
- the diff grows past 2x the original files or non-test LOC without explicit approval to expand scope;
|
||||||
|
- two review-triggered patch cycles have not converged; pause and reclassify every remaining finding before another edit;
|
||||||
|
- the best fix is "define the canonical contract first" rather than another local inference layer;
|
||||||
|
- fixing the accepted finding would make the PR no longer describe the same behavior, issue, or owner boundary.
|
||||||
|
|
||||||
|
After the two-cycle pause, continue only when every remaining accepted finding is still an in-scope blocker. Otherwise preserve the useful analysis, identify the smallest safe landed subset if one exists, and open or request a follow-up for the larger fix. Do not keep committing speculative fixes just to satisfy the reviewer.
|
||||||
|
|
||||||
|
Do not stack or push review-triggered fix commits while scope classification or focused proof is unresolved. Keep exploratory edits local until the cycle is proven in scope; if scope breaks, remove them from the landing lane instead of preserving them as branch history.
|
||||||
|
|
||||||
|
Critical exceptions must be explicit: active data loss, crash, broken install/upgrade, release blocker, or concrete security exposure. If the exception is not one of those, it is not critical enough to blow up scope.
|
||||||
|
|
||||||
|
## Release Branches And Release Process
|
||||||
|
|
||||||
|
On release, beta, stable, hotfix, signing, notarization, appcast, package-publish, or release-check work, use freeze discipline even when the branch name is not release-like:
|
||||||
|
|
||||||
|
- Fix only release blockers, failed release infrastructure, exact backports, install/upgrade breakage, data loss, crashes, or concrete security exposure.
|
||||||
|
- Treat non-blocking autoreview findings as follow-ups for `main`, not reasons to broaden the release branch.
|
||||||
|
- Do not introduce new product behavior, config surface, protocol shape, migration, plugin ownership, docs narrative, or process policy unless it directly unblocks the release.
|
||||||
|
- Keep proof tied to the release target: exact branch/ref, failing check or shipped-risk reason, smallest command/proof, and whether the fix must also forward-port to `main`.
|
||||||
|
- If review discovers a real but non-critical design problem during release closeout, stop with a follow-up issue/PR plan; do not use the release branch as the refactor lane.
|
||||||
|
|
||||||
|
## Pick Target
|
||||||
|
|
||||||
|
Dirty local work:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --mode local
|
||||||
|
```
|
||||||
|
|
||||||
|
Use this only when the patch is actually unstaged/staged/untracked in the
|
||||||
|
current checkout. `--mode uncommitted` is accepted as an alias for `--mode local`.
|
||||||
|
For committed, pushed, or PR work, point the helper at the commit
|
||||||
|
or branch diff instead; do not force dirty modes just
|
||||||
|
because the helper docs mention dirty work first. A clean local review
|
||||||
|
only proves there is no local patch.
|
||||||
|
|
||||||
|
Branch/PR work:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --mode branch --base origin/main
|
||||||
|
```
|
||||||
|
|
||||||
|
Optional review context is first-class:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --mode branch --base origin/main --prompt-file /tmp/review-notes.md --dataset /tmp/evidence.json
|
||||||
|
```
|
||||||
|
|
||||||
|
If an open PR exists, use its actual base:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
base=$(gh pr view --json baseRefName --jq .baseRefName)
|
||||||
|
<autoreview-helper> --mode branch --base "origin/$base"
|
||||||
|
```
|
||||||
|
|
||||||
|
Committed single change:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --mode commit --commit HEAD
|
||||||
|
```
|
||||||
|
|
||||||
|
or with the helper:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/Users/steipete/Projects/agent-scripts/skills/autoreview/scripts/autoreview --mode commit --commit HEAD
|
||||||
|
```
|
||||||
|
|
||||||
|
Use commit review for already-landed or already-pushed work on `main`. Reviewing
|
||||||
|
clean `main` against `origin/main` is usually an empty diff after push. For a
|
||||||
|
small stack, review each commit explicitly or review the branch before merging
|
||||||
|
with `--base`.
|
||||||
|
|
||||||
|
## Parallel Closeout
|
||||||
|
|
||||||
|
Format first if formatting can change line locations. Then it is OK to run tests and review in parallel:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
scripts/autoreview --parallel-tests "<focused test command>"
|
||||||
|
```
|
||||||
|
|
||||||
|
On Windows, the default `--parallel-tests` shell preserves the platform `cmd.exe`
|
||||||
|
semantics used by Python `shell=True`. Use `--parallel-tests-shell powershell`
|
||||||
|
or `--parallel-tests-shell pwsh` when the focused test command is PowerShell-specific.
|
||||||
|
|
||||||
|
Tradeoff: tests may force code changes that stale the review. If tests or review lead to code edits, rerun the affected tests and rerun review until no accepted/actionable findings remain. Once that rerun exits cleanly, stop; do not spend another long review cycle on redundant confirmation.
|
||||||
|
|
||||||
|
## Review Panels
|
||||||
|
|
||||||
|
Run multiple reviewers against one frozen bundle:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --reviewers codex,claude
|
||||||
|
```
|
||||||
|
|
||||||
|
`--panel` is shorthand for Codex plus Claude unless `--engine` changes the first reviewer:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --panel
|
||||||
|
```
|
||||||
|
|
||||||
|
Set reviewer models and thinking/effort explicitly:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --reviewers codex,claude --model codex=gpt-5.1 --thinking codex=high --model claude=sonnet --thinking claude=max
|
||||||
|
```
|
||||||
|
|
||||||
|
Inline syntax is also supported:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
<autoreview-helper> --reviewers codex:gpt-5.1:high,claude:sonnet:max
|
||||||
|
```
|
||||||
|
|
||||||
|
Codex maps thinking to `model_reasoning_effort` and accepts `low`, `medium`,
|
||||||
|
`high`, or `xhigh`. Claude maps thinking to `--effort` and also accepts `max`.
|
||||||
|
Engines without a real thinking knob reject `--thinking`.
|
||||||
|
|
||||||
|
## Context Efficiency
|
||||||
|
|
||||||
|
Run the helper directly so target selection, engine choice, structured validation, and exit status all stay in one path. If output is noisy, summarize the completed helper output after it returns; do not ask another agent or reviewer to rerun the review.
|
||||||
|
|
||||||
|
## Helper
|
||||||
|
|
||||||
|
OpenClaw repo-local helper:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
.agents/skills/autoreview/scripts/autoreview --help
|
||||||
|
```
|
||||||
|
|
||||||
|
On native Windows, invoke the extensionless Python helper through Python:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
python .agents\skills\autoreview\scripts\autoreview --help
|
||||||
|
```
|
||||||
|
|
||||||
|
The smoke harness has thin shell wrappers over a shared Python implementation:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
.agents/skills/autoreview/scripts/test-review-harness --fixture benign --engine codex
|
||||||
|
```
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
.agents\skills\autoreview\scripts\test-review-harness.ps1 -Fixture benign -Engine codex
|
||||||
|
```
|
||||||
|
|
||||||
|
`agent-scripts` checkout helper:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
skills/autoreview/scripts/autoreview --help
|
||||||
|
```
|
||||||
|
|
||||||
|
Global helper from `agent-scripts`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
~/.codex/skills/agent-scripts/autoreview/scripts/autoreview --help
|
||||||
|
```
|
||||||
|
|
||||||
|
If installed from `agent-scripts`, path is:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/Users/steipete/Projects/agent-scripts/skills/autoreview/scripts/autoreview --help
|
||||||
|
```
|
||||||
|
|
||||||
|
The helper:
|
||||||
|
|
||||||
|
- chooses dirty local changes first
|
||||||
|
- accepts `--mode uncommitted` as an alias for `--mode local`
|
||||||
|
- otherwise uses current PR base if `gh pr view` works
|
||||||
|
- otherwise uses `origin/main` for non-main branches
|
||||||
|
- supports `--engine codex`, `claude`, `droid`, `copilot`, and `cursor-agent`; default is `AUTOREVIEW_ENGINE` or `codex`; Codex should remain the default when nothing is set
|
||||||
|
- resolves bare `git`, `gh`, reviewer, and PowerShell shell commands from absolute `PATH` entries only, never from the reviewed checkout; explicit relative `--*-bin` paths are resolved from the reviewed repository root
|
||||||
|
- use `--mode commit --commit <ref>` for already-committed work, especially clean `main` after landing
|
||||||
|
- should be left in `--mode auto` or forced to `--mode branch` for PR/branch work; do not force `--mode local` after committing
|
||||||
|
- writes only to stdout unless `--output`, `--json-output`, or live streamed engine stderr is set
|
||||||
|
- supports `--dry-run`, `--parallel-tests`, `--parallel-tests-shell`, `--prompt`, `--prompt-file`, `--dataset`, `--no-tools`, `--no-web-search`, and commit refs
|
||||||
|
- supports `--stream-engine-output` or `AUTOREVIEW_STREAM_ENGINE_OUTPUT=1` for live engine text while preserving structured validation; Codex, Claude, and cursor-agent hide tool/file event details, emit compact activity summaries, and report usage at turn completion
|
||||||
|
- supports opt-in review panels with `--panel` / `--reviewers`, plus per-engine `--model` and `--thinking`
|
||||||
|
- allows read-only tools and web search by default where the selected CLI supports them; forbids nested review in the prompt; Codex is run through `codex exec` with read-only sandbox and structured output; cursor-agent is run through headless `--print` in ask mode with sandboxing enabled from a helper-owned temporary workspace
|
||||||
|
- rejects `--no-web-search` for cursor-agent because the Cursor CLI does not expose a CLI-level web-search disable switch
|
||||||
|
- prints `review still running: <engine> elapsed=<seconds>s pid=<pid>` to stderr at long-running intervals while waiting for the selected review engine, unless streamed output or compact Codex activity has been visible recently
|
||||||
|
- prints `autoreview clean: no accepted/actionable findings reported` when the selected review command exits 0
|
||||||
|
- exits nonzero when accepted/actionable findings are present
|
||||||
|
|
||||||
|
## Final Report
|
||||||
|
|
||||||
|
Include:
|
||||||
|
|
||||||
|
- review command used
|
||||||
|
- tests/proof run
|
||||||
|
- findings accepted/rejected, briefly why
|
||||||
|
- the clean review result from the final helper/review run, or why a remaining finding was consciously rejected
|
||||||
|
|
||||||
|
Do not run another review solely to improve the final report wording. If the final helper run exited 0 and produced no accepted/actionable findings, report that exact run as clean.
|
||||||
1399
.agents/skills/autoreview/scripts/autoreview
Executable file
1399
.agents/skills/autoreview/scripts/autoreview
Executable file
File diff suppressed because it is too large
Load Diff
16
.agents/skills/autoreview/scripts/test-review-harness
Executable file
16
.agents/skills/autoreview/scripts/test-review-harness
Executable file
@@ -0,0 +1,16 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
|
||||||
|
harness="$script_dir/test-review-harness.py"
|
||||||
|
|
||||||
|
if command -v python3 >/dev/null 2>&1; then
|
||||||
|
exec python3 "$harness" "$@"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if command -v python >/dev/null 2>&1; then
|
||||||
|
exec python "$harness" "$@"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "Python 3 is required to run test-review-harness." >&2
|
||||||
|
exit 127
|
||||||
45
.agents/skills/autoreview/scripts/test-review-harness.ps1
Normal file
45
.agents/skills/autoreview/scripts/test-review-harness.ps1
Normal file
@@ -0,0 +1,45 @@
|
|||||||
|
[CmdletBinding()]
|
||||||
|
param(
|
||||||
|
[ValidateSet('malicious', 'benign')]
|
||||||
|
[string] $Fixture,
|
||||||
|
|
||||||
|
[ValidateSet('codex', 'claude', 'droid', 'copilot', 'cursor-agent')]
|
||||||
|
[string[]] $Engine,
|
||||||
|
|
||||||
|
[Alias('h')]
|
||||||
|
[switch] $Help
|
||||||
|
)
|
||||||
|
|
||||||
|
$ErrorActionPreference = 'Stop'
|
||||||
|
|
||||||
|
$Harness = Join-Path $PSScriptRoot 'test-review-harness.py'
|
||||||
|
$ForwardedArgs = @()
|
||||||
|
|
||||||
|
if ($Help) {
|
||||||
|
$ForwardedArgs += '--help'
|
||||||
|
}
|
||||||
|
|
||||||
|
if ($PSBoundParameters.ContainsKey('Fixture')) {
|
||||||
|
$ForwardedArgs += @('--fixture', $Fixture)
|
||||||
|
}
|
||||||
|
|
||||||
|
if ($PSBoundParameters.ContainsKey('Engine')) {
|
||||||
|
foreach ($SelectedEngine in $Engine) {
|
||||||
|
$ForwardedArgs += @('--engine', $SelectedEngine)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
$PyLauncher = Get-Command py -ErrorAction SilentlyContinue
|
||||||
|
if ($null -ne $PyLauncher) {
|
||||||
|
& $PyLauncher.Source -3 $Harness @ForwardedArgs
|
||||||
|
exit $LASTEXITCODE
|
||||||
|
}
|
||||||
|
|
||||||
|
$Python = Get-Command python -ErrorAction SilentlyContinue
|
||||||
|
if ($null -ne $Python) {
|
||||||
|
& $Python.Source $Harness @ForwardedArgs
|
||||||
|
exit $LASTEXITCODE
|
||||||
|
}
|
||||||
|
|
||||||
|
Write-Error 'Python 3 is required to run test-review-harness.'
|
||||||
|
exit 127
|
||||||
215
.agents/skills/autoreview/scripts/test-review-harness.py
Normal file
215
.agents/skills/autoreview/scripts/test-review-harness.py
Normal file
@@ -0,0 +1,215 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import os
|
||||||
|
import runpy
|
||||||
|
import shutil
|
||||||
|
import stat
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
from collections.abc import Callable
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
ENGINES = ("codex", "claude", "droid", "copilot", "cursor-agent")
|
||||||
|
DEFAULT_ENGINES = ("codex", "claude")
|
||||||
|
|
||||||
|
MALICIOUS_INITIAL = """export function uploadPath(name) {
|
||||||
|
return `uploads/${name.replaceAll("/", "")}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function publicUser(user) {
|
||||||
|
return { id: user.id, name: user.name };
|
||||||
|
}
|
||||||
|
"""
|
||||||
|
|
||||||
|
BENIGN_INITIAL = r"""import path from "node:path";
|
||||||
|
|
||||||
|
const uploadsRoot = path.resolve("uploads");
|
||||||
|
|
||||||
|
function safeChildPath(root, name) {
|
||||||
|
if (name.includes("/") || name.includes("\\") || name.includes("\0")) {
|
||||||
|
throw new Error("invalid child name");
|
||||||
|
}
|
||||||
|
const resolved = path.resolve(root, name);
|
||||||
|
if (!resolved.startsWith(root + path.sep)) throw new Error("path escaped root");
|
||||||
|
return resolved;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function uploadPath(name) {
|
||||||
|
return safeChildPath(uploadsRoot, name);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function publicUser(user) {
|
||||||
|
return { id: user.id, name: user.name };
|
||||||
|
}
|
||||||
|
"""
|
||||||
|
|
||||||
|
MALICIOUS_CHANGED = """import { execSync } from "node:child_process";
|
||||||
|
|
||||||
|
export function uploadPath(name) {
|
||||||
|
return `uploads/${name}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function deleteUpload(name) {
|
||||||
|
return execSync(`rm -rf uploads/${name}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function publicUser(user) {
|
||||||
|
return { id: user.id, name: user.name, password: user.password };
|
||||||
|
}
|
||||||
|
"""
|
||||||
|
|
||||||
|
BENIGN_CHANGED = r"""import { execFile } from "node:child_process";
|
||||||
|
import path from "node:path";
|
||||||
|
import { promisify } from "node:util";
|
||||||
|
|
||||||
|
const uploadsRoot = path.resolve("uploads");
|
||||||
|
const reposRoot = path.resolve("repos");
|
||||||
|
const execFileAsync = promisify(execFile);
|
||||||
|
|
||||||
|
function safeChildPath(root, name) {
|
||||||
|
if (name.includes("/") || name.includes("\\") || name.includes("\0")) {
|
||||||
|
throw new Error("invalid child name");
|
||||||
|
}
|
||||||
|
const resolved = path.resolve(root, name);
|
||||||
|
if (!resolved.startsWith(root + path.sep)) throw new Error("path escaped root");
|
||||||
|
return resolved;
|
||||||
|
}
|
||||||
|
|
||||||
|
function repoChildPath(root, name) {
|
||||||
|
if (!/^[A-Za-z0-9._-]+$/.test(name)) throw new Error("invalid repo name");
|
||||||
|
return safeChildPath(root, name);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function uploadPath(name) {
|
||||||
|
return safeChildPath(uploadsRoot, name);
|
||||||
|
}
|
||||||
|
|
||||||
|
export async function repoProbe(repoName) {
|
||||||
|
const { stdout } = await execFileAsync(process.execPath, ["--version"], {
|
||||||
|
cwd: repoChildPath(reposRoot, repoName),
|
||||||
|
encoding: "utf8",
|
||||||
|
maxBuffer: 16 * 1024 * 1024,
|
||||||
|
});
|
||||||
|
return stdout;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function publicUser(user) {
|
||||||
|
return { id: user.id, name: user.name };
|
||||||
|
}
|
||||||
|
|
||||||
|
export function accountSettingsForOwner(user, requesterId) {
|
||||||
|
if (String(requesterId) !== String(user.id)) throw new Error("forbidden");
|
||||||
|
return { id: user.id, hasPassword: Boolean(user.passwordHash) };
|
||||||
|
}
|
||||||
|
"""
|
||||||
|
|
||||||
|
MALICIOUS_PROMPT = "This is an acceptance test fixture. The changed app.js patch contains real security bugs. Review normally and report only actionable defects from the patch."
|
||||||
|
BENIGN_PROMPT = "Security calibration fixture: this patch intentionally uses filesystem paths, async execFile, and owner-gated password-adjacent state safely. Do not flag legitimate shell/filesystem/auth-adjacent functionality unless there is a concrete exploitable risk in the diff."
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args(argv: list[str]) -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
prog="test-review-harness",
|
||||||
|
description=(
|
||||||
|
"Creates a temporary git repo with either a deliberately unsafe patch "
|
||||||
|
"or a security-sensitive-but-safe patch, then verifies each selected "
|
||||||
|
"engine through autoreview."
|
||||||
|
),
|
||||||
|
epilog="Default engines: codex, claude.",
|
||||||
|
)
|
||||||
|
parser.add_argument("--fixture", choices=("malicious", "benign"), default="malicious")
|
||||||
|
parser.add_argument("--engine", action="append", choices=ENGINES, dest="engines")
|
||||||
|
return parser.parse_args(argv)
|
||||||
|
|
||||||
|
|
||||||
|
def write_fixture_file(repo: Path, content: str) -> None:
|
||||||
|
with (repo / "app.js").open("w", encoding="utf-8", newline="\n") as handle:
|
||||||
|
handle.write(content)
|
||||||
|
|
||||||
|
|
||||||
|
def run(command: list[str], cwd: Path) -> None:
|
||||||
|
subprocess.run(command, cwd=cwd, check=True)
|
||||||
|
|
||||||
|
|
||||||
|
def create_fixture_repo(repo: Path, fixture: str) -> None:
|
||||||
|
run(["git", "init", "--quiet"], repo)
|
||||||
|
run(["git", "config", "user.name", "Review Fixture"], repo)
|
||||||
|
run(["git", "config", "user.email", "review-fixture@example.com"], repo)
|
||||||
|
|
||||||
|
write_fixture_file(repo, MALICIOUS_INITIAL if fixture == "malicious" else BENIGN_INITIAL)
|
||||||
|
run(["git", "add", "app.js"], repo)
|
||||||
|
run(["git", "commit", "--quiet", "-m", "initial safe version"], repo)
|
||||||
|
write_fixture_file(repo, MALICIOUS_CHANGED if fixture == "malicious" else BENIGN_CHANGED)
|
||||||
|
|
||||||
|
|
||||||
|
def validate_prompt_policy(repo: Path, autoreview: Path) -> None:
|
||||||
|
namespace = runpy.run_path(str(autoreview))
|
||||||
|
prompt = namespace["build_prompt"](repo, "local", None, "fixture diff", "", "")
|
||||||
|
required = (
|
||||||
|
"This helper is a closeout gate.",
|
||||||
|
"Do not turn a narrow patch into a broad",
|
||||||
|
"If this is release-branch or release-process work",
|
||||||
|
"Non-blocking design,",
|
||||||
|
)
|
||||||
|
missing = [needle for needle in required if needle not in prompt]
|
||||||
|
if missing:
|
||||||
|
raise RuntimeError(f"autoreview prompt missing scope policy: {missing}")
|
||||||
|
|
||||||
|
|
||||||
|
def run_reviews(repo: Path, script_dir: Path, fixture: str, engines: list[str]) -> None:
|
||||||
|
autoreview = script_dir / "autoreview"
|
||||||
|
validate_prompt_policy(repo, autoreview)
|
||||||
|
for engine in engines:
|
||||||
|
print(f"== {engine} ==", flush=True)
|
||||||
|
command = [
|
||||||
|
sys.executable,
|
||||||
|
str(autoreview),
|
||||||
|
"--mode",
|
||||||
|
"local",
|
||||||
|
"--engine",
|
||||||
|
engine,
|
||||||
|
"--prompt",
|
||||||
|
MALICIOUS_PROMPT if fixture == "malicious" else BENIGN_PROMPT,
|
||||||
|
]
|
||||||
|
if fixture == "malicious":
|
||||||
|
command.extend(["--require-finding", "command", "--expect-findings"])
|
||||||
|
run(command, repo)
|
||||||
|
|
||||||
|
|
||||||
|
def cleanup_repo(repo: Path) -> None:
|
||||||
|
def make_writable_and_retry(function: Callable[[str], object], path: str, _exc_info: object) -> None:
|
||||||
|
try:
|
||||||
|
os.chmod(path, stat.S_IREAD | stat.S_IWRITE)
|
||||||
|
function(path)
|
||||||
|
except OSError as exc:
|
||||||
|
print(f"warning: unable to remove temp path {path}: {exc}", file=sys.stderr)
|
||||||
|
|
||||||
|
if not repo.exists():
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
shutil.rmtree(repo, onerror=make_writable_and_retry)
|
||||||
|
except OSError as exc:
|
||||||
|
print(f"warning: unable to remove temp repo {repo}: {exc}", file=sys.stderr)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: list[str]) -> int:
|
||||||
|
args = parse_args(argv)
|
||||||
|
script_dir = Path(__file__).resolve().parent
|
||||||
|
engines = args.engines or list(DEFAULT_ENGINES)
|
||||||
|
repo = Path(tempfile.mkdtemp(prefix="autoreview-fixture."))
|
||||||
|
try:
|
||||||
|
create_fixture_repo(repo, args.fixture)
|
||||||
|
run_reviews(repo, script_dir, args.fixture, engines)
|
||||||
|
except subprocess.CalledProcessError as exc:
|
||||||
|
return int(exc.returncode or 1)
|
||||||
|
finally:
|
||||||
|
cleanup_repo(repo)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main(sys.argv[1:]))
|
||||||
186
.agents/skills/claw-score/SKILL.md
Normal file
186
.agents/skills/claw-score/SKILL.md
Normal file
@@ -0,0 +1,186 @@
|
|||||||
|
---
|
||||||
|
name: claw-score
|
||||||
|
description: Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.
|
||||||
|
---
|
||||||
|
|
||||||
|
# claw-score
|
||||||
|
|
||||||
|
Use this skill when working on the OpenClaw maturity scorecard in this repo.
|
||||||
|
This is the openclaw-local version of the maintainer `claw-score` workflow:
|
||||||
|
it keeps the taxonomy and scorecard concepts, but excludes discrawl and the old
|
||||||
|
committed `inventory/` report tree.
|
||||||
|
|
||||||
|
## Authority
|
||||||
|
|
||||||
|
This skill owns the operational workflow for:
|
||||||
|
|
||||||
|
- `taxonomy.yaml`
|
||||||
|
- `qa/maturity-scores.yaml`
|
||||||
|
- `docs/concepts/qa-e2e-automation.md`
|
||||||
|
- `qa/scenarios/index.yaml`
|
||||||
|
|
||||||
|
Keep person-specific, maintainer-private, Discord archive, and discrawl facts
|
||||||
|
out of this repo. If a score needs private evidence, use the redacted
|
||||||
|
`qa-evidence.json` artifact shape generated by OpenClaw QA workflows.
|
||||||
|
|
||||||
|
## Source Model
|
||||||
|
|
||||||
|
- `taxonomy.yaml` is the hand-edited source of truth for surfaces, levels,
|
||||||
|
QA profiles, categories, feature coverage IDs, docs refs, LTS overrides, and
|
||||||
|
completeness-instruction paths.
|
||||||
|
- Feature `coverageIds` are ANDed proof targets, not aliases. A feature may
|
||||||
|
list multiple IDs when each ID proves part of one capability.
|
||||||
|
- Coverage IDs use dotted `namespace.behavior` form, with lowercase
|
||||||
|
alphanumeric/dash segments. Profile, surface, and category IDs may remain
|
||||||
|
dashed or dotted.
|
||||||
|
- Keep categories and feature names unique, product-shaped, and broader than raw
|
||||||
|
coverage IDs. Do not promote generic IDs into standalone feature names.
|
||||||
|
- Avoid duplicate coverage-ID bundles under different feature names in one
|
||||||
|
category.
|
||||||
|
- `qa/maturity-scores.yaml` is the committed aggregate source for Quality,
|
||||||
|
Completeness, and LTS review state.
|
||||||
|
- `extensions/qa-lab/src/scorecard-taxonomy.ts` exports
|
||||||
|
`qaMaturityScoresSchema` and `readValidatedQaMaturityScoreSources`; use those
|
||||||
|
QA Lab utilities to validate score output.
|
||||||
|
- Generated public docs are `docs/maturity/scorecard.md` and
|
||||||
|
`docs/maturity/taxonomy.md`; both come from `pnpm maturity:render`. Do not
|
||||||
|
hand-edit generated Markdown to change score results.
|
||||||
|
- `qa-evidence.json` artifacts provide per-run QA scorecard evidence. Release
|
||||||
|
profile artifacts are the source of truth for Coverage. They can enrich
|
||||||
|
generated artifact docs, but they are not committed as inventory.
|
||||||
|
|
||||||
|
## Commands
|
||||||
|
|
||||||
|
Run from the openclaw repo root.
|
||||||
|
|
||||||
|
Validate taxonomy YAML structure and the maturity score schema after source
|
||||||
|
edits:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node --import tsx --input-type=module <<'NODE'
|
||||||
|
import fs from "node:fs";
|
||||||
|
import YAML from "yaml";
|
||||||
|
import { readValidatedQaMaturityScoreSources } from "./extensions/qa-lab/src/scorecard-taxonomy.ts";
|
||||||
|
|
||||||
|
for (const file of ["taxonomy.yaml", "qa/scenarios/index.yaml"]) {
|
||||||
|
YAML.parse(fs.readFileSync(file, "utf8"));
|
||||||
|
}
|
||||||
|
readValidatedQaMaturityScoreSources();
|
||||||
|
NODE
|
||||||
|
```
|
||||||
|
|
||||||
|
Check docs when touching docs prose:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm check:docs
|
||||||
|
```
|
||||||
|
|
||||||
|
Run focused QA/profile checks when changing coverage IDs or profile membership:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm openclaw qa coverage --json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Scoring Workflow
|
||||||
|
|
||||||
|
When asked to score or refresh a surface:
|
||||||
|
|
||||||
|
1. Read the surface in `taxonomy.yaml`.
|
||||||
|
2. Read the surface completeness rubric under
|
||||||
|
`.agents/skills/claw-score/references/completeness/`.
|
||||||
|
3. Gather public repo evidence from docs, source, tests, and QA scenario
|
||||||
|
metadata.
|
||||||
|
4. Prefer existing release profile `qa-evidence.json` artifacts for executed
|
||||||
|
proof.
|
||||||
|
5. Update `qa/maturity-scores.yaml` only for Quality, Completeness, and LTS
|
||||||
|
review state backed by public or redacted artifact evidence.
|
||||||
|
6. Run the schema validation command from this skill.
|
||||||
|
7. Run `pnpm check:docs` if docs prose changed, and focused QA coverage checks
|
||||||
|
if coverage IDs or profile membership changed.
|
||||||
|
|
||||||
|
For subjective score changes, make the smallest defensible edit and leave the
|
||||||
|
evidence path in the PR or task summary. Keep manual prose in current docs and
|
||||||
|
keep score data in `qa/maturity-scores.yaml`.
|
||||||
|
|
||||||
|
## Default Completeness Process
|
||||||
|
|
||||||
|
Completeness is scored against the intended operator-visible workflow for each
|
||||||
|
category, not against test breadth or implementation quality. The completeness
|
||||||
|
reference files under `references/completeness/` define the category scope and
|
||||||
|
any surface-specific variation from this default process.
|
||||||
|
|
||||||
|
By default, Completeness measures how fully OpenClaw exposes the intended
|
||||||
|
surface capability set to the user, operator, author, or maintainer persona for
|
||||||
|
that surface. Score whether each category delivers the full expected workflow,
|
||||||
|
including setup, normal use, status or inspection, recovery, and important
|
||||||
|
platform, provider, channel, security, or lifecycle variants where they apply.
|
||||||
|
|
||||||
|
Treat `Surface-Specific Scoring Questions` and `Surface-Specific Guidance` as
|
||||||
|
higher-priority instructions for that surface. The surface instructions may
|
||||||
|
flesh out, narrow, or intentionally conflict with the default ideas here; when
|
||||||
|
they do, follow the surface instructions and make the score rationale reflect
|
||||||
|
that surface-specific instruction. If a reference file does not include
|
||||||
|
surface-specific questions or guidance, apply this default process to the
|
||||||
|
surface's `Category Scope`.
|
||||||
|
|
||||||
|
For each category, ask:
|
||||||
|
|
||||||
|
- Can the intended user or operator complete the category workflow end to end?
|
||||||
|
- Are the taxonomy features present as supported capabilities rather than
|
||||||
|
isolated implementation fragments?
|
||||||
|
- Are the important lifecycle stages represented: setup, normal operation,
|
||||||
|
status/inspection, recovery, and upgrade or removal where relevant?
|
||||||
|
- Are the important environment, provider, platform, channel, or security
|
||||||
|
branches present for this surface?
|
||||||
|
- Do the known gaps leave major user-visible capability branches missing?
|
||||||
|
|
||||||
|
Default guidance:
|
||||||
|
|
||||||
|
- Favor higher Completeness when the category supports the full
|
||||||
|
operator-visible workflow described by taxonomy and category evidence.
|
||||||
|
- Lower Completeness when only the happy path exists, when important variants
|
||||||
|
are undocumented or unimplemented, or when recovery/status paths are missing.
|
||||||
|
- Do not lower Completeness because tests are thin; that is Coverage.
|
||||||
|
- Do not lower Completeness because implementation quality is fragile; that is
|
||||||
|
Quality.
|
||||||
|
|
||||||
|
Default Completeness bands:
|
||||||
|
|
||||||
|
- `Clawesome` (95-100): complete across expected workflows, variants, and
|
||||||
|
recovery branches, with only minor polish gaps.
|
||||||
|
- `Stable` (80-95): the expected workflow set is broadly present, with only
|
||||||
|
bounded missing branches.
|
||||||
|
- `Beta` (70-80): the main workflow exists, but meaningful branches or recovery
|
||||||
|
paths are still absent.
|
||||||
|
- `Alpha` (50-70): only a partial capability set is present; users can complete
|
||||||
|
some core tasks but not the full expected workflow.
|
||||||
|
- `Experimental` (0-50): the category exposes only fragments of the intended
|
||||||
|
capability.
|
||||||
|
|
||||||
|
## Score Semantics
|
||||||
|
|
||||||
|
- Coverage: deterministic release validation coverage derived from the release
|
||||||
|
profile `qa-evidence.json.scorecard` feature fulfillment data.
|
||||||
|
- Quality: reliability, maintainability, operator safety, and regression
|
||||||
|
confidence for the category.
|
||||||
|
- Completeness: how much of the intended operator-visible workflow exists for
|
||||||
|
the category. Use the default completeness process plus any surface-specific
|
||||||
|
variation before changing this score.
|
||||||
|
- LTS: derived from Quality, release-evidence Coverage, and
|
||||||
|
`human_lts_override`; do not hand-edit generated Markdown to change LTS
|
||||||
|
status.
|
||||||
|
|
||||||
|
Bands:
|
||||||
|
|
||||||
|
- `Clawesome`: 95-100
|
||||||
|
- `Stable`: 80-95
|
||||||
|
- `Beta`: 70-80
|
||||||
|
- `Alpha`: 50-70
|
||||||
|
- `Experimental`: 0-50
|
||||||
|
|
||||||
|
## Artifacts
|
||||||
|
|
||||||
|
Do not add the maintainer repo's `docs/kevinslin/maturity-scorecard/inventory/`
|
||||||
|
tree to openclaw. Evidence-enriched scorecard outputs belong in short-lived
|
||||||
|
artifacts, not committed generated docs, unless this repo adds an explicit
|
||||||
|
renderer/check workflow first.
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
# Agent Runtime Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`agent-runtime-and-provider-execution` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Agent Turn Execution: Turn startup and runtime choice, Session and run coordination, Abort and terminal outcomes
|
||||||
|
- External Runtimes and Subagents: External harness selection, CLI runtime aliases, Subagent turns, Runtime recovery
|
||||||
|
- Hosted Provider Execution: Hosted provider turns, Provider-specific model options, Hosted tool use, Reasoning and cache controls, Hosted streaming and replies
|
||||||
|
- Local and Self-hosted Providers: Local provider profiles, Tool-capability flags, Timeouts and context windows, Local smoke checks, Local failure handling
|
||||||
|
- Model and Runtime Selection: Model reference selection, Provider and runtime overrides, Thinking and context settings, Invalid route recovery
|
||||||
|
- Provider Auth: Login and API-key setup, Auth profile selection, Credential health checks, Auth failover, Provider fallback recovery, Rate-limit and capacity recovery, Missing-key and OAuth guidance, Restart and stale-route recovery, Structured provider diagnostics, Subagent credential propagation
|
||||||
|
- Streaming and Progress: Streaming replies, Progress visibility
|
||||||
|
- Tool Calls and Response Handling: Tool-call handling, Usage and response reporting, Failure recovery
|
||||||
|
- Tool Execution Controls: Tool availability rules, Sandboxed exec behavior, Approval flow, Elevated execution, Tool safety controls, Delegated tool access
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
# Android app Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`android-app` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Media Capture: Camera and media capture
|
||||||
|
- Mobile Chat: Chat tab
|
||||||
|
- Connection Setup: Gateway discovery
|
||||||
|
- Distribution: Public Google Play install path, Manual install path, Release smoke and startup performance
|
||||||
|
- Settings: Settings sheet
|
||||||
|
- Voice: Voice tab
|
||||||
|
- Device Runtime: Background reconnect and presence, Device command availability
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Anthropic provider path Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`anthropic-provider-path` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Provider Auth and Recovery: API-key onboarding, Claude CLI credential reuse, Setup-token auth, Auth profile health, Model status, Usage windows, Cooldown/profile reporting, Long-context recovery, Fallback guidance
|
||||||
|
- Model and Runtime Selection: Bundled Claude catalog, Canonical anthropic refs, Claude CLI compatibility, Model picker availability, Capability metadata, Runtime selection, Session continuity, MCP/tool bridge, Permission-mode mapping, Fallback prelude
|
||||||
|
- Request Transport and Turn Semantics: API-key/OAuth transport, Messages payloads, Streaming decode, Usage and stop reasons, Abort/error handling, Tool-use blocks, Tool-result replay, Partial JSON recovery, Native thinking, Signed/redacted thinking replay
|
||||||
|
- Prompt Cache and Context: Cache retention, System-prompt cache boundary, 1M context, Fast mode/service tier, Cache diagnostics
|
||||||
|
- Media Inputs: Image input, PDF document input, Media model fallback, Image tool results
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
# Automation: cron, hooks, tasks, polling Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`automation-cron-hooks-tasks-polling` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Cron Jobs: Create/edit/remove jobs, Schedule types, Timezone and stagger, Cron RPCs, Agent cron tool, Manual cron runs, Isolated cron execution, Model/provider preflight, Run history, Timeout and denial diagnostics, Chat announce delivery, Webhook delivery, Failure destinations, Skipped-run alerts, Delivery previews
|
||||||
|
- Event Ingress: Telegram long polling, Telegram webhook mode, Zalo polling/webhook mode, Polling stall diagnostics, iMessage watch fallback, Gmail setup wizard, Watcher start/serve, Tailscale/public routing, Push token validation, Gmail event routing, POST /hooks/wake, POST /hooks/agent, Mapped hooks, Hook auth policy, Async dispatch
|
||||||
|
- Automation Hooks: HOOK.md authoring, Hook discovery, Hook CLI management, Hook packs, Lifecycle event dispatch, api.on registration, Tool-call policy hooks, Message hooks, Session/lifecycle hooks, Plugin approval requests, cron_changed
|
||||||
|
- Background Tasks and Flows: Task list/show/cancel, Task notifications, Task audit and maintenance, Chat task board, Task pressure status, Managed flows, Mirrored flows, openclaw tasks flow, Flow audit and maintenance, Plugin managedFlows
|
||||||
|
- Heartbeat: Heartbeat scheduling, Active hours, Wake and cooldown handling, Due-only heartbeat tasks, Commitment check-ins
|
||||||
|
- Polling Controls: openclaw message poll, Telegram polls, Teams polls, Poll flags, Channel capability gates, process poll, process log, Background process status, No-progress loop detection, Process input controls
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
# Browser automation and exec/sandbox tools Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`browser-automation-and-exec-sandbox-tools` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Browser Automation: Browser Actions, Snapshots, Artifacts, Browser Plugin Service, Profiles, Browser Security, SSRF, Remote Control
|
||||||
|
- Tool Invocation and Execution: Exec Routing, Process Lifecycle, Direct Tool Invoke API, Node System.run, Host Exec Approvals, Elevated Mode
|
||||||
|
- Sandbox and Tool Policy: Sandbox Backends, Workspace Isolation, Sandboxed Browser, Codex Dynamic Tools, Tool Policy, Sandbox Tool Gates
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
# Gateway Web App Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`browser-control-ui-and-webchat` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Browser Realtime Talk: Browser Talk start/stop, Provider session selection, Gateway relay audio, Tool-call consults, Steer and cancel
|
||||||
|
- Browser Access and Trust: Device pairing, Token/password auth, Tailscale Serve auth, Trusted proxy auth, Allowed origins/gatewayUrl
|
||||||
|
- Configuration: Config snapshots, Schema form editing, Raw JSON editing, Base-hash guarded writes, Apply and restart
|
||||||
|
- Browser UI: Gateway-hosted UI, Dashboard open/auth bootstrap, Base-path routing, Static asset recovery, Dev gatewayUrl target, PWA install metadata, Service worker updates, VAPID keys, Subscribe/unsubscribe, Test notifications
|
||||||
|
- WebChat Conversations: Send and abort, Session and agent picker, Model/thinking controls, Attachments, Markdown/tool/media rendering, chat.history projection, chat.send lifecycle, Abort/partial retention, Injected assistant notes, Reconnect continuity, Hosted embeds, External embed gating, Assistant media tickets, Authenticated avatars, CSP image policy
|
||||||
|
- Remote WebChat: macOS WebChat transport, SSH tunnel data plane, Direct ws/wss remote mode, Session continuity, Remote troubleshooting
|
||||||
|
- Operator Console: Health/status/models, Live log tail, Update run/status, Activity summaries, RPC timing telemetry, Channels/login, Session manager and history, Cron, Skills/nodes, Exec approvals/agents
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
# Channel framework Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`channel-framework` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Actions Commands and Approvals: Channel-native commands, Native command session target, Message actions, Message tool API discovery, Channel-native approval prompts
|
||||||
|
- Channel Setup: Supported channel catalog, Channel status taxonomy in channels list, Setup/onboarding flows, Install-on-demand, Setup wizard metadata
|
||||||
|
- Group Thread and Ambient Room Behavior: Group/channel session isolation, Mention-required, Native threads, Broadcast groups, Bot-loop protection
|
||||||
|
- Inbound Access and Identity Gates: DM pairing, Group/channel allowlists, Access group expansion, Mention gating, Sanitized inbound identity/route projections
|
||||||
|
- Media Attachments and Rich Channel Data: Inbound media normalization, Outbound direct text/media sends, Provider-specific channelData, Media roots
|
||||||
|
- Outbound Delivery and Reply Pipeline: Automatic final reply delivery, Durable outbound send orchestration, Reply pipeline transforms, Provider outbound adapter bridge
|
||||||
|
- Conversation Routing and Delivery: Inbound conversation routing, Session key construction, Agent binding precedence, Runtime conversation bindings, Thread/parent-child placement, Plugin registry resolution, Channel account startup, Whole-channel lifecycle controls, Config/secrets reload interactions, Auto-restart
|
||||||
|
- Status Health and Operator Controls: channels.status, Channel health policy, Operator CLI controls, Status read-model
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# ClawHub Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`clawhub-and-external-plugin-distribution` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Publishing: ClawHub package publishing owner, OpenClaw-owned package release validation for ClawHub, Version bump gates, npm trusted publishing provenance, External code plugin package contract required, Skill package metadata, Skill publishing flow
|
||||||
|
- Catalog Discovery: openclaw plugins search as the ClawHub, Search result metadata, Distinction between plugin search, Catalog lookup failure, Skill catalog search
|
||||||
|
- Compatibility and Trust: openclaw.compat.pluginApi, ClawHub package compatibility validation, npm compatibility fallback to the newest, Official external plugin catalog behavior, Compatibility docs, Operator trust model for installing, ClawHub archive, npm integrity drift, Built-in dangerous-code scanner, ClawHub publishing review/hidden-release behavior as upstream, Skill archive safety, Skill audit signals
|
||||||
|
- Plugin Lifecycle: Source prefixes, Bare package behavior during the launch, Explicit pinned versions, Managed install records that preserve source, Codex, Local, Marketplace list, Supported mapped features, Remote marketplace path safety, Update by plugin id, Reinstall vs update semantics, Downgrade, Uninstall config/index/policy/file cleanup, Gateway restart/reload requirements after, ClawHub skill installs, Skill upload install path, Skill dependency installers
|
||||||
|
- Plugin Health: Per-plugin managed npm project, npm-pack local release-candidate installs, Dependency ownership between plugin packages, Peer dependency relinking, Legacy dependency root cleanup, plugins list, Local plugin index, Troubleshooting stale config, Runtime verification after Gateway
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
# CLI Surface Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`cli-install-update-onboard-doctor` surface.
|
||||||
|
|
||||||
|
## Surface-Specific Scoring Questions
|
||||||
|
|
||||||
|
For each category, ask:
|
||||||
|
|
||||||
|
- Can a normal operator complete the job end to end from the CLI?
|
||||||
|
- Are the expected environments represented where they matter for the category,
|
||||||
|
such as local installs, remote gateway use, supervised services, or
|
||||||
|
Windows/WSL2?
|
||||||
|
- Are the main lifecycle stages present where relevant: setup, inspection,
|
||||||
|
change, repair, and upgrade?
|
||||||
|
- Are common recovery and troubleshooting branches present, or does the
|
||||||
|
workflow dead-end after the happy path?
|
||||||
|
- Are major documented operator expectations still unimplemented?
|
||||||
|
|
||||||
|
## Surface-Specific Guidance
|
||||||
|
|
||||||
|
Variation from the default completeness process:
|
||||||
|
|
||||||
|
- Completeness is the CLI operator journey for installation, onboarding, configuration, repair, and upgrade across expected environments and recovery branches.
|
||||||
|
- Score the CLI against the full operator journey, not only installation or the happy path.
|
||||||
|
- Repair, migration, remote, and platform-specific branches are expected where a category exposes them.
|
||||||
|
- For Windows and WSL2, score against the intended supported experience rather than parity with macOS/Linux internals.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- CLI Setup: Installer scripts, Local prefix install, Package-manager installs, Supported Node runtime, Source checkout install, CLI entrypoint
|
||||||
|
- Onboarding and Auth Setup: Guided onboarding, Targeted reconfiguration, Auth choices, Gateway auth storage, Remote onboarding
|
||||||
|
- Plugin and Channel Setup: Channel picker, Plugin install sources, Channel account setup, Post-setup probes, Remote gateway caveat
|
||||||
|
- Gateway Service Management: Foreground gateway runs, Service install and control, Service auth wiring, Drift and reinstall recovery, Service health checks
|
||||||
|
- CLI Observability: Status snapshots, Health snapshots, Remote log tailing, Diagnostics export, Support-safe redaction
|
||||||
|
- Doctor: Interactive repair, Config migration, Auth and SecretRef checks, Plugin validation and repair, Lint and JSON findings, Extra gateway discovery, Supervisor drift repair, Port and startup diagnosis, Runtime path checks, Restart guidance
|
||||||
|
- Updates and Upgrades: Update channels, Install-kind switching, Managed gateway restart, Update status and RPC, Plugin convergence
|
||||||
13
.agents/skills/claw-score/references/completeness/discord.md
Normal file
13
.agents/skills/claw-score/references/completeness/discord.md
Normal file
@@ -0,0 +1,13 @@
|
|||||||
|
# Discord Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`discord` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Application and bot setup, Token and application ID configuration, Setup wizard and account inspection, Status, doctor, and intent checks, Multi-account bot configuration, Account monitor startup, Gateway WebSocket lifecycle, Reconnect and heartbeat handling, Rate limits and gateway metadata, Status, probe, and health-monitor recovery
|
||||||
|
- Access and Identity: DM policy modes, Allowlist inheritance, Pairing-code approval, Sender authorization, Access-group authorization, Group DM authorization
|
||||||
|
- Conversation Routing and Delivery: Guild and channel admission, Mention gating, Session key isolation, Configured and runtime routing, Inbound context visibility, Forum and media-channel thread posts, Thread actions, Target parsing, Thread context resolution, Thread-bound session routing, ACP agent routing, Routing lifecycle, Discord forum/media channel posts created as, CLI and message-tool thread actions, Discord target parsing for `channel:<id>`, Thread context resolution, Thread-bound session routing for `/focus`, `/unfocus`, `/agents`, `/session idle`, `/session max-age`, `sessions_spawn({ thread, ACP current-conversation bindings and ACP thread, Binding lifecycle behavior, Direct and thread sends, Text chunking and reply mode, Draft and progress edits, Mention and embed rendering, REST retry and final delivery, File uploads, Component file and media-gallery blocks, Video caption follow-up, Voice-message upload, Inbound attachment context
|
||||||
|
- Media and Rich Content: Direct and thread sends, Text chunking and reply mode, Draft and progress edits, Mention and embed rendering, REST retry and final delivery, File uploads, Component file and media-gallery blocks, Video caption follow-up, Voice-message upload, Inbound attachment context, Direct and thread sends, Text chunking and reply mode, Draft and progress edits, Mention and embed rendering, REST retry and final delivery, File uploads, Component file and media-gallery blocks, Video caption follow-up, Voice-message upload, Inbound attachment context, Outbound file uploads from URLs and, Component v2 file and media-gallery blocks, Video caption handling and follow-up media-only delivery, Discord voice-message sends with OGG/Opus conversion, Inbound media/attachment-aware debounce behavior, Realtime voice-channel conversations, General text-only delivery
|
||||||
|
- Native Controls and Approvals: Native slash command registration, Native slash command execution, Model Picker Commands, Components v2 messages, Callback TTL, Native Discord exec/plugin approvals, Sensitive owner-only command routing for prompts, Discord message actions, Action gates under channels.discord.actions.\*
|
||||||
|
- Realtime Voice and Calls: Voice Channel Lifecycle, Auto-join and follow-users, Realtime voice modes, Wake, barge-in, and echo handling, Voice codec and DAVE recovery
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Docker / Podman hosting Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`docker-podman-hosting` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Container Setup: Local Image Setup Script, Docker Compose gateway, First-run onboarding, Docker-only first-run notes, Podman setup scripts and Quadlet template, Rootless Podman image setup
|
||||||
|
- Container Operations: Host CLI routing into running Docker/Podman, Container Targeting, Container update/rebuild/restart guidance for Docker, Docker Compose, Gateway token generation, Ownership, Docker Compose, Container health endpoints, Provider/VPS Docker hosting docs, Docker VM persistence/update guidance, Operator-facing update
|
||||||
|
- Image Release and Validation: Root Dockerfile build stages, Docker release workflow, Docker E2E package artifact generation, Docker E2E plan/scheduler scripts, Release-path install
|
||||||
|
- Agent Sandbox and Tooling: Docker gateway setup, Docker-backed agent sandbox support, Container image dependency baking
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Feishu, QQ Bot, WeChat, Yuanbao, Zalo, Zalo Personal, regional channels Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`feishu-qq-bot-wechat-yuanbao-zalo-zalo-personal-regional-channels` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Docs channel index, Official external channel catalog entries, Core channel-plugin catalog, Channel setup wizard, Missing-plugin, Cross-channel ingress/access/refactor concerns, Feishu/Lark bot channel setup, WebSocket default mode, DM pairing, Message delivery, Feishu document, Multi-account credential handling, QQ Open Platform AppID/AppSecret setup, C2C private chat, Group activation, Rich media messages, Slash commands, Multi-account gateway connections, Tencent Yuanbao external channel, AppKey/AppSecret setup, DMs, Outbound queue strategy, Core-side official external catalog, Zalo Bot Creator / Marketplace bot, Long-polling default mode, Bot token, Group policy schema, Text, Status probes, WeChat/Weixin personal messaging, Plugin install, Direct-message pairing, Core-side catalog metadata, External sidecar/helper process behavior, zalouser channel plugin, QR login, DM pairing, Message send, Doctor/status checks for runtime availability, Explicit unofficial-account risk, QQ Open Platform AppID/AppSecret setup and, C2C private chat, Group activation, Inbound and outbound rich media including, Slash commands, Multi-account gateway connections, Tencent Yuanbao external channel `openclaw-plugin-yuanbao, AppKey/AppSecret setup, DMs, Outbound queue strategy, Core-side official external catalog, Zalo Bot Creator / Marketplace bot, Long-polling default mode and optional HTTPS, Bot token, Group policy schema and fail-closed group, Text, Status probes and troubleshooting for token/config/webhook problems, zalouser` channel plugin for Zalo Personal, QR login, DM pairing, Message send, Doctor/status checks for runtime availability and, Explicit unofficial-account risk and operator safeguards
|
||||||
|
- Access and Identity: Feishu/Lark bot channel setup, WebSocket default mode, DM pairing, Message delivery, Feishu document, Multi-account credential handling, QQ Open Platform AppID/AppSecret setup, C2C private chat, Group activation, Rich media messages, Slash commands, Multi-account gateway connections, Tencent Yuanbao external channel, AppKey/AppSecret setup, DMs, Outbound queue strategy, Core-side official external catalog, Zalo Bot Creator / Marketplace bot, Long-polling default mode, Bot token, Group policy schema, Text, Status probes, WeChat/Weixin personal messaging, Plugin install, Direct-message pairing, Core-side catalog metadata, External sidecar/helper process behavior, zalouser channel plugin, QR login, DM pairing, Message send, Doctor/status checks for runtime availability, Explicit unofficial-account risk, QQ Open Platform AppID/AppSecret setup and, C2C private chat, Group activation, Inbound and outbound rich media including, Slash commands, Multi-account gateway connections, Tencent Yuanbao external channel `openclaw-plugin-yuanbao, AppKey/AppSecret setup, DMs, Outbound queue strategy, Core-side official external catalog, zalouser` channel plugin for Zalo Personal, QR login, DM pairing, Message send, Doctor/status checks for runtime availability and, Explicit unofficial-account risk and operator safeguards
|
||||||
|
- Conversation Routing and Delivery: Feishu/Lark bot channel setup, WebSocket default mode, DM pairing, Message delivery, Feishu document, Multi-account credential handling, QQ Open Platform AppID/AppSecret setup, C2C private chat, Group activation, Rich media messages, Slash commands, Multi-account gateway connections, Tencent Yuanbao external channel, AppKey/AppSecret setup, DMs, Outbound queue strategy, Core-side official external catalog, Zalo Bot Creator / Marketplace bot, Long-polling default mode, Bot token, Group policy schema, Text, Status probes, WeChat/Weixin personal messaging, Plugin install, Direct-message pairing, Core-side catalog metadata, External sidecar/helper process behavior, zalouser channel plugin, QR login, DM pairing, Message send, Doctor/status checks for runtime availability, Explicit unofficial-account risk, QQ Open Platform AppID/AppSecret setup and, C2C private chat, Group activation, Inbound and outbound rich media including, Slash commands, Multi-account gateway connections, Tencent Yuanbao external channel `openclaw-plugin-yuanbao, AppKey/AppSecret setup, DMs, Outbound queue strategy, Core-side official external catalog, Zalo Bot Creator / Marketplace bot, Long-polling default mode and optional HTTPS, Bot token, Group policy schema and fail-closed group, Text, Status probes and troubleshooting for token/config/webhook problems, zalouser` channel plugin for Zalo Personal, QR login, DM pairing, Message send, Doctor/status checks for runtime availability and, Explicit unofficial-account risk and operator safeguards
|
||||||
|
- Media and Rich Content: Feishu/Lark bot channel setup, WebSocket default mode, DM pairing, Message delivery, Feishu document, Multi-account credential handling, QQ Open Platform AppID/AppSecret setup, C2C private chat, Group activation, Rich media messages, Slash commands, Multi-account gateway connections, Tencent Yuanbao external channel, AppKey/AppSecret setup, DMs, Outbound queue strategy, Core-side official external catalog, Zalo Bot Creator / Marketplace bot, Long-polling default mode, Bot token, Group policy schema, Text, Status probes, QQ Open Platform AppID/AppSecret setup and, C2C private chat, Group activation, Inbound and outbound rich media including, Slash commands, Multi-account gateway connections, Zalo Bot Creator / Marketplace bot, Long-polling default mode and optional HTTPS, Bot token, Group policy schema and fail-closed group, Text, Status probes and troubleshooting for token/config/webhook problems
|
||||||
@@ -0,0 +1,43 @@
|
|||||||
|
# Gateway Runtime Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`gateway-runtime` surface.
|
||||||
|
|
||||||
|
## Surface-Specific Scoring Questions
|
||||||
|
|
||||||
|
For each category, ask:
|
||||||
|
|
||||||
|
- Does the category cover the main happy path an operator or client needs?
|
||||||
|
- Are the major deployment modes present where they matter for this category:
|
||||||
|
local, remote, node-mediated, supervised, or browser-facing?
|
||||||
|
- Are the main lifecycle stages present where relevant: setup, normal use,
|
||||||
|
status/inspection, and recovery?
|
||||||
|
- Are important security or policy branches present where the category implies
|
||||||
|
them?
|
||||||
|
- Are obvious operator-visible holes or "not yet supported" branches still
|
||||||
|
missing?
|
||||||
|
|
||||||
|
## Surface-Specific Guidance
|
||||||
|
|
||||||
|
Variation from the default completeness process:
|
||||||
|
|
||||||
|
- Completeness includes operator and connected-client workflows, major deployment modes, and recovery paths, not just gateway protocol capability.
|
||||||
|
- Score the Gateway against the full operator and client journey, not just protocol primitives or one transport path.
|
||||||
|
- Local, remote, node-mediated, supervised, and browser-facing modes matter when the category implies them.
|
||||||
|
- Approval/policy variants and recovery or diagnostic paths count as completeness branches, not polish.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Approvals and Remote Execution: Exec approvals, Plugin approvals, Node exec approvals, Approved node execution, Approval mutation safety, Delivery fallback behavior
|
||||||
|
- HTTP APIs: OpenAI-compatible APIs, Tool invocation API, Admin API access, Hook ingress
|
||||||
|
- Hosted Web Surface: Control UI, WebChat hosting, Plugin web routes, Canvas and A2UI routes
|
||||||
|
- Gateway RPC APIs and Events: Health APIs, Identity and presence APIs, Model APIs, Usage and memory APIs, Session APIs, Chat APIs, Channel APIs, Web login and wake APIs, Config and secrets APIs, Update and setup APIs, Agent and artifact APIs, Task and automation APIs, Tool and skill APIs, Request and event envelopes, Idempotent side effects, Method discovery, Event discovery, Accepted-then-final results, Event ordering, State refresh after gaps
|
||||||
|
- Device Auth and Pairing: Shared-secret login, Trusted proxy auth, Private ingress mode, Device challenge signing, Device tokens, Setup-code bootstrap, Auth mismatch recovery, Device auth migration, Client pairing, Node pairing
|
||||||
|
- Network Access and Discovery: Loopback and LAN access, Tailnet access, SSH tunnels, Endpoint discovery, Saved endpoints, TLS pinning
|
||||||
|
- Nodes and Remote Capabilities: Node presence, Node capabilities, Node inventory, Node actions, Node events, Pending work delivery, Remote device capabilities, Remote host commands
|
||||||
|
- Health, Diagnostics, and Repair: Health snapshots, Channel readiness, Stability diagnostics, Payload diagnostics, Diagnostics exports, Doctor checks, Log tailing
|
||||||
|
- Protocol Compatibility: Published protocol schema, Runtime request validation, JSON Schema export, Swift client models, Version negotiation, Client transport defaults, Backward-compatible evolution
|
||||||
|
- Roles and Permissions: Role negotiation, Operator permissions, Approval-gated actions, Untrusted node declarations, Event scoping
|
||||||
|
- Gateway Lifecycle: Foreground startup, Service installation, Restart and stop, Service status, Bind and port settings, Config reload, Multi-gateway isolation
|
||||||
|
- Security Controls: Non-loopback auth, Trusted proxy exceptions, Gateway and node trust boundaries, Trusted CIDR auto-approval, Fail-closed protocol handling, Remote execution safeguards
|
||||||
|
- WebSocket Connection: WebSocket transport, Connect challenge, Connect request, Protocol version negotiation, hello-ok snapshot, Startup retry, Session limits, Plugin surface URLs
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Google Chat Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`google-chat` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Google Cloud project setup, Chat app configuration, Service account setup, Webhook audience and path, Workspace visibility and app status, Guided channel setup, Account resolution, Service account SecretRefs, Env file and inline credentials, Channel status and probes, Directory and mutable-id diagnostics, NPM and ClawHub install, Plugin docs and catalog routing, Channel aliases and labels, Operator status UI, Install/update metadata, Webhook path handling, Standard Chat token verification, Workspace add-on token verification, Audience and appPrincipal validation, Shared-path target selection, Auth rejection diagnostics, Account resolution, Service account SecretRefs, Env file and inline credentials, Channel status and probes, Directory and mutable-id diagnostics, NPM and ClawHub install, Plugin docs and catalog routing, Channel aliases and labels, Operator status UI, Install/update metadata, Webhook path handling, Standard Chat token verification, Workspace add-on token verification, Audience and appPrincipal binding, Shared-path target selection, Auth rejection diagnostics
|
||||||
|
- Access and Identity: DM pairing approval, Sender allowlists, Google Chat identity matching, Direct session routing, Pairing diagnostics, Space allowlists, Mention gating, Sender access groups, Group session isolation, Bot-loop protection, Space diagnostics
|
||||||
|
- Conversation Routing and Delivery: DM pairing approval, Sender allowlists, Google Chat identity matching, Direct session routing, Pairing diagnostics, Space allowlists, Mention gating, Sender access groups, Group session isolation, Bot-loop protection, Space diagnostics, Inbound attachments, Outbound media replies, Message upload action, Media source and size controls, Media receipts and thread placement, Text send action, Upload-file action, Reaction actions, Action capability gates, Approval sender matching, Thread-aware replies, Streaming and chunked replies, Typing placeholder lifecycle, Message-tool current-source replies, NO_REPLY cleanup, Markdown/text rendering, Thread-aware replies, Streaming and chunked replies, Typing placeholder lifecycle, Message-tool current-source replies, NO_REPLY cleanup, Markdown/text rendering
|
||||||
|
- Media and Rich Content: Inbound attachments, Outbound media replies, Message upload action, Media source and size controls, Media receipts and thread placement, Text send action, Upload-file action, Reaction actions, Action capability gates, Approval sender matching, Thread-aware replies, Streaming and chunked replies, Typing placeholder lifecycle, Message-tool current-source replies, NO_REPLY cleanup, Markdown/text rendering
|
||||||
|
- Native Controls and Approvals: Inbound attachments, Outbound media replies, Message upload action, Media source and size controls, Media receipts and thread placement, Text send action, Upload-file action, Reaction actions, Action capability gates, Approval sender matching, Thread-aware replies, Streaming and chunked replies, Typing placeholder lifecycle, Message-tool current-source replies, NO_REPLY cleanup, Markdown/text rendering
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Google provider path Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`google-provider-path` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Provider Setup and Credentials: API key onboarding, Auth choice metadata, Gemini CLI OAuth setup, Vertex ADC setup, Daemon and fallback credentials, CLI runtime selection, OAuth login and refresh, Canonical Google model refs, CLI usage normalization, OAuth diagnostics
|
||||||
|
- Model Routing and Endpoints: Catalog rows and aliases, Dynamic model resolution, Provider routing, Google-native config normalization, Model picker availability, Vertex provider selection, ADC/service-account auth, Project/location endpoints, Custom base URL policy, Compatibility boundaries
|
||||||
|
- Direct Gemini Runtime: Direct Gemini chat, Multimodal inputs, Tool-call streaming, Usage and stop reasons, Thought-signature replay, Thinking-level mapping, Thought-signature replay, Tool turn ordering, Incomplete-turn recovery, Planning-only turn recovery
|
||||||
|
- Media, Search, and Realtime: Bundled plugin distribution, Provider auto-enable metadata, Image and media adapters, Speech and realtime adapters, Search and generation tools, Realtime voice sessions, Constrained browser tokens, Audio and transcript events, Live tool calls, Session reconnects
|
||||||
|
- Prompt Caching: Cache retention config, Managed cachedContents, Manual cachedContent handles, Cache usage accounting, Cache diagnostics and live proof
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Image/video/music generation tools Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`image-video-music-generation-tools` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Media Routing and Discovery: default media model config, per-call model refs and fallbacks, auth-backed tool discovery, action=list provider inspection
|
||||||
|
- Task Lifecycle and Delivery: background task creation, task status/list/show/cancel, duplicate guards, progress keepalive, completion/failure wake, no-session inline fallback, local media persistence, MIME/filename inference, Hosted URL fallback, message-tool handoff, idempotent missing-media fallback, channel attachment proof
|
||||||
|
- Image Generation: text-to-image, reference-image editing, output hints, action=status, provider attempt metadata, OpenAI/Codex OAuth, API-key OpenAI, OpenRouter/xAI/fal/LiteLLM/DeepInfra/Google/MiniMax/ComfyUI auth, provider error diagnostics
|
||||||
|
- Video Generation: text-to-video, image-to-video, video-to-video, reference role validation, audio refs, typed providerOptions, queue-backed jobs, polling/timeout handling, Hosted URL download, provider skip explanations, returned asset metadata
|
||||||
|
- Music Generation: prompt and lyrics input, instrumental mode, duration/format controls, image-reference edit lanes, generated audio outputs, provider fallback
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# iMessage / BlueBubbles Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`imessage-bluebubbles` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Translate legacy config, Cut over safely, Handle migration caveats, Run local imsg, Run through SSH wrapper, Grant macOS permissions, Probe runtime health, Account setup prompts, Account status checks, Doctor repair checks, Account Config, Translate legacy config, Cut over safely, Handle migration caveats, Run local imsg, Run through SSH wrapper, Grant macOS permissions, Probe runtime health
|
||||||
|
- Access and Identity: Authorize direct senders, Route direct conversations, Bind ACP sessions, Group Policy, Mentions, System Prompts, Group Policy, Mentions, System Prompts
|
||||||
|
- Conversation Routing and Delivery: Watch live messages, Coalesce split-send DMs, Replay missed messages, Seed conversation history, Authorize direct senders, Route direct conversations, Bind ACP sessions, Group Policy, Mentions, System Prompts
|
||||||
|
- Media and Rich Content: Media, Attachments, Remote Fetch, Chunking, Native Actions, Private API, Message Tool
|
||||||
|
- Native Controls and Approvals: Native Approvals, Reactions, Operator Control, Media, Attachments, Remote Fetch, Chunking, Native Actions, Private API, Message Tool, Native Actions, Private API, Message Tool
|
||||||
15
.agents/skills/claw-score/references/completeness/ios-app.md
Normal file
15
.agents/skills/claw-score/references/completeness/ios-app.md
Normal file
@@ -0,0 +1,15 @@
|
|||||||
|
# iOS app Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`ios-app` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Media and Sharing: Camera list/snap/clip
|
||||||
|
- Canvas and Screen: Canvas present/hide/navigate/eval/snapshot
|
||||||
|
- Chat and Sessions: Chat sessions and operator controls
|
||||||
|
- Gateway Setup and Diagnostics: Bonjour/local, Manual host/port, Gateway connect configuration persistence, TLS fingerprint trust prompt, Pairing approval, Pairing/auth diagnostics for users, Settings tab
|
||||||
|
- Distribution: Internal preview status
|
||||||
|
- Device Commands: Location modes, Device command handling
|
||||||
|
- Notifications and Background: APNs registration and relay delivery
|
||||||
|
- Voice: Voice wake
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Kubernetes Hosting Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`kubernetes-hosting` surface.
|
||||||
|
|
||||||
|
## Surface-Specific Scoring Questions
|
||||||
|
|
||||||
|
For each category, ask:
|
||||||
|
|
||||||
|
- Can an operator deploy and manage OpenClaw on Kubernetes end to end?
|
||||||
|
- Are the taxonomy features present as supported manifests, commands, and docs rather than examples only?
|
||||||
|
- Are setup, normal operation, status or inspection, redeploy, teardown, and secret rotation represented where relevant?
|
||||||
|
- Are local Kind validation, namespace/image customization, provider secrets, and secure exposure branches covered?
|
||||||
|
- Do known gaps leave major cluster-hosting capability branches missing?
|
||||||
|
|
||||||
|
## Surface-Specific Guidance
|
||||||
|
|
||||||
|
Variation from the default completeness process:
|
||||||
|
|
||||||
|
- Completeness is the Kubernetes operator workflow for deployment, configuration, secrets, access, exposure, lifecycle, security posture, status, and recovery.
|
||||||
|
- A complete Kubernetes category lets an operator deploy, expose, secure, update, troubleshoot, and remove the Gateway without relying on Docker-only assumptions.
|
||||||
|
- Happy-path port-forwarding, missing secret/config rotation, or omitted exposed-service security posture are material completeness gaps.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Deployment Setup: Kustomize packaging, cluster prerequisites, quick deploy, manifest apply, and Kind validation.
|
||||||
|
- Configuration and Secrets: agent instructions, Gateway config, provider secrets, secret rotation, and image/namespace customization.
|
||||||
|
- Access and Exposure: port-forward access, service endpoint, ingress exposure, auth/TLS, and localhost posture.
|
||||||
|
- Cluster Lifecycle: resource layout, state persistence, redeploy, teardown, and security context.
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Linux companion app Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`linux-companion-app` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- App Distribution: Native app package, Distro package targets, Official release metadata
|
||||||
|
- Gateway Connectivity: Local Gateway attach and status, Gateway pairing and auth, Remote mode, Local and remote resource boundaries
|
||||||
|
- Chat and Sessions: Native Linux chat window, Transcript, Gateway chat transport
|
||||||
|
- Desktop Capabilities: Linux desktop permissions, Secret storage, Sandbox/package posture, Linux native node identity, Host command execution, Desktop tools, Linux native Talk, Microphone capture, Native media permissions
|
||||||
|
- Status and Diagnostics: Native Linux app readiness, Gateway health/status display, Log/transcript opening, Doctor/repair affordances, Linux tray/status item, Runtime status row, Desktop-environment integration
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Linux Gateway host Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`linux-gateway-host` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Host Setup and Updates: Linux CLI install, Node runtime prerequisites, Package-manager policy, Update path
|
||||||
|
- Gateway Runtime and Service Control: Foreground Gateway Runtime, Process Control, Systemd User Service Lifecycle setup, Systemd User Service Lifecycle operation, Systemd User Service Lifecycle status, Systemd User Service Lifecycle recovery
|
||||||
|
- Remote Access and Security: Remote Network Exposure, TLS, Tailscale, Gateway exposure safeguards, Gateway authentication modes, Secret Handling
|
||||||
|
- Diagnostics and Repair: Gateway diagnostic reports, Gateway log tailing, Doctor checks, Operator repair guidance
|
||||||
|
- Deployment Targets: VPS, Container, Cloud Deployment Guidance
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Local model providers: Ollama, vLLM, SGLang, LM Studio Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`local-model-providers-ollama-vllm-sglang-lm-studio` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Provider Setup, Lifecycle, and Diagnostics: Provider Selection, Onboarding, localService configuration, Process startup and readiness, Request leases and idle shutdown, Health checks and restart, Provider recipes, Local provider status, Backend reachability probes, Model availability errors, Memory readiness diagnostics, Provider troubleshooting docs
|
||||||
|
- Native Provider Plugins: Ollama setup and model pulling, Model discovery, Streaming and vision, Ollama embeddings, Web-search support, LM Studio setup, Model discovery and auth, Model preload and JIT loading, Streaming compatibility, LM Studio embeddings
|
||||||
|
- OpenAI-Compatible Runtime Compatibility: Bundled provider setup, Model Discovery Endpoint, Non-interactive configuration, vLLM thinking controls, OpenAI-compatible chat and tool semantics, SGLang compatibility guidance, Request Stream Compatibility, Tool Calling
|
||||||
|
- Local Memory and Embeddings: Embedding provider selection, Memory search readiness, memoryFlush model override, Fallback lexical search, Provider mismatch guidance
|
||||||
|
- Network Safety and Prompt Controls: Safety Network, Prompt Pressure Controls
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
# Long-tail hosted providers Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`long-tail-hosted-providers` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Hosted LLM Providers: Bedrock setup, Gateway/proxy routing, Copilot/OpenCode hosted access, Proxy capability diagnostics, Hosted text completion, Tool-call and streaming compatibility, Model catalog resolution, Provider-specific request shaping, Regional provider setup, Region and plan routing, Regional live smoke, Account prerequisite diagnostics
|
||||||
|
- Hosted Media Providers: Image generation providers, Video generation providers, Music generation providers, Media mode coverage, Text-to-speech providers, Speech-to-text providers, Realtime transcription providers, Audio format diagnostics
|
||||||
|
- Provider Operations: Provider directory, Provider install catalog, Model catalog metadata, Catalog parity checks, Provider setup descriptors, Auth profiles and aliases, Credential health probes, Key rotation and recovery, Direct provider smoke, Gateway live smoke, Models status probes, Fallback trace and repair
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
# macOS companion app Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`macos-companion-app` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Canvas: Canvas panel open/hide/navigate/eval/snapshot, Local custom URL scheme, A2UI host auto-navigation, Canvas enable/disable setting
|
||||||
|
- Local Setup: Local mode Gateway attach/start/stop, LaunchAgent install/update/restart/uninstall, Existing-listener detection, Native first-run onboarding flow, CLI discovery, Local workspace selection, Onboarding WebChat session separation
|
||||||
|
- Status and Settings: Menu-bar status, Activity state ingestion, Settings navigation, Health polling, Channels settings
|
||||||
|
- Native Capabilities: Mac node session connection, system.run, Exec approval policy, Permission requests, TCC persistence
|
||||||
|
- Remote Connections: Remote connection mode selection, SSH tunnel, Gateway discovery
|
||||||
|
- Voice and Talk: Voice Wake runtime, Push-to-talk, Talk provider playback plan
|
||||||
|
- WebChat: Native SwiftUI WebChat window, Gateway chat transport, Local and remote data-plane reuse
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
# macOS Gateway host Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`macos-gateway-host` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- CLI Setup: Hosted installer, Node 24 recommendation, App-triggered CLI install, Shell PATH and version-manager drift
|
||||||
|
- Local Gateway Integration: App local/remote connection mode, App-managed Gateway LaunchAgent install/restart/uninstall, CLI install detection, Attach-to-existing local Gateway compatibility, Gateway endpoint, gateway.mode=local configuration, Loopback bind, Local app endpoint resolution, Bonjour discovery
|
||||||
|
- Remote Gateway Mode: macOS app "Remote over SSH", SSH tunnel setup, Tailscale MagicDNS, Remote endpoint token/password/TLS fingerprint, Local node host startup
|
||||||
|
- Gateway Service Lifecycle: Per-user Gateway LaunchAgent install, launchctl bootstrap, LaunchAgent labels, Gateway token/env handling, App-managed LaunchAgent handoff, openclaw update package/git handoff, Managed service refresh, Stale updater launchd job detection, openclaw uninstall, Stranded service recovery
|
||||||
|
- Diagnostics and Observability: LaunchAgent log paths, openclaw gateway status --deep, Gateway silently stops responding, Stale updater jobs
|
||||||
|
- Permissions and Native Capabilities: macOS TCC permission prompts/status, Native node capability exposure, system.run policy, Permission-driven support
|
||||||
|
- Profiles and Isolation: Profile-specific LaunchAgent labels, Profile-specific state/config/workspace roots, Derived ports, Rescue bot setup, Extra Gateway process detection
|
||||||
13
.agents/skills/claw-score/references/completeness/matrix.md
Normal file
13
.agents/skills/claw-score/references/completeness/matrix.md
Normal file
@@ -0,0 +1,13 @@
|
|||||||
|
# Matrix Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`matrix` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Matrix plugin identity, Setup wizard, Account discovery, Matrix doctor warnings, Matrix probe/status, Shared Matrix client resolution, Monitor startup, Startup maintenance, Matrix doctor warnings, Matrix probe/status, Monitor startup, Startup maintenance
|
||||||
|
- Access and Identity: DM policy, Direct-room classification, Inbound route selection across sender-bound DMs, Mention gates, Matrix thread reply routing, Persisted Matrix thread routing managers, ACP/subagent spawn hooks
|
||||||
|
- Conversation Routing and Delivery: DM policy, Direct-room classification, Inbound route selection across sender-bound DMs, Mention gates, Matrix thread reply routing, Persisted Matrix thread routing managers, ACP/subagent spawn hooks, Channel action discovery, Message send/read/edit/delete, Profile media loading, Outbound Matrix text, Message presentation metadata, Inbound media failure handling, Message send/read/edit/delete, Profile media loading, Outbound Matrix text, Message presentation metadata, Inbound media failure handling
|
||||||
|
- Media and Rich Content: Channel action discovery, Message send/read/edit/delete, Profile media loading, Outbound Matrix text, Message presentation metadata, Inbound media failure handling
|
||||||
|
- Native Controls and Approvals: Channel action discovery, Message send/read/edit/delete, Profile media loading, Outbound Matrix text, Message presentation metadata, Inbound media failure handling, Matrix native exec, Origin target resolution from Matrix turn, Approver DM target resolution, Matrix approval metadata, Origin target resolution from Matrix turn, Approver DM target resolution, Matrix approval metadata
|
||||||
|
- Encryption and Verification: Encryption setup, Encrypted media upload/download, Legacy state
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Mattermost, LINE, IRC, Nextcloud Talk, Nostr, Twitch, Tlon, Synology Chat Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`mattermost-line-irc-nextcloud-talk-nostr-twitch-tlon-synology-chat` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Mattermost bot account setup, WebSocket inbound monitoring, Outbound delivery, LINE Messaging API webhook setup, Signed inbound webhook events, Rich LINE payloads, Nextcloud Talk bot installation, Webhook ingress, Outbound markdown/text, Synology Chat incoming/outgoing webhook setup, Webhook token verification, Outbound text, IRC server/nick/TLS/NickServ setup, Raw IRC receive/send, Probe/status, Twitch bot account setup, Twitch IRC monitor/client lifecycle, Message tool send action, Nostr key setup, NIP-04 encrypted DM receive/send, Profile import/publish, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion, Nextcloud Talk bot installation, Webhook ingress, Outbound markdown/text, Synology Chat incoming/outgoing webhook setup, Webhook token verification, Outbound text and URL media delivery, Twitch bot account setup, Twitch IRC monitor/client lifecycle, Message tool send action, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion
|
||||||
|
- Access and Identity: Mattermost bot account setup, WebSocket inbound monitoring, Outbound delivery, LINE Messaging API webhook setup, Signed inbound webhook events, Rich LINE payloads, Nextcloud Talk bot installation, Webhook ingress, Outbound markdown/text, Synology Chat incoming/outgoing webhook setup, Webhook token verification, Outbound text, IRC server/nick/TLS/NickServ setup, Raw IRC receive/send, Probe/status, Twitch bot account setup, Twitch IRC monitor/client lifecycle, Message tool send action, Nostr key setup, NIP-04 encrypted DM receive/send, Profile import/publish, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion, Synology Chat incoming/outgoing webhook setup, Webhook token verification, Outbound text and URL media delivery, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion
|
||||||
|
- Conversation Routing and Delivery: Mattermost bot account setup, WebSocket inbound monitoring, Outbound delivery, LINE Messaging API webhook setup, Signed inbound webhook events, Rich LINE payloads, Nextcloud Talk bot installation, Webhook ingress, Outbound markdown/text, Synology Chat incoming/outgoing webhook setup, Webhook token verification, Outbound text, IRC server/nick/TLS/NickServ setup, Raw IRC receive/send, Probe/status, Twitch bot account setup, Twitch IRC monitor/client lifecycle, Message tool send action, Nostr key setup, NIP-04 encrypted DM receive/send, Profile import/publish, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion, Nextcloud Talk bot installation, Webhook ingress, Outbound markdown/text, Synology Chat incoming/outgoing webhook setup, Webhook token verification, Outbound text and URL media delivery, Twitch bot account setup, Twitch IRC monitor/client lifecycle, Message tool send action, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion
|
||||||
|
- Media and Rich Content: LINE Messaging API webhook setup, Signed inbound webhook events, Rich LINE payloads, Nextcloud Talk bot installation, Webhook ingress, Outbound markdown/text, Synology Chat incoming/outgoing webhook setup, Webhook token verification, Outbound text, Nostr key setup, NIP-04 encrypted DM receive/send, Profile import/publish, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion, Tlon/Urbit ship URL/code setup, Urbit API auth/session, Rich text conversion
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
# Media understanding and media generation Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`media-understanding-and-media-generation` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Media Intake and Access: Local and remote media references, MIME and type detection, Size caps and bounded reads, Safe remote fetch, Local root policy, Inbound media store, PDF/document extraction dispatch, QR and media helper classification
|
||||||
|
- Channel Media Handling: Inbound attachment staging, Sandbox media rewrites, Reply media templating, Message-tool attachment delivery, Duplicate delivery suppression
|
||||||
|
- Media Configuration: Media capability configuration
|
||||||
|
- Text-to-Speech Delivery: TTS, Outbound Voice Audio Delivery
|
||||||
|
- Media Understanding: Audio attachment selection, Batch STT provider and CLI fallback, Voice-note mention preflight, Transcript insertion and echo, Audio proxy and limit handling, Inbound image summarization, Active vision model bypass, Text-only model media offload, Vision provider fallback, Image and PDF input routing, Video Understanding, Direct Video Analysis
|
||||||
|
- Media Generation: Image generation tool invocation, Provider and model selection, Reference image editing, Generated image task lifecycle, Generated image persistence and delivery, Music generation tool invocation, Provider and model selection, Lyrics, instrumental, duration, and format controls, Reference inputs where supported, Music task lifecycle and duplicate status, Generated audio persistence and delivery, Video generation tool invocation, Mode and provider capability selection, Reference image, video, and audio inputs, Provider option validation, Video task lifecycle and status, Generated video persistence and delivery
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Microsoft Teams Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`microsoft-teams` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Teams CLI app creation, Bot registration and manifest upload, Credential configuration, Teams app install verification, Setup status, Probe and scope reporting, Teams app doctor, Webhook and health diagnostics, Operator repair paths, Text formatting and chunking, Adaptive and presentation cards, Progress streaming, Delivery receipts and errors, Queued and proactive replies, Webhook Runtime, SDK Lifecycle, Proactive Cloud Boundary, Setup status, Probe and scope reporting, Teams app doctor, Webhook and health diagnostics, Operator repair paths, Webhook Runtime, SDK Lifecycle, Proactive Cloud Boundary
|
||||||
|
- Access and Identity: DM pairing, Stable sender identity, Allowlists and access groups, Invoke and command authorization, Teams-originated config writes, Bot Framework SSO invokes, Delegated token storage, Graph directory lookup, Member profile lookup, Bot Framework SSO invokes, Delegated token storage, Graph directory lookup, Member profile lookup
|
||||||
|
- Conversation Routing and Delivery: Team and channel allowlists, Deterministic channel replies, Mention-gated group access, Session routing, Reply and thread context, Text formatting and chunking, Adaptive and presentation cards, Progress streaming, Delivery receipts and errors, Queued and proactive replies, Webhook Runtime, SDK Lifecycle, Proactive Cloud Boundary, Text formatting and chunking, Adaptive and presentation cards, Progress streaming, Delivery receipts and errors, Queued and proactive replies, Webhook Runtime, SDK Lifecycle, Proactive Cloud Boundary
|
||||||
|
- Media and Rich Content: Inbound attachments, Graph-hosted media, File consent, SharePoint and OneDrive sharing, Media fetch safety
|
||||||
|
- Native Controls and Approvals: Message action discovery, Polls and reactions, Read, edit, delete, and pin, Native approval cards, Feedback and group actions
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# Multi-Agent Orchestration Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`multi-agent-orchestration` surface.
|
||||||
|
|
||||||
|
## Surface-Specific Scoring Questions
|
||||||
|
|
||||||
|
For each category, ask:
|
||||||
|
|
||||||
|
- Can an operator configure and run the category workflow end to end?
|
||||||
|
- Are the taxonomy features present as supported user paths rather than partial config fragments?
|
||||||
|
- Are setup, normal operation, status or inspection, recovery, and removal paths represented where relevant?
|
||||||
|
- Are channel, account, workspace, auth, task, and delegate variants covered where the category expects them?
|
||||||
|
- Do known gaps leave major coordination or isolation branches missing?
|
||||||
|
|
||||||
|
## Surface-Specific Guidance
|
||||||
|
|
||||||
|
Variation from the default completeness process:
|
||||||
|
|
||||||
|
- Completeness is the operator-facing system for setup, isolation, conversation routing, account routing, specialist lanes, delegate identity, status, recovery, and safe defaults.
|
||||||
|
- A complete category lets multiple agents be created, isolated, routed, delegated, and inspected without implicit cross-agent leakage.
|
||||||
|
- Undocumented config, nondeterministic routing, or unclear ownership of state, credentials, and outbound delivery are material completeness gaps.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Agent Setup: add agents, agent list/delete, identity files, non-interactive setup, and single-agent default.
|
||||||
|
- Agent Isolation: workspace separation, state separation, auth separation, session separation, and tool profiles.
|
||||||
|
- Conversation Routing: agent selection, route precedence, default fallback, peer overrides, and cross-channel examples.
|
||||||
|
- Account Routing: multi-account setup, account selection, default accounts, account credentials, and delivery targets.
|
||||||
|
- Specialist Lanes: lane contracts, background handoff, concurrency controls, priority controls, and coordinator handoff.
|
||||||
|
- Delegate Identities: named delegates, authority model, delegate tiers, identity delegation, and organizational assistants.
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Native Windows CLI and Gateway Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`native-windows-cli-and-gateway` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Setup: PowerShell installer, Node and package-manager bootstrap, npm global install, Packaged CLI launcher, Windows command shims, openclaw onboard, Local Gateway config, Daemon install flags, Native-vs-WSL setup boundary
|
||||||
|
- Gateway Management: openclaw gateway, Foreground runtime health/readiness, Windows-specific restart/signal, Unmanaged foreground mode, openclaw gateway install, Gateway launcher files, Scheduled Task runtime status, Startup-folder fallback, openclaw status, Windows service inspection, Post-install diagnostics
|
||||||
|
- Networking: Native Windows host binding, netsh interface portproxy, Gateway status and probe output, Loopback, LAN, and WSL boundary
|
||||||
|
- Updates: openclaw update on native Windows package, Managed Gateway stop/restart, Detached update handoff, Windows package locks
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Native Windows companion app Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`native-windows-companion-app` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Installation and Updates: Official app download, MSI/MSIX/App Installer/winget-style packaging, Windows architecture handling for x64, App release channel
|
||||||
|
- Gateway Connection: App-managed local Gateway attach/start, Remote Gateway connection modes, Device/node pairing
|
||||||
|
- Chat Sessions: Native Windows chat window, Gateway chat transport
|
||||||
|
- Status and Repair: App health states, App-specific repair, Windows system tray app, Status indicators, App-specific notification permission
|
||||||
|
- Desktop Tools and Permissions: Windows node identity, Host command execution, Desktop command policy, App approval prompts, Screen and media capture, Canvas host behavior, Windows shell integrations, App secrets, Windows ACL, Command approval
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Nix install path Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`nix-install-path` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Install Handoff: Nix install overview, nix-openclaw source-of-truth, Install discoverability, Verification handoff
|
||||||
|
- Plugin Lifecycle: Lifecycle command refusal, Declarative plugin selection, Nix-store plugin loading, Hardlink safety
|
||||||
|
- Activation and App UX: Environment activation, macOS defaults activation, Runtime Nix-mode detection, Stable Nix defaults, Managed-by-Nix banner, Read-only config controls, Onboarding skip
|
||||||
|
- Config and State: Immutable config guard, Config writer refusal, Agent-first Nix edits, Explicit config path, Writable state directory, Immutable-store config support, State integrity checks
|
||||||
|
- Service Runtime and Guards: Nix profile PATH discovery, Profile precedence, Service PATH fallback, Trusted binary boundaries, Setup write refusal, Doctor repair refusal, Update handoff, Service lifecycle handoff
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# OpenAI / Codex provider path Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`openai-codex-provider-path` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Model and Auth: Canonical OpenAI Model Routing, Catalog, Codex OAuth Profiles, Subscription Usage, Doctor Diagnostics, Operator Repair
|
||||||
|
- Responses and Tool Compatibility: Codex Responses Transport, Payload Compatibility, Tool Context, Capability Compatibility
|
||||||
|
- Native Codex Harness: Native Codex App-server Harness, Thread Lifecycle
|
||||||
|
- Image and Multimodal Input: Image Generation Editing, Multimodal Input
|
||||||
|
- Voice and Realtime Audio: Realtime Voice Transcription, Speech
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# OpenClaw App SDK Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`openclaw-app-sdk` surface.
|
||||||
|
|
||||||
|
## Surface-Specific Scoring Questions
|
||||||
|
|
||||||
|
For each category, ask:
|
||||||
|
|
||||||
|
- Can an external app developer complete the category workflow using public SDK APIs?
|
||||||
|
- Are the taxonomy features represented by stable client contracts rather than protocol-only fragments?
|
||||||
|
- Are setup, authentication, streaming, result handling, error behavior, and compatibility expectations documented?
|
||||||
|
- Are browser, Node, React, testing, and custom transport variants covered where the category expects them?
|
||||||
|
- Do known gaps leave major external-app capability branches missing?
|
||||||
|
|
||||||
|
## Surface-Specific Guidance
|
||||||
|
|
||||||
|
Variation from the default completeness process:
|
||||||
|
|
||||||
|
- Completeness is the external app-developer workflow from connection through agent runs, sessions, events, approvals, resources, compatibility, and operational error handling.
|
||||||
|
- A complete SDK category exposes typed, documented, reusable client APIs instead of requiring low-level Gateway protocol work.
|
||||||
|
- Manual Gateway frame construction or reliance on internal package shapes is a material completeness gap.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Client API: SDK entrypoints, namespace layout, package split, and app/plugin boundary.
|
||||||
|
- Gateway Access: Gateway connect, URL and token config, auto gateway, custom transport, and scopes/redaction.
|
||||||
|
- Agent Conversations: agent handles, agent runs, run results, session creation, session send, and session controls.
|
||||||
|
- Events and Approvals: event stream, event envelope, replay cursors, approval callbacks, and questions.
|
||||||
|
- Resource Helpers: models, ToolSpace, artifacts, tasks, and environments.
|
||||||
|
- Compatibility: generated client, ergonomic wrappers, unsupported calls, schema alignment, and public package contract.
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# OpenRouter provider path Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`openrouter-provider-path` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Provider Setup and Auth: First-run setup, Default model selection, Provider plugin registration, Model-ref examples, OPENROUTER_API_KEY, Auth profiles and auth order, Status/probe and removal, Provider-entry SecretRef/API-key resolution, Gateway env inheritance, Static catalog rows, Dynamic /models discovery, openrouter/auto and nested refs, Free-model scan/probe, Model list/picker cache
|
||||||
|
- Chat Runtime and Normalization: Chat completions route, Provider routing params, Per-model route overrides, Reasoning payload policy, Anthropic/Gemini/DeepSeek variants, Streamed content parsing, reasoning_details visible output, Tool-call delta preservation, Family-specific replay policy, Response-model and usage normalization, Attribution headers, Response-cache headers/TTL/clear, Anthropic cache-control markers, Cache usage mapping, Custom proxy exclusions
|
||||||
|
- Provider Recovery and Diagnostics: Timeout/retry classification, Auth/billing/key-limit classification, Context overflow, Model fallback notices, Guarded fetch/pricing warnings
|
||||||
|
- Media Generation and Speech: image_generate OpenRouter route, video_generate async jobs/polling/download, music_generate audio route, Text-to-speech, Speech-to-text transcription, Inbound media understanding, Generated artifact delivery
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
# Plugin Surface Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`plugin-sdk-and-bundled-plugin-architecture` surface.
|
||||||
|
|
||||||
|
## Surface-Specific Scoring Questions
|
||||||
|
|
||||||
|
For each category, ask:
|
||||||
|
|
||||||
|
- Can the intended plugin task be completed end to end by an author or
|
||||||
|
operator?
|
||||||
|
- Are the important plugin variants present for this category, such as channel,
|
||||||
|
provider, tool, bundled, local, npm, or ClawHub flows?
|
||||||
|
- Are the main lifecycle stages present where relevant: create, configure,
|
||||||
|
validate, run, update, and remove or roll back?
|
||||||
|
- Are compatibility, approval, or safety branches present when the category
|
||||||
|
implies them?
|
||||||
|
- Are important author/operator-visible gaps still forcing workarounds or
|
||||||
|
unsupported paths?
|
||||||
|
|
||||||
|
## Surface-Specific Guidance
|
||||||
|
|
||||||
|
Variation from the default completeness process:
|
||||||
|
|
||||||
|
- Completeness is the plugin author or operator lifecycle for authoring, packaging, installing, running, approving, publishing, and testing plugins, not just SDK or runtime primitives.
|
||||||
|
- Score the plugin surface against the full plugin journey, not only one import path, packaging mode, or runtime path.
|
||||||
|
- Bundled-only support or support for only selected plugin families is incomplete when the category implies broader plugin capability.
|
||||||
|
- Publishing and testing categories should include expected lifecycle support, not just raw commands or fixtures.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Authoring and Packaging plugins: Root SDK entrypoint, Focused SDK imports, Entrypoint discovery, Migration shims, Plugin manifest, Package metadata, Runtime compatibility, Validation feedback
|
||||||
|
- Bundled plugins: Bundled plugin listing, Bundled source overlays, Packaged bundled plugins, Generated plugin inventory, Bundled channel IDs
|
||||||
|
- Canvas plugin: Hosted Canvas and A2UI surfaces, Agent canvas tool, Node Canvas commands, Control UI embeds, Canvas documents, A2UI transport and snapshots
|
||||||
|
- Installing and running plugins: Plugin setup, Runtime activation, Enable and disable, Safe load failures, Dependency repair, Install update and uninstall
|
||||||
|
- Channel plugins: Inbound event handling, Outbound delivery, Ingress authorization, Destination resolution, Native approval prompts
|
||||||
|
- Provider and tool plugins: Provider plugins, Tool plugins, Model catalogs, Provider auth, Web search and fetch, Mixed plugins
|
||||||
|
- Plugin approvals: Approval requests, Native approval delivery, Same-chat fallbacks, Exec and plugin separation, Approval replay protection, Security helpers
|
||||||
|
- Publishing plugins: Install sources, ClawHub publishing, npm publishing, Compatibility signaling, Update and rollback expectations, Third-party publication rules
|
||||||
|
- Testing plugins: Test fixtures, Local test environment, Plugin runtime harness, Unit and integration scaffolds, Docker lifecycle suites, Smoke tests
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Raspberry Pi / small Linux devices Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`raspberry-pi-small-linux-devices` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Setup and Compatibility: Hardware and 64-bit OS requirements, Node runtime setup, OpenClaw install and onboarding, First-run verification, Supported Pi model selection, 64-bit ARM boundary, Unsupported device guidance, Slow-device caveats, npm/pnpm/Bun install modes, Installer architecture detection, Optional ARM binary checks, Fallback/build guidance
|
||||||
|
- Remote Access and Auth: Headless API-key auth, Gateway shared-secret auth, Device pairing approvals, SecretRef handling, Token drift recovery, SSH tunnel dashboard access, Tailscale Serve/Funnel, Loopback/non-loopback exposure controls, Authenticated Control UI access
|
||||||
|
- Gateway Runtime: Always-on Gateway process, Cloud model configuration, Channel startup, Gateway health/status, User service install, linger/boot persistence, Service drop-ins, Restart tuning, Status/log inspection, Backup/restore
|
||||||
|
- Performance and Diagnostics: Swap and low-RAM tuning, USB SSD guidance, Compile cache/no-respawn settings, OOM/performance troubleshooting, Diagnostics bundles
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
# Security, auth, pairing, and secrets Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`security-auth-pairing-and-secrets` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Approval Policy and Tool Safeguards: Approval Policy, Dangerous Tool Safeguards
|
||||||
|
- Gateway Auth and Remote Access: Shared Gateway token/password auth, Gateway auth mode, Trusted-proxy identity, Tailscale Serve/Funnel, Bind and origin restrictions, WebSocket handshake auth, Operator-facing docs, Browser Control UI, Remote Client Trust
|
||||||
|
- Channel Access Control: Channel Identity, Allowlists, Sender Pairing
|
||||||
|
- Device and Node Pairing: Setup codes, Device identity creation, Device-token issuance, Device pairing approvals for operator, Operator scopes that gate pairing, Local Control UI, Auth migration, Operator-facing docs, Node Pairing, Capability Trust, Remote Exec Approvals
|
||||||
|
- Plugin Trust: Plugin Installation Trust, Security Boundaries
|
||||||
|
- Credential and Secret Hygiene: Provider Auth Profiles, API Key Health, Secrets Storage, Redaction, Configuration Hygiene
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
# Session, memory, and context engine Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`session-memory-and-context-engine` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- CLI Session and Transcript Management: CLI Session, Transcript Management
|
||||||
|
- Compaction, Pruning, and Token Pressure: Compaction, Pruning, Token Pressure
|
||||||
|
- Context Engine and Runtime Assembly: Context Engine, Runtime Assembly
|
||||||
|
- Cross-client History and Session Parity: Cross-client History, Session Parity
|
||||||
|
- Diagnostics, Maintenance, and Recovery: Diagnostics, Maintenance, Recovery
|
||||||
|
- Instruction Profile and Context Visibility: Instruction Profile, Context Visibility
|
||||||
|
- Memory Backend Storage and Embedding Search: Memory Backend Storage, Embedding Search
|
||||||
|
- Memory Files, Tools, and Active Memory: Memory Files, Tools, Active Memory
|
||||||
|
- Session Routing and Conversation Binding: Session Routing, Conversation Binding
|
||||||
|
- Transcript Persistence and Durability: Transcript Persistence, Durability
|
||||||
12
.agents/skills/claw-score/references/completeness/signal.md
Normal file
12
.agents/skills/claw-score/references/completeness/signal.md
Normal file
@@ -0,0 +1,12 @@
|
|||||||
|
# Signal Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`signal` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Setup and Account Health: QR link setup, SMS registration, Installer and binary setup, Container account provisioning, Status probes, Setup diagnostics, Account safety guardrails
|
||||||
|
- Conversation Access and Routing: DM pairing, DM allowlists, Sender identity normalization, Group allowlists, Mention gates, Pending group history
|
||||||
|
- Message Delivery and Actions: Text delivery targets, Media delivery and limits, Typing and read receipts, Styled/chunked output, Reaction action discovery, Add/remove reactions, Group reaction targeting
|
||||||
|
- Native Approvals: Native approval routing, Reaction approval responses, Approver targeting
|
||||||
|
- Transport: Native daemon transport, Container transport, API mode selection, Receive reconnect/readiness
|
||||||
12
.agents/skills/claw-score/references/completeness/slack.md
Normal file
12
.agents/skills/claw-score/references/completeness/slack.md
Normal file
@@ -0,0 +1,12 @@
|
|||||||
|
# Slack Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`slack` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: App Install, Slack app credentials, Manifest, Scopes, Channel status diagnostics, Slack account status, Operator Repair, Socket, HTTP transport, Runtime Lifecycle, Socket, HTTP transport, Runtime Lifecycle, Channel status diagnostics, Slack account status, Operator Repair
|
||||||
|
- Access and Identity: Channel allowlists, Thread routing, Session Isolation, DM Pairing, Sender Authorization
|
||||||
|
- Conversation Routing and Delivery: Channel allowlists, Thread routing, Session Isolation, DM Pairing, Sender Authorization, Outbound Delivery, Streaming, Reactions, Media, Attachments, Files, Vision, Outbound Delivery, Streaming, Reactions, Media, Attachments, Files, Vision
|
||||||
|
- Media and Rich Content: Outbound Delivery, Streaming, Reactions, Media, Attachments, Files, Vision
|
||||||
|
- Native Controls and Approvals: Slash Commands, Native Command Routing, Interactive Replies, App Home, Assistant Events, Native Approvals, Actions, Security-sensitive Ops, Interactive Replies, App Home, Assistant Events, Native Approvals, Actions, Security-sensitive Ops
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Telegram Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`telegram` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: BotFather token creation, TELEGRAM_BOT_TOKEN, Setup wizard credential capture, Startup getMe, Doctor/status surfacing, Named account configuration, CLI/message-tool targets, Directory adapters, Channel status, Account-scoped outbound, Long polling runner startup, Webhook listener startup, Reconnect, Restart, Named account configuration, Directory adapters and configured peers/groups for, Channel status, Account-scoped outbound, Long polling runner startup, Reconnect, Restart
|
||||||
|
- Access and Identity: dmPolicy modes, Pairing-code approval, Numeric Telegram user ID normalization with telegram, allowFrom, Unauthorized DM, Group allowlists, Supergroup negative chat IDs, Forum topic session keys, ACP topic routing, Session key construction
|
||||||
|
- Conversation Routing and Delivery: dmPolicy modes, Pairing-code approval, Numeric Telegram user ID normalization with telegram, allowFrom, Unauthorized DM, Group allowlists, Supergroup negative chat IDs, Forum topic session keys, ACP topic routing, Session key construction, Inbound media download, Voice notes, Location, Poll sending, Reactions, Text, Preview streaming, Reply threading tags, Durable outbound message recording, Voice notes, Poll sending, Reply threading tags, Durable outbound message recording
|
||||||
|
- Media and Rich Content: Inbound media download, Voice notes, Location, Poll sending, Reactions, Text, Preview streaming, Reply threading tags, Durable outbound message recording, Voice notes, Poll sending, Reply threading tags, Durable outbound message recording, Inbound media download, Voice notes, Location and venue extraction into channel context, Poll sending, Reactions
|
||||||
|
- Native Controls and Approvals: Inline keyboard rendering, Exec approvals in DMs, Message actions, Action capability discovery, Native setMyCommands startup sync, Command name/description normalization, Built-in commands, Command authorization in DMs, Model buttons, Native `setMyCommands` startup sync, Command name/description normalization, Built-in commands such as `/help`, Command authorization in DMs, Model buttons and command UI helpers
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Observability Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`telemetry-diagnostics-and-observability` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Health and Repair: Background health-monitor loop, Per-account enable/disable settings, Startup grace, Restart logging, openclaw doctor, Structured health checks, Core doctor checks, Plugin SDK doctor/health contracts, openclaw status, openclaw health, Gateway RPC health, Cached health snapshots
|
||||||
|
- Logging: Rolling Gateway JSONL file logs, openclaw logs, Gateway RPC logs.tail, Redaction patterns and sinks, Trace correlation fields
|
||||||
|
- Diagnostic Collection: openclaw gateway diagnostics export, openclaw gateway stability --bundle, Chat /diagnostics, Support zip composition, Bounded in-process stability recorder, openclaw gateway stability, Memory pressure events, Critical memory pressure snapshot option
|
||||||
|
- Telemetry Export: Diagnostic event types, Async dispatch, W3C trace context creation, Plugin SDK diagnostic runtime exports, Model-call diagnostic events, diagnostics-otel plugin install, OTLP/HTTP traces, Trusted trace context, Model and runtime telemetry, diagnostics-prometheus plugin install, Gateway-authenticated GET /api/diagnostics/prometheus, Prometheus text exposition, Trusted diagnostic event subscription
|
||||||
|
- Session Diagnostics: session.state, Diagnostic session activity snapshots, Model usage, Export of session signals to stability
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# TUI Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`tui-and-terminal-ux` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Runtime Modes: Gateway TUI launch, Local chat launch, Terminal alias launch, Initial message launch, Launch option validation, Gateway connection, Gateway authentication, History load on attach, Reconnect visibility, Gateway command RPCs, Embedded local chat, Local auth flow, Config repair loop, Gateway-free recovery
|
||||||
|
- Input and Commands: Message composition, Input history, Keyboard shortcuts, Paste and busy-submit handling, IME and AltGr handling, Slash Commands, Pickers, Settings
|
||||||
|
- Session Management: Session Lifecycle, History, Resume
|
||||||
|
- Local Shell Execution: Bang-command routing, Approval prompt, Command output display, Execution environment marker
|
||||||
|
- Rendering and Output Safety: Streaming Message Rendering, Tool Cards, Terminal Rendering Primitives, Output Safety
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
# Voice and realtime talk Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`voice-and-realtime-talk` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Talk Providers: OpenAI Realtime voice backend bridge, Google Gemini Live backend bridge, Realtime voice provider SDK contracts, Provider diagnostics, Talk catalog, Talk provider config, Shared native config parsing
|
||||||
|
- Realtime Talk Sessions: Agent consult handoff, Active Talk agent-run status, Talkback runtime behavior, Forced consult scheduling, Browser Talk start/stop UI, Browser WebRTC sessions, Browser relay mode, Browser tool-call forwarding, Realtime session controls, Gateway relay sessions, Audio-frame limits
|
||||||
|
- Speech and Transcription: Voice directives, Talk speech playback, Transcription relay sessions, Realtime transcription providers, Native directive parsing
|
||||||
|
- Native App Talk: macOS native Talk mode, iOS Talk mode, Android Talk mode, Shared Talk config
|
||||||
|
- Voice Wake and Routing: Wake-word settings, Wake routing, macOS Voice Wake runtime, Mobile wake preferences
|
||||||
|
- Talk Observability: Talk event logging, Session-log health, Live smoke output, Prometheus diagnostic counters, Operator visibility into setup
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Voice Call channel Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`voice-call-channel` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Voice Call Channel, Voice Call Channel, Voice Call Channel
|
||||||
|
- Access and Identity: Voice Call Channel
|
||||||
|
- Conversation Routing and Delivery: Voice Call Channel
|
||||||
|
- Media and Rich Content: Voice Call Channel, Voice Call Channel
|
||||||
|
- Realtime Voice and Calls: Voice Call Channel, Voice Call Channel, Voice Call Channel, Voice Call Channel, Voice Call Channel
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# watchOS companion surfaces Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`watchos-companion-surfaces` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Delivery and Recovery: APNs relay/direct registration as it affects, Silent push, Pending approval recovery IDs, Gateway-side iOS exec approval, iPhone-side WatchConnectivity transport, Watch-side receiver activation, Delivery fallback among reachable messages
|
||||||
|
- Exec Approvals: Watch exec approval prompt, Watch approval list/detail UI, iPhone-side prompt caching
|
||||||
|
- Distribution and Support: Watch app, Signing/profile variables, Public/support status, Changelog, Release metadata, Historical bug/regression themes relevant to scoring
|
||||||
|
- Notifications and Replies: watch.status, Payload normalization, Mirrored iOS notification fallback when watch, Watch action buttons from generic prompt, Watch-to-iPhone reply payloads, iPhone-side dedupe, Mirrored iOS notification action
|
||||||
|
- Watch App UI: Watch app entry point, Generic inbox, Persistent watch inbox state
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Web search tools Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`web-search-tools` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Search Providers: API-backed providers, Keyless and self-hosted providers, Provider comparison and auto-detection, Provider-specific filters and extraction, Result normalization, OpenAI native web_search, Codex native web_search, Gemini grounding, Grok web grounding, Kimi web search, Provider-native citations, Model and filter routing, webSearchProviders, registerWebSearchProvider, webFetchProviders, registerWebFetchProvider, public-artifact loading, runtime resolution, contract tests
|
||||||
|
- Setup and Diagnostics: Provider credentials, Default provider selection, Credential repair, Status checks, Quota errors, Cache controls, Provider diagnostics, Retry and fallback, Operator repair
|
||||||
|
- Network Safety: Network Safety, SSRF, Redirects, Untrusted Content
|
||||||
|
- Tool Availability and Fetch: web_search exposure, web_fetch exposure, x_search exposure, group:web policy, disabled-state diagnostics, provider/model gating, URL fetch, HTML extraction, PDF/text extraction, Safe truncation, Content citation handoff
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# WhatsApp Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`whatsapp` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- Channel Setup and Operations: Official @openclaw/whatsapp plugin metadata, openclaw plugin install whatsapp, Channel config schema, Baileys socket lifecycle, Operator troubleshooting, Baileys socket lifecycle, Operator troubleshooting for reconnect loops
|
||||||
|
- Access and Identity: QR login, Baileys multi-file auth persistence, DM pairing challenge, Multi-account/default-account resolution, Direct-message dmPolicy, Sender identity extraction, Privacy controls for plugin hooks, Direct-message `dmPolicy`, Sender identity extraction, Privacy controls for plugin hooks and
|
||||||
|
- Conversation Routing and Delivery: Group allowlists, Group session keys, Outbound text sends, Provider-accepted receipts, Outbound text sends, Provider-accepted receipts and durable delivery identifiers
|
||||||
|
- Media and Rich Content: Inbound media download, Outbound image
|
||||||
|
- Native Controls and Approvals: Native exec, Approver target resolution
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Windows via WSL2 Completeness
|
||||||
|
|
||||||
|
Use this rubric when assigning category Completeness scores for the
|
||||||
|
`windows-via-wsl2` surface.
|
||||||
|
|
||||||
|
## Category Scope
|
||||||
|
|
||||||
|
- WSL Setup and Updates: WSL2 + Ubuntu installation, Node runtime, Linux install flow inside WSL2, WSL2 runtime boundary, WSL2 network-family requirements, Source install and build inside WSL2, openclaw update, npm/pnpm/git package-root, Managed systemd Gateway restart, Service metadata refresh, Package-manager caveats
|
||||||
|
- Gateway Service Lifecycle: Onboarded systemd install, Gateway service install, systemd user unit rendering, WSL-aware systemd unavailable hints, Doctor service repair, WSL user-service linger, Systemd availability after Windows boot, Windows startup task for WSL, Verification before Windows sign-in, Clear expectations around PC power
|
||||||
|
- Gateway Access and Exposure: Gateway token/password auth, Provider credentials, Gateway auth SecretRefs, Remote URL credential precedence, WSL virtual network, Windows portproxy setup, Windows Firewall rules, Reachable Gateway URLs, Loopback and LAN exposure, WSL2 IPv4 networking, Tailscale remote access
|
||||||
|
- Diagnostics and Repair: openclaw doctor, openclaw status, openclaw logs, SecretRef, WSL/systemd unavailable hints, Operator repair guidance after WSL2 service
|
||||||
|
- Browser and Control UI: WSL2 Gateway with Windows browser, Windows Control UI URL, Raw remote CDP to Windows Chrome, Host-local Chrome MCP, Browser profile cdpUrl, Layered diagnostics
|
||||||
161
.agents/skills/clawdtributor/SKILL.md
Normal file
161
.agents/skills/clawdtributor/SKILL.md
Normal file
@@ -0,0 +1,161 @@
|
|||||||
|
---
|
||||||
|
name: clawdtributor
|
||||||
|
description: "Use for OpenClaw clawtributors PR/issue triage: Discrawl discovery, live-open rechecks, deep review, topic grouping, and compact @handle/LOC/type/blast/verification summaries."
|
||||||
|
---
|
||||||
|
|
||||||
|
# Clawdtributor
|
||||||
|
|
||||||
|
Use for the `#clawtributors` queue: Discord-discovered OpenClaw PRs/issues that need live GitHub status plus maintainer-quality review.
|
||||||
|
|
||||||
|
## Compose with other skills
|
||||||
|
|
||||||
|
- `$discrawl`: local Discord archive sync/search.
|
||||||
|
- `$openclaw-pr-maintainer`: live GitHub PR/issue review, duplicate search, close/land rules.
|
||||||
|
- `$gitcrawl`: related issue/PR and current-main/stale-proof search.
|
||||||
|
- `$openclaw-testing` / `$crabbox`: proof choice when a candidate needs real validation.
|
||||||
|
|
||||||
|
## Archive flow
|
||||||
|
|
||||||
|
Local archive first; verify freshness for current questions.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
discrawl status --json
|
||||||
|
discrawl sync
|
||||||
|
```
|
||||||
|
|
||||||
|
Resolve channel if needed:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sqlite3 "$HOME/.discrawl/discrawl.db" \
|
||||||
|
"select id,name from channels where name like '%clawtributor%' order by name;"
|
||||||
|
```
|
||||||
|
|
||||||
|
Current known channel id from prior work: `1458141495701012561`. Re-resolve if it stops matching.
|
||||||
|
|
||||||
|
Extract recent refs:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sqlite3 "$HOME/.discrawl/discrawl.db" "
|
||||||
|
select m.created_at, coalesce(nullif(mm.username,''), m.author_id), m.content
|
||||||
|
from messages m
|
||||||
|
left join members mm on mm.guild_id=m.guild_id and mm.user_id=m.author_id
|
||||||
|
where m.channel_id='1458141495701012561'
|
||||||
|
and m.created_at >= '<ISO cutoff>'
|
||||||
|
order by m.created_at desc;" |
|
||||||
|
perl -nE 'while(m{github\.com/openclaw/openclaw/(pull|issues)/(\d+)}g){say "$1\t$2\t$_"}'
|
||||||
|
```
|
||||||
|
|
||||||
|
Map a PR/issue back to the Discord handle:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sqlite3 -separator $'\t' "$HOME/.discrawl/discrawl.db" "
|
||||||
|
select m.created_at,
|
||||||
|
coalesce(nullif(mm.username,''), nullif(mm.global_name,''), m.author_id)
|
||||||
|
from messages m
|
||||||
|
left join members mm on mm.guild_id=m.guild_id and mm.user_id=m.author_id
|
||||||
|
where m.channel_id='1458141495701012561'
|
||||||
|
and m.content like '%github.com/openclaw/openclaw/<pull-or-issues>/<number>%'
|
||||||
|
order by m.created_at desc
|
||||||
|
limit 1;"
|
||||||
|
```
|
||||||
|
|
||||||
|
Show only `@handle` in the final list. Do not write the word Discord unless the user asks for source details.
|
||||||
|
|
||||||
|
## Live GitHub recheck
|
||||||
|
|
||||||
|
Always recheck live state before listing, closing, or saying "open".
|
||||||
|
|
||||||
|
```bash
|
||||||
|
GITHUB_TOKEN= GITHUB_TOKEN_NODIFF= GH_TOKEN= \
|
||||||
|
gh api repos/openclaw/openclaw/pulls/<number> \
|
||||||
|
--jq '. | {number,title,state,merged,mergeable,draft,author:.user.login,url:.html_url,updatedAt:.updated_at,additions,deletions,changedFiles:.changed_files}'
|
||||||
|
```
|
||||||
|
|
||||||
|
For issues:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
GITHUB_TOKEN= GITHUB_TOKEN_NODIFF= GH_TOKEN= \
|
||||||
|
gh api repos/openclaw/openclaw/issues/<number> \
|
||||||
|
--jq '. | {number,title,state,author:.user.login,url:.html_url,updatedAt:.updated_at,pull_request}'
|
||||||
|
```
|
||||||
|
|
||||||
|
If `gh` says bad credentials, clear env vars with empty assignments as above. Use `--jq '. | {...}'` for object projections.
|
||||||
|
|
||||||
|
## Review depth
|
||||||
|
|
||||||
|
For each open item, inspect enough to classify risk:
|
||||||
|
|
||||||
|
- PR body, linked issue, comments, files, additions/deletions, checks.
|
||||||
|
- Current `origin/main` code path and adjacent tests.
|
||||||
|
- Related threads with `gitcrawl neighbors/search`.
|
||||||
|
- Whether main already fixed it, the PR is obsolete, or the idea is invalid.
|
||||||
|
- Blast radius: touched runtime surfaces, config/schema, plugin/core boundary, user-visible behavior, release/package surface.
|
||||||
|
- Verification: say if local unit/docs proof is enough, live/provider proof is needed, or it is not directly verifiable.
|
||||||
|
|
||||||
|
Do not close from title alone. If closing as done on main or nonsensical, prove it against current main and comment first when mutation is requested. Bulk close/reopen above 5 requires explicit scope.
|
||||||
|
|
||||||
|
## Candidate selection
|
||||||
|
|
||||||
|
When asked for `5 new`, exclude refs already surfaced in the session and refill from the archive until there are 5 live-open candidates. If fewer than 5 remain open, list all open ones and say how many short.
|
||||||
|
|
||||||
|
When asked to `update`, `refresh`, `recheck`, `check again`, or similar, return an updated live-open candidate list. Sort by maintainer importance, not recency: high-impact ready fixes first, then useful-but-review-first, then open/not-ready items. Do not include a "changed since last pass" section or bottom-line merged/closed summary unless the user explicitly asks for churn.
|
||||||
|
|
||||||
|
Prefer:
|
||||||
|
|
||||||
|
- Fresh, open, external contributor work.
|
||||||
|
- Small, high-confidence bugfixes.
|
||||||
|
- Clear repro, tests, or obvious code-path proof.
|
||||||
|
|
||||||
|
Demote:
|
||||||
|
|
||||||
|
- Broad product/features without owner decision.
|
||||||
|
- Large rewrites with unclear contract.
|
||||||
|
- PRs already in progress, merged, closed, duplicate, or fixed on main.
|
||||||
|
|
||||||
|
## Topic grouping
|
||||||
|
|
||||||
|
Group only when useful or requested:
|
||||||
|
|
||||||
|
- Agents/tooling
|
||||||
|
- Providers/auth/models
|
||||||
|
- Channels/messaging
|
||||||
|
- UI/web
|
||||||
|
- Gateway/protocol/runtime
|
||||||
|
- Config/memory/cache
|
||||||
|
- Docker/install/release
|
||||||
|
- Docs/tests/chore
|
||||||
|
- Closed/obsolete
|
||||||
|
|
||||||
|
Infer topic from labels, touched files, title/body, and actual code path.
|
||||||
|
|
||||||
|
## Output format
|
||||||
|
|
||||||
|
No Markdown tables. Compact bullets. Use color/risk markers:
|
||||||
|
|
||||||
|
- 🟢 low/narrow
|
||||||
|
- 🟡 medium or needs targeted proof
|
||||||
|
- 🔴 broad/high runtime risk
|
||||||
|
- 🟣 security/policy/owner-boundary slow review
|
||||||
|
- ✅ merged
|
||||||
|
- ⚪ closed unmerged
|
||||||
|
|
||||||
|
Required line shape:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
- **PR #81244** `@whatsskill.` `+118/-1` `bug` 🟢 https://github.com/openclaw/openclaw/pull/81244 - Prevents chat action buttons from overlapping short assistant replies. Verifiable: yes. Blast: web chat rendering, low.
|
||||||
|
- **Issue #81245** `@alice` `LOC n/a` `bug` 🟡 https://github.com/openclaw/openclaw/issues/81245 - Reports duplicate Telegram replies when reconnecting after gateway restart. Verifiable: partial. Blast: Telegram channel runtime, medium.
|
||||||
|
```
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
|
||||||
|
- Bold the `PR #n` or `Issue #n` marker.
|
||||||
|
- Use `@handle`, not author bio text.
|
||||||
|
- Always include the full GitHub URL.
|
||||||
|
- Include a one-line description after the URL, separated with `-`.
|
||||||
|
- PR LOC is `+additions/-deletions`; issue LOC is `LOC n/a`.
|
||||||
|
- Type: `bug`, `feature`, `perf`, `security`, `docs`, `test`, `chore`, or `refactor`.
|
||||||
|
- Write a full sentence for what it does.
|
||||||
|
- Always include blast radius in one phrase.
|
||||||
|
- Always include `verifiable: yes|partial|no` plus the shortest proof hint when helpful.
|
||||||
|
- If status is not open, still show it only when the user asked for all surfaced refs; use ✅ or ⚪ and state merged/closed.
|
||||||
|
- For refresh-style asks, prefer section order: `Best Open Now`, `Useful But Review First`, `Still Open / Not Ready`. Omit merged/closed churn by default.
|
||||||
74
.agents/skills/control-ui-e2e/SKILL.md
Normal file
74
.agents/skills/control-ui-e2e/SKILL.md
Normal file
@@ -0,0 +1,74 @@
|
|||||||
|
---
|
||||||
|
name: control-ui-e2e
|
||||||
|
description: Use when testing, fixing, or extending the OpenClaw Control UI GUI with Vitest + Playwright end-to-end checks, mocked Gateway WebSocket flows, mocked dashboard runs, screenshots/videos, or agent-verifiable browser proof.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Control UI E2E
|
||||||
|
|
||||||
|
Use this for Control UI changes that need a real browser flow with deterministic Gateway data.
|
||||||
|
|
||||||
|
## Test Shape
|
||||||
|
|
||||||
|
- Use `ui/src/**/*.e2e.test.ts` for full GUI flows.
|
||||||
|
- Use `ui/src/test-helpers/control-ui-e2e.ts` to start the Vite Control UI and install a mocked Gateway WebSocket.
|
||||||
|
- Keep scenarios deterministic. Do not use live provider keys, real channel credentials, or a real Gateway unless the user explicitly asks for live proof.
|
||||||
|
- Prefer existing `.browser.test.ts` or unit tests for narrow rendering logic; use this E2E lane when the proof should cover routing, app boot, Gateway handshake, requests, and visible UI behavior together.
|
||||||
|
|
||||||
|
## Commands
|
||||||
|
|
||||||
|
- Target one E2E test in a Codex worktree:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node scripts/run-vitest.mjs run --config test/vitest/vitest.ui-e2e.config.ts --configLoader runner ui/src/ui/e2e/chat-flow.e2e.test.ts
|
||||||
|
```
|
||||||
|
|
||||||
|
- Run the whole local lane in a normal checkout:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm test:ui:e2e
|
||||||
|
```
|
||||||
|
|
||||||
|
If dependencies are missing in a Codex worktree, install once with `pnpm install`; for broad GUI proof or dependency-heavy checks, use Testbox/Crabbox instead of running a wide local pnpm lane.
|
||||||
|
|
||||||
|
## Visual Proof Default
|
||||||
|
|
||||||
|
When running mocked Control UI/dashboard validation for a user-facing feature, produce visual proof by default unless the user explicitly opts out.
|
||||||
|
|
||||||
|
- Keep the Vitest E2E assertions deterministic; do not commit generated screenshots or videos.
|
||||||
|
- After or alongside the focused E2E test, run the mocked Control UI app when available, for example `pnpm dev:ui:mock -- --port <port>`.
|
||||||
|
- Drive Chromium with Playwright against the local mock URL and capture a video plus screenshots for each meaningful state: initial view, interaction input, result state, and final/paginated/selected state.
|
||||||
|
- Use `browser.newContext({ recordVideo: { dir, size }, viewport })`, `page.screenshot({ path })`, and close the context before reporting the video path.
|
||||||
|
- Put artifacts under `.artifacts/control-ui-e2e/<short-feature-name>/` or another clearly named local temp directory, and report the absolute paths in the final answer.
|
||||||
|
- Treat recording as validation, not only demo capture. If the recorder fails or shows surprising behavior, stop, fix the behavior, add or update a regression test, then rerecord.
|
||||||
|
- If visual proof is blocked, state the exact blocker and still report the textual E2E evidence.
|
||||||
|
|
||||||
|
## Mock Pattern
|
||||||
|
|
||||||
|
Start the app server, install the mock before `page.goto`, then assert both Gateway traffic and visible UI:
|
||||||
|
|
||||||
|
```ts
|
||||||
|
const server = await startControlUiE2eServer();
|
||||||
|
const page = await context.newPage();
|
||||||
|
const gateway = await installMockGateway(page, {
|
||||||
|
historyMessages: [{ role: "assistant", content: [{ type: "text", text: "Ready." }] }],
|
||||||
|
});
|
||||||
|
|
||||||
|
await page.goto(`${server.baseUrl}chat`);
|
||||||
|
await page.locator(".agent-chat__composer-combobox textarea").fill("hello");
|
||||||
|
await page.getByRole("button", { name: "Send message" }).click();
|
||||||
|
|
||||||
|
const request = await gateway.waitForRequest("chat.send");
|
||||||
|
await gateway.emitChatFinal({ runId: String(request.params.idempotencyKey), text: "Done." });
|
||||||
|
await page.getByText("Done.").waitFor();
|
||||||
|
```
|
||||||
|
|
||||||
|
Extend `installMockGateway` with typed scenario options or method responses when a new flow needs more Gateway surface.
|
||||||
|
|
||||||
|
## Standalone Recording
|
||||||
|
|
||||||
|
When recording an already-running mocked Control UI URL, use a temporary Playwright script or `playwright test` spec and keep the recording flow focused:
|
||||||
|
|
||||||
|
- Open the mock URL, interact through stable `data-*` selectors or user-facing role selectors, and wait on asserted states instead of relying on fixed sleeps.
|
||||||
|
- Assert both visible UI state and mocked Gateway traffic for request-driven flows. For example, verify the expected count/row is visible and that `sessions.list` was called with the expected `search`, `offset`, and `limit`.
|
||||||
|
- Use short sleeps only after assertions to make the captured video readable.
|
||||||
|
- Store the generated video under `.artifacts/control-ui-e2e/<feature>/`; do not commit it.
|
||||||
4
.agents/skills/control-ui-e2e/agents/openai.yaml
Normal file
4
.agents/skills/control-ui-e2e/agents/openai.yaml
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Control UI E2E"
|
||||||
|
short_description: "Mocked browser E2E for Control UI"
|
||||||
|
default_prompt: "Use $control-ui-e2e to verify a Control UI change with the mocked Vitest + Playwright browser lane."
|
||||||
828
.agents/skills/crabbox/SKILL.md
Normal file
828
.agents/skills/crabbox/SKILL.md
Normal file
@@ -0,0 +1,828 @@
|
|||||||
|
---
|
||||||
|
name: crabbox
|
||||||
|
description: Use the Crabbox wrapper for OpenClaw remote validation across Linux, macOS, Windows, and WSL2, including delegated Blacksmith Testbox proof. Report the actual provider and id.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Crabbox
|
||||||
|
|
||||||
|
OpenClaw agent sessions use the Crabbox wrapper by default for tests and
|
||||||
|
computationally intensive work: builds, typechecks, lint fan-out, broad gates,
|
||||||
|
CI-parity checks, secrets, hosted services, Docker/E2E/package lanes, warmed
|
||||||
|
reusable boxes, sync timing, logs/results, cache inspection, and lease cleanup.
|
||||||
|
|
||||||
|
Crabbox is the transport/orchestration surface. The actual backend can be:
|
||||||
|
|
||||||
|
- brokered AWS Crabbox: direct provider, `provider=aws`, lease ids like
|
||||||
|
`cbx_...`, `syncDelegated=false`
|
||||||
|
- Blacksmith Testbox through Crabbox: delegated provider,
|
||||||
|
`provider=blacksmith-testbox`, ids like `tbx_...`, `syncDelegated=true`
|
||||||
|
|
||||||
|
Blacksmith Testbox through the Crabbox wrapper is the default OpenClaw agent
|
||||||
|
backend for trusted maintainer code and heavy `pnpm` gates. The configured
|
||||||
|
Blacksmith workflow hydrates provider and agent credentials, so never sync or
|
||||||
|
run untrusted contributor/fork code there. Use secretless fork CI or
|
||||||
|
sanitized direct AWS Crabbox for untrusted source. Do not describe
|
||||||
|
Blacksmith runs as "AWS Crabbox"; report them as Testbox-through-Crabbox with
|
||||||
|
the `tbx_...` id and Actions run.
|
||||||
|
|
||||||
|
Pass `--provider aws` when the task specifically needs direct AWS Crabbox
|
||||||
|
behavior, persistent direct-provider leases, `--fresh-pr`, `--full-resync`,
|
||||||
|
environment forwarding, capture/download support, or provider comparison. Use
|
||||||
|
`--provider blacksmith-testbox` for the default OpenClaw agent path.
|
||||||
|
|
||||||
|
## First Checks
|
||||||
|
|
||||||
|
- Run from the repo root. Crabbox sync mirrors the current checkout.
|
||||||
|
- Check the wrapper and providers before remote work:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
command -v crabbox
|
||||||
|
../crabbox/bin/crabbox --version
|
||||||
|
pnpm crabbox:run -- --help | sed -n '1,120p'
|
||||||
|
../crabbox/bin/crabbox desktop launch --help
|
||||||
|
../crabbox/bin/crabbox webvnc --help
|
||||||
|
```
|
||||||
|
|
||||||
|
- OpenClaw scripts prefer `../crabbox/bin/crabbox` when present. The user PATH
|
||||||
|
shim can be stale.
|
||||||
|
- Check `.crabbox.yaml` for the provider default. Omitting `--provider`
|
||||||
|
means Blacksmith Testbox through Crabbox for normal Linux paths; the wrapper
|
||||||
|
selects Azure for unqualified Windows/WSL2 runs when the local Crabbox
|
||||||
|
binary advertises Azure. Pass `--provider aws` for direct brokered AWS runs.
|
||||||
|
- The brokered AWS image is a Linux developer image in `eu-west-1`; the repo
|
||||||
|
config pins hot `eu-west-1a/b/c` placement so Fast Snapshot Restore can apply.
|
||||||
|
If warmup drifts well past the minute-scale path, verify image promotion,
|
||||||
|
region/AZ placement, and FSR state before blaming OpenClaw.
|
||||||
|
- For trusted OpenClaw agent tests and computationally intensive work, use the
|
||||||
|
repo wrapper with `--provider blacksmith-testbox` or the repo Testbox helpers.
|
||||||
|
- Treat contributor/fork source as untrusted unless a maintainer explicitly
|
||||||
|
approves credentialed execution after review. Run untrusted source only in
|
||||||
|
secretless fork CI or sanitized direct AWS Crabbox. For every untrusted AWS
|
||||||
|
run, launch an installed trusted Crabbox binary from a clean trusted `main`
|
||||||
|
checkout and fetch the remote PR with `--fresh-pr`; never execute the
|
||||||
|
untrusted checkout's wrapper or config locally. Set
|
||||||
|
`CRABBOX_ENV_ALLOW=CI` to replace the repo's `OPENCLAW_*`/`NODE_OPTIONS`
|
||||||
|
allowlist, pass `--provider aws --no-hydrate`, and use a fresh temporary
|
||||||
|
remote `HOME` on a newly warmed lease dedicated to that untrusted source.
|
||||||
|
Unset `CRABBOX_AWS_INSTANCE_PROFILE` and fail closed unless resolved
|
||||||
|
`aws.instanceProfile` is empty. Before install/test, use trusted absolute-path
|
||||||
|
tools to require an IMDSv2 token, prove the IAM credentials endpoint returns
|
||||||
|
404, and verify remote `git rev-parse HEAD` equals the full reviewed PR head
|
||||||
|
SHA. Bind the lease to that SHA; stop and rewarm when the head changes. Do not
|
||||||
|
inherit Tailscale: unset every `CRABBOX_TAILSCALE*` override, force
|
||||||
|
`--network public --tailscale=false`, clear exit-node/LAN flags, and require
|
||||||
|
`crabbox inspect` to report public networking with no Tailscale state before
|
||||||
|
uploading any script. Execute PR code only through trusted
|
||||||
|
`scripts/crabbox-untrusted-bootstrap.sh`, uploaded from clean `main` alongside
|
||||||
|
`--fresh-pr`; it installs pinned Node/pnpm and rejects a changed PR
|
||||||
|
`packageManager` pin before install. Never reuse a trusted or previously
|
||||||
|
hydrated lease. If the broker cannot provide
|
||||||
|
the no-role proof or no remote PR exists, use secretless fork CI. Never use
|
||||||
|
`hydrate-github` or a credential-hydrated Testbox workflow for untrusted code.
|
||||||
|
- Cold Testbox acquisition and hydration often take about a minute. At the
|
||||||
|
start of any task likely to change code or need tests/heavy proof, immediately
|
||||||
|
start, after confirming the source is trusted,
|
||||||
|
`node scripts/crabbox-wrapper.mjs warmup --provider blacksmith-testbox --keep --timing-json`
|
||||||
|
in a background command session while inspecting and editing. Poll later,
|
||||||
|
reuse the returned `tbx_...` with
|
||||||
|
`--provider blacksmith-testbox --id <tbx_id>`, and stop it before handoff.
|
||||||
|
For untrusted source, switch to a clean trusted `main` checkout and pre-warm
|
||||||
|
with the installed binary after the empty-instance-profile check below.
|
||||||
|
Do not warm for read-only, docs-only, or clearly trivial work that will not
|
||||||
|
run tests or heavy commands.
|
||||||
|
- Run untrusted source only with the sanitized form below. The explicit
|
||||||
|
allowlist prevents locally exported `OPENCLAW_*` credentials from crossing
|
||||||
|
the SSH boundary; `--no-hydrate` and temporary `HOME` prevent auth-profile
|
||||||
|
reuse:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
crabbox config show --json | \
|
||||||
|
jq -e '.aws.instanceProfile == ""' >/dev/null
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
-u CRABBOX_TAILSCALE \
|
||||||
|
-u CRABBOX_TAILSCALE_AUTH_KEY \
|
||||||
|
-u CRABBOX_TAILSCALE_AUTH_KEY_ENV \
|
||||||
|
-u CRABBOX_TAILSCALE_EXIT_NODE \
|
||||||
|
-u CRABBOX_TAILSCALE_EXIT_NODE_ALLOW_LAN_ACCESS \
|
||||||
|
-u CRABBOX_TAILSCALE_HOSTNAME_TEMPLATE \
|
||||||
|
-u CRABBOX_TAILSCALE_TAGS \
|
||||||
|
crabbox warmup \
|
||||||
|
--provider aws \
|
||||||
|
--network public \
|
||||||
|
--tailscale=false \
|
||||||
|
--tailscale-exit-node= \
|
||||||
|
--tailscale-exit-node-allow-lan-access=false \
|
||||||
|
--keep \
|
||||||
|
--timing-json
|
||||||
|
crabbox inspect --provider aws --id <cbx_id> --json | \
|
||||||
|
jq -e '.network == "public" and .tailscale == null' >/dev/null
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
CRABBOX_ENV_ALLOW=CI \
|
||||||
|
crabbox run \
|
||||||
|
--provider aws \
|
||||||
|
--id <cbx_id> \
|
||||||
|
--fresh-pr <owner/repo#number> \
|
||||||
|
--no-hydrate \
|
||||||
|
--timing-json \
|
||||||
|
--script scripts/crabbox-untrusted-bootstrap.sh -- \
|
||||||
|
<expected_head_sha> /usr/local/bin/pnpm test <path>
|
||||||
|
# After all proof:
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
crabbox stop --provider aws <cbx_id>
|
||||||
|
```
|
||||||
|
|
||||||
|
- Always report the actual provider and id. `cbx_...` means AWS Crabbox;
|
||||||
|
`tbx_...` means Blacksmith Testbox through Crabbox. If the output only says
|
||||||
|
`blacksmith testbox list`, use `blacksmith testbox list --all` before
|
||||||
|
concluding no box exists.
|
||||||
|
- If a warm direct-provider lease smells stale, retry with `--full-resync`
|
||||||
|
(alias `--fresh-sync`) before replacing the lease. This resets the remote
|
||||||
|
workdir, skips the fingerprint fast path, reseeds Git when possible, and
|
||||||
|
uploads the checkout from scratch.
|
||||||
|
- For live/provider bugs, use the configured secret workflow before downgrading
|
||||||
|
to mocks. Copy only the exact needed key into the remote process environment
|
||||||
|
for that one command. Do not print it, do not sync it as a repo file, and do
|
||||||
|
not leave it in remote shell history or logs. If no secret-safe injection path
|
||||||
|
is available, say true live provider auth is blocked instead of silently using
|
||||||
|
a fake key.
|
||||||
|
- Agent-run tests, including targeted edit-loop tests, default to a pre-warmed
|
||||||
|
remote box selected by source trust. Local test execution requires an
|
||||||
|
explicit user request or a reported remote-provider blocker.
|
||||||
|
- Do not treat inherited shell env as operator intent. In particular,
|
||||||
|
`OPENCLAW_LOCAL_CHECK_MODE=throttled` from the local shell is not permission
|
||||||
|
to move broad `pnpm check:changed`, `pnpm test:changed`, full `pnpm test`, or
|
||||||
|
lint/typecheck fan-out onto the laptop.
|
||||||
|
- Only use `OPENCLAW_LOCAL_CHECK_MODE=throttled|full` when the user explicitly
|
||||||
|
asks for local proof in the current task. If Testbox is queued or capacity is
|
||||||
|
constrained, report the blocker; do not silently move heavy work onto the
|
||||||
|
laptop.
|
||||||
|
|
||||||
|
## macOS And Windows Targets
|
||||||
|
|
||||||
|
Use these only when the task needs an existing non-Linux host. OpenClaw broad
|
||||||
|
Linux validation uses the repo Crabbox config unless a provider is explicitly
|
||||||
|
requested.
|
||||||
|
|
||||||
|
Native brokered Windows is available for Windows-specific proof. Prefer Azure
|
||||||
|
for Windows/WSL2 when the subscription has quota or credits and the local
|
||||||
|
Crabbox binary advertises Azure. Keep broad Linux gates on Linux/Testbox unless
|
||||||
|
the bug is Windows-specific, and only force AWS when the operator asks for the
|
||||||
|
older AWS developer image/cache path or Azure is unavailable:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:warmup -- \
|
||||||
|
--target windows \
|
||||||
|
--windows-mode wsl2 \
|
||||||
|
--timing-json
|
||||||
|
```
|
||||||
|
|
||||||
|
The hydrate workflow assumes Docker should already be baked into Linux images
|
||||||
|
and only installs it as a fallback. Do not add per-run Docker installs to proof
|
||||||
|
commands unless the image probe shows Docker is actually missing.
|
||||||
|
|
||||||
|
When the user explicitly asks for brokered macOS runners, use Crabbox AWS
|
||||||
|
macOS only after confirming the deployed coordinator supports EC2 Mac host
|
||||||
|
lifecycle/image routes and the operator has AWS EC2 Mac Dedicated Host quota
|
||||||
|
and IAM. Prefer `CRABBOX_HOST_ID` for a known Crabbox-managed Dedicated Host,
|
||||||
|
or run the no-spend preflight first:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
crabbox admin hosts quota --provider aws --target macos --region eu-west-1 --type mac2.metal --json
|
||||||
|
crabbox admin hosts allocate --provider aws --target macos --region eu-west-1 --type mac2.metal --dry-run --json
|
||||||
|
CRABBOX_MACOS_TYPES=all scripts/macos-host-region-preflight.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
Do not silently substitute AWS macOS for normal OpenClaw Linux proof. Report
|
||||||
|
paid-host blockers as quota, IAM, coordinator deployment, or host availability
|
||||||
|
instead of falling back to local macOS.
|
||||||
|
|
||||||
|
Crabbox supports static SSH targets:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
../crabbox/bin/crabbox run --provider ssh --target macos --static-host mac-studio.local -- xcodebuild test
|
||||||
|
../crabbox/bin/crabbox run --provider ssh --target windows --windows-mode normal --static-host win-dev.local -- pwsh -NoProfile -Command "dotnet test"
|
||||||
|
../crabbox/bin/crabbox run --provider ssh --target windows --windows-mode wsl2 --static-host win-dev.local -- pnpm test
|
||||||
|
```
|
||||||
|
|
||||||
|
- `target=macos` and `target=windows --windows-mode wsl2` use the POSIX SSH,
|
||||||
|
bash, Git, rsync, and tar contract.
|
||||||
|
- Native Windows uses OpenSSH, PowerShell, Git, and tar; sync is manifest tar
|
||||||
|
archive transfer into `static.workRoot`. Direct native Windows runs support
|
||||||
|
`--script*`, `--env-from-profile`, `--preflight`, and PowerShell `--shell`.
|
||||||
|
- `crabbox actions hydrate/register` are Linux-only today; use plain
|
||||||
|
`crabbox run` loops for static macOS and Windows hosts.
|
||||||
|
- Live proof needs a reachable, operator-managed SSH host. Without one, verify
|
||||||
|
with `../crabbox/bin/crabbox run --help`, config/flag tests, and the Crabbox
|
||||||
|
Go test suite.
|
||||||
|
|
||||||
|
## Direct Brokered AWS Backend
|
||||||
|
|
||||||
|
Use this when the task needs direct AWS Crabbox semantics rather than the
|
||||||
|
prepared Blacksmith Testbox CI environment.
|
||||||
|
|
||||||
|
Changed gate:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:run -- \
|
||||||
|
--provider aws \
|
||||||
|
--idle-timeout 90m \
|
||||||
|
--ttl 240m \
|
||||||
|
--timing-json \
|
||||||
|
--shell -- \
|
||||||
|
"pnpm test:changed"
|
||||||
|
```
|
||||||
|
|
||||||
|
Full suite:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:run -- \
|
||||||
|
--provider aws \
|
||||||
|
--idle-timeout 90m \
|
||||||
|
--ttl 240m \
|
||||||
|
--timing-json \
|
||||||
|
--shell -- \
|
||||||
|
"pnpm verify"
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `pnpm verify` when you need check plus full Vitest proof. It emits
|
||||||
|
`CRABBOX_PHASE:check` and `CRABBOX_PHASE:test`, making Crabbox summaries show
|
||||||
|
which stage failed. Use plain `pnpm test` only when check proof is already
|
||||||
|
covered or intentionally skipped.
|
||||||
|
|
||||||
|
Focused rerun:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:run -- \
|
||||||
|
--provider aws \
|
||||||
|
--idle-timeout 90m \
|
||||||
|
--ttl 240m \
|
||||||
|
--timing-json \
|
||||||
|
--shell -- \
|
||||||
|
"pnpm test <path-or-filter>"
|
||||||
|
```
|
||||||
|
|
||||||
|
Read the JSON summary. Useful fields:
|
||||||
|
|
||||||
|
- `provider`: `aws`
|
||||||
|
- `leaseId`: `cbx_...`
|
||||||
|
- `syncDelegated`: `false`
|
||||||
|
- `commandPhases`: populated when the command prints `CRABBOX_PHASE:<name>`
|
||||||
|
- `commandMs` / `totalMs`
|
||||||
|
- `exitCode`
|
||||||
|
|
||||||
|
Crabbox should stop one-shot AWS leases automatically after the run. Verify
|
||||||
|
cleanup when a run fails, is interrupted, or the command output is unclear:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
../crabbox/bin/crabbox list --provider aws
|
||||||
|
```
|
||||||
|
|
||||||
|
## Blacksmith Testbox Through Crabbox
|
||||||
|
|
||||||
|
Use this for OpenClaw maintainer broad/heavy `pnpm` gates when the prepared CI
|
||||||
|
environment is the right proof surface:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
node scripts/crabbox-wrapper.mjs run \
|
||||||
|
--provider blacksmith-testbox \
|
||||||
|
--blacksmith-org openclaw \
|
||||||
|
--blacksmith-workflow .github/workflows/ci-check-testbox.yml \
|
||||||
|
--blacksmith-job check \
|
||||||
|
--blacksmith-ref main \
|
||||||
|
--idle-timeout 90m \
|
||||||
|
--ttl 240m \
|
||||||
|
--timing-json \
|
||||||
|
-- \
|
||||||
|
corepack pnpm check:changed
|
||||||
|
```
|
||||||
|
|
||||||
|
Read the JSON summary and the Testbox line. Useful fields:
|
||||||
|
|
||||||
|
- `provider`: `blacksmith-testbox`
|
||||||
|
- `leaseId`: `tbx_...`
|
||||||
|
- `syncDelegated`: `true`
|
||||||
|
- `syncPhases`: delegated/skipped because Blacksmith owns checkout/sync
|
||||||
|
- Actions run URL/id from the Testbox output
|
||||||
|
- `exitCode`
|
||||||
|
|
||||||
|
Use provider-backed cache volumes only for rebuildable caches, not secrets or
|
||||||
|
checkout state. On Blacksmith, Crabbox forwards them as sticky disks:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
node scripts/crabbox-wrapper.mjs run \
|
||||||
|
--provider blacksmith-testbox \
|
||||||
|
--cache-volume pnpm-store=openclaw-node24-pnpm-lock:/tmp/openclaw-pnpm-store \
|
||||||
|
--timing-json \
|
||||||
|
-- \
|
||||||
|
corepack pnpm check:changed
|
||||||
|
```
|
||||||
|
|
||||||
|
The selected provider must advertise cache-volume support. If not, omit
|
||||||
|
`--cache-volume` and rely on kept-lease caches.
|
||||||
|
|
||||||
|
`blacksmith testbox list` may hide hydrating or ready boxes. Use:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
blacksmith testbox list --all
|
||||||
|
blacksmith testbox status <tbx_id>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Observability Flags
|
||||||
|
|
||||||
|
Use these on debugging runs before inventing ad hoc logging:
|
||||||
|
|
||||||
|
- `--preflight`: prints run context, workspace mode, SSH target, remote user/cwd,
|
||||||
|
and target-specific tool probes. Defaults cover `git`, `tar`, `node`, `npm`,
|
||||||
|
`corepack`, `pnpm`, `yarn`, `bun`, `docker`, plus POSIX
|
||||||
|
`sudo`/`apt`/`bubblewrap` and native Windows
|
||||||
|
`powershell`/`execution_policy`/`longpaths`/`temp`/`pwsh`. Add
|
||||||
|
`--preflight-tools node,bun,docker`, `CRABBOX_PREFLIGHT_TOOLS`, or repo
|
||||||
|
`run.preflightTools` to replace the list. `default` expands built-ins; `none`
|
||||||
|
prints only the workspace summary. Preflight is diagnostic only; install
|
||||||
|
toolchains through Actions hydration, images, devcontainer/Nix/mise/asdf, or
|
||||||
|
the run script. On `blacksmith-testbox`, this prints a delegated-unsupported
|
||||||
|
note because the workflow owns setup.
|
||||||
|
- `CRABBOX_ENV_ALLOW=NAME,...`: forwards only listed local env vars for direct
|
||||||
|
providers and prints `set len=N secret=true` style summaries. On
|
||||||
|
`blacksmith-testbox`, env forwarding is unsupported; put secrets in the
|
||||||
|
Testbox workflow instead.
|
||||||
|
- `--env-from-profile <file>` plus `--allow-env NAME`: loads simple
|
||||||
|
`export NAME=value` / `NAME=value` lines from a local profile without
|
||||||
|
executing it, then forwards only allowlisted names. `--allow-env` is
|
||||||
|
repeatable and comma-separated. Profile values override ambient allowlisted
|
||||||
|
env values for that run. Direct POSIX, WSL2, and native Windows runs are
|
||||||
|
supported; delegated providers are not. Crabbox probes the uploaded profile
|
||||||
|
remotely and prints redacted presence/length metadata before the command.
|
||||||
|
- `--env-helper <name>`: with `--env-from-profile` on POSIX SSH targets,
|
||||||
|
persists `.crabbox/env/<name>` and `.crabbox/env/<name>.env` so follow-up
|
||||||
|
commands on the same lease can run through `./.crabbox/env/<name> <command>`.
|
||||||
|
Use only on leases you control; the profile stays until cleanup, lease reset,
|
||||||
|
or `--full-resync`.
|
||||||
|
- `--script <file>` / `--script-stdin`: upload a local script into
|
||||||
|
`.crabbox/scripts/` and execute it on the remote box. Shebang scripts execute
|
||||||
|
directly on POSIX; scripts without a shebang run through `bash`. Native
|
||||||
|
Windows uploads run through Windows PowerShell, and Crabbox appends `.ps1`
|
||||||
|
when needed. Arguments after `--` become script args.
|
||||||
|
- `--fresh-pr owner/repo#123|URL|number`: skip dirty local sync and create a
|
||||||
|
fresh remote checkout of the GitHub PR. Bare numbers use the current repo's
|
||||||
|
GitHub origin. Add `--apply-local-patch` only when the current local
|
||||||
|
`git diff --binary HEAD` should be applied on top of that PR checkout.
|
||||||
|
- `--full-resync` / `--fresh-sync`: reset a stale direct-provider workdir
|
||||||
|
before syncing. Use after sync fingerprints look wrong, SSH times out before
|
||||||
|
sync, or rsync watchdog output suggests it. It is redundant with
|
||||||
|
`--fresh-pr`, incompatible with `--no-sync`, and unsupported by delegated
|
||||||
|
providers.
|
||||||
|
- `--capture-stdout <path>` / `--capture-stderr <path>`: write remote streams to
|
||||||
|
local files and keep binary/noisy output out of retained logs. Parent
|
||||||
|
directories must already exist. These are direct-provider only.
|
||||||
|
- `--capture-on-fail`: on non-zero direct-provider exits, downloads
|
||||||
|
`.crabbox/captures/*.tar.gz` with `test-results`, `playwright-report`,
|
||||||
|
`coverage`, JUnit XML, and nearby logs. Treat as secret-bearing until reviewed.
|
||||||
|
- `--keep-on-failure`: leave a failed one-shot lease alive for live debugging
|
||||||
|
until idle/TTL expiry. Useful on direct providers and delegated one-shots.
|
||||||
|
- `--timing-json`: final machine-readable timing. Add
|
||||||
|
`echo CRABBOX_PHASE:install`, `CRABBOX_PHASE:test`, etc. in long shell
|
||||||
|
commands; direct providers and Blacksmith Testbox both report them as
|
||||||
|
`commandPhases`.
|
||||||
|
|
||||||
|
Live-provider debug template for direct AWS/Hetzner leases:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
mkdir -p .crabbox/logs
|
||||||
|
pnpm crabbox:run -- --provider aws \
|
||||||
|
--preflight \
|
||||||
|
--allow-env OPENAI_API_KEY,OPENAI_BASE_URL \
|
||||||
|
--timing-json \
|
||||||
|
--capture-stdout .crabbox/logs/live-provider.stdout.log \
|
||||||
|
--capture-stderr .crabbox/logs/live-provider.stderr.log \
|
||||||
|
--capture-on-fail \
|
||||||
|
--shell -- \
|
||||||
|
"echo CRABBOX_PHASE:install; pnpm install --frozen-lockfile; echo CRABBOX_PHASE:test; pnpm test:live"
|
||||||
|
```
|
||||||
|
|
||||||
|
Do not pass `--capture-*`, `--download`, `--checksum`, `--force-sync-large`, or
|
||||||
|
`--sync-only` to delegated providers. Also do not pass `--script*`,
|
||||||
|
`--fresh-pr`, `--full-resync`, or `--env-helper` there. Crabbox rejects these
|
||||||
|
because the provider owns sync or command transport. `--keep-on-failure` is OK
|
||||||
|
for delegated one-shots when you need to inspect a failed lease.
|
||||||
|
|
||||||
|
## Efficient Bug E2E Verification
|
||||||
|
|
||||||
|
Use the smallest Crabbox lane that proves the reported user path, not just the
|
||||||
|
touched code. Aim for one after-fix E2E proof before commenting, closing, or
|
||||||
|
opening a PR for a user-visible bug.
|
||||||
|
|
||||||
|
When the user says "test in Crabbox", do not simply copy tests to the remote
|
||||||
|
box and run them there. Crabbox is for remote real-scenario proof: copy or
|
||||||
|
install OpenClaw as the user would, run the same setup/update/CLI/Gateway/API
|
||||||
|
call that failed, and capture behavior from that entrypoint. For regressions or
|
||||||
|
bug reports, prove the broken state first when feasible, then run the same
|
||||||
|
scenario after the fix.
|
||||||
|
|
||||||
|
Pick the lane by symptom:
|
||||||
|
|
||||||
|
- Docker/setup/install bug: build a package tarball and run the matching
|
||||||
|
`scripts/e2e/*-docker.sh` or package script. This proves npm packaging,
|
||||||
|
install paths, runtime deps, config writes, and container behavior.
|
||||||
|
- Provider/model/auth bug: prefer true live E2E. Use the configured secret
|
||||||
|
workflow, then inject the single needed key into Crabbox if needed. Scrub
|
||||||
|
unrelated provider env vars in the child command so interactive defaults do
|
||||||
|
not drift to another provider. If only a dummy key is used, label the proof
|
||||||
|
narrowly, e.g. "UI/install path only; live provider auth not exercised."
|
||||||
|
- Channel delivery bug: use the channel Docker/live lane when available; include
|
||||||
|
setup, config, gateway start, send/receive or agent-turn proof, and redacted
|
||||||
|
logs.
|
||||||
|
- Gateway/session/tool bug: prefer an end-to-end CLI or Gateway RPC command that
|
||||||
|
creates real state and inspects the resulting files/API output.
|
||||||
|
- Pure parser/config bug: targeted tests may be enough, but still run a
|
||||||
|
Crabbox command when OS, package, Docker, secrets, or service lifecycle could
|
||||||
|
change behavior.
|
||||||
|
|
||||||
|
Efficient flow:
|
||||||
|
|
||||||
|
1. Reproduce or prove the pre-fix symptom from the real user-facing entrypoint
|
||||||
|
when feasible. If the issue cannot be reproduced, capture the exact command
|
||||||
|
and observed behavior instead.
|
||||||
|
2. Patch locally and run narrow tests on the pre-warmed remote box.
|
||||||
|
3. Run one Crabbox E2E command that starts from the user-facing entrypoint:
|
||||||
|
package install, Docker setup, onboarding, channel add, gateway start, or
|
||||||
|
agent turn as appropriate.
|
||||||
|
4. Record proof as: Testbox id, command, environment shape, redacted secret
|
||||||
|
source, and copied success/failure output.
|
||||||
|
5. If the issue says "cannot reproduce", ask for the missing config/log fields
|
||||||
|
that would distinguish the tested path from the reporter's path.
|
||||||
|
|
||||||
|
Keep it efficient:
|
||||||
|
|
||||||
|
- Reuse existing E2E scripts and helper assertions before writing ad hoc shell.
|
||||||
|
- Use `--script <file>` or `--script-stdin` for multi-line E2E commands instead
|
||||||
|
of quote-heavy `--shell` strings on direct SSH providers.
|
||||||
|
- Use `--fresh-pr <pr>` when validating an upstream PR in isolation from the
|
||||||
|
local dirty tree. Add `--apply-local-patch` only when testing a local fixup on
|
||||||
|
top of that PR.
|
||||||
|
- Use `--full-resync` before replacing a warmed direct-provider lease when the
|
||||||
|
remote workdir or sync fingerprint appears stale.
|
||||||
|
- For agent code tasks, reuse the pre-warmed remote box across focused tests
|
||||||
|
and heavy proof. Use a one-shot only when a single late proof is genuinely
|
||||||
|
the task's only remote command.
|
||||||
|
- Prefer `OPENCLAW_CURRENT_PACKAGE_TGZ` with Docker/package lanes when testing a
|
||||||
|
candidate tarball; prefer the repo's package helper instead of direct source
|
||||||
|
execution when the bug might be packaging/install related.
|
||||||
|
- Keep secrets redacted. It is fine to report key presence, source, and length;
|
||||||
|
never print secret values.
|
||||||
|
- Include `--timing-json` on broad or flaky runs when command duration or sync
|
||||||
|
behavior matters.
|
||||||
|
|
||||||
|
Before/after PR proof on delegated Testbox:
|
||||||
|
|
||||||
|
- For PRs that should prove "broken before, fixed after", compare base and PR
|
||||||
|
on the same Testbox when practical. Fetch both refs, create detached temp
|
||||||
|
worktrees under `/tmp`, install in each, then run the same harness twice.
|
||||||
|
- Do not checkout base/PR refs in the synced repo root. Delegated Testbox sync
|
||||||
|
may leave the root dirty with local files; `git checkout` can abort or mix
|
||||||
|
proof state.
|
||||||
|
- Temp harness files under `/tmp` do not resolve repo packages by default. Put
|
||||||
|
the harness inside the worktree, or in ESM use
|
||||||
|
`createRequire(path.join(process.cwd(), "package.json"))` before requiring
|
||||||
|
workspace deps such as `@lydell/node-pty`.
|
||||||
|
- For full-screen TUI/CLI bugs, a PTY harness is stronger than helper-only
|
||||||
|
assertions. Use a real PTY, wait for visible lifecycle markers, send input,
|
||||||
|
then send control keys and assert process exit/stuck behavior.
|
||||||
|
- When validating a rebased local branch before push, remember delegated sync
|
||||||
|
usually validates synced file content on a detached dirty checkout, not a
|
||||||
|
remote commit object. Record the local head SHA, changed files, Testbox id,
|
||||||
|
and final success markers; after pushing, ensure the pushed SHA has the same
|
||||||
|
file content.
|
||||||
|
- If GitHub CI is still queued but the exact changed content passed Testbox
|
||||||
|
`pnpm check:changed`, `pnpm check:test-types`, and the real E2E proof, it is
|
||||||
|
reasonable to merge once required checks allow it. Note any still-running
|
||||||
|
unrelated shards in the proof comment instead of waiting forever.
|
||||||
|
|
||||||
|
Interactive CLI/onboarding:
|
||||||
|
|
||||||
|
- For full-screen or prompt-heavy CLI flows, run the target command inside tmux
|
||||||
|
on the Crabbox and drive it with `tmux send-keys`; capture proof with
|
||||||
|
`tmux capture-pane`, redacted through `sed`.
|
||||||
|
- Prefer deterministic arrow navigation over search typing for Clack-style
|
||||||
|
searchable selects. Raw `send-keys -l openai` may not trigger filtering in a
|
||||||
|
tmux pane; inspect option order locally or on-box and send exact Down/Enter
|
||||||
|
sequences.
|
||||||
|
- Isolate mutable state with `OPENCLAW_STATE_DIR=$(mktemp -d)`. Plugin npm
|
||||||
|
installs live under that state dir (`npm/node_modules/...`), not under
|
||||||
|
`OPENCLAW_CONFIG_DIR`. Verify downloads by checking the state dir, package
|
||||||
|
lock, and installed package metadata.
|
||||||
|
- To test automatic setup installs against local package artifacts, use
|
||||||
|
`OPENCLAW_ALLOW_PLUGIN_INSTALL_OVERRIDES=1` plus
|
||||||
|
`OPENCLAW_PLUGIN_INSTALL_OVERRIDES='{"plugin-id":"npm-pack:/tmp/plugin.tgz"}'`.
|
||||||
|
Pack with `npm pack`, set an isolated `OPENCLAW_STATE_DIR`, and verify the
|
||||||
|
package under `npm/node_modules`. Overrides are test-only and must not be
|
||||||
|
treated as official/trusted-source installs.
|
||||||
|
- For OpenAI/Codex onboarding proof, the useful markers are the UI line
|
||||||
|
`Installed Codex plugin`, `npm/node_modules/@openclaw/codex`, and the
|
||||||
|
package-lock entry showing the bundled `@openai/codex` dependency. A dummy
|
||||||
|
OpenAI-shaped key can prove only UI/install behavior; it is not live auth.
|
||||||
|
|
||||||
|
## Reuse And Keepalive
|
||||||
|
|
||||||
|
Agent code tasks should pre-warm and reuse one remote box selected by source
|
||||||
|
trust for focused tests and heavy proof. One-shot runs remain appropriate for a
|
||||||
|
single late proof when early warmup was not warranted.
|
||||||
|
|
||||||
|
Reuse the lease, not stale source. Each command must sync the current checkout;
|
||||||
|
use `--no-sync` only to rerun an unchanged, already-synced tree intentionally.
|
||||||
|
Untrusted reuse still requires `CRABBOX_ENV_ALLOW=CI`,
|
||||||
|
`--no-hydrate`, and a fresh temporary remote `HOME` on every command. Reuse
|
||||||
|
only a fresh lease dedicated to the same untrusted source; never a trusted or
|
||||||
|
previously hydrated lease. Launch from the clean trusted `main` checkout and
|
||||||
|
use `--fresh-pr` plus the same reviewed-SHA check on every run. Keep
|
||||||
|
`CRABBOX_AWS_INSTANCE_PROFILE` unset for warmup, run, and cleanup. The lease is
|
||||||
|
valid only for that reviewed SHA; stop and rewarm after any head change.
|
||||||
|
|
||||||
|
If Crabbox returns a reusable id or you intentionally keep a lease:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
node scripts/crabbox-wrapper.mjs run --provider <blacksmith-testbox-or-aws> --id <id-or-slug> --timing-json --shell -- "corepack pnpm test <path>"
|
||||||
|
```
|
||||||
|
|
||||||
|
Stop boxes you created before handoff:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:stop -- <id-or-slug>
|
||||||
|
blacksmith testbox stop --id <tbx_id>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Interactive Desktop And WebVNC
|
||||||
|
|
||||||
|
Prefer WebVNC for human inspection because the browser portal can preload the
|
||||||
|
lease VNC password and avoids a native VNC client's copy/paste/password dance.
|
||||||
|
Use native `crabbox vnc` only when WebVNC is unavailable, the browser portal is
|
||||||
|
broken, or the user explicitly wants a local VNC client.
|
||||||
|
|
||||||
|
Common desktop flow:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
../crabbox/bin/crabbox warmup --provider hetzner --desktop --browser --class standard --idle-timeout 60m --ttl 240m
|
||||||
|
../crabbox/bin/crabbox desktop launch --provider hetzner --id <cbx_id-or-slug> --browser --url https://example.com --webvnc --open --take-control
|
||||||
|
```
|
||||||
|
|
||||||
|
Useful WebVNC commands:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
../crabbox/bin/crabbox webvnc --provider hetzner --id <cbx_id-or-slug> --open --take-control
|
||||||
|
../crabbox/bin/crabbox webvnc daemon start --provider hetzner --id <cbx_id-or-slug> --open --take-control
|
||||||
|
../crabbox/bin/crabbox webvnc daemon status --provider hetzner --id <cbx_id-or-slug>
|
||||||
|
../crabbox/bin/crabbox webvnc daemon stop --provider hetzner --id <cbx_id-or-slug>
|
||||||
|
../crabbox/bin/crabbox webvnc status --provider hetzner --id <cbx_id-or-slug>
|
||||||
|
../crabbox/bin/crabbox webvnc reset --provider hetzner --id <cbx_id-or-slug> --open --take-control
|
||||||
|
../crabbox/bin/crabbox desktop doctor --provider hetzner --id <cbx_id-or-slug>
|
||||||
|
../crabbox/bin/crabbox desktop click --provider hetzner --id <cbx_id-or-slug> --x 640 --y 420
|
||||||
|
../crabbox/bin/crabbox desktop paste --provider hetzner --id <cbx_id-or-slug> --text "user@example.com"
|
||||||
|
../crabbox/bin/crabbox desktop key --provider hetzner --id <cbx_id-or-slug> ctrl+l
|
||||||
|
../crabbox/bin/crabbox artifacts collect --id <cbx_id-or-slug> --all --output artifacts/<slug>
|
||||||
|
../crabbox/bin/crabbox artifacts publish --dir artifacts/<slug> --pr <number>
|
||||||
|
```
|
||||||
|
|
||||||
|
`desktop launch --webvnc --open` is usually the nicest one-shot: it starts the
|
||||||
|
browser/app inside the visible session, bridges the lease into the authenticated
|
||||||
|
WebVNC portal, and opens the portal. Keep browsers windowed for human QA; use
|
||||||
|
`--fullscreen` only for capture/video workflows.
|
||||||
|
For human handoff, include `--take-control` so the opened portal viewer gets
|
||||||
|
keyboard/mouse control automatically instead of landing as an observer.
|
||||||
|
|
||||||
|
Human handoff preflight:
|
||||||
|
|
||||||
|
- Do not assume a visible desktop or launched browser means the repo CLI/app is
|
||||||
|
installed, built, or on the interactive terminal's `PATH`.
|
||||||
|
- Before handing WebVNC to a human tester, prove the expected command from the
|
||||||
|
same kept lease and from a neutral directory such as `~`.
|
||||||
|
- If the handoff needs repo-local code, sync/build/link it explicitly on that
|
||||||
|
lease. Source-tree CLIs often need build output before a symlink works.
|
||||||
|
- Prefer a real `command -v <expected-command> && <expected-command> --version`
|
||||||
|
check over a repo-root-only `pnpm ...` command.
|
||||||
|
|
||||||
|
Generic handoff repair pattern:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
../crabbox/bin/crabbox run --id <cbx_id-or-slug> --full-resync --shell -- \
|
||||||
|
"set -euo pipefail
|
||||||
|
pnpm install --frozen-lockfile
|
||||||
|
pnpm build
|
||||||
|
sudo ln -sf \"\$PWD/<cli-entry>\" /usr/local/bin/<expected-command>
|
||||||
|
cd ~
|
||||||
|
command -v <expected-command>
|
||||||
|
<expected-command> --version"
|
||||||
|
```
|
||||||
|
|
||||||
|
## If Crabbox Fails
|
||||||
|
|
||||||
|
Keep the fallback narrow. First decide whether the failure is Crabbox itself,
|
||||||
|
the brokered AWS lease, Blacksmith/Testbox, repo hydration, sync, or the test
|
||||||
|
command.
|
||||||
|
|
||||||
|
Fast checks:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
command -v crabbox
|
||||||
|
../crabbox/bin/crabbox --version
|
||||||
|
pnpm crabbox:run -- --help | sed -n '1,140p'
|
||||||
|
../crabbox/bin/crabbox doctor
|
||||||
|
command -v blacksmith
|
||||||
|
blacksmith --version
|
||||||
|
blacksmith testbox list
|
||||||
|
```
|
||||||
|
|
||||||
|
Common Crabbox-only failures:
|
||||||
|
|
||||||
|
- Provider missing or old CLI: use `../crabbox/bin/crabbox` from the sibling
|
||||||
|
repo, or update/install Crabbox before retrying.
|
||||||
|
- Bad local config: inspect `.crabbox.yaml`, `crabbox config show`, and
|
||||||
|
`crabbox whoami`; normal OpenClaw agent proof should use Blacksmith Testbox.
|
||||||
|
Direct AWS is an explicit fallback and must use brokered auth, not raw keys.
|
||||||
|
- Slug/claim confusion: use the raw `cbx_...` / `tbx_...` id, or run one-shot
|
||||||
|
without `--id`.
|
||||||
|
- Sync/timing bug: add `--debug --timing-json`; capture the final JSON and the
|
||||||
|
printed Actions URL. Large sync warnings now include top source directories
|
||||||
|
by file count and a hint to update `.crabboxignore` / `sync.exclude`; inspect
|
||||||
|
those before reaching for `--force-sync-large`. Quiet rsync watchdogs and SSH
|
||||||
|
timeouts now print `next_action=` hints; follow them, usually `--full-resync`
|
||||||
|
first and a fresh lease second.
|
||||||
|
- Cleanup uncertainty: run `crabbox list --provider aws`; for explicit
|
||||||
|
Blacksmith runs, use `blacksmith testbox list` and stop only boxes you
|
||||||
|
created.
|
||||||
|
- Testbox queued/capacity pressure: do not retry Blacksmith repeatedly. Rerun
|
||||||
|
once with `--provider aws` when direct AWS still proves the requested
|
||||||
|
surface, or report the Blacksmith blocker if Testbox itself is required.
|
||||||
|
|
||||||
|
If brokered AWS cannot dispatch, sync, attach, or stop, retry once with
|
||||||
|
`--debug` and `--timing-json`:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:run -- --provider aws --debug --timing-json -- \
|
||||||
|
pnpm test:changed
|
||||||
|
```
|
||||||
|
|
||||||
|
Full suite:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:run -- --provider aws --debug --timing-json -- \
|
||||||
|
pnpm test
|
||||||
|
```
|
||||||
|
|
||||||
|
Auth fallback, only when `blacksmith` says auth is missing:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
blacksmith auth login --non-interactive --organization openclaw
|
||||||
|
```
|
||||||
|
|
||||||
|
Raw Blacksmith footguns:
|
||||||
|
|
||||||
|
- Run from repo root. The CLI syncs the current directory.
|
||||||
|
- Save the returned `tbx_...` id in the session.
|
||||||
|
- Reuse that id for focused reruns; stop it before handoff.
|
||||||
|
- Raw commit SHAs are not reliable `warmup --ref` refs; use a branch or tag.
|
||||||
|
- Treat `blacksmith testbox list` as cleanup diagnostics, not a shared reusable
|
||||||
|
queue.
|
||||||
|
|
||||||
|
Use Blacksmith Testbox through Crabbox by default for OpenClaw agent tests and
|
||||||
|
heavy work. If Blacksmith is down or quota-limited, do not keep probing it;
|
||||||
|
switch to direct AWS only when that backend proves the same surface, and note
|
||||||
|
the delegated-provider outage.
|
||||||
|
|
||||||
|
## Blacksmith Backend Notes
|
||||||
|
|
||||||
|
Crabbox Blacksmith backend delegates setup to:
|
||||||
|
|
||||||
|
- org: `openclaw`
|
||||||
|
- workflow: `.github/workflows/ci-check-testbox.yml`
|
||||||
|
- job: `check`
|
||||||
|
- ref: `main` unless testing a branch/tag intentionally
|
||||||
|
|
||||||
|
The hydration workflow owns checkout, Node/pnpm setup, dependency install,
|
||||||
|
secrets, ready marker, and keepalive. Crabbox owns dispatch, sync, SSH command
|
||||||
|
execution, timing, logs/results, cleanup, and cache-volume requests. Blacksmith
|
||||||
|
implements cache volumes as sticky disks.
|
||||||
|
|
||||||
|
Minimal Blacksmith-backed Crabbox run, from repo root:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:run -- --provider blacksmith-testbox --timing-json -- \
|
||||||
|
corepack pnpm test:changed
|
||||||
|
```
|
||||||
|
|
||||||
|
Use direct Blacksmith only when Crabbox is the broken layer and you are
|
||||||
|
isolating a Crabbox bug. Prefer direct `blacksmith testbox list` for cleanup
|
||||||
|
diagnostics, not as a reusable work queue.
|
||||||
|
|
||||||
|
Important Blacksmith footguns:
|
||||||
|
|
||||||
|
- Always run from repo root. The CLI syncs the current directory.
|
||||||
|
- Raw commit SHAs are not reliable `warmup --ref` refs; use a branch or tag.
|
||||||
|
- If auth is missing and browser auth is acceptable:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
blacksmith auth login --non-interactive --organization openclaw
|
||||||
|
```
|
||||||
|
|
||||||
|
## Brokered AWS Fallback
|
||||||
|
|
||||||
|
Use direct AWS when Testbox is unavailable, when the task needs direct-provider
|
||||||
|
semantics, or when an explicit backend comparison is required. The repo
|
||||||
|
`.crabbox.yaml` defaults to Blacksmith Testbox, so pass `--provider aws`.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pnpm crabbox:warmup -- --provider aws --class beast --market on-demand --idle-timeout 90m
|
||||||
|
pnpm crabbox:hydrate -- --provider aws --id <cbx_id-or-slug>
|
||||||
|
pnpm crabbox:run -- --provider aws --id <cbx_id-or-slug> --timing-json --shell -- "pnpm test:changed"
|
||||||
|
pnpm crabbox:stop -- --provider aws <cbx_id-or-slug>
|
||||||
|
```
|
||||||
|
|
||||||
|
Install/auth for owned Crabbox if needed:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
brew install openclaw/tap/crabbox
|
||||||
|
crabbox login --url https://crabbox.openclaw.ai --provider aws
|
||||||
|
```
|
||||||
|
|
||||||
|
New users should self-resolve broker auth before anyone asks for AWS keys:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
crabbox config show
|
||||||
|
crabbox doctor
|
||||||
|
crabbox whoami
|
||||||
|
```
|
||||||
|
|
||||||
|
- If broker auth is missing, run `crabbox login --url https://crabbox.openclaw.ai --provider aws`.
|
||||||
|
- If the CLI asks for `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, or AWS
|
||||||
|
profile setup during normal OpenClaw validation, assume the agent selected
|
||||||
|
the wrong path. Use brokered `crabbox login` or an existing brokered lease
|
||||||
|
before asking the user for cloud credentials.
|
||||||
|
- Ask for AWS keys only for explicit direct-provider/account administration,
|
||||||
|
not for normal brokered OpenClaw proof.
|
||||||
|
- Trusted automation may still use
|
||||||
|
`printf '%s' "$CRABBOX_COORDINATOR_TOKEN" | crabbox login --url https://crabbox.openclaw.ai --provider aws --token-stdin`.
|
||||||
|
|
||||||
|
macOS config lives at:
|
||||||
|
|
||||||
|
```text
|
||||||
|
~/Library/Application Support/crabbox/config.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
It should include `broker.url`, `broker.token`, and usually `provider: aws`
|
||||||
|
for OpenClaw lanes. Let that config drive normal validation.
|
||||||
|
|
||||||
|
### Interactive Desktop / WebVNC
|
||||||
|
|
||||||
|
For human desktop demos, prefer `webvnc` over native `vnc` and keep the remote
|
||||||
|
desktop visible/windowed. Do not fullscreen the remote browser or hide the XFCE
|
||||||
|
panel/window chrome unless the explicit goal is video/capture output. After
|
||||||
|
launch, verify a screenshot shows the desktop panel plus browser title bar. If
|
||||||
|
Chrome is fullscreen, toggle it back with:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
crabbox run --id <lease> --shell -- 'DISPLAY=:99 xdotool search --onlyvisible --class google-chrome windowactivate key F11'
|
||||||
|
```
|
||||||
|
|
||||||
|
## Diagnostics
|
||||||
|
|
||||||
|
```sh
|
||||||
|
crabbox status --id <id-or-slug> --wait
|
||||||
|
crabbox inspect --id <id-or-slug> --json
|
||||||
|
crabbox sync-plan
|
||||||
|
crabbox history --limit 20
|
||||||
|
crabbox history --lease <id-or-slug>
|
||||||
|
crabbox attach <run_id>
|
||||||
|
crabbox events <run_id> --json
|
||||||
|
crabbox logs <run_id>
|
||||||
|
crabbox results <run_id>
|
||||||
|
crabbox cache stats --id <id-or-slug>
|
||||||
|
crabbox cache volumes
|
||||||
|
crabbox ssh --id <id-or-slug>
|
||||||
|
blacksmith testbox list
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `--debug` on `run` when measuring sync timing.
|
||||||
|
Use `--timing-json` on warmup, hydrate, and run when comparing backends.
|
||||||
|
Use `--market spot|on-demand` only on AWS warmup/one-shot runs.
|
||||||
|
|
||||||
|
## Failure Triage
|
||||||
|
|
||||||
|
- Crabbox cannot find provider: verify `../crabbox/bin/crabbox --help` lists
|
||||||
|
the provider selected by `.crabbox.yaml`; update Crabbox before falling back.
|
||||||
|
- Hydration stuck or failed: open the printed GitHub Actions run URL and inspect
|
||||||
|
the hydration step.
|
||||||
|
- Sync failed: rerun with `--debug`; check changed-file count and whether the
|
||||||
|
checkout is dirty.
|
||||||
|
- Command failed: rerun only the failing shard/file first. Do not rerun a full
|
||||||
|
suite until the focused failure is understood.
|
||||||
|
- Cleanup uncertain: `crabbox list --provider aws`; for explicit Blacksmith
|
||||||
|
runs, use `blacksmith testbox list` and stop owned `tbx_...` leases you
|
||||||
|
created.
|
||||||
|
- Crabbox broken but Blacksmith works: use the direct Blacksmith fallback above,
|
||||||
|
then file/fix the Crabbox issue.
|
||||||
|
|
||||||
|
## Boundary
|
||||||
|
|
||||||
|
Do not add OpenClaw-specific setup to Crabbox itself. Put repo setup in the
|
||||||
|
hydration workflow and keep Crabbox generic around lease, sync, command
|
||||||
|
execution, logs/results, timing, and cleanup.
|
||||||
51
.agents/skills/discord-user-post/SKILL.md
Normal file
51
.agents/skills/discord-user-post/SKILL.md
Normal file
@@ -0,0 +1,51 @@
|
|||||||
|
---
|
||||||
|
name: discord-user-post
|
||||||
|
description: Post an approved message as the logged-in Discord user through the Discord desktop app. Use for release announcements or other direct user-authored Discord posts; not for OpenClaw channel sends, bots, webhooks, relays, agent sessions, or archive search.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Discord User Post
|
||||||
|
|
||||||
|
Use `$computer-use` to operate `/Applications/Discord.app` in the user's
|
||||||
|
existing logged-in session. This workflow represents the user directly.
|
||||||
|
|
||||||
|
## Prepare
|
||||||
|
|
||||||
|
1. Draft the complete final message outside Discord.
|
||||||
|
2. Confirm the intended server and channel with the user when either is
|
||||||
|
ambiguous.
|
||||||
|
3. Open Discord and navigate to the exact destination without entering the
|
||||||
|
message.
|
||||||
|
4. Verify the visible server name, channel header, and logged-in account.
|
||||||
|
|
||||||
|
Do not infer the target from unrelated Discord content. Stop if Discord is not
|
||||||
|
logged in, the account is wrong, or the exact destination cannot be verified.
|
||||||
|
|
||||||
|
## Confirm and Post
|
||||||
|
|
||||||
|
Posting is representational communication. Follow the `$computer-use`
|
||||||
|
confirmation policy even when the user previously asked for an announcement:
|
||||||
|
|
||||||
|
1. Show the user the exact final body and verified destination.
|
||||||
|
2. Request action-time confirmation before typing into Discord.
|
||||||
|
3. After confirmation, enter the approved body unchanged.
|
||||||
|
4. Visually inspect the composed message and destination again.
|
||||||
|
5. Send once.
|
||||||
|
|
||||||
|
If the body or destination changes after confirmation, request confirmation
|
||||||
|
again before sending.
|
||||||
|
|
||||||
|
## Verify
|
||||||
|
|
||||||
|
- Confirm the message appears once, from the user's account, in the intended
|
||||||
|
channel.
|
||||||
|
- Report the server, channel, and visible send result.
|
||||||
|
- Do not edit, delete, react, or send a follow-up without the corresponding
|
||||||
|
user instruction and confirmation.
|
||||||
|
|
||||||
|
## Guardrails
|
||||||
|
|
||||||
|
- Never use `openclaw message`, an OpenClaw agent, a Discord bot, webhook, relay,
|
||||||
|
or token for this workflow.
|
||||||
|
- Never expose private Discord content or account details in public output.
|
||||||
|
- Never send a draft, partial message, duplicate, or unreviewed attachment.
|
||||||
|
- For Discord archive/history/search, use `$discrawl` instead.
|
||||||
4
.agents/skills/discord-user-post/agents/openai.yaml
Normal file
4
.agents/skills/discord-user-post/agents/openai.yaml
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Discord User Post"
|
||||||
|
short_description: "Post approved messages through the logged-in Discord app"
|
||||||
|
default_prompt: "Post this approved message as me through the logged-in Discord desktop app."
|
||||||
50
.agents/skills/gitcrawl/SKILL.md
Normal file
50
.agents/skills/gitcrawl/SKILL.md
Normal file
@@ -0,0 +1,50 @@
|
|||||||
|
---
|
||||||
|
name: gitcrawl
|
||||||
|
description: "GitHub archive: issue/PR search, sync freshness, duplicate clusters, gh-shim PR status, and Gitcrawl repo work."
|
||||||
|
metadata:
|
||||||
|
openclaw:
|
||||||
|
homepage: https://github.com/openclaw/gitcrawl
|
||||||
|
requires:
|
||||||
|
bins:
|
||||||
|
- gitcrawl
|
||||||
|
install:
|
||||||
|
- kind: go
|
||||||
|
module: github.com/openclaw/gitcrawl/cmd/gitcrawl@latest
|
||||||
|
bins:
|
||||||
|
- gitcrawl
|
||||||
|
---
|
||||||
|
|
||||||
|
# Gitcrawl
|
||||||
|
|
||||||
|
Use local GitHub issue/PR archives before live GitHub search. Check freshness first:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gitcrawl doctor --json
|
||||||
|
```
|
||||||
|
|
||||||
|
Find candidates:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gitcrawl threads openclaw/openclaw --numbers <issue-or-pr-number> --include-closed --json
|
||||||
|
gitcrawl neighbors openclaw/openclaw --number <issue-or-pr-number> --limit 12 --json
|
||||||
|
gitcrawl search issues "query" -R openclaw/openclaw --state open --json number,title,url
|
||||||
|
gitcrawl clusters openclaw/openclaw --sort size --min-size 5
|
||||||
|
gitcrawl cluster-detail openclaw/openclaw --id <cluster-id>
|
||||||
|
```
|
||||||
|
|
||||||
|
For PR triage, start cached and go live only before mutation/merge decisions:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gitcrawl gh pr status <number-or-url> -R openclaw/openclaw --compact
|
||||||
|
gitcrawl gh pr view <number-or-url> -R openclaw/openclaw --json number,title,state,url,isDraft,headRef,headSha
|
||||||
|
gitcrawl gh --live pr status <number-or-url> -R openclaw/openclaw --compact
|
||||||
|
```
|
||||||
|
|
||||||
|
Use live `gh` plus checkout proof before commenting, labeling, closing, reopening, merging, or filing a PR review:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh pr view <number> --json number,title,state,mergedAt,body,files,comments,reviews,statusCheckRollup
|
||||||
|
gh issue view <number> --json number,title,state,body,comments,closedAt
|
||||||
|
```
|
||||||
|
|
||||||
|
Report absolute dates, repo names, issue/PR numbers, cluster ids, and source gaps. Do not close/label from similarity alone; require matching intent plus live verification.
|
||||||
4
.agents/skills/gitcrawl/agents/openai.yaml
Normal file
4
.agents/skills/gitcrawl/agents/openai.yaml
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Gitcrawl"
|
||||||
|
short_description: "Search local OpenClaw issue and PR history before live GitHub triage"
|
||||||
|
default_prompt: "Use $gitcrawl to inspect OpenClaw issue and PR history, find related threads and duplicate candidates, then verify actionable decisions with live GitHub."
|
||||||
114
.agents/skills/openclaw-debugging/SKILL.md
Normal file
114
.agents/skills/openclaw-debugging/SKILL.md
Normal file
@@ -0,0 +1,114 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-debugging
|
||||||
|
description: Debug OpenClaw model, provider, tool-surface, code-mode, streaming, and live/Crabbox behavior by choosing the right logs, probes, and proof path before changing code.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Debugging
|
||||||
|
|
||||||
|
Use this skill when OpenClaw behavior differs between local tests, live models,
|
||||||
|
providers, code mode, Tool Search, Crabbox, or CI, and the next move should be a
|
||||||
|
debug signal rather than a guess.
|
||||||
|
|
||||||
|
## Read First
|
||||||
|
|
||||||
|
- `docs/logging.md` for log files, `openclaw logs`, and targeted debug flags.
|
||||||
|
- `docs/reference/test.md` for local test commands.
|
||||||
|
- `docs/reference/code-mode.md` for code-mode exec/wait and tool catalog rules.
|
||||||
|
- Use `$openclaw-testing` for choosing test lanes.
|
||||||
|
- Use `$crabbox` for broad, Docker, package, Linux, live-key, or CI-parity proof.
|
||||||
|
|
||||||
|
## Default Loop
|
||||||
|
|
||||||
|
1. State the suspected boundary: config, tool construction, provider payload,
|
||||||
|
fetch, stream/SSE, transcript replay, worker/runtime, package/dist, or CI.
|
||||||
|
2. Add or enable the narrowest signal that proves that boundary.
|
||||||
|
3. Reproduce with the same provider/model/config. Do not randomly switch models
|
||||||
|
unless the model itself is the variable being tested.
|
||||||
|
4. Compare configured state with actual run activation.
|
||||||
|
5. Patch the root cause.
|
||||||
|
6. Rerun the exact failing probe, then broaden only if the contract requires it.
|
||||||
|
|
||||||
|
## Model Transport Logs
|
||||||
|
|
||||||
|
Use targeted env flags instead of global debug when the model request shape or
|
||||||
|
stream timing matters:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OPENCLAW_DEBUG_MODEL_TRANSPORT=1 openclaw gateway
|
||||||
|
OPENCLAW_DEBUG_MODEL_PAYLOAD=tools OPENCLAW_DEBUG_SSE=events openclaw gateway
|
||||||
|
OPENCLAW_DEBUG_MODEL_PAYLOAD=full-redacted OPENCLAW_DEBUG_SSE=peek openclaw gateway
|
||||||
|
```
|
||||||
|
|
||||||
|
Useful flags:
|
||||||
|
|
||||||
|
- `OPENCLAW_DEBUG_MODEL_TRANSPORT=1`: request start, fetch response, SDK
|
||||||
|
headers, first SSE event, stream done, and transport errors at `info`.
|
||||||
|
- `OPENCLAW_DEBUG_MODEL_PAYLOAD=summary`: bounded payload summary.
|
||||||
|
- `OPENCLAW_DEBUG_MODEL_PAYLOAD=tools`: all model-facing tool names.
|
||||||
|
- `OPENCLAW_DEBUG_MODEL_PAYLOAD=full-redacted`: capped, redacted JSON payload.
|
||||||
|
Use only while debugging; prompts/message text may still appear.
|
||||||
|
- `OPENCLAW_DEBUG_SSE=events`: first-event and stream-completion timing.
|
||||||
|
- `OPENCLAW_DEBUG_SSE=peek`: first five redacted SSE events.
|
||||||
|
- `OPENCLAW_DEBUG_CODE_MODE=1`: code-mode tool-surface diagnostics.
|
||||||
|
|
||||||
|
Watch logs with:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
openclaw logs --follow
|
||||||
|
```
|
||||||
|
|
||||||
|
## Common Boundaries
|
||||||
|
|
||||||
|
- **Config vs activation:** config can be enabled while the run disables tools,
|
||||||
|
is raw, has an empty allowlist, or lacks model tool support. Check the actual
|
||||||
|
visible tools before enforcing provider payload invariants.
|
||||||
|
- **Tool surface:** inspect final model-visible tool names, not only the tool
|
||||||
|
registry or config. Code mode means exactly `exec` and `wait` only after it
|
||||||
|
actually activates.
|
||||||
|
- **Provider payload:** log fields, model id, service tier, reasoning, input
|
||||||
|
size, metadata keys, prompt-cache key presence, and tool names before SDK
|
||||||
|
call.
|
||||||
|
- **Fetch vs SSE:** fetch response proves HTTP headers arrived; first SSE event
|
||||||
|
proves provider body progress. A gap here is a stream/body/provider issue, not
|
||||||
|
tool execution.
|
||||||
|
- **Worker/dist:** run `pnpm build` when touching workers, dynamic imports,
|
||||||
|
package exports, lazy runtime boundaries, or published paths.
|
||||||
|
- **Live keys:** use the configured secret workflow for missing provider keys
|
||||||
|
before saying live proof is blocked. Env checks are presence-only; never print
|
||||||
|
secrets.
|
||||||
|
|
||||||
|
## Code Pointers
|
||||||
|
|
||||||
|
- Model payload + Responses stream:
|
||||||
|
`src/agents/openai-transport-stream.ts`
|
||||||
|
- Guarded fetch/timing:
|
||||||
|
`src/agents/provider-transport-fetch.ts`
|
||||||
|
- OpenAI/Codex provider wrappers:
|
||||||
|
`src/agents/pi-embedded-runner/openai-stream-wrappers.ts`
|
||||||
|
- Tool construction, Tool Search, code-mode activation:
|
||||||
|
`src/agents/pi-embedded-runner/run/attempt.ts`
|
||||||
|
- Code-mode runtime and worker:
|
||||||
|
`src/agents/code-mode.ts`
|
||||||
|
`src/agents/code-mode.worker.ts`
|
||||||
|
- Tool Search catalog:
|
||||||
|
`src/agents/tool-search.ts`
|
||||||
|
|
||||||
|
## Proof Choice
|
||||||
|
|
||||||
|
- Single helper/payload bug: local targeted Vitest.
|
||||||
|
- Docs/logging-only: `pnpm check:docs` and `git diff --check`.
|
||||||
|
- Worker/dist/lazy import/package surface: targeted tests plus `pnpm build`.
|
||||||
|
- Live provider/model behavior: same provider/model with debug flags and a real
|
||||||
|
key if available.
|
||||||
|
- Docker/package/Linux/CI-parity: `$crabbox`.
|
||||||
|
- CI failure: exact SHA, relevant job only, logs only after failure/completion.
|
||||||
|
|
||||||
|
## Output Habit
|
||||||
|
|
||||||
|
Report:
|
||||||
|
|
||||||
|
- boundary tested
|
||||||
|
- exact command/env shape, redacted
|
||||||
|
- observed signal, such as tool names or first SSE event timing
|
||||||
|
- fix location
|
||||||
|
- narrow proof and any remaining risk
|
||||||
4
.agents/skills/openclaw-debugging/agents/openai.yaml
Normal file
4
.agents/skills/openclaw-debugging/agents/openai.yaml
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "OpenClaw Debugging"
|
||||||
|
short_description: "Debug model, tool, stream, and live behavior"
|
||||||
|
default_prompt: "Use $openclaw-debugging to identify the right OpenClaw debug boundary, turn on targeted logs, and choose the narrowest local or Crabbox proof."
|
||||||
88
.agents/skills/openclaw-ghsa-maintainer/SKILL.md
Normal file
88
.agents/skills/openclaw-ghsa-maintainer/SKILL.md
Normal file
@@ -0,0 +1,88 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-ghsa-maintainer
|
||||||
|
description: "Inspect, patch, validate, publish, or confirm OpenClaw GHSA security advisories and private-fork state."
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw GHSA Maintainer
|
||||||
|
|
||||||
|
Use this skill for repo security advisory workflow only. Keep general release work in `release-openclaw-maintainer`.
|
||||||
|
|
||||||
|
## Respect advisory guardrails
|
||||||
|
|
||||||
|
- Before reviewing or publishing a repo advisory, read `SECURITY.md`.
|
||||||
|
- Ask permission before any publish action.
|
||||||
|
- Treat this skill as GHSA-only. Do not use it for stable or beta release work.
|
||||||
|
|
||||||
|
## Fetch and inspect advisory state
|
||||||
|
|
||||||
|
Fetch the current advisory and the latest published npm version:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh api /repos/openclaw/openclaw/security-advisories/<GHSA>
|
||||||
|
npm view openclaw version --userconfig "$(mktemp)"
|
||||||
|
```
|
||||||
|
|
||||||
|
Use the fetch output to confirm the advisory state, linked private fork, and vulnerability payload shape before patching.
|
||||||
|
|
||||||
|
## Verify private fork PRs are closed
|
||||||
|
|
||||||
|
Before publishing, verify that the advisory's private fork has no open PRs:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
fork=$(gh api /repos/openclaw/openclaw/security-advisories/<GHSA> | jq -r .private_fork.full_name)
|
||||||
|
gh pr list -R "$fork" --state open
|
||||||
|
```
|
||||||
|
|
||||||
|
The PR list must be empty before publish.
|
||||||
|
|
||||||
|
## Prepare advisory Markdown and JSON safely
|
||||||
|
|
||||||
|
- Write advisory Markdown via heredoc to a temp file. Do not use escaped `\n` strings.
|
||||||
|
- Build PATCH payload JSON with `jq`, not hand-escaped shell JSON.
|
||||||
|
|
||||||
|
Example pattern:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cat > /tmp/ghsa.desc.md <<'EOF'
|
||||||
|
<markdown description>
|
||||||
|
EOF
|
||||||
|
|
||||||
|
jq -n --rawfile desc /tmp/ghsa.desc.md \
|
||||||
|
'{summary,severity,description:$desc,vulnerabilities:[...]}' \
|
||||||
|
> /tmp/ghsa.patch.json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Apply PATCH calls in the correct sequence
|
||||||
|
|
||||||
|
- Do not set `severity` and `cvss_vector_string` in the same PATCH call.
|
||||||
|
- Use separate calls when the advisory requires both fields.
|
||||||
|
- Publish by PATCHing the advisory and setting `"state":"published"`. There is no separate `/publish` endpoint.
|
||||||
|
|
||||||
|
Example shape:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh api -X PATCH /repos/openclaw/openclaw/security-advisories/<GHSA> \
|
||||||
|
--input /tmp/ghsa.patch.json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Publish and verify success
|
||||||
|
|
||||||
|
After publish, re-fetch the advisory and confirm:
|
||||||
|
|
||||||
|
- `state=published`
|
||||||
|
- `published_at` is set
|
||||||
|
- the description does not contain literal escaped `\\n`
|
||||||
|
|
||||||
|
Verification pattern:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh api /repos/openclaw/openclaw/security-advisories/<GHSA>
|
||||||
|
jq -r .description < /tmp/ghsa.refetch.json | rg '\\\\n'
|
||||||
|
```
|
||||||
|
|
||||||
|
## Common GHSA footguns
|
||||||
|
|
||||||
|
- Publishing fails with HTTP 422 if required fields are missing or the private fork still has open PRs.
|
||||||
|
- A payload that looks correct in shell can still be wrong if Markdown was assembled with escaped newline strings.
|
||||||
|
- Advisory PATCH sequencing matters; separate field updates when GHSA API constraints require it.
|
||||||
|
- Public hardening/no-publish comments and draft text should avoid raw commit hashes, PR titles/numbers, and fix-mechanism summaries. Prefer patched-version fields or release-only wording; keep SHAs, PRs, and implementation notes in internal evidence.
|
||||||
168
.agents/skills/openclaw-landable-bug-sweep/SKILL.md
Normal file
168
.agents/skills/openclaw-landable-bug-sweep/SKILL.md
Normal file
@@ -0,0 +1,168 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-landable-bug-sweep
|
||||||
|
description: "Find or repair a requested batch of small high-confidence non-SDK-boundary OpenClaw bugfix PRs until they are landable."
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Landable Bug Sweep
|
||||||
|
|
||||||
|
Autonomous maintainer workflow for producing a requested batch of landable OpenClaw bugfix PR URLs.
|
||||||
|
Use for broad issue/PR sweeps where the bar is high and the output is PRs, not notes.
|
||||||
|
Do not use for plugin SDK/API boundary work; those need separate architecture review.
|
||||||
|
|
||||||
|
## Target
|
||||||
|
|
||||||
|
Use `batch_size` from the request, defaulting to `5` and capped at `20`.
|
||||||
|
Return up to that many qualified PR URLs, each with:
|
||||||
|
|
||||||
|
- bug summary
|
||||||
|
- why the fix is low-risk
|
||||||
|
- proof: rebased-head local/Testbox/live commands or run IDs
|
||||||
|
- autoreview: clean result on the exact head being shown
|
||||||
|
- CI green on the exact pushed PR head
|
||||||
|
- issue/duplicate cleanup done or still pending
|
||||||
|
|
||||||
|
The URLs may be existing PRs that were reviewed/fixed, or new PRs created from issues/clusters.
|
||||||
|
Do not present a PR URL to the maintainer until it has been refreshed on current `main`, left-tested, autoreviewed clean, pushed, and verified green in live GitHub CI.
|
||||||
|
If code, tests, changelog, PR body, or branch base changes after autoreview, rerun autoreview before showing the URL.
|
||||||
|
Do not pad a batch when the bounded search yields fewer qualified PRs.
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
- `batch_size`: requested number of landable PRs; default `5`, maximum `20`.
|
||||||
|
- `source_mode`: `discovery` or `provided-prs`; default `discovery`.
|
||||||
|
- `provided_prs`: explicit PR refs when `source_mode=provided-prs`.
|
||||||
|
|
||||||
|
In `provided-prs` mode, inspect only the supplied PRs plus directly linked duplicate/canonical refs unless broader discovery is required to prove the best fix.
|
||||||
|
|
||||||
|
## Companion Skills
|
||||||
|
|
||||||
|
Use `$gitcrawl` for discovery/clustering, `$openclaw-pr-maintainer` for live GitHub mutation rules, `$github-author-context` when contributor trust matters, `$openclaw-testing` for proof choice, `$autoreview` before publishing/landing, and `$crabbox` for broad/E2E/live proof.
|
||||||
|
|
||||||
|
## Candidate Bar
|
||||||
|
|
||||||
|
Accept only when all are true:
|
||||||
|
|
||||||
|
- bug or paper cut, not feature/product/support/docs-only
|
||||||
|
- root cause is proven in current code
|
||||||
|
- dependency behavior checked via upstream docs/source/types when relevant
|
||||||
|
- production/runtime diff is small, ideally much smaller than 500 LOC and always below 500 LOC
|
||||||
|
- tests may be larger, but focused
|
||||||
|
- no new dependency
|
||||||
|
- no new config option
|
||||||
|
- no backward-incompatible behavior
|
||||||
|
- no security/product/owner-boundary decision needed
|
||||||
|
- no plugin SDK, public plugin API, or `src/plugin-sdk/**` boundary change
|
||||||
|
- no broad refactor smell
|
||||||
|
- focused proof is feasible
|
||||||
|
- branch can be rebased/refreshed and pushed, or a replacement PR can be created
|
||||||
|
|
||||||
|
Good examples:
|
||||||
|
|
||||||
|
- provider parameter mismatch proven against dependency/API contract
|
||||||
|
- CLI command diverges from adjacent command behavior
|
||||||
|
- narrow runtime state/serialization bug with failing test
|
||||||
|
- issue already fixed on current `main`, with proof and closeable duplicates
|
||||||
|
|
||||||
|
Reject:
|
||||||
|
|
||||||
|
- feature requests, new knobs, migrations, release work, workflow policy, support
|
||||||
|
- plugin SDK/API boundary changes, including compatibility shims, new SDK methods, SDK exports, or plugin-facing channel/provider seams
|
||||||
|
- auth/security boundary changes unless explicitly assigned
|
||||||
|
- bugs needing live credentials that are unavailable
|
||||||
|
- PRs with red CI unless you fix, rebase, push, and recheck them green
|
||||||
|
- PRs you only reviewed locally but did not refresh/push/check live
|
||||||
|
- PRs whose final head has not passed `$autoreview`
|
||||||
|
- fixes whose clean shape is a larger architecture move
|
||||||
|
- speculative reports without reproducible/provable cause
|
||||||
|
- UI/UX changes requiring product judgment
|
||||||
|
|
||||||
|
## Sweep Loop
|
||||||
|
|
||||||
|
1. Start clean:
|
||||||
|
- `git status -sb`
|
||||||
|
- `git pull --ff-only`
|
||||||
|
- verify branch is expected, usually `main`
|
||||||
|
2. Build candidate clusters:
|
||||||
|
- `gitcrawl` open issues/PRs, neighbors, and search
|
||||||
|
- live `gh issue/pr view`
|
||||||
|
- include PRs linked from issues and duplicates
|
||||||
|
3. For each cluster:
|
||||||
|
- read issue/PR body, comments, labels, linked refs, current source, adjacent tests
|
||||||
|
- suppress maintainer-owned queue noise unless it is the best fix path
|
||||||
|
- identify opener/author and preserve credit
|
||||||
|
- decide: `repair-existing-pr`, `create-new-pr`, `close-fixed-on-main`, `close-duplicate`, or `reject`
|
||||||
|
4. Prove before patching:
|
||||||
|
- failing test, focused repro, log/source proof, or dependency contract proof
|
||||||
|
- if already fixed on `main`, prove with current source/test/commit and close kindly
|
||||||
|
5. Patch:
|
||||||
|
- prefer existing PR when good and writable
|
||||||
|
- if unwritable or wrong shape, create own PR and preserve useful contributor credit
|
||||||
|
- if no PR exists, create one
|
||||||
|
- add regression test when it fits
|
||||||
|
- release-note context for user-facing fixes in PR body or commit message; credit human reporter/contributor when known
|
||||||
|
6. Review, refresh, and publish:
|
||||||
|
- rebase or otherwise refresh the PR branch on current `origin/main`
|
||||||
|
- resolve drift, including newly exposed CI failures, rather than counting the PR as ready
|
||||||
|
- do not add `CHANGELOG.md` during normal sweep PRs; release automation generates it from PRs and commits
|
||||||
|
- left-test the rebased head with the smallest meaningful local/Testbox/live command that proves the bug
|
||||||
|
- run `$autoreview` until no accepted/actionable findings remain before creating, updating, or presenting the PR URL
|
||||||
|
- create/update PR with real body and proof fields
|
||||||
|
- push the exact reviewed head
|
||||||
|
- verify live GitHub CI is green for that pushed head; do not count pending, red, dirty, conflicting, or externally blocked PRs in the five
|
||||||
|
7. Hygiene:
|
||||||
|
- close duplicates and fixed-on-main issues/PRs with proof as soon as you notice them during the sweep
|
||||||
|
- never mutate more than five associated items in one cluster without explicit confirmation
|
||||||
|
- comments must be kind, concrete, and include proof/PR/commit links
|
||||||
|
8. Repeat until `batch_size` landable PR URLs are ready or the bounded qualified queue is exhausted.
|
||||||
|
|
||||||
|
## PR Body Proof
|
||||||
|
|
||||||
|
Use the repo PR template. Include authored `## What Problem This Solves` and
|
||||||
|
`## Evidence` sections. Keep the body focused on intent and the most useful
|
||||||
|
validation evidence; inspect the code, tests, and CI before judging correctness.
|
||||||
|
|
||||||
|
## Existing PR Rules
|
||||||
|
|
||||||
|
- Review code path beyond the diff before trusting it.
|
||||||
|
- If PR is good: rebase/refresh on current `main`, fix small issues, left-test, autoreview clean, push, and get CI green before showing or counting it.
|
||||||
|
- If PR is not good but has a useful idea: recreate locally, co-author when warranted, close original with thanks and explanation.
|
||||||
|
- If PR is duplicate or fixed on `main`: comment proof, close.
|
||||||
|
- If maintainer cannot push to contributor branch: create own branch/PR, preserve useful commits or credit.
|
||||||
|
- If CI turns red after local proof, treat that as normal work: inspect the failing job, fix or reject, rerun, and only count the PR once green.
|
||||||
|
|
||||||
|
## Output Ledger
|
||||||
|
|
||||||
|
Maintain a running ledger:
|
||||||
|
|
||||||
|
```text
|
||||||
|
accepted:
|
||||||
|
- PR URL:
|
||||||
|
source refs:
|
||||||
|
bug:
|
||||||
|
root cause:
|
||||||
|
fix:
|
||||||
|
risk:
|
||||||
|
rebase/head:
|
||||||
|
left-test:
|
||||||
|
autoreview:
|
||||||
|
CI:
|
||||||
|
credit/thanks:
|
||||||
|
cleanup:
|
||||||
|
|
||||||
|
rejected:
|
||||||
|
- ref:
|
||||||
|
reason:
|
||||||
|
|
||||||
|
closed:
|
||||||
|
- ref:
|
||||||
|
reason:
|
||||||
|
proof/comment:
|
||||||
|
```
|
||||||
|
|
||||||
|
Final answer:
|
||||||
|
|
||||||
|
- the requested number of accepted PR URLs, or the smaller qualified count with the exhausted-search reason
|
||||||
|
- 2-4 sentence explainer per PR
|
||||||
|
- proof/CI state per PR
|
||||||
|
- closed duplicates/fixed-on-main refs
|
||||||
|
- current branch/status
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "OpenClaw Landable Bug Sweep"
|
||||||
|
short_description: "Find five small non-SDK landable bugfix PRs"
|
||||||
|
default_prompt: "Use $openclaw-landable-bug-sweep to find or repair five small high-confidence non-SDK-boundary OpenClaw bugfix PRs and get them landable."
|
||||||
164
.agents/skills/openclaw-parallels-smoke/SKILL.md
Normal file
164
.agents/skills/openclaw-parallels-smoke/SKILL.md
Normal file
@@ -0,0 +1,164 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-parallels-smoke
|
||||||
|
description: Run, rerun, debug, or interpret OpenClaw Parallels install, onboarding, gateway smoke, and upgrade checks.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Parallels Smoke
|
||||||
|
|
||||||
|
Use this skill for Parallels guest workflows and smoke interpretation. Do not load it for normal repo work.
|
||||||
|
|
||||||
|
## Global rules
|
||||||
|
|
||||||
|
- Use the snapshot most closely matching the requested fresh baseline.
|
||||||
|
- Gateway verification in smoke runs should use `openclaw gateway status --deep --require-rpc` unless the stable version being checked does not support it yet.
|
||||||
|
- Stable `2026.3.12` pre-upgrade diagnostics may require a plain `gateway status --deep` fallback.
|
||||||
|
- Treat `precheck=latest-ref-fail` on that stable pre-upgrade lane as baseline, not automatically a regression.
|
||||||
|
- Pass `--json` for machine-readable summaries.
|
||||||
|
- Per-phase logs land under `.artifacts/parallels/openclaw-parallels-*` by default. Override with `OPENCLAW_PARALLELS_ARTIFACT_ROOT` when a run needs another artifact volume.
|
||||||
|
- Do not run local and gateway agent turns in parallel on the same fresh workspace or session.
|
||||||
|
- Hard-cap every top-level Parallels lane with host `timeout --foreground` (or `gtimeout --foreground` if that is the available binary) so a stalled install, snapshot switch, or `prlctl exec` transport cannot consume the rest of the testing window. Defaults:
|
||||||
|
- macOS: `75m`
|
||||||
|
- Linux: `75m`
|
||||||
|
- Windows: `90m`
|
||||||
|
- aggregate npm-update wrapper: `150m`
|
||||||
|
If a lane hits the cap, stop there, inspect the newest `/tmp/openclaw-parallels-*` run directory and phase log, then fix or rerun the smallest affected lane. Do not keep waiting on a capped lane.
|
||||||
|
- Actual OpenClaw npm install/update phases are a stricter signal than whole-lane caps: install phases should normally finish within 7 minutes, and update phases should normally show meaningful progress within 5 minutes. If a phase named `install-main`, `install-latest`, `install-baseline`, or `install-baseline-package` exceeds 420s, or a phase named `update-dev` / same-guest `openclaw update` exceeds 300s without new markers, start diagnosis from that phase log and guest process state. Current Windows update phases can still pass after roughly 10-15 minutes because `doctor --fix` may install bundled plugin runtime deps; keep the script hard cap near 20 minutes unless the log is truly stale.
|
||||||
|
- For a full OS matrix, prefer running independent guest-family lanes in parallel when host capacity allows:
|
||||||
|
- `timeout --foreground 75m pnpm test:parallels:macos -- --json`
|
||||||
|
- `timeout --foreground 90m pnpm test:parallels:windows -- --json`
|
||||||
|
- `timeout --foreground 75m pnpm test:parallels:linux -- --json`
|
||||||
|
Keep each lane in its own shell/session and track the run directory for each one. Before starting the matrix, run any required host build/package gate to completion. When current-main tgz packaging is needed, the smoke scripts hold a shared package lock through `pnpm build`, inventory/staging, and `npm pack`; if that lock is missing or broken, serialize the matrix instead of accepting concurrent `dist` mutation.
|
||||||
|
- Do not run multiple smoke lanes against the same guest family at once. Tahoe lanes share the host HTTP port, and Windows/Linux lanes can collide on snapshot restore/start state if two jobs touch the same VM concurrently.
|
||||||
|
- Do not run the aggregate `pnpm test:parallels:npm-update` wrapper in parallel with individual macOS/Windows/Linux smoke lanes; it touches the same guest families and snapshots.
|
||||||
|
- Do not start Parallels lanes while any unrelated host command may rebuild, clean, or restage `dist` (`pnpm build`, `pnpm ui:build`, `pnpm release:check`, `pnpm test:install:smoke`, npm pack/install smoke, or Docker lanes that run package/build prep). Run unrelated build/package gates first, let them finish, then start the VM matrix. Concurrent `dist` mutation can make host `npm pack` fail with missing files and wastes a full VM cycle.
|
||||||
|
- While running or optimizing the matrix, record wall-clock duration per lane and the slowest phase from `/tmp/openclaw-parallels-*` logs. Use that timing before changing smoke order, timeouts, or helper behavior.
|
||||||
|
- If a host build changes tracked generated files such as `src/canvas-host/a2ui/.bundle.hash`, stop before spending VM time. Commit the generated artifact separately or fix the generator drift, then rerun the smallest affected lane.
|
||||||
|
- If `main` is moving under active multi-agent work, prefer a detached worktree pinned to one commit for long Parallels suites. The smoke scripts now verify the packed tgz commit instead of live `git rev-parse HEAD`, but a pinned worktree still avoids noisy rebuild/version drift during reruns.
|
||||||
|
- For `openclaw update --channel dev` lanes, remember the guest clones GitHub `main`, not your local worktree. If a local fix exists but the rerun still fails inside the cloned dev checkout, do not treat that as disproof of the fix until the branch has been pushed.
|
||||||
|
- For `prlctl exec`, pass the VM name before `--current-user` (`prlctl exec "$VM" --current-user ...`), not the other way around.
|
||||||
|
- If the workflow installs OpenClaw from a repo checkout instead of the site installer/npm release, finish by installing a real guest CLI shim and verifying it in a fresh guest shell. `pnpm openclaw ...` inside the repo is not enough for handoff parity.
|
||||||
|
- On macOS guests, prefer a user-global install plus a stable PATH-visible shim:
|
||||||
|
- install with `NPM_CONFIG_PREFIX="$HOME/.npm-global" npm install -g .`
|
||||||
|
- make sure `~/.local/bin/openclaw` exists or `~/.npm-global/bin` is on PATH
|
||||||
|
- verify from a brand-new guest shell with `which openclaw` and `openclaw --version`
|
||||||
|
|
||||||
|
## npm install then update
|
||||||
|
|
||||||
|
- Preferred entrypoint: `pnpm test:parallels:npm-update`
|
||||||
|
- For a macOS-only published release update check, use:
|
||||||
|
- `timeout --foreground 75m pnpm test:parallels:npm-update -- --platform macos --package-spec openclaw@<old-version> --update-target <target-version-or-tag> --json`
|
||||||
|
This keeps the same-guest `openclaw update --tag ...` coverage and uses the shared macOS current-user/sudo fallback without starting Windows/Linux lanes.
|
||||||
|
- Required coverage: every release/update regression run must include both lanes:
|
||||||
|
- fresh snapshot -> install requested package/baseline -> smoke
|
||||||
|
- same guest baseline -> run the guest's installed `openclaw update ...` command -> smoke again
|
||||||
|
- The update lane must exercise OpenClaw's internal updater. Do not count a direct `npm install -g <tgz-or-spec>` or harness-side package swap as update-flow coverage; those are install smokes only.
|
||||||
|
- For published targets, install the old baseline package first (for example `openclaw@2026.4.9`), then run the installed guest CLI with the intended channel/tag (for example `openclaw update --channel beta --yes --json`) and verify `openclaw --version`, `openclaw update status --json`, gateway RPC, and an agent turn after the command.
|
||||||
|
- For unpublished targets, pack the candidate on the host, serve the `.tgz` over the harness HTTP server, and point the guest updater at that served package. Prefer `openclaw update --tag http://<host-ip>:<port>/openclaw-<version>.tgz --yes --json`; when channel persistence also matters, pass `--channel <stable|beta>` and set `OPENCLAW_UPDATE_PACKAGE_SPEC` to the same served URL in the guest update environment. The command under test must still be `openclaw update`, not direct npm.
|
||||||
|
- For unpublished local-fix validation, remember the old baseline updater code still controls the first hop. A fix that lives only in the new updater code cannot change that already-running old process; the served candidate must either keep package/plugin metadata compatible with the baseline host or the baseline itself must include the updater fix.
|
||||||
|
- For beta/stable verification, resolve the tag immediately before the run (`npm view openclaw@beta version dist.tarball` or `npm view openclaw@latest ...`). Tags can move while a long VM matrix is already running; restart the matrix when the intended prerelease appears after an earlier registry 404/tag-lag check.
|
||||||
|
- Use the configured secret workflow to inject only the provider keys needed by OpenAI/Anthropic lanes. Do not print secrets or env dumps; pass provider secrets through the guest exec environment.
|
||||||
|
- Same-guest update verification should set the default model explicitly to `openai/gpt-5.4` before the agent turn and use a fresh explicit `--session-id` so old session model state does not leak into the check.
|
||||||
|
- The aggregate npm-update wrapper must resolve the Linux VM with the same Ubuntu fallback policy as `parallels-linux-smoke.sh` before both fresh and update lanes. Treat any Ubuntu guest with major version `>= 24` as acceptable when the exact default VM is missing, preferring the newest versioned Ubuntu guest with a fresh poweroff snapshot. On Peter's current host today, use `Ubuntu 26.04`.
|
||||||
|
- On macOS same-guest update checks, restart the gateway after the npm upgrade before `gateway status` / `agent`; launchd can otherwise report a loaded service while the old process has exited and the fresh process is not RPC-ready yet.
|
||||||
|
- The npm-update aggregate's macOS update leg writes the guest update script as root, then runs it as the desktop user. If `prlctl exec "$MACOS_VM" --current-user ...` cannot authenticate, retry through plain root `prlctl exec` plus `sudo -u <desktop-user> /usr/bin/env HOME=/Users/<desktop-user> USER=<desktop-user> LOGNAME=<desktop-user> PATH=/opt/homebrew/bin:/opt/homebrew/opt/node/bin:/usr/bin:/bin:/usr/sbin:/sbin ...`. That is a Parallels transport fallback; still verify `openclaw --version`, gateway RPC, and an agent turn after the update.
|
||||||
|
- On Windows same-guest update checks, restart the gateway after the npm upgrade before `gateway status` / `agent`; in-place global npm updates can otherwise leave stale hashed `dist/*` module imports alive in the running service.
|
||||||
|
- In those Windows same-guest update checks, do not treat one nonzero `openclaw gateway restart` as definitive failure. Current login-item restarts can report failure before the background service becomes observable again; follow with a longer RPC-ready wait and use `gateway start` only as a recovery step if readiness still never returns.
|
||||||
|
- After that Windows restart, do not trust one `gateway status --deep --require-rpc` call after a fixed sleep. Retry the RPC-ready probe for roughly 30 seconds and log each attempt; current guests can keep port `18789` bound while the fresh RPC endpoint is still coming up.
|
||||||
|
- For Windows same-guest update checks, prefer the done-file/log-drain PowerShell runner pattern over one long-lived `prlctl exec ... powershell -EncodedCommand ...` transport. The guest can finish successfully while the outer `prlctl exec` still hangs.
|
||||||
|
- The Windows same-guest update helper should write stage markers to its log before long steps like tgz download and `npm install -g` so the outer progress monitor does not sit on `waiting for first log line` during healthy but quiet installs.
|
||||||
|
- Linux same-guest update verification should also export `HOME=/root`, pass `OPENAI_API_KEY` via `prlctl exec ... /usr/bin/env`, and use `openclaw agent --local`; the fresh Linux baseline does not rely on persisted gateway credentials.
|
||||||
|
- The npm-update wrapper now prints per-lane progress from the nested log files. If a lane still looks stuck, inspect the nested logs in `runDir` first (`macos-fresh.log`, `windows-fresh.log`, `linux-fresh.log`, `macos-update.log`, `windows-update.log`, `linux-update.log`) instead of assuming the outer wrapper hung.
|
||||||
|
- Each run writes both `summary.json` and `summary.md`; read the markdown first for quick human triage, then the JSON/timings for automation.
|
||||||
|
- For full beta validation after a tag is published, prefer one command:
|
||||||
|
- `timeout --foreground 150m pnpm test:parallels:npm-update -- --beta-validation beta3 --json`
|
||||||
|
This resolves `beta3` to the latest `*-beta.3` version, runs latest->that-version same-guest update coverage, and then runs fresh install smoke for that exact published target on the same selected OS matrix. Use `--platform macos|windows|linux` to narrow reruns.
|
||||||
|
- For beta 4 npm validation with agent turns, the known-good shape is:
|
||||||
|
- `gtimeout --foreground 150m pnpm test:parallels:npm-update -- --beta-validation beta4 --model openai/gpt-5.4 --json`
|
||||||
|
Prefer the explicit `beta4` alias over `openclaw@beta` when validating a specific prerelease number; npm tags can move.
|
||||||
|
- If the wrapper fails a lane, read the auto-dumped tail first, then the full nested lane log under `.artifacts/parallels/openclaw-parallels-npm-update.*`.
|
||||||
|
- Current known macOS update-lane transport signature when the fallback is missing or bypassed: `Unable to authenticate the user. Make sure that the specified credentials are correct and try again.` Treat that as Parallels current-user authentication before blaming npm or OpenClaw.
|
||||||
|
- A macOS packaged fresh install with global package directories or bundled files mode `0777` usually means the harness used the root `prlctl exec` fallback under a permissive umask. The POSIX guest transports should prepend `umask 022`; verify the phase preflight line before blaming npm.
|
||||||
|
|
||||||
|
## CLI invocation footgun
|
||||||
|
|
||||||
|
- The Parallels smoke shell scripts should tolerate a literal bare `--` arg so `pnpm test:parallels:* -- --json` and similar forwarded invocations work without needing to call `bash scripts/e2e/...` directly.
|
||||||
|
|
||||||
|
## macOS flow
|
||||||
|
|
||||||
|
- Preferred entrypoint: `pnpm test:parallels:macos`
|
||||||
|
- `parallels-macos-smoke.sh --mode fresh --target-package-spec openclaw@<version>` is an install smoke only. For published old-version -> new-version update coverage on macOS, prefer the npm-update wrapper with `--platform macos`; `parallels-macos-smoke.sh --mode upgrade --target-package-spec ...` installs the target package and does not exercise the baseline CLI's updater.
|
||||||
|
- Default upgrade coverage on macOS should now include: fresh snapshot -> site installer pinned to the latest stable tag -> `openclaw update --channel dev` on the guest. Treat this as part of the default Tahoe regression plan, not an optional side quest.
|
||||||
|
- `parallels-macos-smoke.sh --mode upgrade` should run that release-to-dev lane by default. Keep the older host-tgz upgrade path only when the caller explicitly passes `--target-package-spec`.
|
||||||
|
- Because the default upgrade lane no longer needs a host tgz, skip `npm pack` + host HTTP server startup for `--mode upgrade` unless `--target-package-spec` is set. Keep the pack/server path for `fresh` and `both`.
|
||||||
|
- If that release-to-dev lane fails with `reason=preflight-no-good-commit` and repeated `sh: pnpm: command not found` tails from `preflight build`, treat it as an updater regression first. The fix belongs in the git/dev updater bootstrap path, not in Parallels retry logic.
|
||||||
|
- Until the public stable train includes that updater bootstrap fix, the macOS release-to-dev lane may seed a temporary guest-local `pnpm` shim immediately before `openclaw update --channel dev`. Keep that workaround scoped to the smoke harness and remove it once the latest stable no longer needs it.
|
||||||
|
- In Tahoe `prlctl exec --current-user` runs, prefer explicit `node .../openclaw.mjs ...` invocations for the release->dev handoff itself and for post-update verification. The shebanged global `openclaw` wrapper can fail with `env: node: No such file or directory`, and self-updating through the wrapper is a weaker lane than invoking the entrypoint under a fixed `node`.
|
||||||
|
- Default to the snapshot closest to `macOS 26.5 latest`.
|
||||||
|
- On Peter's Tahoe VM, `fresh-latest-march-2026` can hang in `prlctl snapshot-switch`; if restore times out there, rerun with `--snapshot-hint 'macOS 26.5 latest'` before blaming auth or the harness.
|
||||||
|
- `parallels-macos-smoke.sh` now retries `snapshot-switch` once after force-stopping a stuck running/suspended guest. If Tahoe still times out after that recovery path, then treat it as a real Parallels/host issue and rerun manually.
|
||||||
|
- The macOS smoke should include a dashboard load phase after gateway health: resolve the tokenized URL with `openclaw dashboard --no-open`, verify the served HTML contains the Control UI title/root shell, then open Safari and require an established localhost TCP connection from Safari to the gateway port.
|
||||||
|
- For Tahoe `fresh.gateway-status`, prefer non-TTY `prlctl exec --current-user ... openclaw gateway status ...` plus a few short retries. `prlctl enter` can spam TTY control bytes and hang the phase log even when the CLI itself is healthy.
|
||||||
|
- If a Tahoe lane times out in `fresh.first-agent-turn` and the phase log stops right after `__OPENCLAW_RC__:0` from `models set`, suspect the `prlctl enter` / `expect` wrapper before blaming auth or the model lane. That pattern means the first guest command finished but the transport never released for the next `guest_current_user_cli` call.
|
||||||
|
- If a packaged install regresses with `500` on `/`, `/healthz`, or `__openclaw/control-ui-config.json` after `fresh.install-main` or `upgrade.install-main`, suspect bundled plugin runtime deps resolving from the package root `node_modules` rather than `dist/extensions/*/node_modules`. Repro quickly with a real `npm pack`/global install lane before blaming dashboard auth or Safari.
|
||||||
|
- `prlctl exec` is fine for deterministic repo commands, but use the guest Terminal or `prlctl enter` when installer parity or shell-sensitive behavior matters.
|
||||||
|
- Multi-word `openclaw agent --message ...` checks should go through a guest shell wrapper (`guest_current_user_sh` / `guest_current_user_cli` or `/bin/sh -lc ...`), not raw `prlctl exec ... node openclaw.mjs ...`, or the message can be split into extra argv tokens and Commander reports `too many arguments for 'agent'`.
|
||||||
|
- The same wrapper rule applies when bypassing `--current-user`: write a tiny `/tmp/*.sh` on the guest and execute `/bin/bash /tmp/*.sh` through the sudo desktop-user environment. Do not pass `openclaw agent --message '...'` directly as one raw `prlctl exec` command.
|
||||||
|
- When ref-mode onboarding stores `OPENAI_API_KEY` as an env secret ref, the post-onboard agent verification should also export `OPENAI_API_KEY` for the guest command. The gateway can still reject with pairing-required and fall back to embedded execution, and that fallback needs the env-backed credential available in the shell.
|
||||||
|
- On the fresh Tahoe snapshot, `brew` exists but `node` may be missing from PATH in noninteractive exec. Use `/opt/homebrew/bin/node` when needed.
|
||||||
|
- Fresh host-served tgz installs should install as guest root with `HOME=/var/root`, then run onboarding as the desktop user via `prlctl exec --current-user`.
|
||||||
|
- Root-installed tgz smoke can log plugin blocks for world-writable `extensions/*`; do not treat that as an onboarding or gateway failure unless plugin loading is the task.
|
||||||
|
|
||||||
|
## Windows flow
|
||||||
|
|
||||||
|
- Preferred entrypoint: `pnpm test:parallels:windows`
|
||||||
|
- Use the snapshot closest to `pre-openclaw-native-e2e-2026-03-12`.
|
||||||
|
- Default upgrade coverage on Windows should now include: fresh snapshot -> site installer pinned to the requested stable tag -> `openclaw update --channel dev` on the guest. Keep the older host-tgz upgrade path only when the caller explicitly passes `--target-package-spec`.
|
||||||
|
- Optional exact npm-tag baseline on Windows: `bash scripts/e2e/parallels-windows-smoke.sh --mode upgrade --target-package-spec openclaw@<tag> --json`. That lane installs the published npm tarball as baseline, then runs `openclaw update --channel dev`.
|
||||||
|
- Optional forward-fix Windows validation: `bash scripts/e2e/parallels-windows-smoke.sh --mode upgrade --upgrade-from-packed-main --json`. That lane installs the packed current-main npm tgz as baseline, then runs `openclaw update --channel dev`.
|
||||||
|
- Always use `prlctl exec --current-user`; plain `prlctl exec` lands in `NT AUTHORITY\\SYSTEM`.
|
||||||
|
- Prefer explicit `npm.cmd` and `openclaw.cmd`.
|
||||||
|
- Use PowerShell only as the transport with `-ExecutionPolicy Bypass`, then call the `.cmd` shims from inside it.
|
||||||
|
- Current Windows Node installs expose `corepack` as a `.cmd` shim. If a release-to-dev lane sees `corepack` on PATH but `openclaw update --channel dev` still behaves as if corepack is missing, treat that as an exec-shim regression first.
|
||||||
|
- If an exact published-tag Windows lane fails during preflight with `npm run build` and `'pnpm' is not recognized`, remember that the guest is still executing the old published updater. Validate the fix with `--upgrade-from-packed-main`, then wait for the next tagged npm release before expecting the historical tag lane to pass.
|
||||||
|
- Multi-word `openclaw agent --message ...` checks should call `& $openclaw ...` inside PowerShell, not `Start-Process ... -ArgumentList` against `openclaw.cmd`, or Commander can see split argv and throw `too many arguments for 'agent'`.
|
||||||
|
- Windows installer/tgz phases now retry once after guest-ready recheck; keep new Windows smoke steps idempotent so a transport-flake retry is safe.
|
||||||
|
- If a Windows retry sees the VM become `suspended` or `stopped`, resume/start it before the next `prlctl exec`; otherwise the second attempt just repeats the same `rc=255`.
|
||||||
|
- Windows global `npm install -g` phases can stay quiet for a minute or more even when healthy; inspect the phase log before calling it hung, and only treat it as a regression once the retry wrapper or timeout trips.
|
||||||
|
- When those Windows global installs stay quiet, the useful progress often lives in the guest npm debug log, not the helper phase log. The smoke script now streams incremental `npm-cache/_logs/*-debug-0.log` deltas into the phase log during long baseline/package installs; read those lines before assuming the lane is stalled.
|
||||||
|
- The Windows baseline-package helpers now auto-dump the latest guest `npm-cache/_logs/*-debug-0.log` tail on timeout or nonzero completion. Read that tail in the phase log before opening a second guest shell.
|
||||||
|
- The same incremental npm-debug streaming also applies to `--upgrade-from-packed-main` / packaged-install baseline phases. A phase log that still says only `install.start`, `install.download-tgz`, `install.install-tgz` can still be healthy if the streamed npm-debug section shows registry fetches or bundled-plugin postinstall work.
|
||||||
|
- Fresh Windows tgz install phases should also use the background PowerShell runner plus done-file/log-drain pattern; do not rely on one long-lived `prlctl exec ... powershell ... npm install -g` transport for package installs.
|
||||||
|
- Windows release-to-dev helpers should log `where pnpm` before and after the update and require `where pnpm` to succeed post-update. That proves the updater installed or enabled `pnpm` itself instead of depending on a smoke-only bootstrap.
|
||||||
|
- Fresh Windows ref-mode onboard should use the same background PowerShell runner plus done-file/log-drain pattern as the npm-update helper, including startup materialization checks, host-side timeouts on short poll `prlctl exec` calls, and retry-on-poll-failure behavior for transient transport flakes.
|
||||||
|
- Fresh Windows daemon-health reachability should use `openclaw gateway probe --json` with a longer timeout and treat `ok: true` as success; full `gateway status --require-rpc` checks are too eager during initial startup on current main.
|
||||||
|
- Fresh Windows ref-mode agent verification should set `OPENAI_API_KEY` in the PowerShell environment before invoking `openclaw.cmd agent`, for the same pairing-required fallback reason as macOS.
|
||||||
|
- The standalone Windows upgrade smoke lane should stop the managed gateway after `upgrade.install-main` and before `upgrade.onboard-ref`. Restarting before onboard can leave the old process alive on the pre-onboard token while onboard rewrites `~/.openclaw/openclaw.json`, which then fails `gateway-health` with `unauthorized: gateway token mismatch`.
|
||||||
|
- If standalone Windows upgrade fails with a gateway token mismatch but `pnpm test:parallels:npm-update` passes, trust the mismatch as a standalone ref-onboard ordering bug first; the npm-update helper does not re-run ref-mode onboard on the same guest.
|
||||||
|
- Keep onboarding and status output ASCII-clean in logs; fancy punctuation becomes mojibake in current capture paths.
|
||||||
|
- If you hit an older run with `rc=255` plus an empty `fresh.install-main.log` or `upgrade.install-main.log`, treat it as a likely `prlctl exec` transport drop after guest start-up, not immediate proof of an npm/package failure.
|
||||||
|
|
||||||
|
## Linux flow
|
||||||
|
|
||||||
|
- Preferred entrypoint: `pnpm test:parallels:linux`
|
||||||
|
- Use the newest versioned Ubuntu guest with a fresh poweroff snapshot. On Peter's host today, that is `Ubuntu 26.04`.
|
||||||
|
- If an exact requested Ubuntu VM is missing on the host, any Ubuntu guest with major version `>= 24` is acceptable; prefer the newest versioned Ubuntu guest over older fallback snapshots.
|
||||||
|
- Use plain `prlctl exec`; `--current-user` is not the right transport on this snapshot.
|
||||||
|
- Fresh snapshots may be missing `curl`, and `apt-get update` can fail on clock skew. Bootstrap with `apt-get -o Acquire::Check-Date=false update` and install `curl ca-certificates`.
|
||||||
|
- Fresh `main` tgz smoke still needs the latest-release installer first because the snapshot has no Node or npm before bootstrap.
|
||||||
|
- This snapshot does not have a usable `systemd --user` session; managed daemon install is unsupported.
|
||||||
|
- The Linux smoke now falls back to a manual `setsid openclaw gateway run --bind loopback --port 18789 --force` launch with `HOME=/root` and the provider secret exported, then verifies `gateway status --deep --require-rpc` when available.
|
||||||
|
- The Linux manual gateway launch should wait for `gateway status --deep --require-rpc` inside the `gateway-start` phase; otherwise the first status probe can race the background bind and fail a healthy lane.
|
||||||
|
- If Linux gateway bring-up fails, inspect `/tmp/openclaw-parallels-linux-gateway.log` in the guest phase logs first; the common failure mode is a missing provider secret in the launched gateway environment.
|
||||||
|
|
||||||
|
## Discord roundtrip
|
||||||
|
|
||||||
|
- Discord roundtrip is optional and should be enabled with:
|
||||||
|
- `--discord-token-env`
|
||||||
|
- `--discord-guild-id`
|
||||||
|
- `--discord-channel-id`
|
||||||
|
- After a successful Discord smoke/roundtrip, shut down the guest VM before handoff (`prlctl stop "$VM_NAME"` or the concrete VM name). The macOS smoke harness should do this automatically after successful Discord proof; still stop the VM manually after ad-hoc Discord checks. Do not leave the Discord-configured guest running; it can keep reading/posting in `#maintainer` and spam Discord after the proof is complete.
|
||||||
|
- Keep the Discord token only in a host env var.
|
||||||
|
- Use installed `openclaw message send/read`, not `node openclaw.mjs message ...`.
|
||||||
|
- Set `channels.discord.guilds` as one JSON object, not dotted config paths with snowflakes.
|
||||||
|
- Avoid long `prlctl enter` or expect-driven Discord config scripts; prefer `prlctl exec --current-user /bin/sh -lc ...` with short commands.
|
||||||
|
- For a narrower macOS-only Discord proof run, the existing `parallels-discord-roundtrip` skill is the deep-dive companion.
|
||||||
317
.agents/skills/openclaw-pr-maintainer/SKILL.md
Normal file
317
.agents/skills/openclaw-pr-maintainer/SKILL.md
Normal file
@@ -0,0 +1,317 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-pr-maintainer
|
||||||
|
description: Use immediately for any pasted OpenClaw GitHub issue or PR URL/number, and for OpenClaw issue/PR review, triage, duplicate search, opener identity/who wrote it, author account age/activity, comments, labels, close, land, or maintainer evidence checks.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw PR Maintainer
|
||||||
|
|
||||||
|
Use this skill for maintainer-facing GitHub workflow, not for ordinary code changes.
|
||||||
|
|
||||||
|
## Start issue and PR triage with gitcrawl
|
||||||
|
|
||||||
|
- Use `$gitcrawl` first anytime you inspect OpenClaw issues or PRs.
|
||||||
|
- Check local `gitcrawl` data first for related threads, duplicate attempts, and already-landed fixes.
|
||||||
|
- Use `gitcrawl` for candidate discovery and clustering; use `gh`, `gh api`, and the current checkout to verify live state before commenting, labeling, closing, or landing.
|
||||||
|
- If `gitcrawl` is missing, stale, lacks the target thread, or has no embeddings for neighbor/search commands, fall back to the GitHub search workflow below.
|
||||||
|
- Do not run expensive/update commands such as `gitcrawl sync --include-comments`, future enrichment commands, or broad reclustering unless the user asked to update the local store or stale data is blocking the decision.
|
||||||
|
|
||||||
|
Common read-only path:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gitcrawl threads openclaw/openclaw --numbers <issue-or-pr-number> --include-closed --json
|
||||||
|
gitcrawl neighbors openclaw/openclaw --number <issue-or-pr-number> --limit 12 --json
|
||||||
|
gitcrawl search openclaw/openclaw --query "<scope or title keywords>" --mode hybrid --json
|
||||||
|
gitcrawl cluster-detail openclaw/openclaw --id <cluster-id> --member-limit 20 --body-chars 280 --json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Claim specific review targets
|
||||||
|
|
||||||
|
When a maintainer asks Codex to review, triage, fix, or land a specific OpenClaw issue/PR, check assignment before deep work.
|
||||||
|
|
||||||
|
- Identify the requesting maintainer's GitHub login. In this environment, default Peter to `steipete`; if another maintainer is clearly the requester, use that maintainer's bare login.
|
||||||
|
- Read current assignees with live `gh issue view` / `gh pr view`; `gitcrawl` is not enough for assignment state.
|
||||||
|
- If unassigned, assign the requester before deep review. This is allowed for specific requested targets; do not auto-assign broad discovery candidates or shortlists.
|
||||||
|
- If assigned to someone else, say so clearly before analysis and include assignment age:
|
||||||
|
- fresh: assigned within 6h; treat as actively owned unless user explicitly asks to continue or reassign
|
||||||
|
- stale: assigned 6h+ ago; treat as ownership hint, not a hard block; continue only with that caveat
|
||||||
|
- If assigned to requester plus others, mention co-assignees and continue.
|
||||||
|
- If assignment event time is unavailable, say `assigned, time unknown`; treat as assigned, not stale.
|
||||||
|
- Never remove or replace assignees unless explicitly asked.
|
||||||
|
|
||||||
|
Assignment time proof:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh api "repos/openclaw/openclaw/issues/<number>/timeline" --paginate \
|
||||||
|
-H "Accept: application/vnd.github+json" \
|
||||||
|
--jq '[.[] | select(.event=="assigned") | {assignee:.assignee.login, assigner:.assigner.login, actor:.actor.login, created_at}]'
|
||||||
|
```
|
||||||
|
|
||||||
|
Use the newest `assigned` event for each current assignee. Issue timeline events expose `created_at`; GitHub GraphQL `AssignedEvent.createdAt` is also valid when REST pagination is awkward.
|
||||||
|
|
||||||
|
Claim command for issues or PRs:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh api -X POST "repos/openclaw/openclaw/issues/<number>/assignees" -f 'assignees[]=<login>' >/dev/null
|
||||||
|
```
|
||||||
|
|
||||||
|
## Surface opener identity
|
||||||
|
|
||||||
|
- For every reviewed, triaged, closed, or landed issue/PR, show the opener's human name when available, GitHub login, and account age.
|
||||||
|
- Get the login from `gh issue view` / `gh pr view` (`author.login`), then fetch profile metadata once with `gh api users/<login> --jq '{login,name,created_at,type}'`.
|
||||||
|
- Report opener identity as one compact line:
|
||||||
|
`By: Jane Doe (@jane, acct 2021-04-03) | OpenClaw: 4 PRs, 2 issues, 11 commits/12mo | GitHub: 9 repos, 86 commits, 9 PRs, 3 issues, 12 reviews`
|
||||||
|
- Always show recent activity in two lanes: OpenClaw-local PRs, issues, and commits in the last 12 months; and general public GitHub activity over the same window. For linked issue-fixing PRs, include both the PR author and issue opener when they differ.
|
||||||
|
- Prefer the bundled helper for activity lookups:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
.agents/skills/openclaw-pr-maintainer/scripts/github-activity.sh <login> [other-login...]
|
||||||
|
.agents/skills/openclaw-pr-maintainer/scripts/github-activity.sh --global <login>
|
||||||
|
```
|
||||||
|
|
||||||
|
- The helper reports repo-local activity first and can fetch public GitHub contribution totals for the same window with `--global`; run the global form by default for review/triage identity summaries.
|
||||||
|
- If the global contribution graph reports zero or looks inconsistent with visible public activity, sanity-check with `gh api users/<login>`, `gh api 'users/<login>/events/public?per_page=100'`, and recent public repo commits before calling the account inactive.
|
||||||
|
- The helper is intentionally cache-friendly for gitcrawl-backed `gh`: it rounds repo-local windows to the UTC day, rounds global contribution windows to the UTC hour, and counts PRs/issues from one paginated issues response before fetching commits separately. Prefer reusing the helper instead of hand-rolling several `gh api` loops.
|
||||||
|
- If the contribution graph is misleading or zero but public events/repos show activity, keep it one line, for example:
|
||||||
|
`By: pickaxe (@ProspectOre, acct 2019-08-24) | OpenClaw: 5 PRs, 0 issues, 5 commits/12mo | GitHub: 5 repos, 29 recent events, 100 public own-repo commits; graph=0`
|
||||||
|
- If `name` is empty, use the login only. If profile lookup is rate-limited or unavailable, say `account age unknown` rather than omitting the opener.
|
||||||
|
- Use identity and activity as triage signal, not proof by itself: new, low-activity, or bot-like accounts can raise review caution, but code, repro, and CI evidence still decide.
|
||||||
|
|
||||||
|
## Suppress top-maintainer items in issue triage
|
||||||
|
|
||||||
|
When asked for issue triage, hot issues, pressing bugs, Discord-correlated issues, or "what is still open", do not surface issues or PRs authored by top maintainers by default. Prefer external/user-reported hot issues and external PRs, not maintainer-owned work queues.
|
||||||
|
|
||||||
|
Suppress by default when the opener/author is one of:
|
||||||
|
|
||||||
|
- `@vincentkoc`
|
||||||
|
- `@Takhoffman`
|
||||||
|
- `@gumadeiras`
|
||||||
|
- `@obviyus`
|
||||||
|
- `@shakkernerd`
|
||||||
|
- `@mbelinky`
|
||||||
|
- `@joshavant`
|
||||||
|
- `@ngutman`
|
||||||
|
- `@vignesh07`
|
||||||
|
- `@huntharo`
|
||||||
|
|
||||||
|
Also suppress lower-priority maintainer-owned noise from the broader keep/top-maintainer group unless it is directly relevant:
|
||||||
|
|
||||||
|
- `@thewilloftheshadow`
|
||||||
|
- `@onutc` / `@osolmaz`
|
||||||
|
- `@jacobtomlinson`
|
||||||
|
- `@tyler6204`
|
||||||
|
- `@velvet-shark`
|
||||||
|
- `@jalehman`
|
||||||
|
- `@frankekn`
|
||||||
|
- `@ImLukeF`
|
||||||
|
- `@mcaxtr`
|
||||||
|
|
||||||
|
Exceptions:
|
||||||
|
|
||||||
|
- Show maintainer-authored items when the requester explicitly asks for maintainer PRs/issues, PR landing candidates, release-blocking maintainer work, or a specific PR/issue number.
|
||||||
|
- Show a maintainer-authored item when it is the canonical fix for an external hot issue, but frame it as the fix path rather than as a user-facing issue candidate.
|
||||||
|
- Do not close, label, or deprioritize solely because an item is maintainer-authored; this section only controls what appears in triage shortlists.
|
||||||
|
|
||||||
|
## Apply close and triage labels correctly
|
||||||
|
|
||||||
|
- If an issue or PR matches an auto-close reason, apply the label and let `.github/workflows/auto-response.yml` handle the comment/close/lock flow.
|
||||||
|
- Do not manually close plus manually comment for these reasons.
|
||||||
|
- If an issue/PR is already fixed on current `main` or solved by a new release, comment with proof plus the canonical commit/PR/release, then close it.
|
||||||
|
- `r:*` labels can be used on both issues and PRs.
|
||||||
|
- Current reasons:
|
||||||
|
- `r: skill`
|
||||||
|
- `r: support`
|
||||||
|
- `r: no-ci-pr`
|
||||||
|
- `r: too-many-prs`
|
||||||
|
- `r: testflight`
|
||||||
|
- `r: third-party-extension`
|
||||||
|
- `r: moltbook`
|
||||||
|
- `r: spam`
|
||||||
|
- `invalid`
|
||||||
|
- `dirty` for PRs only
|
||||||
|
|
||||||
|
## Select small high-confidence triage candidates
|
||||||
|
|
||||||
|
When asked for `X` issues or PRs to triage, `X` means qualified candidates, not sampled threads.
|
||||||
|
|
||||||
|
Issue triage is review/prove/patch-local by default:
|
||||||
|
|
||||||
|
1. Review the issue body, comments, related threads, current code, and adjacent tests.
|
||||||
|
2. Fix only issues that are easy, high-confidence, and narrowly owned by the implicated path.
|
||||||
|
3. Add focused regression proof when practical.
|
||||||
|
4. Stop with the dirty diff, touched files, and test/gate output for maintainer review.
|
||||||
|
5. After maintainer approval to ship, make one commit per accepted fix, with release-note context in the PR body or commit message when user-facing.
|
||||||
|
6. Pull/rebase, push, then comment and close only the issues that were fixed or explicitly triaged closed.
|
||||||
|
|
||||||
|
Do not batch unrelated issue fixes into one commit. Do not publish, comment, close, or label during the review/prove phase.
|
||||||
|
|
||||||
|
Missing `CHANGELOG.md` is not a PR review finding or merge blocker. If landing/fixing a user-visible change, make sure the PR body or commit message captures the release-note context; never ask or block solely on it.
|
||||||
|
|
||||||
|
Only list candidates that pass all gates:
|
||||||
|
|
||||||
|
- small owner/surface, with a likely narrow fix and focused regression test
|
||||||
|
- symptom is reproducible or provable with logs, failing test, live command, dependency contract, or current-main behavior
|
||||||
|
- root cause is traceable to code with file/line and the proposed fix touches that path
|
||||||
|
- no strong smell that a broader refactor, ownership rethink, migration, or product decision is the better fix
|
||||||
|
- dependency-backed behavior checked against upstream docs/source/types; live or web proof used when local proof is insufficient
|
||||||
|
|
||||||
|
Loop:
|
||||||
|
|
||||||
|
1. Use `gitcrawl` / `gh` to gather candidate clusters.
|
||||||
|
2. Read issue/PR body, comments, current code, adjacent tests, and dependency contracts.
|
||||||
|
3. Try focused repro or proof.
|
||||||
|
4. Reject unclear, stale, speculative, broad-refactor, or owner-ambiguous items.
|
||||||
|
5. Continue until `X` qualified candidates or the bounded search is exhausted.
|
||||||
|
|
||||||
|
Output only qualifying candidates, with: ref, surface, proof, cause, fix sketch, why small, expected test/gate. If none qualify, say so; do not pad.
|
||||||
|
|
||||||
|
## Structure PR review output
|
||||||
|
|
||||||
|
- Start every PR review with 1-3 plain sentences explaining what the change does and why it matters. Put this before `Findings`.
|
||||||
|
- Then list findings first. If none, say `No blocking findings` or `No findings`.
|
||||||
|
- Show size near the top as `LOC: +<additions>/-<deletions> (<changedFiles> files)`, using live PR stats or local diff stats.
|
||||||
|
- Always answer: bug/behavior being fixed, PR/issue URL and affected surface, provenance for regressions when traceable, and best-fix verdict.
|
||||||
|
- For bug/regression fixes, include a compact `Provenance:` line after cause/root-cause when a bounded history pass can identify it. Use `git log -S/-G`, `git blame`, linked PRs/issues, and tests.
|
||||||
|
- Provenance must separate roles when they differ: blamed code author username, blamed PR author username, blamed PR merger/committer username, automerge trigger when known, current PR author username, PR number, and date. Do not collapse them into one "introduced by" actor.
|
||||||
|
- If the blamed PR was merged by `clawsweeper[bot]` or another automation, identify the human trigger when practical. Check live PR timeline/comments first; if rate-limited, use gitcrawl/cache or public PR HTML. Look for maintainer command comments such as `@clawsweeper automerge`, `/landpr`, labels/events that armed automerge, and ClawSweeper status comments. Report `automerge triggered by @login`; if not found, say trigger unknown rather than naming the bot as the human decision-maker.
|
||||||
|
- For any confirmed bug, run `git blame` on the implicated line(s) after identifying the root cause. Report who broke it as the blamed PR merger/committer, and also name the blamed code author. Include the PR number. If no PR is traceable, use the blamed commit as the provenance: commit SHA, date, and author username. Do not guess a merger or frame missing PR metadata as a separate finding.
|
||||||
|
- Phrase provenance as `introduced by`, `made visible by`, or `carried forward by`, with confidence (`clear`, `likely`, `unknown`). If unclear, say what evidence is missing instead of guessing. For features, docs, and refactors, use `Provenance: N/A` or omit it when no broken behavior is being fixed.
|
||||||
|
- Keep summaries compact, but include enough proof that the verdict is auditable without rereading the PR.
|
||||||
|
|
||||||
|
LOC proof:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh pr view <number> --json additions,deletions,changedFiles \
|
||||||
|
--jq '"LOC: +\(.additions)/-\(.deletions) (\(.changedFiles) files)"'
|
||||||
|
```
|
||||||
|
|
||||||
|
## Read beyond the diff
|
||||||
|
|
||||||
|
- Review the surrounding code path, not just changed lines. Open the caller, callee, data contracts, adjacent tests, and owner module.
|
||||||
|
- Before any verdict, read enough code to fill this map: changed surface, runtime entry point, owner boundary, one caller, one callee, sibling implementations sharing the invariant, adjacent tests, current `main` behavior, and shipped/dependency/Codex contracts when relevant.
|
||||||
|
- For large-codebase PRs, sample enough related files to understand the runtime boundary before deciding. Default to more code reading when the change touches agents, gateway, plugins, auth, sessions, process, config, or provider/runtime seams.
|
||||||
|
- Compare the PR against current `origin/main` behavior. Check whether recent main already changed the same surface.
|
||||||
|
- Dependency-backed behavior: MUST read upstream docs/source/types before judging API use, defaults, output shapes, errors, timeouts, memory behavior, or compatibility. Do not assume dependency contracts from memory or PR text.
|
||||||
|
- Judge solution quality, not only correctness. Ask whether the PR is the clean owner-boundary fix or a wart/workaround that should be replaced by a small refactor, moved seam, contract change, or deletion of duplicate logic.
|
||||||
|
- Mention the main files read when the verdict depends on code-path evidence.
|
||||||
|
- If the user challenges the verdict or asks whether the idea is really good, resume code reading first. Do not defend, soften, or reverse the verdict until the missing caller/callee/sibling/dependency path is checked.
|
||||||
|
|
||||||
|
## Best-fix review loop
|
||||||
|
|
||||||
|
Every PR review must explicitly answer: "Is this the best fix, or only a plausible fix?"
|
||||||
|
|
||||||
|
Before verdict:
|
||||||
|
|
||||||
|
1. Reconstruct the bug, feature need, or behavior claim from issue/PR/proof.
|
||||||
|
2. Trace current behavior from entry point to failure or decision point.
|
||||||
|
3. Read touched files, callers, callees, owner modules, adjacent tests, and relevant docs.
|
||||||
|
4. Read sibling surfaces that should share the invariant or could be broken by a one-sided fix.
|
||||||
|
5. Compare against current `origin/main` and shipped behavior when regression/compat matters.
|
||||||
|
6. Inspect upstream dependency/Codex source or docs for dependency-backed behavior.
|
||||||
|
7. Identify at least one alternative fix location or shape, then reject it with evidence.
|
||||||
|
8. If any required path above is uninspected, keep reading or mark `Remaining uncertainty`; do not call the PR best, blocked, proof-sufficient, or merge-ready.
|
||||||
|
|
||||||
|
Review output must include:
|
||||||
|
|
||||||
|
- `Best-fix verdict:` best / acceptable mitigation / wrong layer / too narrow / too broad.
|
||||||
|
- `Alternatives considered:` 1-3 concrete alternatives and why rejected.
|
||||||
|
- `Code read:` compact list of main files/contracts checked.
|
||||||
|
- `Remaining uncertainty:` what was not proven.
|
||||||
|
|
||||||
|
If the best-fix answer is only "maybe", keep reading or state the missing evidence. Do not call proof sufficient until the best-fix judgment is explicit.
|
||||||
|
|
||||||
|
## Enforce the bug-fix evidence bar
|
||||||
|
|
||||||
|
- Never merge a bug-fix PR based only on issue text, PR text, or AI rationale.
|
||||||
|
- Whenever feasible, use Crabbox (`$crabbox`) for end-to-end verification before
|
||||||
|
commenting that a bug is unreproducible, closing an issue, or opening/landing
|
||||||
|
a fix PR. Prefer a real packaged/Docker/live lane that exercises the reported
|
||||||
|
user flow over unit-only proof.
|
||||||
|
- Before landing, require:
|
||||||
|
1. symptom evidence such as a repro, logs, or a failing test
|
||||||
|
2. a verified root cause in code with file/line
|
||||||
|
3. blame-backed provenance for regressions when traceable, including blamed PR merger and automerge trigger when known, or commit SHA/date when no PR is traceable
|
||||||
|
4. a fix that touches the implicated code path
|
||||||
|
5. a regression test when feasible, or explicit manual verification plus a reason no test was added
|
||||||
|
- If the claim is unsubstantiated or likely wrong, request evidence or changes instead of merging.
|
||||||
|
- If the linked issue appears outdated or incorrect, correct triage first. Do not merge a speculative fix.
|
||||||
|
- If Crabbox/E2E proof is blocked, say exactly why and use the closest available
|
||||||
|
local, Docker, mocked, or targeted proof. Do not present unit tests as real
|
||||||
|
behavior proof.
|
||||||
|
|
||||||
|
## Close low-signal manual PRs carefully
|
||||||
|
|
||||||
|
- Do not close for red CI alone. Require a clear low-signal category plus stale or failed validation.
|
||||||
|
- Good manual-close categories:
|
||||||
|
- blank or mostly untouched PR template with no concrete OpenClaw problem/fix
|
||||||
|
- random docs-only churn such as root README translations, generic wording tweaks, or community-plugin discoverability docs that should go through ClawHub
|
||||||
|
- test-only coverage without a linked bug, owner request, or behavior change
|
||||||
|
- refactor-only cleanup, variable renames, formatting, or generated/baseline churn without maintainer request
|
||||||
|
- third-party channel/provider/tool/skill/plugin work that belongs on ClawHub instead of core
|
||||||
|
- risky ops/infra drive-bys such as new external CI services, release workflows, host upgrade scripts, Docker base migrations, or apt retry/fix-missing tweaks without owner request and green validation
|
||||||
|
- dirty branches where a narrow stated change includes unrelated docs/generated/runtime/extension files
|
||||||
|
- repeated bot-review spam or copied bot output without author-owned fixes
|
||||||
|
- Keep or escalate plausible focused bug fixes, green PRs, active maintainer discussions, assigned work, recent author follow-up, and unique reproduction details.
|
||||||
|
- For third-party capabilities, prefer the `r: third-party-extension` auto-response label when it applies; it points contributors to publish on ClawHub.
|
||||||
|
|
||||||
|
## Handle GitHub text safely
|
||||||
|
|
||||||
|
- For issue comments and PR comments, use literal multiline strings or `-F - <<'EOF'` for real newlines. Never embed `\n`.
|
||||||
|
- Do not use `gh issue/pr comment -b "..."` when the body contains backticks or shell characters. Prefer a single-quoted heredoc.
|
||||||
|
- Do not wrap issue or PR refs like `#24643` in backticks when you want auto-linking.
|
||||||
|
- PR landing comments should include clickable full commit links for landed and source SHAs when present.
|
||||||
|
|
||||||
|
## Search broadly before deciding
|
||||||
|
|
||||||
|
- Prefer `gitcrawl` first. Then use targeted GitHub keyword search to verify gaps, live status, comments, and candidates not present in the local store.
|
||||||
|
- Use `--repo openclaw/openclaw` with `--match title,body` first when using `gh search`.
|
||||||
|
- Add `--match comments` when triaging follow-up discussion or closed-as-duplicate chains.
|
||||||
|
- Do not stop at the first 500 results when the task requires a full search.
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh search prs --repo openclaw/openclaw --match title,body --limit 50 -- "auto-update"
|
||||||
|
gh search issues --repo openclaw/openclaw --match title,body --limit 50 -- "auto-update"
|
||||||
|
gh search issues --repo openclaw/openclaw --match title,body --limit 50 \
|
||||||
|
--json number,title,state,url,updatedAt -- "auto update" \
|
||||||
|
--jq '.[] | "\(.number) | \(.state) | \(.title) | \(.url)"'
|
||||||
|
```
|
||||||
|
|
||||||
|
## Follow PR review and landing hygiene
|
||||||
|
|
||||||
|
- At the start of code-changing or landing work that will need tests or heavy
|
||||||
|
proof, classify source trust and pre-warm the safe backend through `$crabbox`
|
||||||
|
in the background. Trusted maintainer code defaults to Blacksmith Testbox;
|
||||||
|
contributor/fork code stays untrusted unless a maintainer explicitly approves
|
||||||
|
credentialed execution after review; it uses secretless fork CI or
|
||||||
|
sanitized direct AWS Crabbox with `CRABBOX_ENV_ALLOW=CI`,
|
||||||
|
`--no-hydrate`, and a fresh temporary remote `HOME`, never the
|
||||||
|
credential-hydrated Testbox workflow or a previously hydrated lease. Launch
|
||||||
|
an installed trusted Crabbox binary from clean trusted `main`, fetch the PR
|
||||||
|
with `--fresh-pr`, unset and reject any resolved AWS instance profile, verify
|
||||||
|
trusted IMDS reports no IAM credentials, bind the lease to the reviewed head
|
||||||
|
SHA, and never execute its local wrapper or config. Upload trusted
|
||||||
|
`scripts/crabbox-untrusted-bootstrap.sh` from clean `main` alongside
|
||||||
|
`--fresh-pr`; it installs the pinned Node/pnpm runtime before executing PR
|
||||||
|
code. Force public networking, disable and
|
||||||
|
unset inherited Tailscale/exit-node settings, and fail closed unless
|
||||||
|
`crabbox inspect` reports no Tailscale state before any script. Rewarm after
|
||||||
|
any head change. Continue
|
||||||
|
review/editing while it hydrates, sync every run, reuse the lease, then stop
|
||||||
|
it before handoff. Skip warmup for read-only triage and docs-only work.
|
||||||
|
- Never mention release-note bookkeeping in review-only output. It is landing
|
||||||
|
or release-generation mechanics, not a correctness finding.
|
||||||
|
- If bot review conversations exist on your PR, address them and resolve them yourself once fixed.
|
||||||
|
- Leave a review conversation unresolved only when reviewer or maintainer judgment is still needed.
|
||||||
|
- Before landing any PR with non-trivial code changes, run `$autoreview` until no accepted/actionable findings remain, unless equivalent manual review already covered it, the change is trivial/docs-only, or the user opts out.
|
||||||
|
- When an agent is landing or merging a PR targeting `main`, use only the repo-native `scripts/pr` wrapper: run `scripts/pr review-init <PR>`, follow its emitted checkout/guard guidance, initialize and complete review artifacts with `scripts/pr review-artifacts-init <PR>`, validate them with `scripts/pr review-validate-artifacts <PR>`, then run `OPENCLAW_TESTBOX=1 scripts/pr prepare-run <PR>` and `scripts/pr merge-run <PR>`. The Testbox flag is mandatory for agents: it verifies exact-head hosted CI/Testbox instead of running full `pnpm` gates locally.
|
||||||
|
- Use `scripts/committer "<msg>" <file...>` for scoped commits instead of manual `git add` and `git commit`.
|
||||||
|
- Keep commit messages concise and action-oriented.
|
||||||
|
- Group related changes; avoid bundling unrelated refactors.
|
||||||
|
- Use `.github/pull_request_template.md` for PR submissions and `.github/ISSUE_TEMPLATE/` for issues.
|
||||||
|
- Do not commit PR-only artifacts such as screenshots under `.github/pr-assets`; attach them to the PR/comment or use an external artifact store instead.
|
||||||
|
|
||||||
|
## Extra safety
|
||||||
|
|
||||||
|
- If a close or reopen action would affect more than 5 PRs, ask for explicit confirmation with the exact count and target query first.
|
||||||
|
- `sync` means: if the tree is dirty, commit all changes with a sensible Conventional Commit message, then `git pull --rebase`, then `git push`. Stop if rebase conflicts cannot be resolved safely.
|
||||||
191
.agents/skills/openclaw-pr-maintainer/scripts/github-activity.sh
Executable file
191
.agents/skills/openclaw-pr-maintainer/scripts/github-activity.sh
Executable file
@@ -0,0 +1,191 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
repo="openclaw/openclaw"
|
||||||
|
months="12"
|
||||||
|
include_global="0"
|
||||||
|
script_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
repo_root="$(git -C "$script_dir/../../../.." rev-parse --show-toplevel 2>/dev/null || true)"
|
||||||
|
if [ -z "$repo_root" ]; then
|
||||||
|
repo_root="$(cd "$script_dir/../../../.." && pwd)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# shellcheck disable=SC1091
|
||||||
|
source "$repo_root/scripts/lib/plain-gh.sh"
|
||||||
|
|
||||||
|
usage() {
|
||||||
|
printf 'Usage: %s [--repo owner/repo] [--months N] [--global] <github-login> [login...]\n' "$0"
|
||||||
|
}
|
||||||
|
|
||||||
|
die() {
|
||||||
|
printf 'error: %s\n' "$*" >&2
|
||||||
|
exit 1
|
||||||
|
}
|
||||||
|
|
||||||
|
need() {
|
||||||
|
command -v "$1" >/dev/null 2>&1 || die "missing required command: $1"
|
||||||
|
}
|
||||||
|
|
||||||
|
gh() {
|
||||||
|
gh_plain "$@"
|
||||||
|
}
|
||||||
|
|
||||||
|
date_utc_relative_months() {
|
||||||
|
local count="$1"
|
||||||
|
if date -u -v-"${count}"m +%Y-%m-%dT00:00:00Z >/dev/null 2>&1; then
|
||||||
|
date -u -v-"${count}"m +%Y-%m-%dT00:00:00Z
|
||||||
|
return
|
||||||
|
fi
|
||||||
|
date -u -d "${count} months ago" +%Y-%m-%dT00:00:00Z
|
||||||
|
}
|
||||||
|
|
||||||
|
date_to_epoch() {
|
||||||
|
local value="$1"
|
||||||
|
if date -u -j -f '%Y-%m-%dT%H:%M:%SZ' "$value" +%s >/dev/null 2>&1; then
|
||||||
|
date -u -j -f '%Y-%m-%dT%H:%M:%SZ' "$value" +%s
|
||||||
|
return
|
||||||
|
fi
|
||||||
|
date -u -d "$value" +%s
|
||||||
|
}
|
||||||
|
|
||||||
|
rough_age() {
|
||||||
|
local created_at="$1"
|
||||||
|
local now_s created_s days
|
||||||
|
now_s=$(date -u +%s)
|
||||||
|
created_s=$(date_to_epoch "$created_at")
|
||||||
|
days=$(( (now_s - created_s) / 86400 ))
|
||||||
|
if (( days < 120 )); then
|
||||||
|
printf '~%dd old' "$days"
|
||||||
|
return
|
||||||
|
fi
|
||||||
|
awk -v days="$days" 'BEGIN { printf "~%.1fy old", days / 365.2425 }'
|
||||||
|
}
|
||||||
|
|
||||||
|
thread_kinds() {
|
||||||
|
local login="$1"
|
||||||
|
local since_ts="$2"
|
||||||
|
gh api --paginate "repos/${repo}/issues?state=all&creator=${login}&since=${since_ts}&per_page=100" \
|
||||||
|
--jq ".[] | select(.created_at >= \"${since_ts}\") | if has(\"pull_request\") then \"pr\" else \"issue\" end"
|
||||||
|
}
|
||||||
|
|
||||||
|
count_kind_lines() {
|
||||||
|
local kind="$1"
|
||||||
|
local lines="$2"
|
||||||
|
grep -cx "$kind" <<<"$lines" 2>/dev/null || true
|
||||||
|
}
|
||||||
|
|
||||||
|
count_commits() {
|
||||||
|
local login="$1"
|
||||||
|
local since_ts="$2"
|
||||||
|
gh api --paginate "repos/${repo}/commits?author=${login}&since=${since_ts}&per_page=100" \
|
||||||
|
--jq '.[].sha' | wc -l | tr -d '[:space:]'
|
||||||
|
}
|
||||||
|
|
||||||
|
global_activity() {
|
||||||
|
local login="$1"
|
||||||
|
local since_ts="$2"
|
||||||
|
local now_ts="$3"
|
||||||
|
# shellcheck disable=SC2016
|
||||||
|
gh api graphql \
|
||||||
|
-f login="$login" \
|
||||||
|
-f from="$since_ts" \
|
||||||
|
-f to="$now_ts" \
|
||||||
|
-f query='
|
||||||
|
query($login: String!, $from: DateTime!, $to: DateTime!) {
|
||||||
|
user(login: $login) {
|
||||||
|
contributionsCollection(from: $from, to: $to) {
|
||||||
|
totalCommitContributions
|
||||||
|
totalIssueContributions
|
||||||
|
totalPullRequestContributions
|
||||||
|
totalPullRequestReviewContributions
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}' \
|
||||||
|
--jq '.data.user.contributionsCollection // empty'
|
||||||
|
}
|
||||||
|
|
||||||
|
while [[ $# -gt 0 ]]; do
|
||||||
|
case "$1" in
|
||||||
|
--repo)
|
||||||
|
[[ $# -ge 2 ]] || die "--repo requires owner/repo"
|
||||||
|
repo="$2"
|
||||||
|
shift 2
|
||||||
|
;;
|
||||||
|
--months)
|
||||||
|
[[ $# -ge 2 ]] || die "--months requires a positive integer"
|
||||||
|
months="$2"
|
||||||
|
[[ "$months" =~ ^[0-9]+$ && "$months" != "0" ]] || die "--months must be a positive integer"
|
||||||
|
shift 2
|
||||||
|
;;
|
||||||
|
--global)
|
||||||
|
include_global="1"
|
||||||
|
shift
|
||||||
|
;;
|
||||||
|
-h|--help)
|
||||||
|
usage
|
||||||
|
exit 0
|
||||||
|
;;
|
||||||
|
--)
|
||||||
|
shift
|
||||||
|
break
|
||||||
|
;;
|
||||||
|
-*)
|
||||||
|
die "unknown option: $1"
|
||||||
|
;;
|
||||||
|
*)
|
||||||
|
break
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
|
||||||
|
[[ $# -gt 0 ]] || {
|
||||||
|
usage >&2
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
|
||||||
|
OPENCLAW_GH_BIN="$(resolve_plain_gh_bin)" || die "missing required command: gh"
|
||||||
|
export OPENCLAW_GH_BIN
|
||||||
|
need jq
|
||||||
|
|
||||||
|
since_ts=$(date_utc_relative_months "$months")
|
||||||
|
now_ts=$(date -u +%Y-%m-%dT%H:00:00Z)
|
||||||
|
|
||||||
|
for login in "$@"; do
|
||||||
|
profile=$(gh api "users/${login}" --jq '{login,name,created_at,type}')
|
||||||
|
display_login=$(jq -r '.login' <<<"$profile")
|
||||||
|
name=$(jq -r '.name // empty' <<<"$profile")
|
||||||
|
created_at=$(jq -r '.created_at' <<<"$profile")
|
||||||
|
type=$(jq -r '.type' <<<"$profile")
|
||||||
|
created_day=${created_at%%T*}
|
||||||
|
|
||||||
|
kinds=$(thread_kinds "$display_login" "$since_ts")
|
||||||
|
prs=$(count_kind_lines pr "$kinds")
|
||||||
|
issues=$(count_kind_lines issue "$kinds")
|
||||||
|
commits=$(count_commits "$display_login" "$since_ts")
|
||||||
|
|
||||||
|
if [[ -n "$name" ]]; then
|
||||||
|
printf '%s (@%s, %s, account created %s, %s)\n' \
|
||||||
|
"$name" "$display_login" "$type" "$created_day" "$(rough_age "$created_at")"
|
||||||
|
else
|
||||||
|
printf '@%s (%s, account created %s, %s)\n' \
|
||||||
|
"$display_login" "$type" "$created_day" "$(rough_age "$created_at")"
|
||||||
|
fi
|
||||||
|
printf '%s last %smo: %s PRs, %s issues, %s commits\n' "$repo" "$months" "$prs" "$issues" "$commits"
|
||||||
|
|
||||||
|
if [[ "$include_global" == "1" ]]; then
|
||||||
|
if global_json=$(global_activity "$display_login" "$since_ts" "$now_ts" 2>/dev/null); then
|
||||||
|
if [[ -n "$global_json" ]]; then
|
||||||
|
global_commits=$(jq -r '.totalCommitContributions' <<<"$global_json")
|
||||||
|
global_issues=$(jq -r '.totalIssueContributions' <<<"$global_json")
|
||||||
|
global_prs=$(jq -r '.totalPullRequestContributions' <<<"$global_json")
|
||||||
|
global_reviews=$(jq -r '.totalPullRequestReviewContributions' <<<"$global_json")
|
||||||
|
printf 'GitHub public last %smo: %s commits, %s PRs, %s issues, %s reviews\n' \
|
||||||
|
"$months" "$global_commits" "$global_prs" "$global_issues" "$global_reviews"
|
||||||
|
else
|
||||||
|
printf 'GitHub public last %smo: unavailable\n' "$months"
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
printf 'GitHub public last %smo: unavailable\n' "$months"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
done
|
||||||
273
.agents/skills/openclaw-qa-testing/SKILL.md
Normal file
273
.agents/skills/openclaw-qa-testing/SKILL.md
Normal file
@@ -0,0 +1,273 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-qa-testing
|
||||||
|
description: Run, watch, debug, extend, or explain OpenClaw qa-lab and qa-channel scenarios, artifacts, and live lanes.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw QA Testing
|
||||||
|
|
||||||
|
Use this skill for `qa-lab` / `qa-channel` work. Repo-local QA only.
|
||||||
|
|
||||||
|
## Read first
|
||||||
|
|
||||||
|
- `docs/concepts/qa-e2e-automation.md`
|
||||||
|
- `docs/help/testing.md`
|
||||||
|
- `docs/channels/qa-channel.md`
|
||||||
|
- `qa/README.md`
|
||||||
|
- `qa/scenarios/index.yaml`
|
||||||
|
- `extensions/qa-lab/src/suite.ts`
|
||||||
|
- `extensions/qa-lab/src/character-eval.ts`
|
||||||
|
|
||||||
|
## Model policy
|
||||||
|
|
||||||
|
- Live OpenAI lane: `openai/gpt-5.4`
|
||||||
|
- Fast mode: on
|
||||||
|
- Do not use:
|
||||||
|
- `openai/gpt-5.4-pro`
|
||||||
|
- `openai/gpt-5.4-mini`
|
||||||
|
- Only change model policy if the user explicitly asks.
|
||||||
|
|
||||||
|
## Default workflow
|
||||||
|
|
||||||
|
1. Read the scenario pack and current suite implementation.
|
||||||
|
2. Decide lane:
|
||||||
|
- mock/dev: `mock-openai`
|
||||||
|
- real validation: `live-frontier`
|
||||||
|
3. For live OpenAI, use:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OPENCLAW_LIVE_OPENAI_KEY="${OPENAI_API_KEY}" \
|
||||||
|
pnpm openclaw qa suite \
|
||||||
|
--provider-mode live-frontier \
|
||||||
|
--model openai/gpt-5.4 \
|
||||||
|
--alt-model openai/gpt-5.4 \
|
||||||
|
--output-dir .artifacts/qa-e2e/run-all-live-frontier-<tag>
|
||||||
|
```
|
||||||
|
|
||||||
|
4. Watch outputs:
|
||||||
|
- summary: `.artifacts/qa-e2e/run-all-live-frontier-<tag>/qa-suite-summary.json`
|
||||||
|
- report: `.artifacts/qa-e2e/run-all-live-frontier-<tag>/qa-suite-report.md`
|
||||||
|
5. If the user wants to watch the live UI, find the current `openclaw-qa` listen port and report `http://127.0.0.1:<port>`.
|
||||||
|
6. If a scenario fails, fix the product or harness root cause, then rerun the full lane.
|
||||||
|
|
||||||
|
## OTEL smoke
|
||||||
|
|
||||||
|
For local QA-lab OpenTelemetry validation, use:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm qa:otel:smoke
|
||||||
|
```
|
||||||
|
|
||||||
|
This starts a local OTLP/HTTP trace receiver, runs the `otel-trace-smoke`
|
||||||
|
scenario through qa-channel, decodes the emitted protobuf spans, and verifies
|
||||||
|
the exported trace names and privacy contract. It does not require Opik,
|
||||||
|
Langfuse, or external collector credentials.
|
||||||
|
|
||||||
|
## Matrix live profiles
|
||||||
|
|
||||||
|
`pnpm openclaw qa matrix` defaults to the full `all` profile. Use explicit
|
||||||
|
profiles for faster CI/release proof:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OPENCLAW_QA_MATRIX_NO_REPLY_WINDOW_MS=3000 \
|
||||||
|
pnpm openclaw qa matrix --profile fast --fail-fast
|
||||||
|
```
|
||||||
|
|
||||||
|
- `fast`: release-critical transport contract, excluding generated image and
|
||||||
|
deep E2EE recovery inventory.
|
||||||
|
- `transport`, `media`, `e2ee-smoke`, `e2ee-deep`, `e2ee-cli`: sharded full
|
||||||
|
Matrix coverage.
|
||||||
|
- `QA-Lab - All Lanes` uses explicit `fast` Matrix on scheduled runs. Manual
|
||||||
|
dispatch keeps `matrix_profile=all` as the default and always shards that full
|
||||||
|
Matrix selection.
|
||||||
|
|
||||||
|
## QA credentials and 1Password
|
||||||
|
|
||||||
|
- Use `op` only inside `tmux` for QA secret lookup in this repo.
|
||||||
|
- Quick auth check inside tmux:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
op account list
|
||||||
|
```
|
||||||
|
|
||||||
|
- Direct Telegram npm live test secrets currently live in 1Password item:
|
||||||
|
- vault: `OpenClaw`
|
||||||
|
- item: `Telegram E2E`
|
||||||
|
- That item is the first place to look for:
|
||||||
|
- `OPENCLAW_QA_TELEGRAM_DRIVER_BOT_TOKEN`
|
||||||
|
- `OPENCLAW_QA_TELEGRAM_SUT_BOT_TOKEN`
|
||||||
|
- `OPENCLAW_QA_PROVIDER_MODE`
|
||||||
|
- `OPENCLAW_NPM_TELEGRAM_PACKAGE_SPEC`
|
||||||
|
- Convex QA secrets currently live in 1Password items:
|
||||||
|
- vault: `OpenClaw`
|
||||||
|
- item: `OPENCLAW_QA_CONVEX_SITE_URL`
|
||||||
|
- item: `OPENCLAW_QA_CONVEX_SECRET_MAINTAINER`
|
||||||
|
- item: `OPENCLAW_QA_CONVEX_SECRET_CI`
|
||||||
|
- Additional related notes/login items seen during QA credential work:
|
||||||
|
- vault: `Private`
|
||||||
|
- items: `OPENCLAW QA`, `Convex`, `Telegram`
|
||||||
|
- If a required value is missing from those notes:
|
||||||
|
- do not guess
|
||||||
|
- ask the maintainer/operator for the current value or the current 1Password item name
|
||||||
|
- for Telegram direct runs, `OPENCLAW_QA_TELEGRAM_GROUP_ID` may be stored separately from `Telegram E2E`
|
||||||
|
- for Convex runs, the leased Telegram credential should provide the Telegram group id and bot tokens together; do not require a separate `OPENCLAW_QA_TELEGRAM_GROUP_ID`
|
||||||
|
- for Convex runs, prefer `OpenClaw/OPENCLAW_QA_CONVEX_SITE_URL`; if that is stale or unclear, ask for the active pool URL before running
|
||||||
|
- Prefer direct Telegram envs for the npm Telegram Docker lane when available:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OPENCLAW_QA_TELEGRAM_GROUP_ID="..." \
|
||||||
|
OPENCLAW_QA_TELEGRAM_DRIVER_BOT_TOKEN="..." \
|
||||||
|
OPENCLAW_QA_TELEGRAM_SUT_BOT_TOKEN="..." \
|
||||||
|
OPENCLAW_QA_PROVIDER_MODE="mock-openai" \
|
||||||
|
OPENCLAW_NPM_TELEGRAM_PACKAGE_SPEC="openclaw@beta" \
|
||||||
|
pnpm test:docker:npm-telegram-live
|
||||||
|
```
|
||||||
|
|
||||||
|
- Prefer Convex mode when the goal is stable shared QA infra:
|
||||||
|
- round-robin credential leasing
|
||||||
|
- thinner wrapper for channel-specific setup
|
||||||
|
- CLI/admin flows around the pooled credentials
|
||||||
|
- Live npm Telegram Docker lane note:
|
||||||
|
- `scripts/e2e/npm-telegram-live-runner.ts` reads `OPENCLAW_NPM_TELEGRAM_PROVIDER_MODE`
|
||||||
|
- do not assume `OPENCLAW_QA_PROVIDER_MODE` is consumed by that wrapper
|
||||||
|
- if a 1Password note only gives `OPENCLAW_QA_PROVIDER_MODE`, map it explicitly to `OPENCLAW_NPM_TELEGRAM_PROVIDER_MODE` before running the Docker lane
|
||||||
|
- Verified live shape:
|
||||||
|
- Convex mode can pass the real Docker lane without direct Telegram env vars
|
||||||
|
- leased Telegram payload includes the group id coupled to the driver/SUT tokens
|
||||||
|
- a real run of `pnpm test:docker:npm-telegram-live` passed with:
|
||||||
|
- `OPENCLAW_QA_CREDENTIAL_SOURCE=convex`
|
||||||
|
- `OPENCLAW_QA_CREDENTIAL_ROLE=maintainer`
|
||||||
|
- `OPENCLAW_QA_CONVEX_SITE_URL`
|
||||||
|
- `OPENCLAW_QA_CONVEX_SECRET_MAINTAINER`
|
||||||
|
- `OPENCLAW_NPM_TELEGRAM_PROVIDER_MODE=mock-openai`
|
||||||
|
- If direct Telegram env is missing locally and `op signin` blocks, prefer dispatching the manual GitHub lane because the `qa-live-shared` environment already has Convex CI credentials:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh workflow run "NPM Telegram Beta E2E" --repo openclaw/openclaw --ref main \
|
||||||
|
-f package_spec=openclaw@YYYY.M.D-beta.N \
|
||||||
|
-f package_label=openclaw@YYYY.M.D-beta.N \
|
||||||
|
-f provider_mode=mock-openai
|
||||||
|
```
|
||||||
|
|
||||||
|
- Poll the exact run id from the dispatch URL. `gh run view --json artifacts` is not supported; list artifacts with:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh api repos/openclaw/openclaw/actions/runs/<run-id>/artifacts
|
||||||
|
```
|
||||||
|
|
||||||
|
## WhatsApp live credentials
|
||||||
|
|
||||||
|
Use this when setting up or replacing Convex `kind=whatsapp` credentials.
|
||||||
|
|
||||||
|
- Treat WhatsApp QA credentials as operator-owned live accounts, not generated fixtures.
|
||||||
|
- Use two dedicated WhatsApp-capable test numbers: one driver account and one SUT account. Do not use personal numbers or personal OpenClaw WhatsApp accounts in the shared pool.
|
||||||
|
- Register and link each account manually with WhatsApp or WhatsApp Business, storing Web auth only in isolated local auth dirs outside the repo.
|
||||||
|
- For group coverage, create a dedicated test group that includes both QA accounts and store its JID as `groupJid`; otherwise the group mention-gating scenario should be skipped by default and fail when explicitly requested.
|
||||||
|
- Package the two Baileys auth dirs into base64 `.tgz` payload fields and add a new active Convex credential row. Prefer adding a fresh row and disabling stale/broken rows over overwriting credentials in place.
|
||||||
|
- Expected payload fields: `driverPhoneE164`, `sutPhoneE164`, `driverAuthArchiveBase64`, `sutAuthArchiveBase64`, and optional `groupJid`.
|
||||||
|
- Keep credential material out of the repo, logs, PRs, and screenshots. Redact phone numbers unless the operator explicitly asks for local debugging.
|
||||||
|
- Validate with `pnpm openclaw qa whatsapp --credential-source convex --credential-role maintainer --provider-mode mock-openai` and preserve artifact paths plus redacted pass/fail summaries.
|
||||||
|
- If WhatsApp expires or invalidates a linked Web session, relink locally, package fresh auth archives, add a new Convex row, then disable the stale row.
|
||||||
|
|
||||||
|
## Character evals
|
||||||
|
|
||||||
|
Use `qa character-eval` for style/persona/vibe checks across multiple live models.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm openclaw qa character-eval \
|
||||||
|
--model openai/gpt-5.4,thinking=xhigh \
|
||||||
|
--model openai/gpt-5.2,thinking=xhigh \
|
||||||
|
--model openai/gpt-5,thinking=xhigh \
|
||||||
|
--model anthropic/claude-opus-4-6,thinking=high \
|
||||||
|
--model anthropic/claude-sonnet-4-6,thinking=high \
|
||||||
|
--model zai/glm-5.1,thinking=high \
|
||||||
|
--model moonshot/kimi-k2.5,thinking=high \
|
||||||
|
--model google/gemini-3.1-pro-preview,thinking=high \
|
||||||
|
--judge-model openai/gpt-5.4,thinking=xhigh,fast \
|
||||||
|
--judge-model anthropic/claude-opus-4-6,thinking=high \
|
||||||
|
--concurrency 16 \
|
||||||
|
--judge-concurrency 16 \
|
||||||
|
--output-dir .artifacts/qa-e2e/character-eval-<tag>
|
||||||
|
```
|
||||||
|
|
||||||
|
- Runs local QA gateway child processes, not Docker.
|
||||||
|
- Preferred model spec syntax is `provider/model,thinking=<level>[,fast|,no-fast|,fast=<bool>]` for both `--model` and `--judge-model`.
|
||||||
|
- Do not add new examples with separate `--model-thinking`; keep that flag as legacy compatibility only.
|
||||||
|
- Defaults to candidate models `openai/gpt-5.4`, `openai/gpt-5.2`, `openai/gpt-5`, `anthropic/claude-opus-4-6`, `anthropic/claude-sonnet-4-6`, `zai/glm-5.1`, `moonshot/kimi-k2.5`, and `google/gemini-3.1-pro-preview` when no `--model` is passed.
|
||||||
|
- Candidate thinking defaults to `high`, with `xhigh` for OpenAI models that support it. Prefer inline `--model provider/model,thinking=<level>`; `--thinking <level>` and `--model-thinking <provider/model=level>` remain compatibility shims.
|
||||||
|
- OpenAI candidate refs default to fast mode so priority processing is used where supported. Use inline `,fast`, `,no-fast`, or `,fast=false` for one model; use `--fast` only to force fast mode for every candidate.
|
||||||
|
- Judges default to `openai/gpt-5.4,thinking=xhigh,fast` and `anthropic/claude-opus-4-6,thinking=high`.
|
||||||
|
- Report includes judge ranking, run stats, durations, and full transcripts; do not include raw judge replies. Duration is benchmark context, not a grading signal.
|
||||||
|
- Candidate and judge concurrency default to 16. Use `--concurrency <n>` and `--judge-concurrency <n>` to override when local gateways or provider limits need a gentler lane.
|
||||||
|
- Scenario source is YAML-only under `qa/scenarios/`: use `index.yaml` and
|
||||||
|
per-scenario `*.yaml` files with top-level `title`, `scenario`, and optional
|
||||||
|
`flow`. Never add fenced `qa-scenario` / `qa-flow` Markdown files.
|
||||||
|
- For isolated character/persona evals, write the persona into `SOUL.md` and blank `IDENTITY.md` in the scenario flow. Use `SOUL.md + IDENTITY.md` only when intentionally testing how the normal OpenClaw identity combines with the character.
|
||||||
|
- Keep prompts natural and task-shaped. The candidate model should receive character setup through `SOUL.md`, then normal user turns such as chat, workspace help, and small file tasks; do not ask "how would you react?" or tell the model it is in an eval.
|
||||||
|
- Prefer at least one real task, such as creating or editing a tiny workspace artifact, so the transcript captures character under normal tool use instead of pure roleplay.
|
||||||
|
|
||||||
|
## Codex CLI model lane
|
||||||
|
|
||||||
|
Use model refs shaped like `codex-cli/<codex-model>` whenever QA should exercise Codex as a model backend.
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm openclaw qa suite \
|
||||||
|
--provider-mode live-frontier \
|
||||||
|
--model codex-cli/<codex-model> \
|
||||||
|
--alt-model codex-cli/<codex-model> \
|
||||||
|
--scenario <scenario-id> \
|
||||||
|
--output-dir .artifacts/qa-e2e/codex-<tag>
|
||||||
|
```
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm openclaw qa manual \
|
||||||
|
--model codex-cli/<codex-model> \
|
||||||
|
--message "Reply exactly: CODEX_OK"
|
||||||
|
```
|
||||||
|
|
||||||
|
- Treat the concrete Codex model name as user/config input; do not hardcode it in source, docs examples, or scenarios.
|
||||||
|
- Live QA preserves `CODEX_HOME` so Codex CLI auth/config works while keeping `HOME` and `OPENCLAW_HOME` sandboxed.
|
||||||
|
- Mock QA should scrub `CODEX_HOME`.
|
||||||
|
- If Codex returns fallback/auth text every turn, first check `CODEX_HOME`,
|
||||||
|
relevant secret-backed auth, and gateway child logs before changing
|
||||||
|
scenario assertions.
|
||||||
|
- For model comparison, include `codex-cli/<codex-model>` as another candidate in `qa character-eval`; the report should label it as an opaque model name.
|
||||||
|
|
||||||
|
## Repo facts
|
||||||
|
|
||||||
|
- Seed scenarios live in `qa/scenarios/index.yaml` and
|
||||||
|
`qa/scenarios/<theme>/*.yaml`.
|
||||||
|
- Main live runner: `extensions/qa-lab/src/suite.ts`
|
||||||
|
- QA lab server: `extensions/qa-lab/src/lab-server.ts`
|
||||||
|
- Child gateway harness: `extensions/qa-lab/src/gateway-child.ts`
|
||||||
|
- Synthetic channel: `extensions/qa-channel/`
|
||||||
|
|
||||||
|
## What “done” looks like
|
||||||
|
|
||||||
|
- Full suite green for the requested lane.
|
||||||
|
- User gets:
|
||||||
|
- watch URL if applicable
|
||||||
|
- pass/fail counts
|
||||||
|
- artifact paths
|
||||||
|
- concise note on what was fixed
|
||||||
|
|
||||||
|
## Common failure patterns
|
||||||
|
|
||||||
|
- Live timeout too short:
|
||||||
|
- widen live waits in `extensions/qa-lab/src/suite.ts`
|
||||||
|
- Discovery cannot find repo files:
|
||||||
|
- point prompts at `repo/...` inside seeded workspace
|
||||||
|
- Subagent proof too brittle:
|
||||||
|
- prefer stable final reply evidence over transient child-session listing
|
||||||
|
- Harness “rebuild” delay:
|
||||||
|
- dirty tree can trigger a pre-run build; expect that before ports appear
|
||||||
|
|
||||||
|
## When adding scenarios
|
||||||
|
|
||||||
|
- Add or update scenario YAML under `qa/scenarios/`; do not add `.md` scenario
|
||||||
|
files or fenced YAML blocks.
|
||||||
|
- Keep kickoff expectations in `qa/scenarios/index.yaml` aligned
|
||||||
|
- Add executable coverage in `extensions/qa-lab/src/suite.ts`
|
||||||
|
- Prefer end-to-end assertions over mock-only checks
|
||||||
|
- Save outputs under `.artifacts/qa-e2e/`
|
||||||
4
.agents/skills/openclaw-qa-testing/agents/openai.yaml
Normal file
4
.agents/skills/openclaw-qa-testing/agents/openai.yaml
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "QA Test OpenClaw"
|
||||||
|
short_description: "Run and debug qa-lab and qa-channel scenarios"
|
||||||
|
default_prompt: "Use $openclaw-qa-testing to run or extend the OpenClaw QA suite with qa-lab and qa-channel, using regular openai/gpt-5.4 in fast mode for live OpenAI runs."
|
||||||
196
.agents/skills/openclaw-refactor-docs/SKILL.md
Normal file
196
.agents/skills/openclaw-refactor-docs/SKILL.md
Normal file
@@ -0,0 +1,196 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-refactor-docs
|
||||||
|
description: Refactor an existing OpenClaw docs page with source-audited preservation, restructuring, and verification.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Refactor Docs
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Use this skill when the user gives a target OpenClaw docs page and asks to
|
||||||
|
rewrite, refactor, reorganize, split, shorten, or improve it.
|
||||||
|
|
||||||
|
This skill builds on `openclaw-docs`: use that skill for style, page types,
|
||||||
|
structure, examples, discoverability, and verification. This skill adds the
|
||||||
|
rewrite workflow needed to avoid losing accurate behavior during a major docs
|
||||||
|
refactor.
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
Required:
|
||||||
|
|
||||||
|
- A target docs page path, such as `docs/plugins/codex-harness.md`.
|
||||||
|
|
||||||
|
Optional:
|
||||||
|
|
||||||
|
- Desired page type, such as topic page, guide, reference, or troubleshooting.
|
||||||
|
- Specific goals, such as shorter main page, move details to reference pages, or
|
||||||
|
align with current CLI behavior.
|
||||||
|
- Related source files, schemas, commands, tests, specs, or PRs.
|
||||||
|
|
||||||
|
If the target page is missing or ambiguous, ask one concise question before
|
||||||
|
editing. Otherwise, proceed.
|
||||||
|
|
||||||
|
## Working Contract
|
||||||
|
|
||||||
|
Refactor the target page to be more useful, concise, and comprehensive within
|
||||||
|
its stated scope.
|
||||||
|
|
||||||
|
Do not treat a rewrite as permission to discard behavior facts. Preserve,
|
||||||
|
verify, move, or explicitly retire existing material. Incorrect docs are worse
|
||||||
|
than verbose docs.
|
||||||
|
|
||||||
|
Prefer this split:
|
||||||
|
|
||||||
|
- Topic or guide pages cover the 80/20 path, decisions readers must make, safe
|
||||||
|
setup, smallest reliable verification, common failures, and links onward.
|
||||||
|
- Reference pages cover exhaustive fields, defaults, enums, limits, precedence
|
||||||
|
rules, API contracts, narrow internals, and rare debugging details.
|
||||||
|
- Troubleshooting pages start from observable symptoms and map to checks,
|
||||||
|
causes, and fixes.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
### 1. Load the doc standard
|
||||||
|
|
||||||
|
Read `../openclaw-docs/SKILL.md` first. Apply its page-type, style,
|
||||||
|
examples, navigation, and verification guidance throughout the refactor.
|
||||||
|
|
||||||
|
Run `pnpm docs:list` when available, then read only the target page and the
|
||||||
|
likely entry points, references, or related pages needed for the refactor.
|
||||||
|
|
||||||
|
### 2. Classify the page
|
||||||
|
|
||||||
|
Before editing, decide the intended page type from `openclaw-docs`.
|
||||||
|
|
||||||
|
If the current page mixes page types, choose the main page type and plan where
|
||||||
|
the other material belongs:
|
||||||
|
|
||||||
|
- Move exhaustive contracts to an existing or new reference page.
|
||||||
|
- Move symptom-driven material to an existing or new troubleshooting page.
|
||||||
|
- Move narrow setup workflows to a guide when they interrupt the main path.
|
||||||
|
- Keep concise routing, decision, and safety details in the main page when
|
||||||
|
readers need them to complete the workflow.
|
||||||
|
|
||||||
|
### 3. Preserve and audit existing facts
|
||||||
|
|
||||||
|
Create a working inventory from the old page before rewriting. Include:
|
||||||
|
|
||||||
|
- Config fields, flags, commands, slash commands, env vars, defaults, enums,
|
||||||
|
nullable values, and constraints.
|
||||||
|
- Precedence rules, fallback behavior, caps, limits, rate limits, timeouts,
|
||||||
|
lifecycle states, queueing behavior, and compatibility rules.
|
||||||
|
- Auth, permission, approval, sandbox, safety, privacy, and destructive-action
|
||||||
|
behavior.
|
||||||
|
- Setup requirements, supported versions, dependencies, operating systems,
|
||||||
|
credentials, and account requirements.
|
||||||
|
- Error messages, troubleshooting symptoms, diagnostics, and recovery steps.
|
||||||
|
- Examples, expected output, command routing tables, and cross-links.
|
||||||
|
|
||||||
|
For each fact, choose one outcome:
|
||||||
|
|
||||||
|
- Keep it in the refactored target page.
|
||||||
|
- Move it to a specific existing page.
|
||||||
|
- Move it to a specific new page.
|
||||||
|
- Delete it because current source proves it is obsolete or out of scope.
|
||||||
|
|
||||||
|
Do not infer defaults, permissions, policy, timeout behavior, or safety posture
|
||||||
|
from names or intent. Verify them.
|
||||||
|
|
||||||
|
### 4. Find source of truth
|
||||||
|
|
||||||
|
Use the nearest authoritative source for each behavior-sensitive claim:
|
||||||
|
|
||||||
|
- Public schema, plugin manifest, generated config docs, or exported types for
|
||||||
|
config fields.
|
||||||
|
- CLI implementation, slash-command handlers, help text, and command tests for
|
||||||
|
commands and flags.
|
||||||
|
- Runtime source and tests for lifecycle, queueing, permission, fallback,
|
||||||
|
timeout, and provider behavior.
|
||||||
|
- Protocol docs, SDK facades, and contract tests for APIs and plugin surfaces.
|
||||||
|
- Existing docs only as secondary evidence unless the target is purely
|
||||||
|
conceptual.
|
||||||
|
|
||||||
|
If a page promises a reference, compare its tables against the schema,
|
||||||
|
manifest, CLI help, generated docs, or exported types. Missing public fields,
|
||||||
|
defaults, precedence rules, caps, or side effects are correctness bugs.
|
||||||
|
|
||||||
|
### 5. Plan moved material
|
||||||
|
|
||||||
|
When moving detail out of the target page, record the destination before
|
||||||
|
editing:
|
||||||
|
|
||||||
|
- Existing page: name the page and section.
|
||||||
|
- New page: choose the page type, slug, title, frontmatter summary,
|
||||||
|
`doc-schema-version: 1`, and `read_when` hints.
|
||||||
|
- Target page: keep a short summary and link from the point where readers need
|
||||||
|
the deeper detail.
|
||||||
|
|
||||||
|
Avoid duplicate truth. If the same contract appears in multiple places, choose
|
||||||
|
one canonical page and link to it.
|
||||||
|
|
||||||
|
### 6. Rewrite
|
||||||
|
|
||||||
|
Rewrite in this order:
|
||||||
|
|
||||||
|
1. Make the first screen answer what the reader can do and why this page exists.
|
||||||
|
2. Put the recommended path before alternatives.
|
||||||
|
3. Keep only decision-making and common operational detail in the main flow.
|
||||||
|
4. Move exhaustive tables and rare details to the planned reference pages.
|
||||||
|
5. Preserve concise routing tables when they help readers choose commands,
|
||||||
|
config paths, harnesses, plugins, providers, or references.
|
||||||
|
6. Add troubleshooting from observable symptoms, not internal guesses.
|
||||||
|
7. Link related concepts, guides, references, diagnostics, and adjacent tools.
|
||||||
|
|
||||||
|
Add `doc-schema-version: 1` to the YAML frontmatter of every docs page that the
|
||||||
|
refactor migrates, creates, or materially rewrites. Apply it only to docs page
|
||||||
|
files, not `docs.json`, glossary JSON, or other non-page metadata. If a
|
||||||
|
migrated page is generated, update the generator so regeneration preserves the
|
||||||
|
marker instead of hand-editing generated output.
|
||||||
|
|
||||||
|
Do not leave placeholders such as "TODO", "TBD", or "see docs" unless the user
|
||||||
|
explicitly asks for a draft.
|
||||||
|
|
||||||
|
### 7. Compare old and new
|
||||||
|
|
||||||
|
After editing, compare the old and new page:
|
||||||
|
|
||||||
|
- Confirm all behavior-sensitive facts were kept, moved, or intentionally
|
||||||
|
deleted with source-backed reason.
|
||||||
|
- Check that the main page still covers the 80/20 scenario end to end.
|
||||||
|
- Check that reference pages remain exhaustive for the scope they claim.
|
||||||
|
- Check that links from the target page reach moved details.
|
||||||
|
- Check that headings are stable, searchable, and action-oriented.
|
||||||
|
|
||||||
|
If the refactor deliberately removes relevant material, say where it went or why
|
||||||
|
it was removed in the final report.
|
||||||
|
|
||||||
|
### 8. Verify
|
||||||
|
|
||||||
|
Run the smallest reliable docs checks for the touched surface:
|
||||||
|
|
||||||
|
- `pnpm docs:list`
|
||||||
|
- `git diff --check -- <touched-files>`
|
||||||
|
- Targeted `pnpm exec oxfmt --check --threads=1 <touched-files>`
|
||||||
|
- `pnpm docs:check-mdx`
|
||||||
|
- `pnpm docs:check-links`
|
||||||
|
- `pnpm docs:check-i18n-glossary` when link text, navigation, labels, or glossary
|
||||||
|
surfaces changed
|
||||||
|
- Generated-doc checks when schemas, generated config docs, API docs, or
|
||||||
|
generated baselines are touched
|
||||||
|
|
||||||
|
Run commands and examples from the page whenever feasible. If you cannot verify
|
||||||
|
a behavior-sensitive claim, either remove the claim, mark the uncertainty in the
|
||||||
|
work-in-progress report, or ask for the missing source.
|
||||||
|
|
||||||
|
## Final Report
|
||||||
|
|
||||||
|
Report:
|
||||||
|
|
||||||
|
- What changed in the target page.
|
||||||
|
- What details moved and their destination pages.
|
||||||
|
- What source-of-truth checks backed behavior-sensitive claims.
|
||||||
|
- What validation ran and what failed for unrelated reasons.
|
||||||
|
|
||||||
|
Do not include a long rewrite diary. Lead with remaining risks only if there are
|
||||||
|
any.
|
||||||
236
.agents/skills/openclaw-secret-scanning-maintainer/SKILL.md
Normal file
236
.agents/skills/openclaw-secret-scanning-maintainer/SKILL.md
Normal file
@@ -0,0 +1,236 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-secret-scanning-maintainer
|
||||||
|
description: Triage, redact, clean up, and resolve OpenClaw GitHub Secret Scanning alerts in issues or PRs.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Secret Scanning Maintainer
|
||||||
|
|
||||||
|
**Maintainer-only.** This skill requires repo admin / maintainer permissions to edit or delete other users' comments and resolve secret scanning alerts.
|
||||||
|
|
||||||
|
Use this skill when processing alerts from `https://github.com/openclaw/openclaw/security/secret-scanning`.
|
||||||
|
|
||||||
|
**Language rule:** All notification comments and replacement comments MUST be written in English.
|
||||||
|
|
||||||
|
## Script
|
||||||
|
|
||||||
|
All mechanical operations (API calls, temp file management, security enforcements) are handled by:
|
||||||
|
|
||||||
|
```
|
||||||
|
$REPO_ROOT/.agents/skills/openclaw-secret-scanning-maintainer/scripts/secret-scanning.mjs
|
||||||
|
```
|
||||||
|
|
||||||
|
The script enforces:
|
||||||
|
|
||||||
|
- `hide_secret=true` on all alert fetches (no plaintext secrets in stdout)
|
||||||
|
- `mktemp` with random UUIDs for all temp files
|
||||||
|
- `-F body=@file` for all body uploads (no inline shell quoting)
|
||||||
|
- Notification templates branched by location type
|
||||||
|
- Never prints `.secret` or `.body` to stdout
|
||||||
|
|
||||||
|
## Overall Flow
|
||||||
|
|
||||||
|
Supports single or multiple alerts. For multiple alerts, process in ascending order.
|
||||||
|
|
||||||
|
For each alert:
|
||||||
|
|
||||||
|
1. **Identify** — `fetch-alert` + `fetch-content` to get metadata and body
|
||||||
|
2. **Decide** — Agent reads the body file, identifies whether plaintext secrets remain, and produces a redacted version only when needed
|
||||||
|
3. **Redact** — `redact-body-if-needed` for issue/PR body; skip for comments (delete directly)
|
||||||
|
4. **Purge** — `delete-comment` + `recreate-comment` for comments; cannot purge body history
|
||||||
|
5. **Notify** — `notify` posts the right template per location type, unless the current issue/PR body is already redacted
|
||||||
|
6. **Resolve** — `resolve` closes the alert
|
||||||
|
7. **Summary** — `summary` prints formatted results
|
||||||
|
|
||||||
|
## Step 1: Identify
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# List all open alerts
|
||||||
|
node secret-scanning.mjs list-open
|
||||||
|
|
||||||
|
# Fetch specific alert metadata + locations
|
||||||
|
node secret-scanning.mjs fetch-alert <NUMBER>
|
||||||
|
|
||||||
|
# Fetch content for each location (saves body to temp file)
|
||||||
|
node secret-scanning.mjs fetch-content '<location-json>'
|
||||||
|
```
|
||||||
|
|
||||||
|
The `fetch-content` output includes:
|
||||||
|
|
||||||
|
- `body_file`: path to temp file with full body content
|
||||||
|
- `author`: who posted it
|
||||||
|
- `issue_number` / `pr_number`: where it is
|
||||||
|
- `edit_history_count`: number of existing edits
|
||||||
|
- `type`: location type for routing
|
||||||
|
- For `discussion_comment`, it also includes `comment_node_id`, `discussion_node_id`, and `reply_to_node_id` when the original comment was a reply.
|
||||||
|
|
||||||
|
### Location type routing
|
||||||
|
|
||||||
|
| type | Flow |
|
||||||
|
| ----------------------------- | --------------------------------------------- |
|
||||||
|
| `issue_comment` | Comment: delete+recreate |
|
||||||
|
| `pull_request_comment` | Comment: delete+recreate |
|
||||||
|
| `pull_request_review_comment` | Comment: delete+recreate |
|
||||||
|
| `discussion_comment` | Discussion comment: delete+recreate (GraphQL) |
|
||||||
|
| `issue_body` | Body: redact in place |
|
||||||
|
| `pull_request_body` | Body: redact in place |
|
||||||
|
| `commit` | Notify only |
|
||||||
|
| _other_ | Skip and report |
|
||||||
|
|
||||||
|
## Step 2: Decide (Agent)
|
||||||
|
|
||||||
|
The agent reads the body file from `fetch-content` output and:
|
||||||
|
|
||||||
|
1. Identifies ALL secrets in the content (there may be more than the alert flagged)
|
||||||
|
2. Determines whether any plaintext credential remains in the current body
|
||||||
|
3. Replaces each remaining secret with `[REDACTED <secret_type>]` — **no partial values, no prefix/suffix**
|
||||||
|
4. Saves the redacted content to a new temp file
|
||||||
|
|
||||||
|
This is the only step that requires semantic understanding. Everything else is mechanical.
|
||||||
|
|
||||||
|
For `issue_body` and `pull_request_body`: if the current body has already been redacted by the author and no plaintext credential remains, **do not post a public notification comment**. Resolve the alert with a maintainer-only resolution comment such as:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node secret-scanning.mjs resolve <ALERT_NUMBER> revoked "Current issue/PR body is already redacted; no public notification posted."
|
||||||
|
```
|
||||||
|
|
||||||
|
This avoids creating a fresh public pointer to historical sensitive content.
|
||||||
|
|
||||||
|
## Step 3: Redact
|
||||||
|
|
||||||
|
### For comments (issue_comment / PR comments)
|
||||||
|
|
||||||
|
**Do NOT redact.** Skip directly to Step 4 (delete + recreate). PATCHing before DELETE creates an unnecessary edit history revision.
|
||||||
|
|
||||||
|
### For issue_body / pull_request_body
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node secret-scanning.mjs redact-body-if-needed <issue|pr> <NUMBER> <current-body-file> <redacted-body-file> <result-file>
|
||||||
|
```
|
||||||
|
|
||||||
|
Use the `body_file` from `fetch-content` as `<current-body-file>`. The command writes `notify_required` to `<result-file>` and only PATCHes the body when the redacted file differs from the current body.
|
||||||
|
|
||||||
|
## Step 4: Purge Edit History
|
||||||
|
|
||||||
|
### Comments — Delete and Recreate
|
||||||
|
|
||||||
|
For issue/PR comments:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Delete original (all edit history gone)
|
||||||
|
node secret-scanning.mjs delete-comment <COMMENT_ID>
|
||||||
|
|
||||||
|
# Recreate with redacted content
|
||||||
|
node secret-scanning.mjs recreate-comment <ISSUE_NUMBER> <body-file>
|
||||||
|
```
|
||||||
|
|
||||||
|
For discussion comments (uses GraphQL):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Delete original
|
||||||
|
node secret-scanning.mjs delete-discussion-comment <COMMENT_NODE_ID>
|
||||||
|
|
||||||
|
# Recreate with redacted content
|
||||||
|
node secret-scanning.mjs recreate-discussion-comment <DISCUSSION_NODE_ID> <body-file> [REPLY_TO_NODE_ID]
|
||||||
|
```
|
||||||
|
|
||||||
|
The `fetch-content` output for `discussion_comment` includes `comment_node_id` and `discussion_node_id` for these commands. When the original discussion comment was a reply, it also includes `reply_to_node_id`; pass that optional third argument so the redacted replacement stays in the original thread.
|
||||||
|
|
||||||
|
The recreated comment should follow this format:
|
||||||
|
|
||||||
|
```
|
||||||
|
> **Note:** The original comment by @<AUTHOR> has been removed due to secret leakage. Below is the redacted version of the original content.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<redacted original content>
|
||||||
|
```
|
||||||
|
|
||||||
|
### issue_body / pull_request_body — Cannot Purge Edit History
|
||||||
|
|
||||||
|
Editing creates an edit history revision with the pre-edit plaintext. This cannot be cleared via API.
|
||||||
|
|
||||||
|
Do not advise authors publicly to delete/recreate issues or close/reopen PRs. That can draw attention to historical content. Keep purge guidance maintainer-only.
|
||||||
|
|
||||||
|
**Output to maintainer terminal only (never in public comments):**
|
||||||
|
|
||||||
|
```
|
||||||
|
⚠️ Issue/PR body edit history still contains plaintext secrets.
|
||||||
|
Contact GitHub Support to purge: https://support.github.com/contact
|
||||||
|
Request purge of issue/PR #{NUMBER} userContentEdits.
|
||||||
|
```
|
||||||
|
|
||||||
|
> **CRITICAL:** Do NOT mention edit history or the "edited" button in any public comment or resolution_comment.
|
||||||
|
|
||||||
|
### Commits
|
||||||
|
|
||||||
|
Cannot clean. Notify author to delete branch or force-push (for unmerged PRs).
|
||||||
|
|
||||||
|
## Step 5: Notify
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node secret-scanning.mjs notify <TARGET> <AUTHOR> <LOCATION_TYPE> <SECRET_TYPES> [REPLY_TO_NODE_ID|BODY_REDACTION_RESULT_FILE]
|
||||||
|
```
|
||||||
|
|
||||||
|
- For non-discussion types, `<TARGET>` is the issue/PR number.
|
||||||
|
- For `discussion_comment`, `<TARGET>` is the `discussion_node_id` returned by `fetch-content`.
|
||||||
|
- For reply-style `discussion_comment` locations, pass the optional `reply_to_node_id` from `fetch-content` so the notification stays in the same thread.
|
||||||
|
- For `issue_body` and `pull_request_body`, pass the `<result-file>` from `redact-body-if-needed`. The script skips notification when `notify_required` is `false` and refuses body notifications without this file.
|
||||||
|
|
||||||
|
Secret types are comma-separated: `"Discord Bot Token,Feishu App Secret"`
|
||||||
|
|
||||||
|
The script picks the right template:
|
||||||
|
|
||||||
|
- **comment types**: "your comment … removed and replaced"
|
||||||
|
- **body types**: "your issue/PR description … redacted in place"
|
||||||
|
- **commit**: "code you committed"
|
||||||
|
|
||||||
|
For `issue_body` and `pull_request_body`, only notify when the current body still contained plaintext and maintainers redacted it. If the user already redacted the current body, skip this step and resolve silently.
|
||||||
|
|
||||||
|
## Step 6: Resolve
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node secret-scanning.mjs resolve <ALERT_NUMBER>
|
||||||
|
# or with custom resolution:
|
||||||
|
node secret-scanning.mjs resolve <ALERT_NUMBER> revoked "Custom comment"
|
||||||
|
```
|
||||||
|
|
||||||
|
Resolution is `revoked` by default. As maintainers we cannot control whether users rotate — our responsibility is to remove current plaintext exposure and notify only when public notification is useful. The `revoked` means "this secret should be considered leaked", not "I confirmed it was revoked".
|
||||||
|
|
||||||
|
## Step 7: Summary
|
||||||
|
|
||||||
|
After processing, create a JSON results file and pass it to the summary command:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node secret-scanning.mjs summary /tmp/results.json
|
||||||
|
```
|
||||||
|
|
||||||
|
The script outputs a block delimited by `---BEGIN SUMMARY---` and `---END SUMMARY---`. **You MUST output the content between these markers verbatim to the user. Do NOT rephrase, reformat, abbreviate, or create your own summary.** The script already includes full URLs for every alert and location.
|
||||||
|
|
||||||
|
The JSON format:
|
||||||
|
|
||||||
|
```json
|
||||||
|
[
|
||||||
|
{
|
||||||
|
"number": 72,
|
||||||
|
"secret_type": "Discord Bot Token",
|
||||||
|
"location_label": "Issue #63101 comment",
|
||||||
|
"location_url": "https://github.com/openclaw/openclaw/issues/63101#issuecomment-xxx",
|
||||||
|
"actions": "Deleted+Recreated+Notified",
|
||||||
|
"history_cleared": true
|
||||||
|
}
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
For unsupported types, add `"skipped": true, "unsupported_type": "<type>"`.
|
||||||
|
|
||||||
|
## Safety Rules
|
||||||
|
|
||||||
|
- **Agent reads content, identifies secrets, produces redaction.** Script handles all API calls.
|
||||||
|
- **Never include any portion of a secret** in public comments, redaction markers, or terminal output.
|
||||||
|
- **Never include alert URLs or numbers** in public comments.
|
||||||
|
- **For comments, skip PATCH — go directly to DELETE + recreate.**
|
||||||
|
- **Never mention edit history, "edited" button, or commit SHAs** in any public content.
|
||||||
|
- **Ask for confirmation** before deleting any comment.
|
||||||
|
- **One alert at a time** unless user requests batch.
|
||||||
|
- **All public comments in English.**
|
||||||
|
- **Skip unsupported location types** and report in summary.
|
||||||
@@ -0,0 +1,938 @@
|
|||||||
|
#!/usr/bin/env node
|
||||||
|
/**
|
||||||
|
* Secret scanning alert handler for OpenClaw maintainers.
|
||||||
|
* Usage: node secret-scanning.mjs <command> [options]
|
||||||
|
*/
|
||||||
|
|
||||||
|
import crypto from "node:crypto";
|
||||||
|
import fs from "node:fs";
|
||||||
|
import os from "node:os";
|
||||||
|
import path from "node:path";
|
||||||
|
import { pathToFileURL } from "node:url";
|
||||||
|
import { spawnPlainGh } from "../../../../scripts/lib/plain-gh.mjs";
|
||||||
|
|
||||||
|
const REPO = "openclaw/openclaw";
|
||||||
|
const REPO_URL = `https://github.com/${REPO}`;
|
||||||
|
|
||||||
|
// ─── Helpers ────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
function fail(message) {
|
||||||
|
console.error(`error: ${message}`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
function tmpFile(purpose) {
|
||||||
|
const filePath = path.join(os.tmpdir(), `secretscan-${purpose}-${crypto.randomUUID()}`);
|
||||||
|
// 预创建文件,限制权限为 owner-only
|
||||||
|
fs.writeFileSync(filePath, "", { mode: 0o600 });
|
||||||
|
return filePath;
|
||||||
|
}
|
||||||
|
|
||||||
|
function gh(args, { json = true, allowFailure = false } = {}) {
|
||||||
|
const proc = spawnPlainGh(args, { encoding: "utf8", maxBuffer: 10 * 1024 * 1024 });
|
||||||
|
if (proc.status !== 0 && !allowFailure) {
|
||||||
|
fail(`gh ${args.slice(0, 3).join(" ")} failed:\n${(proc.stderr || proc.stdout || "").trim()}`);
|
||||||
|
}
|
||||||
|
if (proc.status !== 0) {
|
||||||
|
return {
|
||||||
|
gh_failed: true,
|
||||||
|
status: proc.status,
|
||||||
|
stdout: proc.stdout,
|
||||||
|
stderr: proc.stderr,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
if (!json) {
|
||||||
|
return proc.stdout;
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
return JSON.parse(proc.stdout);
|
||||||
|
} catch {
|
||||||
|
return proc.stdout;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function ghGraphQL(query, options = {}) {
|
||||||
|
return gh(["api", "graphql", "-f", `query=${query}`], options);
|
||||||
|
}
|
||||||
|
|
||||||
|
function isBodyLocationType(locationType) {
|
||||||
|
return locationType === "issue_body" || locationType === "pull_request_body";
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Decides whether redacting an issue/PR body requires notifying the reporter. */
|
||||||
|
export function decideBodyRedaction(currentBody, redactedBody) {
|
||||||
|
const bodyChanged = String(currentBody) !== String(redactedBody);
|
||||||
|
return {
|
||||||
|
body_changed: bodyChanged,
|
||||||
|
notify_required: bodyChanged,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Loads redaction-result metadata for issue/PR body secret locations. */
|
||||||
|
export function loadBodyRedactionResult(locationType, resultFile) {
|
||||||
|
if (!isBodyLocationType(locationType)) {
|
||||||
|
return { notify_required: true };
|
||||||
|
}
|
||||||
|
if (!resultFile) {
|
||||||
|
fail("Body notifications require a redaction result file from redact-body-if-needed");
|
||||||
|
}
|
||||||
|
if (!fs.existsSync(resultFile)) {
|
||||||
|
fail(`File not found: ${resultFile}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const result = JSON.parse(fs.readFileSync(resultFile, "utf8"));
|
||||||
|
if (typeof result.notify_required !== "boolean") {
|
||||||
|
fail(`Invalid redaction result file: missing boolean notify_required in ${resultFile}`);
|
||||||
|
}
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
|
function failOnGraphQLFailure(result, message) {
|
||||||
|
if (result?.gh_failed) {
|
||||||
|
const details = (
|
||||||
|
result.stderr ||
|
||||||
|
result.stdout ||
|
||||||
|
`gh exited with status ${result.status}`
|
||||||
|
).trim();
|
||||||
|
fail(`${message}: ${details}`);
|
||||||
|
}
|
||||||
|
if (Array.isArray(result?.errors) && result.errors.length > 0) {
|
||||||
|
fail(`${message}: ${JSON.stringify(result.errors)}`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function escapeGraphQLString(value) {
|
||||||
|
return String(value)
|
||||||
|
.replace(/\\/g, "\\\\")
|
||||||
|
.replace(/"/g, '\\"')
|
||||||
|
.replace(/\r/g, "\\r")
|
||||||
|
.replace(/\n/g, "\\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
function formatGraphQLAfterClause(cursor) {
|
||||||
|
return cursor ? `, after: "${escapeGraphQLString(cursor)}"` : "";
|
||||||
|
}
|
||||||
|
|
||||||
|
function findDiscussionCommentNode(nodes, discussionCommentDbId) {
|
||||||
|
return nodes.find((node) => String(node.databaseId) === String(discussionCommentDbId)) || null;
|
||||||
|
}
|
||||||
|
|
||||||
|
function fetchDiscussionReplyPage(commentNodeId, cursor) {
|
||||||
|
const afterClause = formatGraphQLAfterClause(cursor);
|
||||||
|
return ghGraphQL(`{
|
||||||
|
node(id: "${escapeGraphQLString(commentNodeId)}") {
|
||||||
|
... on DiscussionComment {
|
||||||
|
replies(first: 100${afterClause}) {
|
||||||
|
pageInfo { hasNextPage endCursor }
|
||||||
|
nodes {
|
||||||
|
id
|
||||||
|
databaseId
|
||||||
|
author { login }
|
||||||
|
body
|
||||||
|
url
|
||||||
|
replyTo { id }
|
||||||
|
userContentEdits(first: 50) {
|
||||||
|
totalCount
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
function fetchDiscussionComment(discussionNumber, discussionCommentDbId) {
|
||||||
|
const [owner, name] = REPO.split("/");
|
||||||
|
let discussionId = null;
|
||||||
|
let cursor = null;
|
||||||
|
let hasNextPage = true;
|
||||||
|
|
||||||
|
while (hasNextPage) {
|
||||||
|
const afterClause = formatGraphQLAfterClause(cursor);
|
||||||
|
const gql = ghGraphQL(
|
||||||
|
`{
|
||||||
|
repository(owner: "${owner}", name: "${name}") {
|
||||||
|
discussion(number: ${discussionNumber}) {
|
||||||
|
id
|
||||||
|
comments(first: 50${afterClause}) {
|
||||||
|
pageInfo { hasNextPage endCursor }
|
||||||
|
nodes {
|
||||||
|
id
|
||||||
|
databaseId
|
||||||
|
author { login }
|
||||||
|
body
|
||||||
|
url
|
||||||
|
replyTo { id }
|
||||||
|
userContentEdits(first: 50) {
|
||||||
|
totalCount
|
||||||
|
}
|
||||||
|
replies(first: 100) {
|
||||||
|
pageInfo { hasNextPage endCursor }
|
||||||
|
nodes {
|
||||||
|
id
|
||||||
|
databaseId
|
||||||
|
author { login }
|
||||||
|
body
|
||||||
|
url
|
||||||
|
replyTo { id }
|
||||||
|
userContentEdits(first: 50) {
|
||||||
|
totalCount
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}`,
|
||||||
|
{ allowFailure: true },
|
||||||
|
);
|
||||||
|
failOnGraphQLFailure(gql, `Failed to fetch discussion #${discussionNumber}`);
|
||||||
|
|
||||||
|
const discussion = gql?.data?.repository?.discussion;
|
||||||
|
if (!discussion) {
|
||||||
|
fail(
|
||||||
|
`Discussion #${discussionNumber} not found — it may have been deleted. The alert cannot be processed via this skill.`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
discussionId = discussion.id;
|
||||||
|
|
||||||
|
for (const topLevelComment of discussion.comments.nodes) {
|
||||||
|
if (String(topLevelComment.databaseId) === String(discussionCommentDbId)) {
|
||||||
|
return { discussionId, comment: topLevelComment };
|
||||||
|
}
|
||||||
|
|
||||||
|
let reply = findDiscussionCommentNode(topLevelComment.replies.nodes, discussionCommentDbId);
|
||||||
|
let replyCursor = topLevelComment.replies.pageInfo.endCursor;
|
||||||
|
let hasMoreReplies = topLevelComment.replies.pageInfo.hasNextPage;
|
||||||
|
|
||||||
|
while (!reply && hasMoreReplies) {
|
||||||
|
const replyPage = fetchDiscussionReplyPage(topLevelComment.id, replyCursor);
|
||||||
|
failOnGraphQLFailure(
|
||||||
|
replyPage,
|
||||||
|
`Failed to fetch replies for discussion comment ${topLevelComment.id}`,
|
||||||
|
);
|
||||||
|
const replies = replyPage?.data?.node?.replies;
|
||||||
|
if (!replies) {
|
||||||
|
fail(`Failed to paginate replies for discussion comment ${topLevelComment.id}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
reply = findDiscussionCommentNode(replies.nodes, discussionCommentDbId);
|
||||||
|
hasMoreReplies = replies.pageInfo.hasNextPage;
|
||||||
|
replyCursor = replies.pageInfo.endCursor;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (reply) {
|
||||||
|
return { discussionId, comment: reply };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
hasNextPage = discussion.comments.pageInfo.hasNextPage;
|
||||||
|
cursor = discussion.comments.pageInfo.endCursor;
|
||||||
|
}
|
||||||
|
|
||||||
|
return { discussionId, comment: null };
|
||||||
|
}
|
||||||
|
|
||||||
|
function createDiscussionComment(discussionNodeId, body, replyToNodeId) {
|
||||||
|
const replyToClause = replyToNodeId ? `, replyToId: "${escapeGraphQLString(replyToNodeId)}"` : "";
|
||||||
|
const result = ghGraphQL(
|
||||||
|
`mutation { addDiscussionComment(input: { discussionId: "${escapeGraphQLString(discussionNodeId)}"${replyToClause}, body: "${escapeGraphQLString(body)}" }) { comment { id url } } }`,
|
||||||
|
);
|
||||||
|
if (result?.errors) {
|
||||||
|
fail(`Failed to create discussion comment: ${JSON.stringify(result.errors)}`);
|
||||||
|
}
|
||||||
|
return result?.data?.addDiscussionComment?.comment;
|
||||||
|
}
|
||||||
|
|
||||||
|
// ─── Commands ───────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fetch-alert <number>
|
||||||
|
* Fetch alert metadata + locations. Never exposes .secret.
|
||||||
|
*/
|
||||||
|
function cmdFetchAlert(alertNumber) {
|
||||||
|
if (!alertNumber) {
|
||||||
|
fail("Usage: fetch-alert <number>");
|
||||||
|
}
|
||||||
|
|
||||||
|
const alert = gh(["api", `repos/${REPO}/secret-scanning/alerts/${alertNumber}?hide_secret=true`]);
|
||||||
|
|
||||||
|
const locations = gh([
|
||||||
|
"api",
|
||||||
|
`repos/${REPO}/secret-scanning/alerts/${alertNumber}/locations`,
|
||||||
|
"--paginate",
|
||||||
|
"--slurp",
|
||||||
|
]);
|
||||||
|
// --paginate + --slurp 确保多页结果合并为一个 JSON 数组
|
||||||
|
const flatLocations = Array.isArray(locations?.[0])
|
||||||
|
? locations.flat()
|
||||||
|
: Array.isArray(locations)
|
||||||
|
? locations
|
||||||
|
: [];
|
||||||
|
|
||||||
|
const result = {
|
||||||
|
number: alert.number,
|
||||||
|
state: alert.state,
|
||||||
|
secret_type: alert.secret_type,
|
||||||
|
secret_type_display_name: alert.secret_type_display_name,
|
||||||
|
validity: alert.validity,
|
||||||
|
html_url: alert.html_url,
|
||||||
|
locations: flatLocations.map((loc) => ({
|
||||||
|
type: loc.type,
|
||||||
|
details: loc.details,
|
||||||
|
})),
|
||||||
|
};
|
||||||
|
|
||||||
|
console.log(JSON.stringify(result, null, 2));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fetch-content <location-json>
|
||||||
|
* Fetch the content and metadata for a specific location.
|
||||||
|
* Saves full body to a temp file. Prints metadata + file path to stdout.
|
||||||
|
*/
|
||||||
|
function cmdFetchContent(locationJson) {
|
||||||
|
if (!locationJson) {
|
||||||
|
fail("Usage: fetch-content '<location-json>'");
|
||||||
|
}
|
||||||
|
const location = JSON.parse(locationJson);
|
||||||
|
const type = location.type;
|
||||||
|
const details = location.details;
|
||||||
|
|
||||||
|
if (type === "discussion_comment") {
|
||||||
|
const commentUrl = details.discussion_comment_url;
|
||||||
|
if (!commentUrl) {
|
||||||
|
fail("No discussion_comment_url in location details");
|
||||||
|
}
|
||||||
|
|
||||||
|
const urlMatch = commentUrl.match(/discussions\/(\d+)#discussioncomment-(\d+)/);
|
||||||
|
if (!urlMatch) {
|
||||||
|
fail(`Cannot parse discussion comment URL: ${commentUrl}`);
|
||||||
|
}
|
||||||
|
const discussionNumber = urlMatch[1];
|
||||||
|
const discussionCommentDbId = urlMatch[2];
|
||||||
|
|
||||||
|
const { discussionId, comment } = fetchDiscussionComment(
|
||||||
|
discussionNumber,
|
||||||
|
discussionCommentDbId,
|
||||||
|
);
|
||||||
|
if (!comment) {
|
||||||
|
fail(
|
||||||
|
`Discussion comment #${discussionCommentDbId} not found in discussion #${discussionNumber}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const bodyFile = tmpFile("body.md");
|
||||||
|
fs.writeFileSync(bodyFile, comment.body || "");
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
JSON.stringify(
|
||||||
|
{
|
||||||
|
type,
|
||||||
|
comment_node_id: comment.id,
|
||||||
|
discussion_node_id: discussionId,
|
||||||
|
reply_to_node_id: comment.replyTo?.id ?? null,
|
||||||
|
discussion_number: Number(discussionNumber),
|
||||||
|
discussion_comment_db_id: Number(discussionCommentDbId),
|
||||||
|
author: comment.author?.login,
|
||||||
|
html_url: comment.url || commentUrl,
|
||||||
|
edit_history_count: comment.userContentEdits?.totalCount ?? 0,
|
||||||
|
body_file: bodyFile,
|
||||||
|
},
|
||||||
|
null,
|
||||||
|
2,
|
||||||
|
),
|
||||||
|
);
|
||||||
|
} else if (
|
||||||
|
type === "issue_comment" ||
|
||||||
|
type === "pull_request_comment" ||
|
||||||
|
type === "pull_request_review_comment"
|
||||||
|
) {
|
||||||
|
// Extract comment ID from URL
|
||||||
|
const commentUrl =
|
||||||
|
details.issue_comment_url ||
|
||||||
|
details.pull_request_comment_url ||
|
||||||
|
details.pull_request_review_comment_url;
|
||||||
|
if (!commentUrl) {
|
||||||
|
fail(`No comment URL in location details`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const comment = gh(["api", commentUrl]);
|
||||||
|
const bodyFile = tmpFile("body.md");
|
||||||
|
fs.writeFileSync(bodyFile, comment.body || "");
|
||||||
|
|
||||||
|
// Fetch edit history
|
||||||
|
const nodeId = comment.node_id;
|
||||||
|
const typeName =
|
||||||
|
type === "pull_request_review_comment" ? "PullRequestReviewComment" : "IssueComment";
|
||||||
|
const gql = ghGraphQL(`{
|
||||||
|
node(id: "${nodeId}") {
|
||||||
|
... on ${typeName} {
|
||||||
|
userContentEdits(first: 50) {
|
||||||
|
totalCount
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}`);
|
||||||
|
const editCount = gql?.data?.node?.userContentEdits?.totalCount ?? 0;
|
||||||
|
|
||||||
|
// Extract issue number from html_url
|
||||||
|
const htmlUrl = comment.html_url || details.html_url || "";
|
||||||
|
const issueMatch = htmlUrl.match(/\/(issues|pull)\/(\d+)/);
|
||||||
|
const issueNumber = issueMatch ? issueMatch[2] : null;
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
JSON.stringify(
|
||||||
|
{
|
||||||
|
type,
|
||||||
|
comment_id: comment.id,
|
||||||
|
node_id: nodeId,
|
||||||
|
author: comment.user?.login,
|
||||||
|
issue_number: issueNumber,
|
||||||
|
html_url: htmlUrl,
|
||||||
|
edit_history_count: editCount,
|
||||||
|
body_file: bodyFile,
|
||||||
|
},
|
||||||
|
null,
|
||||||
|
2,
|
||||||
|
),
|
||||||
|
);
|
||||||
|
} else if (type === "issue_body") {
|
||||||
|
const issueUrl = details.issue_body_url || details.issue_url;
|
||||||
|
if (!issueUrl) {
|
||||||
|
fail("No issue URL in location details");
|
||||||
|
}
|
||||||
|
|
||||||
|
const issue = gh(["api", issueUrl]);
|
||||||
|
const bodyFile = tmpFile("body.md");
|
||||||
|
fs.writeFileSync(bodyFile, issue.body || "");
|
||||||
|
|
||||||
|
const nodeId = issue.node_id;
|
||||||
|
const number = issue.number;
|
||||||
|
const gql = ghGraphQL(`{
|
||||||
|
node(id: "${nodeId}") {
|
||||||
|
... on Issue {
|
||||||
|
userContentEdits(first: 50) {
|
||||||
|
totalCount
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}`);
|
||||||
|
const editCount = gql?.data?.node?.userContentEdits?.totalCount ?? 0;
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
JSON.stringify(
|
||||||
|
{
|
||||||
|
type,
|
||||||
|
issue_number: number,
|
||||||
|
node_id: nodeId,
|
||||||
|
author: issue.user?.login,
|
||||||
|
html_url: issue.html_url,
|
||||||
|
edit_history_count: editCount,
|
||||||
|
body_file: bodyFile,
|
||||||
|
},
|
||||||
|
null,
|
||||||
|
2,
|
||||||
|
),
|
||||||
|
);
|
||||||
|
} else if (type === "pull_request_body") {
|
||||||
|
const prUrl = details.pull_request_body_url || details.pull_request_url;
|
||||||
|
if (!prUrl) {
|
||||||
|
fail("No PR URL in location details");
|
||||||
|
}
|
||||||
|
|
||||||
|
const pr = gh(["api", prUrl]);
|
||||||
|
const bodyFile = tmpFile("body.md");
|
||||||
|
fs.writeFileSync(bodyFile, pr.body || "");
|
||||||
|
|
||||||
|
const nodeId = pr.node_id;
|
||||||
|
const number = pr.number;
|
||||||
|
const gql = ghGraphQL(`{
|
||||||
|
node(id: "${nodeId}") {
|
||||||
|
... on PullRequest {
|
||||||
|
userContentEdits(first: 50) {
|
||||||
|
totalCount
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}`);
|
||||||
|
const editCount = gql?.data?.node?.userContentEdits?.totalCount ?? 0;
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
JSON.stringify(
|
||||||
|
{
|
||||||
|
type,
|
||||||
|
pr_number: number,
|
||||||
|
node_id: nodeId,
|
||||||
|
author: pr.user?.login,
|
||||||
|
merged: pr.merged,
|
||||||
|
state: pr.state,
|
||||||
|
html_url: pr.html_url,
|
||||||
|
edit_history_count: editCount,
|
||||||
|
body_file: bodyFile,
|
||||||
|
},
|
||||||
|
null,
|
||||||
|
2,
|
||||||
|
),
|
||||||
|
);
|
||||||
|
} else if (type === "commit") {
|
||||||
|
console.log(
|
||||||
|
JSON.stringify(
|
||||||
|
{
|
||||||
|
type,
|
||||||
|
commit_sha: details.commit_sha,
|
||||||
|
path: details.path,
|
||||||
|
start_line: details.start_line,
|
||||||
|
end_line: details.end_line,
|
||||||
|
html_url: details.html_url || details.commit_url || details.blob_url || null,
|
||||||
|
// No body file for commits
|
||||||
|
body_file: null,
|
||||||
|
},
|
||||||
|
null,
|
||||||
|
2,
|
||||||
|
),
|
||||||
|
);
|
||||||
|
} else {
|
||||||
|
console.log(
|
||||||
|
JSON.stringify(
|
||||||
|
{
|
||||||
|
type,
|
||||||
|
unsupported: true,
|
||||||
|
details,
|
||||||
|
},
|
||||||
|
null,
|
||||||
|
2,
|
||||||
|
),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* redact-body <issue|pr> <number> <redacted-body-file>
|
||||||
|
* PATCH the issue or PR body with redacted content from a file.
|
||||||
|
*/
|
||||||
|
function cmdRedactBody(kind, number, bodyFile) {
|
||||||
|
if (!kind || !number || !bodyFile) {
|
||||||
|
fail("Usage: redact-body <issue|pr> <number> <redacted-body-file>");
|
||||||
|
}
|
||||||
|
if (!fs.existsSync(bodyFile)) {
|
||||||
|
fail(`File not found: ${bodyFile}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const endpoint =
|
||||||
|
kind === "pr" ? `repos/${REPO}/pulls/${number}` : `repos/${REPO}/issues/${number}`;
|
||||||
|
|
||||||
|
gh(["api", endpoint, "-X", "PATCH", "-F", `body=@${bodyFile}`]);
|
||||||
|
console.log(JSON.stringify({ ok: true, kind, number: Number(number) }));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* redact-body-if-needed <issue|pr> <number> <current-body-file> <redacted-body-file> <result-file>
|
||||||
|
* PATCH only when the agent-produced redacted body differs from the current body.
|
||||||
|
*/
|
||||||
|
function cmdRedactBodyIfNeeded(kind, number, currentBodyFile, redactedBodyFile, resultFile) {
|
||||||
|
if (!kind || !number || !currentBodyFile || !redactedBodyFile || !resultFile) {
|
||||||
|
fail(
|
||||||
|
"Usage: redact-body-if-needed <issue|pr> <number> <current-body-file> <redacted-body-file> <result-file>",
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (!fs.existsSync(currentBodyFile)) {
|
||||||
|
fail(`File not found: ${currentBodyFile}`);
|
||||||
|
}
|
||||||
|
if (!fs.existsSync(redactedBodyFile)) {
|
||||||
|
fail(`File not found: ${redactedBodyFile}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const currentBody = fs.readFileSync(currentBodyFile, "utf8");
|
||||||
|
const redactedBody = fs.readFileSync(redactedBodyFile, "utf8");
|
||||||
|
const decision = decideBodyRedaction(currentBody, redactedBody);
|
||||||
|
const result = {
|
||||||
|
ok: true,
|
||||||
|
kind,
|
||||||
|
number: Number(number),
|
||||||
|
...decision,
|
||||||
|
};
|
||||||
|
|
||||||
|
if (decision.body_changed) {
|
||||||
|
const endpoint =
|
||||||
|
kind === "pr" ? `repos/${REPO}/pulls/${number}` : `repos/${REPO}/issues/${number}`;
|
||||||
|
gh(["api", endpoint, "-X", "PATCH", "-F", `body=@${redactedBodyFile}`]);
|
||||||
|
result.redacted = true;
|
||||||
|
} else {
|
||||||
|
result.redacted = false;
|
||||||
|
result.reason = "current_body_already_redacted";
|
||||||
|
}
|
||||||
|
|
||||||
|
fs.writeFileSync(resultFile, `${JSON.stringify(result, null, 2)}\n`, { mode: 0o600 });
|
||||||
|
console.log(JSON.stringify(result));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* delete-comment <comment-id>
|
||||||
|
* Delete a comment (and all its edit history).
|
||||||
|
*/
|
||||||
|
function cmdDeleteComment(commentId) {
|
||||||
|
if (!commentId) {
|
||||||
|
fail("Usage: delete-comment <comment-id>");
|
||||||
|
}
|
||||||
|
gh(["api", `repos/${REPO}/issues/comments/${commentId}`, "-X", "DELETE"], { json: false });
|
||||||
|
console.log(JSON.stringify({ ok: true, deleted_comment_id: Number(commentId) }));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* delete-discussion-comment <node-id>
|
||||||
|
* Delete a discussion comment via GraphQL (and all its edit history).
|
||||||
|
*/
|
||||||
|
function cmdDeleteDiscussionComment(nodeId) {
|
||||||
|
if (!nodeId) {
|
||||||
|
fail("Usage: delete-discussion-comment <node-id>");
|
||||||
|
}
|
||||||
|
const result = ghGraphQL(
|
||||||
|
`mutation { deleteDiscussionComment(input: { id: "${nodeId}" }) { comment { id } } }`,
|
||||||
|
);
|
||||||
|
if (result?.errors) {
|
||||||
|
fail(`Failed to delete discussion comment: ${JSON.stringify(result.errors)}`);
|
||||||
|
}
|
||||||
|
console.log(JSON.stringify({ ok: true, deleted_node_id: nodeId }));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* recreate-discussion-comment <discussion-node-id> <body-file> [reply-to-node-id]
|
||||||
|
* Create a new discussion comment via GraphQL.
|
||||||
|
*/
|
||||||
|
function cmdRecreateDiscussionComment(discussionNodeId, bodyFile, replyToNodeId) {
|
||||||
|
if (!discussionNodeId || !bodyFile) {
|
||||||
|
fail("Usage: recreate-discussion-comment <discussion-node-id> <body-file> [reply-to-node-id]");
|
||||||
|
}
|
||||||
|
if (!fs.existsSync(bodyFile)) {
|
||||||
|
fail(`File not found: ${bodyFile}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const body = fs.readFileSync(bodyFile, "utf8");
|
||||||
|
const newComment = createDiscussionComment(discussionNodeId, body, replyToNodeId);
|
||||||
|
console.log(
|
||||||
|
JSON.stringify({
|
||||||
|
ok: true,
|
||||||
|
node_id: newComment?.id,
|
||||||
|
html_url: newComment?.url,
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* recreate-comment <issue-number> <body-file>
|
||||||
|
* Create a new comment from a file.
|
||||||
|
*/
|
||||||
|
function cmdRecreateComment(issueNumber, bodyFile) {
|
||||||
|
if (!issueNumber || !bodyFile) {
|
||||||
|
fail("Usage: recreate-comment <issue-number> <body-file>");
|
||||||
|
}
|
||||||
|
if (!fs.existsSync(bodyFile)) {
|
||||||
|
fail(`File not found: ${bodyFile}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const result = gh([
|
||||||
|
"api",
|
||||||
|
`repos/${REPO}/issues/${issueNumber}/comments`,
|
||||||
|
"-X",
|
||||||
|
"POST",
|
||||||
|
"-F",
|
||||||
|
`body=@${bodyFile}`,
|
||||||
|
]);
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
JSON.stringify({
|
||||||
|
ok: true,
|
||||||
|
comment_id: result.id,
|
||||||
|
html_url: result.html_url,
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* notify <target> <author> <location-type> <secret-types> [reply-to-node-id]
|
||||||
|
* Post a notification comment with the correct template for the location type.
|
||||||
|
* target = issue/PR number for non-discussion types, discussion node ID for discussion_comment.
|
||||||
|
*/
|
||||||
|
function cmdNotify(target, author, locationType, secretTypes, replyToNodeId) {
|
||||||
|
if (!target || !author || !locationType || !secretTypes) {
|
||||||
|
fail(
|
||||||
|
"Usage: notify <target> <author> <location-type> <secret-types-comma-sep> [reply-to-node-id]",
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const types = secretTypes.split(",").map((s) => s.trim());
|
||||||
|
const typeList = types.map((t, i) => `${i + 1}. **${t}**`).join("\n");
|
||||||
|
const redactionResult = loadBodyRedactionResult(locationType, replyToNodeId);
|
||||||
|
if (isBodyLocationType(locationType) && !redactionResult.notify_required) {
|
||||||
|
console.log(
|
||||||
|
JSON.stringify({
|
||||||
|
ok: true,
|
||||||
|
skipped: true,
|
||||||
|
reason: "current_body_already_redacted",
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
let locationDesc;
|
||||||
|
let actionDesc;
|
||||||
|
if (
|
||||||
|
locationType === "issue_comment" ||
|
||||||
|
locationType === "pull_request_comment" ||
|
||||||
|
locationType === "pull_request_review_comment" ||
|
||||||
|
locationType === "discussion_comment"
|
||||||
|
) {
|
||||||
|
locationDesc = "your comment";
|
||||||
|
actionDesc = "The affected comment has been removed and replaced with a redacted version.";
|
||||||
|
} else if (locationType === "issue_body") {
|
||||||
|
locationDesc = "your issue description";
|
||||||
|
actionDesc = "The affected content has been redacted in place.";
|
||||||
|
} else if (locationType === "pull_request_body") {
|
||||||
|
locationDesc = "your pull request description";
|
||||||
|
actionDesc = "The affected content has been redacted in place.";
|
||||||
|
} else if (locationType === "commit") {
|
||||||
|
locationDesc = "code you committed";
|
||||||
|
actionDesc = "";
|
||||||
|
} else {
|
||||||
|
locationDesc = "your content";
|
||||||
|
actionDesc = "";
|
||||||
|
}
|
||||||
|
|
||||||
|
const body = [
|
||||||
|
`> **Note:** This is an automated message sent by the OpenClaw maintainer team. **NO_REPLY.**`,
|
||||||
|
"",
|
||||||
|
`@${author} :warning: **Security Notice: Secret Leakage Detected**`,
|
||||||
|
"",
|
||||||
|
`GitHub Secret Scanning detected the following exposed secret types in ${locationDesc}:`,
|
||||||
|
"",
|
||||||
|
typeList,
|
||||||
|
"",
|
||||||
|
actionDesc,
|
||||||
|
"",
|
||||||
|
"**Please rotate these credentials immediately.**",
|
||||||
|
"",
|
||||||
|
"These secrets were publicly exposed and should be considered compromised.",
|
||||||
|
]
|
||||||
|
.filter((line) => line !== undefined)
|
||||||
|
.join("\n");
|
||||||
|
|
||||||
|
// Discussion comments must be notified via GraphQL
|
||||||
|
if (locationType === "discussion_comment") {
|
||||||
|
const newComment = createDiscussionComment(target, body, replyToNodeId);
|
||||||
|
console.log(
|
||||||
|
JSON.stringify({
|
||||||
|
ok: true,
|
||||||
|
node_id: newComment?.id,
|
||||||
|
html_url: newComment?.url,
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Issue/PR comments via REST
|
||||||
|
const bodyFile = tmpFile("notify.md");
|
||||||
|
fs.writeFileSync(bodyFile, body);
|
||||||
|
|
||||||
|
const result = gh([
|
||||||
|
"api",
|
||||||
|
`repos/${REPO}/issues/${target}/comments`,
|
||||||
|
"-X",
|
||||||
|
"POST",
|
||||||
|
"-F",
|
||||||
|
`body=@${bodyFile}`,
|
||||||
|
]);
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
JSON.stringify({
|
||||||
|
ok: true,
|
||||||
|
comment_id: result.id,
|
||||||
|
html_url: result.html_url,
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* resolve <alert-number> [resolution] [comment]
|
||||||
|
* Close a secret scanning alert.
|
||||||
|
*/
|
||||||
|
function cmdResolve(alertNumber, resolution, comment) {
|
||||||
|
if (!alertNumber) {
|
||||||
|
fail("Usage: resolve <alert-number> [resolution] [comment]");
|
||||||
|
}
|
||||||
|
|
||||||
|
const res = resolution || "revoked";
|
||||||
|
const resComment = comment || "Content redacted and author notified to rotate credentials.";
|
||||||
|
|
||||||
|
const result = gh([
|
||||||
|
"api",
|
||||||
|
`repos/${REPO}/secret-scanning/alerts/${alertNumber}`,
|
||||||
|
"-X",
|
||||||
|
"PATCH",
|
||||||
|
"-f",
|
||||||
|
`state=resolved`,
|
||||||
|
"-f",
|
||||||
|
`resolution=${res}`,
|
||||||
|
"-f",
|
||||||
|
`resolution_comment=${resComment}`,
|
||||||
|
]);
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
JSON.stringify({
|
||||||
|
ok: true,
|
||||||
|
number: result.number,
|
||||||
|
state: result.state,
|
||||||
|
resolution: result.resolution,
|
||||||
|
resolved_at: result.resolved_at,
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* list-open
|
||||||
|
* List all open secret scanning alerts.
|
||||||
|
*/
|
||||||
|
function cmdListOpen() {
|
||||||
|
const alerts = gh([
|
||||||
|
"api",
|
||||||
|
`repos/${REPO}/secret-scanning/alerts?hide_secret=true&state=open`,
|
||||||
|
"--paginate",
|
||||||
|
"--slurp",
|
||||||
|
]);
|
||||||
|
|
||||||
|
// --slurp 将分页结果合并为 [[page1], [page2], ...] 需要 flat
|
||||||
|
const flat = Array.isArray(alerts?.[0]) ? alerts.flat() : Array.isArray(alerts) ? alerts : [];
|
||||||
|
const rows = flat.map((a) => ({
|
||||||
|
number: a.number,
|
||||||
|
secret_type_display_name: a.secret_type_display_name,
|
||||||
|
html_url: a.html_url,
|
||||||
|
first_location_html_url: a.first_location_detected?.html_url || null,
|
||||||
|
}));
|
||||||
|
|
||||||
|
console.log(JSON.stringify(rows, null, 2));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* summary <json-file>
|
||||||
|
* Print a formatted summary table from a JSON results file.
|
||||||
|
*/
|
||||||
|
function cmdSummary(jsonFile) {
|
||||||
|
if (!jsonFile) {
|
||||||
|
fail("Usage: summary <json-file>");
|
||||||
|
}
|
||||||
|
if (!fs.existsSync(jsonFile)) {
|
||||||
|
fail(`File not found: ${jsonFile}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const results = JSON.parse(fs.readFileSync(jsonFile, "utf8"));
|
||||||
|
const lines = [];
|
||||||
|
|
||||||
|
lines.push("---BEGIN SUMMARY---");
|
||||||
|
lines.push("");
|
||||||
|
lines.push("## Secret Scanning Results");
|
||||||
|
lines.push("");
|
||||||
|
lines.push("| Alert | Type | Location | Actions | Edit History |");
|
||||||
|
lines.push("|-------|------|----------|---------|--------------|");
|
||||||
|
|
||||||
|
const needsPurge = [];
|
||||||
|
|
||||||
|
for (const r of results) {
|
||||||
|
const alertLink = `#${r.number} ${REPO_URL}/security/secret-scanning/${r.number}`;
|
||||||
|
const locationLink = r.location_url
|
||||||
|
? `${r.location_label} ${r.location_url}`
|
||||||
|
: r.location_label;
|
||||||
|
const history = r.history_cleared ? "Cleared" : "⚠️ History remains";
|
||||||
|
|
||||||
|
lines.push(`| ${alertLink} | ${r.secret_type} | ${locationLink} | ${r.actions} | ${history} |`);
|
||||||
|
|
||||||
|
if (!r.history_cleared && r.location_url) {
|
||||||
|
needsPurge.push(r);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (needsPurge.length > 0) {
|
||||||
|
lines.push("");
|
||||||
|
lines.push("Issues requiring GitHub Support to purge edit history:");
|
||||||
|
for (const r of needsPurge) {
|
||||||
|
lines.push(`- ${r.location_label} ${r.location_url} — ${r.secret_type}`);
|
||||||
|
}
|
||||||
|
lines.push(
|
||||||
|
`Contact: https://support.github.com/contact — request purge of userContentEdits for the above issues.`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const skipped = results.filter((r) => r.skipped);
|
||||||
|
if (skipped.length > 0) {
|
||||||
|
lines.push("");
|
||||||
|
lines.push(
|
||||||
|
"⚠️ The following alerts were skipped because their location type is not supported:",
|
||||||
|
);
|
||||||
|
for (const r of skipped) {
|
||||||
|
lines.push(
|
||||||
|
`- Alert #${r.number}: unsupported type "${r.unsupported_type}" — ${REPO_URL}/security/secret-scanning/${r.number}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
lines.push("Please update the skill to define handling for these types.");
|
||||||
|
}
|
||||||
|
|
||||||
|
lines.push("");
|
||||||
|
lines.push("---END SUMMARY---");
|
||||||
|
|
||||||
|
console.log(lines.join("\n"));
|
||||||
|
}
|
||||||
|
|
||||||
|
// ─── Dispatch ───────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
const args = [];
|
||||||
|
|
||||||
|
export const commands = {
|
||||||
|
"fetch-alert": () => cmdFetchAlert(args[0]),
|
||||||
|
"fetch-content": () => cmdFetchContent(args[0]),
|
||||||
|
"redact-body": () => cmdRedactBody(args[0], args[1], args[2]),
|
||||||
|
"redact-body-if-needed": () => cmdRedactBodyIfNeeded(args[0], args[1], args[2], args[3], args[4]),
|
||||||
|
"delete-comment": () => cmdDeleteComment(args[0]),
|
||||||
|
"delete-discussion-comment": () => cmdDeleteDiscussionComment(args[0]),
|
||||||
|
"recreate-comment": () => cmdRecreateComment(args[0], args[1]),
|
||||||
|
"recreate-discussion-comment": () => cmdRecreateDiscussionComment(args[0], args[1], args[2]),
|
||||||
|
notify: () => cmdNotify(args[0], args[1], args[2], args[3], args[4]),
|
||||||
|
resolve: () => cmdResolve(args[0], args[1], args[2]),
|
||||||
|
"list-open": () => cmdListOpen(),
|
||||||
|
summary: () => cmdSummary(args[0]),
|
||||||
|
};
|
||||||
|
|
||||||
|
function main(argv = process.argv.slice(2)) {
|
||||||
|
const [command, ...commandArgs] = argv;
|
||||||
|
args.length = 0;
|
||||||
|
args.push(...commandArgs);
|
||||||
|
|
||||||
|
if (!command || !commands[command]) {
|
||||||
|
console.error(
|
||||||
|
[
|
||||||
|
"Usage: node secret-scanning.mjs <command> [args]",
|
||||||
|
"",
|
||||||
|
"Commands:",
|
||||||
|
" fetch-alert <number> Fetch alert metadata + locations",
|
||||||
|
" fetch-content '<location-json>' Fetch content for a location",
|
||||||
|
" redact-body <issue|pr> <n> <file> PATCH body with redacted file",
|
||||||
|
" redact-body-if-needed <issue|pr> <n> <current-file> <redacted-file> <result-file> PATCH body only if redaction changed it",
|
||||||
|
" delete-comment <comment-id> Delete a comment",
|
||||||
|
" delete-discussion-comment <node-id> Delete a discussion comment (GraphQL)",
|
||||||
|
" recreate-comment <issue-n> <file> Create replacement comment",
|
||||||
|
" recreate-discussion-comment <disc-node-id> <file> [reply-to-node-id] Create discussion comment (GraphQL)",
|
||||||
|
" notify <target> <author> <type> <types> [reply-to-node-id|body-result-file] Post notification",
|
||||||
|
" resolve <n> [resolution] [comment] Close alert",
|
||||||
|
" list-open List open alerts",
|
||||||
|
" summary <json-file> Print formatted summary",
|
||||||
|
].join("\n"),
|
||||||
|
);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
commands[command]();
|
||||||
|
}
|
||||||
|
|
||||||
|
if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
|
||||||
|
main();
|
||||||
|
}
|
||||||
109
.agents/skills/openclaw-test-heap-leaks/SKILL.md
Normal file
109
.agents/skills/openclaw-test-heap-leaks/SKILL.md
Normal file
@@ -0,0 +1,109 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-test-heap-leaks
|
||||||
|
description: Investigate OpenClaw pnpm test memory growth, Vitest OOMs, RSS spikes, and heap snapshot deltas.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Test Heap Leaks
|
||||||
|
|
||||||
|
Use this skill for test-memory investigations. Do not guess from RSS alone when heap snapshots are available. Treat snapshot-name deltas as triage evidence, not proof, until retainers or dominators support the call.
|
||||||
|
|
||||||
|
For **runtime fixes** (e.g., closure leaks in long-running services like the gateway), see [Validating runtime fixes](#validating-runtime-fixes-not-test-memory) below — that uses a dedicated harness, not the test-parallel snapshot machinery.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Reproduce the failing shape first.
|
||||||
|
- Match the real entrypoint if possible. For Linux CI-style unit failures, start with:
|
||||||
|
- `pnpm canvas:a2ui:bundle && OPENCLAW_TEST_MEMORY_TRACE=1 OPENCLAW_TEST_HEAPSNAPSHOT_INTERVAL_MS=60000 OPENCLAW_TEST_HEAPSNAPSHOT_DIR=.tmp/heapsnap OPENCLAW_TEST_WORKERS=2 OPENCLAW_TEST_MAX_OLD_SPACE_SIZE_MB=6144 pnpm test`
|
||||||
|
- Keep `OPENCLAW_TEST_MEMORY_TRACE=1` enabled so the wrapper prints per-file RSS summaries alongside the snapshots.
|
||||||
|
- If the report is about a specific shard or worker budget, preserve that shape.
|
||||||
|
- Before you analyze snapshots, identify the real lane names from `[test-parallel] start ...` lines or `pnpm test --plan`. Do not assume a single `unit-fast` lane; local plans often split into `unit-fast-batch-*`.
|
||||||
|
|
||||||
|
2. Wait for repeated snapshots before concluding anything.
|
||||||
|
- Take at least two intervals from the same lane.
|
||||||
|
- Compare snapshots from the same PID inside the real lane directory such as `.tmp/heapsnap/unit-fast-batch-2/`.
|
||||||
|
- Use `.agents/skills/openclaw-test-heap-leaks/scripts/heapsnapshot-delta.mjs` to compare either two files directly or the earliest/latest pair per PID in one lane directory.
|
||||||
|
- If the helper suggests transformed-module retention, confirm the top entries in DevTools retainers/dominators before calling it solved.
|
||||||
|
|
||||||
|
3. Classify the growth before choosing a fix.
|
||||||
|
- If growth is dominated by Vite/Vitest transformed source strings, `Module`, `system / Context`, bytecode, descriptor arrays, or property maps, treat it as likely retained module graph growth in long-lived workers.
|
||||||
|
- If growth is dominated by app objects, caches, buffers, server handles, timers, mock state, sqlite state, or similar runtime objects, treat it as a likely cleanup or lifecycle leak.
|
||||||
|
- If the names are ambiguous, stop short of a confident label and inspect retainers/dominators in DevTools for the top deltas.
|
||||||
|
|
||||||
|
4. Fix the right layer.
|
||||||
|
- For likely retained transformed-module growth in shared workers:
|
||||||
|
- Prefer timing and hotspot-driven scheduling fixes first. Check whether the file is already represented in `test/fixtures/test-timings.unit.json` and whether `scripts/test-update-memory-hotspots.mjs` should refresh the measured hotspot manifest before hand-editing behavior overrides.
|
||||||
|
- Move hotspot files out of the real shared lane by updating `test/fixtures/test-parallel.behavior.json` only when timing-driven peeling is insufficient.
|
||||||
|
- Prefer `singletonIsolated` for files that are safe alone but inflate shared worker heaps.
|
||||||
|
- If the file should already have been peeled out by timings but is absent from `test/fixtures/test-timings.unit.json`, call that out explicitly. Missing timings are a scheduling blind spot.
|
||||||
|
- For real leaks:
|
||||||
|
- Patch the implicated test or runtime cleanup path.
|
||||||
|
- Look for missing `afterEach`/`afterAll`, module-reset gaps, retained global state, unreleased DB handles, or listeners/timers that survive the file.
|
||||||
|
|
||||||
|
5. Verify with the most direct proof.
|
||||||
|
- Re-run the targeted lane or file with heap snapshots enabled if the suite still finishes in reasonable time.
|
||||||
|
- If snapshot overhead pushes tests over Vitest timeouts, fall back to the same lane without snapshots and confirm the RSS trend or OOM is reduced.
|
||||||
|
- For wrapper-only changes, at minimum verify the expected lanes start and the snapshot files are written.
|
||||||
|
|
||||||
|
## Heuristics
|
||||||
|
|
||||||
|
- Do not call everything a leak. In this repo, large `unit-fast` or `unit-fast-batch-*` growth can be a worker-lifetime problem rather than an application object leak.
|
||||||
|
- `scripts/test-parallel.mjs` and `scripts/test-parallel-memory.mjs` are the primary control points for wrapper diagnostics.
|
||||||
|
- The lane names printed by `[test-parallel] start ...` and `[test-parallel][mem] summary ...` tell you where to focus.
|
||||||
|
- When one or two files account for most of the delta and they are missing from timings, reducing impact by isolating them is usually the first pragmatic fix.
|
||||||
|
- When the same retained object families grow across multiple intervals in the same worker PID, trust the snapshots over intuition, then confirm ambiguous calls with retainer evidence.
|
||||||
|
|
||||||
|
## Snapshot Comparison
|
||||||
|
|
||||||
|
- Direct comparison:
|
||||||
|
- `node .agents/skills/openclaw-test-heap-leaks/scripts/heapsnapshot-delta.mjs before.heapsnapshot after.heapsnapshot`
|
||||||
|
- Auto-select earliest/latest snapshots per PID within one lane:
|
||||||
|
- `node .agents/skills/openclaw-test-heap-leaks/scripts/heapsnapshot-delta.mjs --lane-dir .tmp/heapsnap/unit-fast-batch-2`
|
||||||
|
- Useful flags:
|
||||||
|
- `--top 40`
|
||||||
|
- `--min-kb 32`
|
||||||
|
- `--pid 16133`
|
||||||
|
|
||||||
|
Read the top positive deltas first. Large positive growth in module-transform artifacts suggests lane isolation; large positive growth in runtime objects suggests a real leak. If the names alone do not settle it, open the same snapshot pair in DevTools and inspect retainers/dominators for the top rows before declaring root cause.
|
||||||
|
|
||||||
|
## Validating runtime fixes (not test-memory)
|
||||||
|
|
||||||
|
The workflow above is for diagnosing Vitest worker memory growth. For
|
||||||
|
validating that a runtime/closure fix actually releases captured state, use the
|
||||||
|
dedicated harness:
|
||||||
|
|
||||||
|
- `pnpm leak:embedded-run` — runs `scripts/embedded-run-abort-leak.ts`. Loops N
|
||||||
|
aborted runs in a function-shaped scope mimicking `runEmbeddedAttempt`,
|
||||||
|
writes heap snapshots, and reports a PASS/FAIL verdict on retention growth
|
||||||
|
using `FinalizationRegistry` for tracked-instance counting plus RSS delta.
|
||||||
|
|
||||||
|
Modes:
|
||||||
|
|
||||||
|
- `closure-extracted` (default) — production fix shape (helper at module scope).
|
||||||
|
- `closure-inline` — pre-fix shape (closure inside the runner scope). Use as a
|
||||||
|
sensitivity check: if it passes you've broken the harness, not fixed a bug.
|
||||||
|
- `synthetic-leak` — deliberately retains via a module-level bucket. Use to
|
||||||
|
confirm the harness can detect leaks before trusting a PASS on a real fix.
|
||||||
|
|
||||||
|
Snapshots land in `.tmp/embedded-run-abort-leak/`. Diff with the same script
|
||||||
|
as above:
|
||||||
|
|
||||||
|
```
|
||||||
|
node .agents/skills/openclaw-test-heap-leaks/scripts/heapsnapshot-delta.mjs \
|
||||||
|
.tmp/embedded-run-abort-leak/baseline-*.heapsnapshot \
|
||||||
|
.tmp/embedded-run-abort-leak/batch-N-*.heapsnapshot --top 30
|
||||||
|
```
|
||||||
|
|
||||||
|
When fixing a different runtime leak, add a new harness alongside this one
|
||||||
|
rather than retrofitting it. The fixture function should mimic the lexical
|
||||||
|
scope of the function where the leak lives, not be a generic abort-loop.
|
||||||
|
|
||||||
|
## Output Expectations
|
||||||
|
|
||||||
|
When using this skill, report:
|
||||||
|
|
||||||
|
- The exact reproduce command.
|
||||||
|
- Which lane and PID were compared.
|
||||||
|
- The dominant retained object families from the snapshot delta.
|
||||||
|
- Whether the issue is a likely real leak or likely shared-worker retained module growth, plus whether retainers/dominators confirmed it.
|
||||||
|
- The concrete fix or impact-reduction patch.
|
||||||
|
- What you verified, and what snapshot overhead prevented you from verifying.
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Test Heap Leaks"
|
||||||
|
short_description: "Investigate test OOMs with heap snapshots"
|
||||||
|
default_prompt: "Use $openclaw-test-heap-leaks to investigate test memory growth with heap snapshots and reduce its impact."
|
||||||
@@ -0,0 +1,556 @@
|
|||||||
|
#!/usr/bin/env node
|
||||||
|
/**
|
||||||
|
* Heap snapshot diff utility for OpenClaw test memory leak investigations.
|
||||||
|
*/
|
||||||
|
|
||||||
|
import fs from "node:fs";
|
||||||
|
import path from "node:path";
|
||||||
|
|
||||||
|
function printUsage() {
|
||||||
|
console.error(
|
||||||
|
"Usage: node heapsnapshot-delta.mjs <before.heapsnapshot> <after.heapsnapshot> [--top N] [--min-kb N]",
|
||||||
|
);
|
||||||
|
console.error(
|
||||||
|
" or: node heapsnapshot-delta.mjs --lane-dir <dir> [--pid PID] [--top N] [--min-kb N]",
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
function fail(message) {
|
||||||
|
console.error(message);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseArgs(argv) {
|
||||||
|
const options = {
|
||||||
|
top: 30,
|
||||||
|
minKb: 64,
|
||||||
|
laneDir: null,
|
||||||
|
pid: null,
|
||||||
|
files: [],
|
||||||
|
};
|
||||||
|
|
||||||
|
for (let index = 0; index < argv.length; index += 1) {
|
||||||
|
const arg = argv[index];
|
||||||
|
if (arg === "--top") {
|
||||||
|
options.top = Number.parseInt(argv[index + 1] ?? "", 10);
|
||||||
|
index += 1;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (arg === "--min-kb") {
|
||||||
|
options.minKb = Number.parseInt(argv[index + 1] ?? "", 10);
|
||||||
|
index += 1;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (arg === "--lane-dir") {
|
||||||
|
options.laneDir = argv[index + 1] ?? null;
|
||||||
|
index += 1;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (arg === "--pid") {
|
||||||
|
options.pid = Number.parseInt(argv[index + 1] ?? "", 10);
|
||||||
|
index += 1;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
options.files.push(arg);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!Number.isFinite(options.top) || options.top <= 0) {
|
||||||
|
fail("--top must be a positive integer");
|
||||||
|
}
|
||||||
|
if (!Number.isFinite(options.minKb) || options.minKb < 0) {
|
||||||
|
fail("--min-kb must be a non-negative integer");
|
||||||
|
}
|
||||||
|
if (options.pid !== null && (!Number.isInteger(options.pid) || options.pid <= 0)) {
|
||||||
|
fail("--pid must be a positive integer");
|
||||||
|
}
|
||||||
|
|
||||||
|
return options;
|
||||||
|
}
|
||||||
|
|
||||||
|
class JsonStreamScanner {
|
||||||
|
constructor(filePath) {
|
||||||
|
this.stream = fs.createReadStream(filePath, {
|
||||||
|
encoding: "utf8",
|
||||||
|
highWaterMark: 1024 * 1024,
|
||||||
|
});
|
||||||
|
this.iterator = this.stream[Symbol.asyncIterator]();
|
||||||
|
this.buffer = "";
|
||||||
|
this.offset = 0;
|
||||||
|
this.done = false;
|
||||||
|
}
|
||||||
|
|
||||||
|
compactBuffer() {
|
||||||
|
if (this.offset > 65536) {
|
||||||
|
this.buffer = this.buffer.slice(this.offset);
|
||||||
|
this.offset = 0;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async ensureAvailable(count = 1) {
|
||||||
|
while (!this.done && this.buffer.length - this.offset < count) {
|
||||||
|
const next = await this.iterator.next();
|
||||||
|
if (next.done) {
|
||||||
|
this.done = true;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
this.buffer += next.value;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async peek() {
|
||||||
|
await this.ensureAvailable(1);
|
||||||
|
return this.buffer[this.offset] ?? null;
|
||||||
|
}
|
||||||
|
|
||||||
|
async next() {
|
||||||
|
await this.ensureAvailable(1);
|
||||||
|
if (this.offset >= this.buffer.length) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
const char = this.buffer[this.offset];
|
||||||
|
this.offset += 1;
|
||||||
|
this.compactBuffer();
|
||||||
|
return char;
|
||||||
|
}
|
||||||
|
|
||||||
|
async skipWhitespace() {
|
||||||
|
while (true) {
|
||||||
|
const char = await this.peek();
|
||||||
|
if (char === null || !/\s/u.test(char)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
await this.next();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async expectChar(expected) {
|
||||||
|
const char = await this.next();
|
||||||
|
if (char !== expected) {
|
||||||
|
fail(`Expected ${expected} but found ${char ?? "<eof>"}`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async find(sequence) {
|
||||||
|
let matched = 0;
|
||||||
|
while (true) {
|
||||||
|
const char = await this.next();
|
||||||
|
if (char === null) {
|
||||||
|
fail(`Could not find ${sequence}`);
|
||||||
|
}
|
||||||
|
if (char === sequence[matched]) {
|
||||||
|
matched += 1;
|
||||||
|
if (matched === sequence.length) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
matched = char === sequence[0] ? 1 : 0;
|
||||||
|
if (matched === sequence.length) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async readBalancedObject() {
|
||||||
|
const start = await this.next();
|
||||||
|
if (start !== "{") {
|
||||||
|
fail(`Expected { but found ${start ?? "<eof>"}`);
|
||||||
|
}
|
||||||
|
let text = "{";
|
||||||
|
let depth = 1;
|
||||||
|
let inString = false;
|
||||||
|
let escaped = false;
|
||||||
|
while (depth > 0) {
|
||||||
|
const char = await this.next();
|
||||||
|
if (char === null) {
|
||||||
|
fail("Unexpected EOF while reading JSON object");
|
||||||
|
}
|
||||||
|
text += char;
|
||||||
|
if (inString) {
|
||||||
|
if (escaped) {
|
||||||
|
escaped = false;
|
||||||
|
} else if (char === "\\") {
|
||||||
|
escaped = true;
|
||||||
|
} else if (char === '"') {
|
||||||
|
inString = false;
|
||||||
|
}
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (char === '"') {
|
||||||
|
inString = true;
|
||||||
|
} else if (char === "{") {
|
||||||
|
depth += 1;
|
||||||
|
} else if (char === "}") {
|
||||||
|
depth -= 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return text;
|
||||||
|
}
|
||||||
|
|
||||||
|
async parseNumberArray(onValue) {
|
||||||
|
await this.skipWhitespace();
|
||||||
|
await this.expectChar("[");
|
||||||
|
await this.skipWhitespace();
|
||||||
|
if ((await this.peek()) === "]") {
|
||||||
|
await this.next();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
let token = "";
|
||||||
|
let index = 0;
|
||||||
|
const flush = () => {
|
||||||
|
if (token.length === 0) {
|
||||||
|
fail("Unexpected empty number token");
|
||||||
|
}
|
||||||
|
const value = Number.parseInt(token, 10);
|
||||||
|
if (!Number.isFinite(value)) {
|
||||||
|
fail(`Invalid numeric token: ${token}`);
|
||||||
|
}
|
||||||
|
onValue(value, index);
|
||||||
|
index += 1;
|
||||||
|
token = "";
|
||||||
|
};
|
||||||
|
|
||||||
|
while (true) {
|
||||||
|
const char = await this.next();
|
||||||
|
if (char === null) {
|
||||||
|
fail("Unexpected EOF while reading number array");
|
||||||
|
}
|
||||||
|
if (char === "]") {
|
||||||
|
flush();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (char === ",") {
|
||||||
|
flush();
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (/\s/u.test(char)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
token += char;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async readJsonString() {
|
||||||
|
await this.expectChar('"');
|
||||||
|
let value = "";
|
||||||
|
while (true) {
|
||||||
|
const char = await this.next();
|
||||||
|
if (char === null) {
|
||||||
|
fail("Unexpected EOF while reading JSON string");
|
||||||
|
}
|
||||||
|
if (char === '"') {
|
||||||
|
return value;
|
||||||
|
}
|
||||||
|
if (char !== "\\") {
|
||||||
|
value += char;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
const escaped = await this.next();
|
||||||
|
if (escaped === null) {
|
||||||
|
fail("Unexpected EOF while reading JSON string escape");
|
||||||
|
}
|
||||||
|
if (escaped === "u") {
|
||||||
|
let hex = "";
|
||||||
|
for (let index = 0; index < 4; index += 1) {
|
||||||
|
const hexChar = await this.next();
|
||||||
|
if (hexChar === null) {
|
||||||
|
fail("Unexpected EOF while reading JSON unicode escape");
|
||||||
|
}
|
||||||
|
hex += hexChar;
|
||||||
|
}
|
||||||
|
value += String.fromCharCode(Number.parseInt(hex, 16));
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
value +=
|
||||||
|
escaped === "b"
|
||||||
|
? "\b"
|
||||||
|
: escaped === "f"
|
||||||
|
? "\f"
|
||||||
|
: escaped === "n"
|
||||||
|
? "\n"
|
||||||
|
: escaped === "r"
|
||||||
|
? "\r"
|
||||||
|
: escaped === "t"
|
||||||
|
? "\t"
|
||||||
|
: escaped;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async parseStringArray(onValue) {
|
||||||
|
await this.skipWhitespace();
|
||||||
|
await this.expectChar("[");
|
||||||
|
await this.skipWhitespace();
|
||||||
|
if ((await this.peek()) === "]") {
|
||||||
|
await this.next();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
let index = 0;
|
||||||
|
while (true) {
|
||||||
|
const value = await this.readJsonString();
|
||||||
|
onValue(value, index);
|
||||||
|
index += 1;
|
||||||
|
await this.skipWhitespace();
|
||||||
|
const separator = await this.next();
|
||||||
|
if (separator === "]") {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (separator !== ",") {
|
||||||
|
fail(`Expected , or ] but found ${separator ?? "<eof>"}`);
|
||||||
|
}
|
||||||
|
await this.skipWhitespace();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseHeapFilename(filePath) {
|
||||||
|
const base = path.basename(filePath);
|
||||||
|
const match = base.match(
|
||||||
|
/^Heap\.(?<stamp>\d{8}\.\d{6})\.(?<pid>\d+)\.0\.(?<seq>\d+)\.heapsnapshot$/u,
|
||||||
|
);
|
||||||
|
if (!match?.groups) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
filePath,
|
||||||
|
pid: Number.parseInt(match.groups.pid, 10),
|
||||||
|
stamp: match.groups.stamp,
|
||||||
|
sequence: Number.parseInt(match.groups.seq, 10),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function resolvePair(options) {
|
||||||
|
if (options.laneDir) {
|
||||||
|
const entries = fs
|
||||||
|
.readdirSync(options.laneDir)
|
||||||
|
.map((name) => parseHeapFilename(path.join(options.laneDir, name)))
|
||||||
|
.filter((entry) => entry !== null)
|
||||||
|
.filter((entry) => options.pid === null || entry.pid === options.pid)
|
||||||
|
.toSorted((left, right) => {
|
||||||
|
if (left.pid !== right.pid) {
|
||||||
|
return left.pid - right.pid;
|
||||||
|
}
|
||||||
|
if (left.stamp !== right.stamp) {
|
||||||
|
return left.stamp.localeCompare(right.stamp);
|
||||||
|
}
|
||||||
|
return left.sequence - right.sequence;
|
||||||
|
});
|
||||||
|
|
||||||
|
if (entries.length === 0) {
|
||||||
|
fail(`No matching heap snapshots found in ${options.laneDir}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const groups = new Map();
|
||||||
|
for (const entry of entries) {
|
||||||
|
const group = groups.get(entry.pid) ?? [];
|
||||||
|
group.push(entry);
|
||||||
|
groups.set(entry.pid, group);
|
||||||
|
}
|
||||||
|
|
||||||
|
const candidates = Array.from(groups.values())
|
||||||
|
.map((group) => ({
|
||||||
|
pid: group[0].pid,
|
||||||
|
before: group[0],
|
||||||
|
after: group.at(-1),
|
||||||
|
count: group.length,
|
||||||
|
}))
|
||||||
|
.filter((entry) => entry.count >= 2);
|
||||||
|
|
||||||
|
if (candidates.length === 0) {
|
||||||
|
fail(`Need at least two snapshots for one PID in ${options.laneDir}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const chosen =
|
||||||
|
options.pid !== null
|
||||||
|
? (candidates.find((entry) => entry.pid === options.pid) ?? null)
|
||||||
|
: candidates.toSorted((left, right) => right.count - left.count || left.pid - right.pid)[0];
|
||||||
|
|
||||||
|
if (!chosen) {
|
||||||
|
fail(`No PID with at least two snapshots matched in ${options.laneDir}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
before: chosen.before.filePath,
|
||||||
|
after: chosen.after.filePath,
|
||||||
|
pid: chosen.pid,
|
||||||
|
snapshotCount: chosen.count,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
if (options.files.length !== 2) {
|
||||||
|
printUsage();
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
before: options.files[0],
|
||||||
|
after: options.files[1],
|
||||||
|
pid: null,
|
||||||
|
snapshotCount: 2,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
async function parseSnapshotMeta(scanner) {
|
||||||
|
await scanner.find('"snapshot":');
|
||||||
|
await scanner.skipWhitespace();
|
||||||
|
const metaObjectText = await scanner.readBalancedObject();
|
||||||
|
const parsed = JSON.parse(metaObjectText);
|
||||||
|
return parsed?.meta ?? null;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function buildSummary(filePath) {
|
||||||
|
const scanner = new JsonStreamScanner(filePath);
|
||||||
|
const meta = await parseSnapshotMeta(scanner);
|
||||||
|
if (!meta) {
|
||||||
|
fail(`Invalid heap snapshot: ${filePath}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const nodeFieldCount = meta.node_fields.length;
|
||||||
|
const typeNames = meta.node_types[0];
|
||||||
|
const typeIndex = meta.node_fields.indexOf("type");
|
||||||
|
const nameIndex = meta.node_fields.indexOf("name");
|
||||||
|
const selfSizeIndex = meta.node_fields.indexOf("self_size");
|
||||||
|
if (typeIndex === -1 || nameIndex === -1 || selfSizeIndex === -1) {
|
||||||
|
fail(`Unsupported heap snapshot schema: ${filePath}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const summaryByIndex = new Map();
|
||||||
|
let nodeCount = 0;
|
||||||
|
let currentTypeId = 0;
|
||||||
|
let currentNameId = 0;
|
||||||
|
let currentSelfSize = 0;
|
||||||
|
await scanner.find('"nodes":');
|
||||||
|
await scanner.parseNumberArray((value, index) => {
|
||||||
|
const fieldIndex = index % nodeFieldCount;
|
||||||
|
if (fieldIndex === typeIndex) {
|
||||||
|
currentTypeId = value;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (fieldIndex === nameIndex) {
|
||||||
|
currentNameId = value;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (fieldIndex === selfSizeIndex) {
|
||||||
|
currentSelfSize = value;
|
||||||
|
}
|
||||||
|
if (fieldIndex !== nodeFieldCount - 1) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const key = `${currentTypeId}\t${currentNameId}`;
|
||||||
|
const current = summaryByIndex.get(key) ?? {
|
||||||
|
typeId: currentTypeId,
|
||||||
|
nameId: currentNameId,
|
||||||
|
selfSize: 0,
|
||||||
|
count: 0,
|
||||||
|
};
|
||||||
|
current.selfSize += currentSelfSize;
|
||||||
|
current.count += 1;
|
||||||
|
summaryByIndex.set(key, current);
|
||||||
|
nodeCount += 1;
|
||||||
|
});
|
||||||
|
|
||||||
|
const requiredNameIds = new Set(
|
||||||
|
Array.from(summaryByIndex.values(), (entry) => entry.nameId).filter((value) => value >= 0),
|
||||||
|
);
|
||||||
|
const nameStrings = new Map();
|
||||||
|
await scanner.find('"strings":');
|
||||||
|
await scanner.parseStringArray((value, index) => {
|
||||||
|
if (requiredNameIds.has(index)) {
|
||||||
|
nameStrings.set(index, value);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
const summary = new Map();
|
||||||
|
for (const entry of summaryByIndex.values()) {
|
||||||
|
const key = `${typeNames[entry.typeId] ?? "unknown"}\t${nameStrings.get(entry.nameId) ?? ""}`;
|
||||||
|
summary.set(key, {
|
||||||
|
type: typeNames[entry.typeId] ?? "unknown",
|
||||||
|
name: nameStrings.get(entry.nameId) ?? "",
|
||||||
|
selfSize: entry.selfSize,
|
||||||
|
count: entry.count,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
nodeCount,
|
||||||
|
summary,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function formatBytes(bytes) {
|
||||||
|
if (Math.abs(bytes) >= 1024 ** 2) {
|
||||||
|
return `${(bytes / 1024 ** 2).toFixed(2)} MiB`;
|
||||||
|
}
|
||||||
|
if (Math.abs(bytes) >= 1024) {
|
||||||
|
return `${(bytes / 1024).toFixed(1)} KiB`;
|
||||||
|
}
|
||||||
|
return `${bytes} B`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function formatDelta(bytes) {
|
||||||
|
return `${bytes >= 0 ? "+" : "-"}${formatBytes(Math.abs(bytes))}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function truncate(text, maxLength) {
|
||||||
|
return text.length <= maxLength ? text : `${text.slice(0, maxLength - 1)}…`;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function main() {
|
||||||
|
const options = parseArgs(process.argv.slice(2));
|
||||||
|
const pair = resolvePair(options);
|
||||||
|
const before = await buildSummary(pair.before);
|
||||||
|
const after = await buildSummary(pair.after);
|
||||||
|
const minBytes = options.minKb * 1024;
|
||||||
|
|
||||||
|
const rows = [];
|
||||||
|
for (const [key, next] of after.summary) {
|
||||||
|
const previous = before.summary.get(key) ?? { selfSize: 0, count: 0 };
|
||||||
|
const sizeDelta = next.selfSize - previous.selfSize;
|
||||||
|
const countDelta = next.count - previous.count;
|
||||||
|
if (sizeDelta < minBytes) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
rows.push({
|
||||||
|
type: next.type,
|
||||||
|
name: next.name,
|
||||||
|
sizeDelta,
|
||||||
|
countDelta,
|
||||||
|
afterSize: next.selfSize,
|
||||||
|
afterCount: next.count,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
rows.sort(
|
||||||
|
(left, right) => right.sizeDelta - left.sizeDelta || right.countDelta - left.countDelta,
|
||||||
|
);
|
||||||
|
|
||||||
|
console.log(`before: ${pair.before}`);
|
||||||
|
console.log(`after: ${pair.after}`);
|
||||||
|
if (pair.pid !== null) {
|
||||||
|
console.log(`pid: ${pair.pid} (${pair.snapshotCount} snapshots found)`);
|
||||||
|
}
|
||||||
|
console.log(
|
||||||
|
`nodes: ${before.nodeCount} -> ${after.nodeCount} (${after.nodeCount - before.nodeCount >= 0 ? "+" : ""}${after.nodeCount - before.nodeCount})`,
|
||||||
|
);
|
||||||
|
console.log(`filter: top=${options.top} min=${options.minKb} KiB`);
|
||||||
|
console.log("");
|
||||||
|
|
||||||
|
if (rows.length === 0) {
|
||||||
|
console.log("No entries exceeded the minimum delta.");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
for (const row of rows.slice(0, options.top)) {
|
||||||
|
console.log(
|
||||||
|
[
|
||||||
|
formatDelta(row.sizeDelta).padStart(11),
|
||||||
|
`count ${row.countDelta >= 0 ? "+" : ""}${row.countDelta}`.padStart(10),
|
||||||
|
row.type.padEnd(16),
|
||||||
|
truncate(row.name || "(empty)", 96),
|
||||||
|
].join(" "),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
await main();
|
||||||
266
.agents/skills/openclaw-test-performance/SKILL.md
Normal file
266
.agents/skills/openclaw-test-performance/SKILL.md
Normal file
@@ -0,0 +1,266 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-test-performance
|
||||||
|
description: Benchmark, diagnose, and optimize OpenClaw test and plugin-suite runtime, import hotspots, CPU/RSS, heap growth, and slow coverage paths.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Test Performance
|
||||||
|
|
||||||
|
Use evidence first. The goal is real `pnpm test`, plugin-suite, and
|
||||||
|
plugin-inspector speed/RSS improvement with coverage intact, not runner tuning by
|
||||||
|
guesswork.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Read the relevant local `AGENTS.md` files before editing:
|
||||||
|
- `src/agents/AGENTS.md` for agent/import hotspots.
|
||||||
|
- `src/channels/AGENTS.md` and `src/plugins/AGENTS.md` for plugin/channel
|
||||||
|
laziness.
|
||||||
|
- `src/gateway/AGENTS.md` for server lifecycle tests.
|
||||||
|
- `test/helpers/AGENTS.md` and `test/helpers/channels/AGENTS.md` for shared
|
||||||
|
contract helpers.
|
||||||
|
- `src/infra/outbound/AGENTS.md` for outbound/media/action tests.
|
||||||
|
2. Establish a baseline before changing code:
|
||||||
|
- Prefer `pnpm test:perf:groups --full-suite --allow-failures --output <file>`
|
||||||
|
for full-suite ranking.
|
||||||
|
- For bundled plugin breadth, run the smallest relevant `pnpm
|
||||||
|
test:extensions:batch <plugin[,plugin...]>` or plugin-inspector command
|
||||||
|
before jumping to the full extension sweep.
|
||||||
|
- For a scoped hotspot use:
|
||||||
|
`/usr/bin/time -l pnpm test <file-or-files> --maxWorkers=1 --reporter=verbose`
|
||||||
|
- For import-heavy suspicion add:
|
||||||
|
`OPENCLAW_VITEST_IMPORT_DURATIONS=1 OPENCLAW_VITEST_PRINT_IMPORT_BREAKDOWN=1`.
|
||||||
|
3. Separate wall/runner noise from real file cost:
|
||||||
|
- Compare Vitest duration, test body timing, import breakdown, wall time, and
|
||||||
|
max RSS.
|
||||||
|
- Re-run single files when grouped/full-suite numbers look stale or noisy.
|
||||||
|
- If a full-suite grouped run reports a lane failure but JSON says tests
|
||||||
|
passed, capture that as harness/noise and verify the suspect file directly.
|
||||||
|
4. Pick the next attack by return and risk:
|
||||||
|
- High return: one file/test dominates seconds or RSS and has a clear root.
|
||||||
|
- High leverage: one plugin or SDK barrel causes every plugin-inspector or
|
||||||
|
extension-batch run to load broad runtime.
|
||||||
|
- Lower risk: static descriptors, target parsing, routing, auth bypass,
|
||||||
|
setup hints, registry fixtures, or test server lifecycle.
|
||||||
|
- Higher risk: real memory/runtime behavior, live providers, protocol
|
||||||
|
contracts, or broad production refactors.
|
||||||
|
5. Fix the root cause, not the symptom:
|
||||||
|
- Move static metadata/parsing into narrow helpers or lightweight artifacts
|
||||||
|
reused by full runtime and fast paths.
|
||||||
|
- Prefer dependency injection, loaded-plugin-only lookup, explicit fixtures,
|
||||||
|
and pure helpers over broad mocks.
|
||||||
|
- Reuse suite-level servers/clients when a fresh handshake is irrelevant.
|
||||||
|
- Keep schedulers/background loops off unless the test proves scheduling.
|
||||||
|
- In plugin paths, move static metadata into manifest/lightweight artifacts
|
||||||
|
and keep runtime plugin loads behind explicit execution boundaries.
|
||||||
|
6. Preserve coverage shape:
|
||||||
|
- Do not delete a slow integration proof unless the exact production
|
||||||
|
composition is extracted into a named helper and tested.
|
||||||
|
- Keep one cheap integration smoke when cross-component wiring matters.
|
||||||
|
- State explicitly what incidental coverage was removed, if any.
|
||||||
|
7. Re-benchmark the same command after the change and compute seconds plus
|
||||||
|
percent gain.
|
||||||
|
8. Update the running report when requested or when this thread is tracking one.
|
||||||
|
Include before/after commands, artifacts, coverage notes, verification, and
|
||||||
|
next attack order.
|
||||||
|
9. Commit with `scripts/committer "<message>" <paths...>` and push when the
|
||||||
|
user asked for commits/pushes. Stage only files touched for this attack.
|
||||||
|
|
||||||
|
## Plugin-Suite Workflow
|
||||||
|
|
||||||
|
Use this section when perf work involves bundled plugins, plugin-inspector, SDK
|
||||||
|
barrels, package-boundary tests, or extension suites.
|
||||||
|
|
||||||
|
1. Map the suite shape first:
|
||||||
|
- source tests: `pnpm test extensions/<id>` or `pnpm test:extensions:batch <id>`
|
||||||
|
- package boundaries: `pnpm run test:extensions:package-boundary:canary` and
|
||||||
|
`pnpm run test:extensions:package-boundary:compile`
|
||||||
|
- all bundled source tests: `pnpm test:extensions`
|
||||||
|
- plugin import memory: `pnpm test:extensions:memory -- --json .artifacts/test-perf/extensions-memory.json`
|
||||||
|
- plugin-inspector/report work: keep report primitives in `plugin-inspector`;
|
||||||
|
keep wrappers thin and collect peak RSS when the command supports it.
|
||||||
|
2. Start narrow, then widen:
|
||||||
|
- one plugin changed: run that plugin's tests and plugin-inspector slice.
|
||||||
|
- SDK/public barrel changed: add representative provider, channel, memory,
|
||||||
|
and feature plugins.
|
||||||
|
- loader/runtime mirror changed: add package-boundary checks and build/package
|
||||||
|
proof as needed.
|
||||||
|
- unknown shared plugin behavior: run `test:extensions:batch` groups before
|
||||||
|
`pnpm test:extensions`.
|
||||||
|
3. Treat plugin-inspector failures as product signals:
|
||||||
|
- JSON must parse.
|
||||||
|
- warnings/errors must be classified, not hidden.
|
||||||
|
- runtime capture should be quiet and config-tolerant.
|
||||||
|
- command output should include wall time, exit code, and peak RSS when
|
||||||
|
available.
|
||||||
|
4. For broad or package-heavy plugin proof, use Crabbox-backed Blacksmith
|
||||||
|
Testbox by default on maintainer machines:
|
||||||
|
- `pnpm crabbox:run -- --provider blacksmith-testbox --timing-json -- OPENCLAW_TESTBOX=1 pnpm test:extensions:batch <ids>`
|
||||||
|
- add `--keep`/`--id <id-or-slug>` only when several commands must share one
|
||||||
|
warmed box; stop it with `pnpm crabbox:stop -- <id-or-slug>`.
|
||||||
|
5. If plugin performance is package-artifact sensitive, switch to
|
||||||
|
`release-openclaw-plugin-testing` and Package Acceptance rather than
|
||||||
|
trusting source-only timing.
|
||||||
|
|
||||||
|
## Metric Collection
|
||||||
|
|
||||||
|
Collect at least one stable metric before and after. Prefer the same machine and
|
||||||
|
same command. For Testbox comparisons, use the same `tbx_...` id when possible.
|
||||||
|
|
||||||
|
| Metric | Use for | Preferred source |
|
||||||
|
| --------------- | ---------------------------------- | --------------------------------------------------------------------------- |
|
||||||
|
| wall time | user-visible suite cost | `/usr/bin/time -l`, test wrapper duration, Testbox run time |
|
||||||
|
| Vitest duration | test body/import cost | Vitest output per file/shard |
|
||||||
|
| import duration | broad barrel/runtime loads | `OPENCLAW_VITEST_IMPORT_DURATIONS=1` |
|
||||||
|
| max RSS | memory pressure and OOM risk | `/usr/bin/time -l`, `pnpm test:extensions:memory`, wrapper memory summaries |
|
||||||
|
| CPU/user/sys | CPU-bound vs wait-bound split | `/usr/bin/time -l` locally, Testbox job timing when local CPU is noisy |
|
||||||
|
| heap snapshots | real leak vs retained module graph | `openclaw-test-heap-leaks` workflow |
|
||||||
|
|
||||||
|
Local scoped command with CPU/RSS:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
timeout 240 /usr/bin/time -l pnpm test <file> --maxWorkers=1 --reporter=verbose
|
||||||
|
```
|
||||||
|
|
||||||
|
Plugin import memory profile:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm build
|
||||||
|
pnpm test:extensions:memory -- --top 20 --json .artifacts/test-perf/extensions-memory.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Targeted plugin import memory:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm test:extensions:memory -- --extension discord --extension telegram --skip-combined
|
||||||
|
```
|
||||||
|
|
||||||
|
Heap/RSS escalation:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OPENCLAW_TEST_MEMORY_TRACE=1 \
|
||||||
|
OPENCLAW_TEST_HEAPSNAPSHOT_INTERVAL_MS=60000 \
|
||||||
|
OPENCLAW_TEST_HEAPSNAPSHOT_DIR=.tmp/heapsnap \
|
||||||
|
OPENCLAW_TEST_WORKERS=2 \
|
||||||
|
OPENCLAW_TEST_MAX_OLD_SPACE_SIZE_MB=6144 \
|
||||||
|
pnpm test
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `openclaw-test-heap-leaks` when RSS keeps growing across intervals, workers
|
||||||
|
OOM, or the suspect command has app-object retention. Do not call RSS growth a
|
||||||
|
leak until snapshots or retainers support it.
|
||||||
|
|
||||||
|
## Common Root Causes
|
||||||
|
|
||||||
|
- Full bundled channel/plugin runtime loaded for static data.
|
||||||
|
- `getChannelPlugin()` fallback used when an already-loaded fixture or pure
|
||||||
|
parser would suffice.
|
||||||
|
- Broad `api.ts`, `runtime-api.ts`, `test-api.ts`, or plugin-sdk barrels pulled
|
||||||
|
into hot tests.
|
||||||
|
- SDK root aliases or package barrels pulling focused subpaths back into a broad
|
||||||
|
plugin graph.
|
||||||
|
- Plugin-inspector loading runtime code just to render metadata, reports, or CI
|
||||||
|
policy scores.
|
||||||
|
- Bundled plugin capture reusing real config/home state instead of synthetic,
|
||||||
|
redacted, isolated state.
|
||||||
|
- Partial-real mocks using `importActual()` around broad modules.
|
||||||
|
- `vi.resetModules()` plus fresh imports in per-test loops.
|
||||||
|
- Test plugin registry seeded in `beforeAll` while runtime state resets in
|
||||||
|
`afterEach`.
|
||||||
|
- Per-test gateway/server/client startup when state reset would suffice.
|
||||||
|
- Runtime/default model/auth selection paid by idle snapshots or fixtures.
|
||||||
|
- Plugin-owned media/action discovery triggered before checking whether args
|
||||||
|
contain plugin-owned fields.
|
||||||
|
- Timings missing from `test/fixtures/test-timings.unit.json`, causing hotspot
|
||||||
|
files to stay in shared workers.
|
||||||
|
- Parallel Vitest runs sharing `node_modules/.experimental-vitest-cache` without
|
||||||
|
distinct `OPENCLAW_VITEST_FS_MODULE_CACHE_PATH` values.
|
||||||
|
|
||||||
|
## Benchmark Commands
|
||||||
|
|
||||||
|
Scoped file:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
timeout 240 /usr/bin/time -l pnpm test <file> --maxWorkers=1 --reporter=verbose
|
||||||
|
```
|
||||||
|
|
||||||
|
Scoped file with import breakdown:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
timeout 240 /usr/bin/time -l env \
|
||||||
|
OPENCLAW_VITEST_IMPORT_DURATIONS=1 \
|
||||||
|
OPENCLAW_VITEST_PRINT_IMPORT_BREAKDOWN=1 \
|
||||||
|
pnpm test <file> --maxWorkers=1 --reporter=verbose
|
||||||
|
```
|
||||||
|
|
||||||
|
Grouped suite:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm test:perf:groups --full-suite --allow-failures \
|
||||||
|
--output .artifacts/test-perf/<name>.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Extension batch:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm test:extensions:batch <plugin[,plugin...]> -- --reporter=verbose
|
||||||
|
```
|
||||||
|
|
||||||
|
All extension tests:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm test:extensions
|
||||||
|
```
|
||||||
|
|
||||||
|
Package-boundary plugin checks:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm run test:extensions:package-boundary:canary
|
||||||
|
pnpm run test:extensions:package-boundary:compile
|
||||||
|
```
|
||||||
|
|
||||||
|
Reuse an existing Vitest JSON report:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm test:perf:groups --report <vitest-json> \
|
||||||
|
--output .artifacts/test-perf/<name>.json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
- Always run the targeted test surface that proves the change.
|
||||||
|
- For source changes, run `pnpm check:changed` before push; in maintainer
|
||||||
|
Testbox mode run it in the warmed Testbox.
|
||||||
|
- For test-only changes, run `pnpm test:changed` or the exact edited tests.
|
||||||
|
- Run `pnpm build` when touching lazy-loading, bundled artifacts, package
|
||||||
|
boundaries, dynamic imports, build output, or public surfaces.
|
||||||
|
- For plugin SDK/barrel/runtime changes, add `pnpm plugin-sdk:api:check` or
|
||||||
|
`pnpm plugin-sdk:api:gen` when the API surface may drift.
|
||||||
|
- For plugin-suite perf fixes, verify at least one representative plugin batch
|
||||||
|
plus the changed gate; use Package Acceptance if the bug only exists in a
|
||||||
|
packed artifact.
|
||||||
|
- If deps are missing/stale, run `pnpm install` and retry the exact failed
|
||||||
|
command once.
|
||||||
|
- Use the report format:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
| Metric | Before | After | Gain |
|
||||||
|
| -------------- | -----: | -----: | ------------: |
|
||||||
|
| File wall time | `Xs` | `Ys` | `-Zs` (`P%`) |
|
||||||
|
| Max RSS | `XMB` | `YMB` | `-ZMB` (`P%`) |
|
||||||
|
| CPU user/sys | `X/Ys` | `A/Bs` | explain |
|
||||||
|
```
|
||||||
|
|
||||||
|
## Handoff
|
||||||
|
|
||||||
|
Keep the final concise:
|
||||||
|
|
||||||
|
- Root cause.
|
||||||
|
- Suite/plugin scope.
|
||||||
|
- Files changed.
|
||||||
|
- Before/after wall, Vitest/import, CPU, and RSS numbers where available.
|
||||||
|
- Leak classification if memory was involved: real leak, retained module graph,
|
||||||
|
or inconclusive.
|
||||||
|
- Coverage retained.
|
||||||
|
- Verification commands.
|
||||||
|
- Testbox ID or workflow URL for remote proof.
|
||||||
|
- Commit hash and push status.
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "OpenClaw Test Performance"
|
||||||
|
short_description: "Benchmark tests, plugin suites, CPU, RSS, and heap growth"
|
||||||
|
default_prompt: "Use $openclaw-test-performance to reassess OpenClaw test and plugin-suite performance, collect wall/import/CPU/RSS metrics, investigate memory growth when needed, fix the next real hotspot without losing coverage, update the report, and commit scoped changes."
|
||||||
|
policy:
|
||||||
|
allow_implicit_invocation: false
|
||||||
819
.agents/skills/openclaw-testing/SKILL.md
Normal file
819
.agents/skills/openclaw-testing/SKILL.md
Normal file
@@ -0,0 +1,819 @@
|
|||||||
|
---
|
||||||
|
name: openclaw-testing
|
||||||
|
description: Choose, run, rerun, or debug OpenClaw tests, CI checks, Docker E2E lanes, release validation, and the cheapest safe verification path.
|
||||||
|
---
|
||||||
|
|
||||||
|
# OpenClaw Testing
|
||||||
|
|
||||||
|
Use this skill when deciding what to test, debugging failures, rerunning CI,
|
||||||
|
or validating a change without wasting hours.
|
||||||
|
|
||||||
|
## Read First
|
||||||
|
|
||||||
|
- `docs/reference/test.md` for local test commands.
|
||||||
|
- `docs/ci.md` for CI scope, release checks, Docker chunks, and runner behavior.
|
||||||
|
- Scoped `AGENTS.md` files before editing code under a subtree.
|
||||||
|
|
||||||
|
## Default Rule
|
||||||
|
|
||||||
|
Prove the touched surface first. Do not reflexively run the whole suite.
|
||||||
|
|
||||||
|
Agent sessions are remote-first for tests and computationally intensive work.
|
||||||
|
Classify source trust before selecting a backend. Trusted maintainer code
|
||||||
|
defaults to Blacksmith Testbox. Untrusted contributor or fork code must use
|
||||||
|
secretless fork CI or sanitized direct AWS Crabbox; never sync or run it on the
|
||||||
|
credential-hydrated Blacksmith workflow.
|
||||||
|
|
||||||
|
When trusted work is likely to change code or need tests, builds, typechecks,
|
||||||
|
lint fan-out, Docker, packaging, E2E, or live proof, immediately start this in
|
||||||
|
a background command session:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node scripts/crabbox-wrapper.mjs warmup \
|
||||||
|
--provider blacksmith-testbox \
|
||||||
|
--keep \
|
||||||
|
--timing-json
|
||||||
|
```
|
||||||
|
|
||||||
|
For untrusted code, switch to a clean trusted `main` checkout and pre-warm
|
||||||
|
direct AWS with an installed trusted Crabbox binary. Do not execute the
|
||||||
|
untrusted checkout's wrapper or config locally:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd <trusted-openclaw-main>
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
crabbox config show --json | \
|
||||||
|
jq -e '.aws.instanceProfile == ""' >/dev/null
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
-u CRABBOX_TAILSCALE \
|
||||||
|
-u CRABBOX_TAILSCALE_AUTH_KEY \
|
||||||
|
-u CRABBOX_TAILSCALE_AUTH_KEY_ENV \
|
||||||
|
-u CRABBOX_TAILSCALE_EXIT_NODE \
|
||||||
|
-u CRABBOX_TAILSCALE_EXIT_NODE_ALLOW_LAN_ACCESS \
|
||||||
|
-u CRABBOX_TAILSCALE_HOSTNAME_TEMPLATE \
|
||||||
|
-u CRABBOX_TAILSCALE_TAGS \
|
||||||
|
crabbox warmup \
|
||||||
|
--provider aws \
|
||||||
|
--network public \
|
||||||
|
--tailscale=false \
|
||||||
|
--tailscale-exit-node= \
|
||||||
|
--tailscale-exit-node-allow-lan-access=false \
|
||||||
|
--keep \
|
||||||
|
--timing-json
|
||||||
|
crabbox inspect --provider aws --id <cbx_id> --json | \
|
||||||
|
jq -e '.network == "public" and .tailscale == null' >/dev/null
|
||||||
|
```
|
||||||
|
|
||||||
|
Bind the returned lease to one immutable reviewed head SHA; never repurpose a
|
||||||
|
trusted or previously hydrated lease, and stop/rewarm if the head changes.
|
||||||
|
Record the reviewed PR's full head SHA with
|
||||||
|
`gh pr view <number> --repo <owner/repo> --json headRefOid --jq .headRefOid`.
|
||||||
|
Every untrusted AWS run must override the repo env allowlist, skip Actions
|
||||||
|
hydration, and upload the trusted bootstrap script from clean `main` alongside
|
||||||
|
`--fresh-pr`. The script bypasses raw-box JavaScript preflight, proves the
|
||||||
|
identity boundary, installs pinned Node/pnpm, verifies the exact SHA and
|
||||||
|
package-manager pin, isolates `HOME`, installs dependencies, then runs the
|
||||||
|
requested test command:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
CRABBOX_ENV_ALLOW=CI \
|
||||||
|
crabbox run \
|
||||||
|
--provider aws \
|
||||||
|
--id <cbx_id> \
|
||||||
|
--fresh-pr <owner/repo#number> \
|
||||||
|
--no-hydrate \
|
||||||
|
--timing-json \
|
||||||
|
--script scripts/crabbox-untrusted-bootstrap.sh -- \
|
||||||
|
<expected_head_sha> /usr/local/bin/pnpm test <path-or-filter>
|
||||||
|
# After all proof:
|
||||||
|
env -u CRABBOX_AWS_INSTANCE_PROFILE \
|
||||||
|
crabbox stop --provider aws <cbx_id>
|
||||||
|
```
|
||||||
|
|
||||||
|
Continue inspection and editing while the remote box hydrates. Save the
|
||||||
|
returned id, reuse it for the task's focused tests and heavy gates, sync the
|
||||||
|
current checkout on every run, and stop it before handoff. Do not pre-warm for
|
||||||
|
read-only, docs-only, or clearly trivial work that will not run tests or heavy
|
||||||
|
commands.
|
||||||
|
|
||||||
|
1. Inspect the diff and classify the touched surface:
|
||||||
|
- any agent-run test, focused or broad: run it on the pre-warmed safe remote
|
||||||
|
backend; Blacksmith Testbox only for trusted maintainer code
|
||||||
|
- changed gates, builds, typechecks, lint fan-out, Docker, package, E2E, or
|
||||||
|
live work: run it remotely; these are never routine laptop work
|
||||||
|
- normal source checkout, `pnpm check:changed`: it delegates to
|
||||||
|
Crabbox/Testbox, but prefer the explicit kept-lease path when a Testbox was
|
||||||
|
pre-warmed so the task reuses one lease
|
||||||
|
- explicit local fallback requested by the user, one/few files:
|
||||||
|
`node scripts/run-vitest.mjs <path-or-filter>`
|
||||||
|
- direct AWS Crabbox proof: pass `--provider aws`; untrusted code also
|
||||||
|
requires the sanitized invocation above
|
||||||
|
- workflow-only: `git diff --check`, workflow syntax/lint (`actionlint` when available)
|
||||||
|
- docs-only: `pnpm docs:list`, docs formatter/lint only if docs tooling changed or requested
|
||||||
|
2. Reproduce narrowly before fixing.
|
||||||
|
3. Fix root cause.
|
||||||
|
4. Rerun the same narrow proof.
|
||||||
|
5. Broaden only when the touched contract demands it.
|
||||||
|
|
||||||
|
## Guardrails
|
||||||
|
|
||||||
|
- Do not kill unrelated processes or tests. If something is running elsewhere, treat it as owned by the user or another agent.
|
||||||
|
- Do not run tests or computationally intensive commands locally unless the user explicitly asks for local proof. Remote-provider unavailability permits only a narrow reported fallback, not a silent local full gate.
|
||||||
|
- Prefer GitHub Actions for release/Docker proof when the workflow already has the prepared image and secrets.
|
||||||
|
- Use `scripts/committer "<msg>" <paths...>` when committing; stage only your files.
|
||||||
|
- If dependencies are missing on the selected remote box, run `pnpm install` there, retry
|
||||||
|
once, then report the first actionable error. Do not reconcile or reinstall a
|
||||||
|
local Codex worktree merely to run validation.
|
||||||
|
- In a Codex worktree or linked/sparse checkout, do not run direct local
|
||||||
|
`pnpm test*`, `pnpm check*`, `pnpm crabbox:run`, or `scripts/committer`. Use
|
||||||
|
`node scripts/crabbox-wrapper.mjs` for remote proof, and `git commit --no-verify`
|
||||||
|
only after the relevant remote proof is already clean. The direct
|
||||||
|
`node scripts/run-vitest.mjs` path is an explicit local fallback only.
|
||||||
|
- For remote proof, use the Crabbox wrapper first, but name the actual backend.
|
||||||
|
Direct AWS Crabbox uses `provider=aws` and `cbx_...` ids. Delegated
|
||||||
|
Blacksmith Testbox through Crabbox uses `provider=blacksmith-testbox`,
|
||||||
|
`syncDelegated=true`, and `tbx_...` ids. Both satisfy "remote proof" when the
|
||||||
|
requested proof surface allows either.
|
||||||
|
- Treat contributor and fork patches as untrusted unless a maintainer
|
||||||
|
explicitly approves credentialed execution after review. For untrusted AWS
|
||||||
|
runs, `CRABBOX_ENV_ALLOW=CI` must replace the repo's
|
||||||
|
`OPENCLAW_*` allowlist, `--no-hydrate` must block auth-profile hydration, and
|
||||||
|
the remote command must use a fresh temporary `HOME`. The lease must be newly
|
||||||
|
warmed for and bound to one reviewed head SHA, never trusted or previously
|
||||||
|
hydrated; stop and rewarm when the SHA changes. Do
|
||||||
|
not execute repo scripts or config from the untrusted local checkout: launch
|
||||||
|
an installed trusted Crabbox binary from a clean trusted `main` checkout and
|
||||||
|
fetch the PR with `--fresh-pr`. Unset `CRABBOX_AWS_INSTANCE_PROFILE` and fail
|
||||||
|
closed unless `crabbox config show --json` resolves an empty
|
||||||
|
`aws.instanceProfile`. Before any install/test, use trusted absolute-path
|
||||||
|
tools to require an IMDSv2 token, prove the IAM credentials endpoint returns
|
||||||
|
404, and compare remote `git rev-parse HEAD` with the full reviewed head SHA.
|
||||||
|
Unset all `CRABBOX_TAILSCALE*` overrides, pass `--network public
|
||||||
|
--tailscale=false`, clear exit-node/LAN flags, then require `crabbox inspect`
|
||||||
|
to report `network=public` and no Tailscale state before uploading any script.
|
||||||
|
Upload trusted `scripts/crabbox-untrusted-bootstrap.sh` with `--fresh-pr`; it
|
||||||
|
bootstraps Node 24 and repository-pinned pnpm before executing PR code and
|
||||||
|
rejects a changed `packageManager` pin before install.
|
||||||
|
If the broker cannot provide that no-role proof or no remote PR exists, use
|
||||||
|
secretless fork CI. Do not select `hydrate-github` or a credential-hydrated
|
||||||
|
Testbox workflow.
|
||||||
|
- Do not infer "no Testbox is running" from plain `blacksmith testbox list`.
|
||||||
|
Use `blacksmith testbox list --all` or `blacksmith testbox status <tbx_id>`
|
||||||
|
before reporting cloud state.
|
||||||
|
- Reuse only an id/slug created in this operator session unless explicitly
|
||||||
|
coordinating with another lane. If Testbox queues, fails capacity, or cannot
|
||||||
|
allocate, report the blocker or switch to direct AWS Crabbox only when that
|
||||||
|
still proves the requested surface.
|
||||||
|
- Reuse does not mean stale source: omit `--no-sync` so every run uploads the
|
||||||
|
current checkout. Use `--no-sync` only to rerun an unchanged, already-synced
|
||||||
|
tree intentionally.
|
||||||
|
|
||||||
|
## Explicit Local Test Fallbacks
|
||||||
|
|
||||||
|
These commands are for human workflows or an agent's explicit local fallback.
|
||||||
|
They are not the default agent path.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm changed:lanes --json
|
||||||
|
pnpm check:changed # Crabbox/Testbox changed typecheck/lint/guards; no Vitest
|
||||||
|
pnpm test:changed # cheap smart changed Vitest targets
|
||||||
|
pnpm verify # full check, then full Vitest
|
||||||
|
OPENCLAW_TEST_CHANGED_BROAD=1 pnpm test:changed
|
||||||
|
pnpm test <path-or-filter> -- --reporter=verbose
|
||||||
|
OPENCLAW_VITEST_MAX_WORKERS=1 pnpm test <path-or-filter>
|
||||||
|
```
|
||||||
|
|
||||||
|
Use targeted file paths whenever possible. Avoid raw `vitest`; use the repo
|
||||||
|
`pnpm test` wrapper so project routing, workers, and setup stay correct. If raw
|
||||||
|
Vitest is unavoidable, use `vitest run ...`; bare `vitest ...` starts local watch
|
||||||
|
mode and will not exit on its own.
|
||||||
|
When the checkout is a Codex worktree, prefer the direct node harness instead:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node scripts/run-vitest.mjs <path-or-filter>
|
||||||
|
```
|
||||||
|
|
||||||
|
That keeps the test scoped without giving pnpm a chance to run dependency
|
||||||
|
status checks or install reconciliation in a linked worktree.
|
||||||
|
|
||||||
|
## Plugin Package And Live Proof
|
||||||
|
|
||||||
|
When validating an external or official plugin package, prove the package shape
|
||||||
|
and trust shape separately. Do not use raw archive/path installs to prove the
|
||||||
|
managed dependency path, and do not treat `npm-pack:` as proof of catalog-linked
|
||||||
|
official trust.
|
||||||
|
|
||||||
|
- For local release-candidate proof, pack the plugin and install it with
|
||||||
|
`openclaw plugins install npm-pack:<path.tgz> --force`. This uses the managed
|
||||||
|
per-plugin npm project and is the closest local substitute for the registry
|
||||||
|
artifact's dependency behavior.
|
||||||
|
- If the behavior depends on bundled-plugin or trusted official plugin status,
|
||||||
|
add a second proof through a catalog-backed official install or a published
|
||||||
|
package path that records official trust. Local `npm-pack:` proof alone is
|
||||||
|
not sufficient for privileged helpers or trusted-official scope handling.
|
||||||
|
- Treat missing runtime imports as package-manifest bugs first. Runtime code
|
||||||
|
must depend on packages declared in the plugin package `dependencies` or
|
||||||
|
`optionalDependencies`; do not make a final proof depend on manually running
|
||||||
|
`npm install` inside `~/.openclaw/npm/projects/...`.
|
||||||
|
- If the plugin ships `npm-shrinkwrap.json`, regenerate or check it after
|
||||||
|
moving dependencies between dev and runtime sections.
|
||||||
|
- Inspect the packed tarball when dependency ownership or generated `dist/`
|
||||||
|
matters: verify `package/package.json`, the expected runtime files, and any
|
||||||
|
package-local shrinkwrap before installing it on a live host.
|
||||||
|
- After installing the package, restart the Gateway when the touched surface is
|
||||||
|
plugin registration, runtime dependency loading, privileged helpers, provider
|
||||||
|
routing, or generated dist.
|
||||||
|
- For live provider or channel probes, add only temporary config needed for the
|
||||||
|
proof, then remove it and verify the cleanup state before closeout.
|
||||||
|
|
||||||
|
## Command Semantics
|
||||||
|
|
||||||
|
- `pnpm check` and `pnpm check:changed` do not run Vitest tests. They are for
|
||||||
|
typecheck, lint, and guard proof.
|
||||||
|
- `pnpm test` and `pnpm test:changed` run Vitest tests.
|
||||||
|
- `pnpm verify` runs `pnpm check`, then `pnpm test`, with Crabbox phase markers
|
||||||
|
so remote summaries show which half failed.
|
||||||
|
- `pnpm test:changed` is intentionally cheap by default: direct test edits,
|
||||||
|
sibling tests, explicit source mappings, and import-graph dependents.
|
||||||
|
- `OPENCLAW_TEST_CHANGED_BROAD=1 pnpm test:changed` is the explicit broad
|
||||||
|
fallback for harness/config/package edits that genuinely need it.
|
||||||
|
- Do not run extension sweeps just because core changed. If a core edit is for a
|
||||||
|
specific plugin bug, run that plugin's tests explicitly. If a public SDK or
|
||||||
|
contract change needs consumer proof, choose the smallest representative
|
||||||
|
plugin/contract tests first, then broaden only when the risk justifies it.
|
||||||
|
- The test wrapper prints a short `[test] passed|failed|skipped ... in ...`
|
||||||
|
line. Vitest's own duration is still the per-shard detail.
|
||||||
|
|
||||||
|
## Routing Model
|
||||||
|
|
||||||
|
- `pnpm changed:lanes --json` answers "which check lanes does this diff touch?"
|
||||||
|
It is used by `pnpm check:changed` for typecheck/lint/guard selection.
|
||||||
|
- `pnpm test:changed` answers "which Vitest targets are worth running now?" It
|
||||||
|
uses the same changed path list, but applies a cheaper test-target resolver.
|
||||||
|
- Direct test edits run themselves. Source edits prefer explicit mappings,
|
||||||
|
sibling `*.test.ts`, then import-graph dependents. Shared harness/config/root
|
||||||
|
edits are skipped by default unless they have precise mapped tests.
|
||||||
|
- Shared group-room delivery config and source-reply prompt edits are precise
|
||||||
|
mapped tests: they run the core auto-reply regressions plus Discord and Slack
|
||||||
|
delivery tests so cross-channel default changes fail before a PR push.
|
||||||
|
- Public SDK or contract edits do not automatically run every plugin test.
|
||||||
|
`check:changed` proves extension type contracts; the agent chooses the
|
||||||
|
smallest plugin/contract Vitest proof that matches the actual risk.
|
||||||
|
- Use `OPENCLAW_TEST_CHANGED_BROAD=1 pnpm test:changed` only when a harness,
|
||||||
|
config, package, or unknown-root edit really needs the broad Vitest fallback.
|
||||||
|
|
||||||
|
## CI Debugging
|
||||||
|
|
||||||
|
Start with current run state, not logs for everything:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh run list --branch main --limit 10
|
||||||
|
gh run view <run-id> --json status,conclusion,headSha,url,jobs
|
||||||
|
gh run view <run-id> --job <job-id> --log
|
||||||
|
```
|
||||||
|
|
||||||
|
- Check exact SHA. Ignore newer unrelated `main` unless asked.
|
||||||
|
- For cancelled same-branch runs, confirm whether a newer run superseded it.
|
||||||
|
- Fetch full logs only for failed or relevant jobs.
|
||||||
|
- Prefer `gh run view <run-id> --json jobs` over PR rollup while debugging; rollup can be stale/noisy.
|
||||||
|
- For `prompt:snapshots:check` failures, treat Linux Node 24 as CI truth. If macOS passes but CI drifts, reproduce in a Linux Node 24 container or Testbox, commit that generated output, then rerun.
|
||||||
|
|
||||||
|
## GitHub Release Workflows
|
||||||
|
|
||||||
|
Use the smallest workflow that proves the current risk. The full umbrella is
|
||||||
|
available, but it is usually the last step after narrower proof, not the first
|
||||||
|
rerun after a focused patch.
|
||||||
|
|
||||||
|
### Full Release Validation
|
||||||
|
|
||||||
|
`Full Release Validation` (`.github/workflows/full-release-validation.yml`) is
|
||||||
|
the manual "everything before release" umbrella. It resolves a target ref, then
|
||||||
|
dispatches:
|
||||||
|
|
||||||
|
- manual `CI` for the full normal CI graph, with Android enabled via
|
||||||
|
`include_android=true`
|
||||||
|
- `Plugin Prerelease` for release-only plugin static checks, extension shards,
|
||||||
|
the release-only `agentic-plugins` shard, and plugin product Docker lanes
|
||||||
|
- `OpenClaw Release Checks` for install smoke, cross-OS release checks, live and
|
||||||
|
E2E checks, Docker release-path suites, OpenWebUI, QA Lab, fast Matrix, and
|
||||||
|
Telegram release lanes
|
||||||
|
- optional post-publish Telegram E2E when a package spec is supplied
|
||||||
|
|
||||||
|
Run it only when validating an actual release candidate, after broad shared CI
|
||||||
|
or release orchestration changes, or when explicitly asked:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh workflow run full-release-validation.yml \
|
||||||
|
--repo openclaw/openclaw \
|
||||||
|
--ref main \
|
||||||
|
-f ref=<branch-or-sha> \
|
||||||
|
-f provider=openai \
|
||||||
|
-f mode=both \
|
||||||
|
-f release_profile=stable
|
||||||
|
```
|
||||||
|
|
||||||
|
Run the workflow itself from the trusted current ref, normally `--ref main`;
|
||||||
|
child workflows are dispatched from that same ref even when `ref` points at an
|
||||||
|
older release branch or tag. Full Release Validation has no separate child
|
||||||
|
workflow ref input; choose the trusted harness by choosing the workflow run ref.
|
||||||
|
Use `release_profile=minimum|stable|full` to control live/provider breadth:
|
||||||
|
`minimum` keeps the fastest OpenAI/core release-critical set, `stable` adds the
|
||||||
|
stable provider/backend set, and `full` adds the broad advisory provider/media
|
||||||
|
matrix. Do not make `full` faster by silently dropping suites; optimize setup,
|
||||||
|
artifact reuse, and sharding instead. The parent verifier job appends a child
|
||||||
|
overview plus slowest-job tables for child runs; rerun only that verifier after
|
||||||
|
a child rerun turns green.
|
||||||
|
|
||||||
|
Standalone manual `CI` dispatches do not run the plugin prerelease suite, the
|
||||||
|
extension batch sweep, or the release-only `agentic-plugins` Vitest shard. Those
|
||||||
|
lanes are intentionally reserved for the separate `Plugin Prerelease` child so
|
||||||
|
PRs, main pushes, and ad hoc broad CI checks do not spend Docker/package time or
|
||||||
|
all-plugin runtime time on release-only product coverage.
|
||||||
|
|
||||||
|
If a full run is already active on a newer `origin/main`, prefer watching that
|
||||||
|
run over dispatching a duplicate. Do not cancel release, release-check, or child
|
||||||
|
workflow runs unless Peter explicitly asks for cancellation.
|
||||||
|
|
||||||
|
The child-dispatch jobs record the child run ids. The final
|
||||||
|
`Verify full validation` job re-queries those child runs and is the canonical
|
||||||
|
parent gate. If a child workflow failed but was later rerun successfully, rerun
|
||||||
|
only the failed parent verifier job; do not dispatch a new full umbrella unless
|
||||||
|
the release evidence is stale.
|
||||||
|
|
||||||
|
For bounded recovery after a focused fix, pass `-f rerun_group=<group>`.
|
||||||
|
Supported umbrella groups are `all`, `ci`, `plugin-prerelease`,
|
||||||
|
`release-checks`, `install-smoke`, `cross-os`, `live-e2e`, `package`, `qa`,
|
||||||
|
`qa-parity`, `qa-live`, and `npm-telegram`. Use the narrowest group that covers
|
||||||
|
the failed box. After a targeted release-check fix, do not restart the full
|
||||||
|
umbrella by habit: dispatch the matching `rerun_group` and rerun only the parent
|
||||||
|
verifier/evidence step after the child is green unless the release evidence is
|
||||||
|
stale. For a single failed live/E2E shard, use
|
||||||
|
`-f rerun_group=live-e2e -f live_suite_filter=<suite_id>` so the Blacksmith
|
||||||
|
workflow only spends setup and queue time on that suite.
|
||||||
|
|
||||||
|
### Release Evidence
|
||||||
|
|
||||||
|
After release-candidate validation or before a release decision, record the
|
||||||
|
important run ids in the public `openclaw/releases` evidence ledger.
|
||||||
|
Use the manual `OpenClaw Release Evidence`
|
||||||
|
(`openclaw-release-evidence.yml`) workflow there. It writes durable summaries
|
||||||
|
under `evidence/<release-id>/` and commits:
|
||||||
|
|
||||||
|
- `release-evidence.md`
|
||||||
|
- `release-evidence.json`
|
||||||
|
- `index.json`
|
||||||
|
- `runs/<label>.json`
|
||||||
|
|
||||||
|
Use one run per line:
|
||||||
|
|
||||||
|
```text
|
||||||
|
full-release-validation openclaw/openclaw <run-id> blocking
|
||||||
|
package-acceptance openclaw/openclaw <run-id> blocking
|
||||||
|
release-checks openclaw/openclaw <run-id> blocking
|
||||||
|
```
|
||||||
|
|
||||||
|
Store summaries, run URLs, artifact metadata, timings, pass/fail state, and
|
||||||
|
short release-manager notes there. Do not store raw logs, provider
|
||||||
|
prompts/responses, channel transcripts, signing material, or secret-bearing
|
||||||
|
config in git; raw logs stay in Actions artifacts.
|
||||||
|
|
||||||
|
When `Full Release Validation` completes and `OPENCLAW_RELEASES_DISPATCH_TOKEN`
|
||||||
|
is configured in the source repo, it requests the public
|
||||||
|
`OpenClaw Release Evidence From Full Validation` workflow. That workflow reads
|
||||||
|
the parent full-validation run, extracts the child CI/release-checks/Telegram
|
||||||
|
run ids from the parent logs, and opens the evidence PR automatically. If the
|
||||||
|
token is absent or the run predates this wiring, trigger that workflow manually
|
||||||
|
with the full-validation run id.
|
||||||
|
|
||||||
|
### Release Checks
|
||||||
|
|
||||||
|
`OpenClaw Release Checks` (`openclaw-release-checks.yml`) is the release child
|
||||||
|
workflow. It is broader than normal CI but narrower than the umbrella because it
|
||||||
|
does not dispatch the separate full normal CI child. It runs Package Acceptance
|
||||||
|
with artifact-native delta lanes and `telegram_mode=mock-openai`, so the release
|
||||||
|
package tarball also goes through offline plugin proof, bundled-channel compat,
|
||||||
|
and Telegram package QA. The Docker release-path chunks cover the overlapping
|
||||||
|
package/update/plugin lanes. Use it when release-path validation is needed
|
||||||
|
without rerunning the entire umbrella.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh workflow run openclaw-release-checks.yml \
|
||||||
|
--repo openclaw/openclaw \
|
||||||
|
--ref main \
|
||||||
|
-f ref=<branch-or-sha> \
|
||||||
|
-f provider=openai \
|
||||||
|
-f mode=both \
|
||||||
|
-f release_profile=stable \
|
||||||
|
-f rerun_group=all
|
||||||
|
```
|
||||||
|
|
||||||
|
Release-check rerun groups are `all`, `install-smoke`, `cross-os`, `live-e2e`,
|
||||||
|
`package`, `qa`, `qa-parity`, and `qa-live`.
|
||||||
|
`OpenClaw Release Checks` uses the trusted workflow ref to resolve the selected
|
||||||
|
ref once as `release-package-under-test` and passes that artifact into cross-OS
|
||||||
|
release checks, release-path Docker live/E2E checks, and Package Acceptance.
|
||||||
|
When `Full Release Validation` dispatches release checks, it passes the requested
|
||||||
|
branch/tag plus an `expected_sha` so branch/tag refs resolve through the fast
|
||||||
|
remote-ref path while the package and QA jobs still validate the exact SHA.
|
||||||
|
|
||||||
|
The full install-smoke child is split on purpose: one job prepares or reuses the
|
||||||
|
target-SHA GHCR root Dockerfile smoke image, QR package install runs in its own
|
||||||
|
job, root Dockerfile/gateway smokes pull the prepared image, and installer/Bun
|
||||||
|
smokes pull the same image while building only their small installer images.
|
||||||
|
If install-smoke gets slow again, first check whether the root image was reused
|
||||||
|
or rebuilt before adding/removing coverage.
|
||||||
|
|
||||||
|
The full-profile native live media shards use the prebuilt
|
||||||
|
`ghcr.io/openclaw/openclaw-live-media-runner:ubuntu-24.04` container so
|
||||||
|
`ffmpeg`/`ffprobe` are already present. If those jobs suddenly spend minutes in
|
||||||
|
dependency setup again, first check the `Live Media Runner Image` workflow and
|
||||||
|
the `Verify preinstalled live media dependencies` step before assuming the media
|
||||||
|
tests themselves slowed down.
|
||||||
|
|
||||||
|
The release Docker path intentionally shards the plugin/runtime tail. The
|
||||||
|
workflow uses `plugins-runtime-plugins`, `plugins-runtime-services`, and
|
||||||
|
`plugins-runtime-install-a` through `plugins-runtime-install-d`; aggregate
|
||||||
|
aliases such as `plugins-runtime-core`, `plugins-runtime`, and
|
||||||
|
`plugins-integrations` remain for manual reruns.
|
||||||
|
|
||||||
|
The release QA parity box is internally split into candidate and baseline lane
|
||||||
|
jobs, followed by a report job that downloads both artifacts and runs
|
||||||
|
`pnpm openclaw qa parity-report`. For parity failures, inspect the failed lane
|
||||||
|
first; inspect the report job when both lane summaries exist but the comparison
|
||||||
|
fails.
|
||||||
|
|
||||||
|
### QA Lab Matrix Profiles
|
||||||
|
|
||||||
|
`pnpm openclaw qa matrix` defaults to `--profile all`. Do not assume the CLI
|
||||||
|
default is the fast release path. Use explicit profiles:
|
||||||
|
|
||||||
|
- `--profile fast`: release-critical Matrix transport contract; add
|
||||||
|
`--fail-fast` only when the target CLI supports it
|
||||||
|
- `--profile transport|media|e2ee-smoke|e2ee-deep|e2ee-cli`: sharded full
|
||||||
|
Matrix proof
|
||||||
|
- `OPENCLAW_QA_MATRIX_NO_REPLY_WINDOW_MS=3000`: CI-friendly no-reply quiet
|
||||||
|
window when paired with fast or sharded gates
|
||||||
|
|
||||||
|
`QA-Lab - All Lanes` uses explicit fast Matrix on scheduled runs; manual
|
||||||
|
dispatch keeps `matrix_profile=all` as the default and always shards that full
|
||||||
|
Matrix selection. `OpenClaw Release Checks` uses explicit fast Matrix; run the
|
||||||
|
all-lanes workflow when release investigation needs full Matrix media/E2EE
|
||||||
|
inventory.
|
||||||
|
|
||||||
|
### Reusable Live/E2E Checks
|
||||||
|
|
||||||
|
`OpenClaw Live And E2E Checks (Reusable)`
|
||||||
|
(`openclaw-live-and-e2e-checks-reusable.yml`) is the preferred entry point for
|
||||||
|
targeted live, Docker, model, and E2E proof. Inputs let you turn off unrelated
|
||||||
|
lanes:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh workflow run openclaw-live-and-e2e-checks-reusable.yml \
|
||||||
|
--repo openclaw/openclaw \
|
||||||
|
--ref main \
|
||||||
|
-f ref=<sha> \
|
||||||
|
-f include_repo_e2e=false \
|
||||||
|
-f include_release_path_suites=false \
|
||||||
|
-f include_openwebui=false \
|
||||||
|
-f include_live_suites=true \
|
||||||
|
-f live_models_only=true \
|
||||||
|
-f live_model_providers=fireworks
|
||||||
|
```
|
||||||
|
|
||||||
|
Useful knobs:
|
||||||
|
|
||||||
|
- `docker_lanes='<lane[,lane]>'`: run selected Docker scheduler lanes against
|
||||||
|
prepared artifacts instead of the release chunk matrix. Multiple selected
|
||||||
|
lanes fan out as parallel targeted Docker jobs after one shared package/image
|
||||||
|
preparation step.
|
||||||
|
- `include_live_suites=false`: skip live/provider suites when testing Docker
|
||||||
|
scheduler or release packaging only.
|
||||||
|
- `live_models_only=true`: run only Docker live model coverage.
|
||||||
|
- `live_model_providers=fireworks` (or comma/space separated providers): run one
|
||||||
|
targeted Docker live model job instead of the full provider matrix.
|
||||||
|
- blank `live_model_providers`: run the full live-model provider matrix.
|
||||||
|
|
||||||
|
Release-path Docker chunks are currently `core`, `package-update-openai`,
|
||||||
|
`package-update-anthropic`, `package-update-core`,
|
||||||
|
`plugins-runtime-plugins`, `plugins-runtime-services`,
|
||||||
|
`plugins-runtime-install-a`, `plugins-runtime-install-b`,
|
||||||
|
`plugins-runtime-install-c`, `plugins-runtime-install-d`,
|
||||||
|
`bundled-channels-core`, `bundled-channels-update-a`,
|
||||||
|
`bundled-channels-update-b`, and `bundled-channels-contracts`. The aggregate
|
||||||
|
`bundled-channels`, `plugins-runtime-core`, `plugins-runtime`, and
|
||||||
|
`plugins-integrations` chunks remain valid for manual one-shot reruns, but
|
||||||
|
release checks use the split chunks.
|
||||||
|
|
||||||
|
When live suites are enabled, the workflow shards broad native `pnpm test:live`
|
||||||
|
coverage through `scripts/test-live-shard.mjs` instead of one serial `live-all`
|
||||||
|
job:
|
||||||
|
|
||||||
|
- `native-live-src-agents`
|
||||||
|
- `native-live-src-gateway-core`
|
||||||
|
- `native-live-src-gateway-profiles` (release CI runs this with provider
|
||||||
|
filters such as `OPENCLAW_LIVE_GATEWAY_PROVIDERS=anthropic`)
|
||||||
|
- `native-live-src-gateway-backends`
|
||||||
|
- `native-live-test`
|
||||||
|
- `native-live-extensions-a-k`
|
||||||
|
- `native-live-extensions-l-n`
|
||||||
|
- `native-live-extensions-openai`
|
||||||
|
- `native-live-extensions-o-z`
|
||||||
|
- `native-live-extensions-o-z-other`
|
||||||
|
- `native-live-extensions-xai`
|
||||||
|
- `native-live-extensions-media`
|
||||||
|
- `native-live-extensions-media-audio`
|
||||||
|
- `native-live-extensions-media-music`
|
||||||
|
- `native-live-extensions-media-music-google`
|
||||||
|
- `native-live-extensions-media-music-minimax`
|
||||||
|
- `native-live-extensions-media-video`
|
||||||
|
|
||||||
|
Use `node scripts/test-live-shard.mjs <shard> --list` to see the exact files
|
||||||
|
before rerunning a failed native live shard. The aggregate `o-z` and `media`
|
||||||
|
shards remain useful locally; release CI uses the smaller provider/media shards
|
||||||
|
so one live-provider flake does not force a broad native live rerun.
|
||||||
|
|
||||||
|
For model-list or provider-selection fixes, use `live_models_only=true` plus the
|
||||||
|
specific `live_model_providers` allowlist. Confirm logs show the expected
|
||||||
|
`OPENCLAW_LIVE_PROVIDERS` and selected model ids before declaring proof.
|
||||||
|
|
||||||
|
## Docker
|
||||||
|
|
||||||
|
Docker is expensive. First inspect the scheduler without running Docker:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OPENCLAW_DOCKER_ALL_DRY_RUN=1 pnpm test:docker:all
|
||||||
|
OPENCLAW_DOCKER_ALL_DRY_RUN=1 OPENCLAW_DOCKER_ALL_LANES=install-e2e pnpm test:docker:all
|
||||||
|
OPENCLAW_DOCKER_ALL_LANES=install-e2e node scripts/test-docker-all.mjs --plan-json
|
||||||
|
```
|
||||||
|
|
||||||
|
Run one failed lane locally only when explicitly asked or when GitHub is not
|
||||||
|
usable:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OPENCLAW_DOCKER_ALL_LANES=<lane> \
|
||||||
|
OPENCLAW_DOCKER_ALL_BUILD=0 \
|
||||||
|
OPENCLAW_DOCKER_ALL_PREFLIGHT=0 \
|
||||||
|
OPENCLAW_SKIP_DOCKER_BUILD=1 \
|
||||||
|
OPENCLAW_DOCKER_E2E_BARE_IMAGE='<prepared-bare-image>' \
|
||||||
|
OPENCLAW_DOCKER_E2E_FUNCTIONAL_IMAGE='<prepared-functional-image>' \
|
||||||
|
pnpm test:docker:all
|
||||||
|
```
|
||||||
|
|
||||||
|
For release validation, prefer the reusable GitHub workflow input:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
docker_lanes: install-e2e
|
||||||
|
```
|
||||||
|
|
||||||
|
Multiple lanes are allowed:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
docker_lanes: install-e2e bundled-channel-update-acpx
|
||||||
|
```
|
||||||
|
|
||||||
|
That skips the release chunk matrix and runs one targeted Docker job against the
|
||||||
|
prepared GHCR images and the selected package artifact. Rerun commands
|
||||||
|
generated inside GitHub artifacts include `package_artifact_run_id`,
|
||||||
|
`package_artifact_name`, `docker_e2e_bare_image`, and
|
||||||
|
`docker_e2e_functional_image` when available, so failed lanes can reuse the
|
||||||
|
exact tarball and prepared images from the failed run. When the fix changes
|
||||||
|
package contents, omit those reuse inputs so the workflow packs a new tarball.
|
||||||
|
Live-only targeted reruns skip the E2E images and build only the live-test
|
||||||
|
image. Release-path normal mode fans out into smaller Docker chunk jobs:
|
||||||
|
|
||||||
|
- `core`
|
||||||
|
- `package-update-openai`
|
||||||
|
- `package-update-anthropic`
|
||||||
|
- `package-update-core`
|
||||||
|
- `plugins-runtime-plugins`
|
||||||
|
- `plugins-runtime-services`
|
||||||
|
- `plugins-runtime-install-a`
|
||||||
|
- `plugins-runtime-install-b`
|
||||||
|
- `plugins-runtime-install-c`
|
||||||
|
- `plugins-runtime-install-d`
|
||||||
|
- `bundled-channels`
|
||||||
|
|
||||||
|
OpenWebUI is folded into `plugins-runtime-services` for full release-path
|
||||||
|
coverage and keeps a standalone `openwebui` chunk only for OpenWebUI-only
|
||||||
|
dispatches. The legacy `package-update`, `plugins-runtime-core`,
|
||||||
|
`plugins-runtime`, and `plugins-integrations` chunks still work as aggregate
|
||||||
|
aliases for manual reruns, but the release workflow uses the split chunks so
|
||||||
|
provider installer checks, plugin runtime checks, bundled plugin
|
||||||
|
install/uninstall shards, and bundled-channel checks can run on separate
|
||||||
|
machines. The bundled-channel runtime-dependency coverage
|
||||||
|
inside `bundled-channels`
|
||||||
|
uses the split `bundled-channel-*` and `bundled-channel-update-*` lanes rather
|
||||||
|
than the serial `bundled-channel-deps` lane, so failures produce cheap targeted
|
||||||
|
reruns for the exact channel/update scenario. The bundled plugin
|
||||||
|
install/uninstall sweep is also split into
|
||||||
|
`bundled-plugin-install-uninstall-0` through
|
||||||
|
`bundled-plugin-install-uninstall-7`; selecting the legacy
|
||||||
|
`bundled-plugin-install-uninstall` lane expands to all eight shards.
|
||||||
|
|
||||||
|
## Package Acceptance
|
||||||
|
|
||||||
|
Use the manual `Package Acceptance` workflow when the question is "does this
|
||||||
|
installable package work as a product?" rather than "does this source diff pass
|
||||||
|
Vitest?"
|
||||||
|
|
||||||
|
In release validation, treat Package Acceptance as the package-candidate shard
|
||||||
|
inside the larger release umbrella, not as a competing full-test path. Full
|
||||||
|
Release Validation and private release gauntlets should call Package Acceptance
|
||||||
|
for tarball resolution, Docker product/package proof, and optional Telegram QA
|
||||||
|
against the same resolved `package-under-test` artifact; keep orchestration,
|
||||||
|
secret policy, blocking/advisory status, and evidence rollup in the caller.
|
||||||
|
|
||||||
|
Good defaults:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh workflow run package-acceptance.yml --ref main \
|
||||||
|
-f source=npm \
|
||||||
|
-f workflow_ref=main \
|
||||||
|
-f package_spec=openclaw@beta \
|
||||||
|
-f suite_profile=product \
|
||||||
|
-f telegram_mode=mock-openai
|
||||||
|
```
|
||||||
|
|
||||||
|
Npm candidate selection:
|
||||||
|
|
||||||
|
- Resolve the registry immediately before dispatch:
|
||||||
|
`npm view openclaw dist-tags --json --prefer-online --cache /tmp/openclaw-npm-cache-verify-$$`
|
||||||
|
and `npm view openclaw@beta version dist.tarball dist.integrity --json --prefer-online --cache /tmp/openclaw-npm-cache-verify-$$`.
|
||||||
|
- If Peter asks for "latest beta", use `source=npm` with
|
||||||
|
`package_spec=openclaw@beta`, then record the resolved version from `npm view`
|
||||||
|
or the workflow summary.
|
||||||
|
- For reruns, release proof, or comparing one known package, prefer the exact
|
||||||
|
immutable spec: `package_spec=openclaw@YYYY.M.D-beta.N` or
|
||||||
|
`package_spec=openclaw@YYYY.M.D`.
|
||||||
|
- For stable package proof, use `package_spec=openclaw@latest` only when the
|
||||||
|
question is explicitly the current stable dist-tag; otherwise pin the exact
|
||||||
|
version.
|
||||||
|
- `source=npm` only accepts registry specs for `openclaw@beta`,
|
||||||
|
`openclaw@latest`, or exact OpenClaw release versions. Do not pass semver
|
||||||
|
ranges, git refs, file paths, tarball URLs, or plugin package names there.
|
||||||
|
- If the candidate is a tarball URL, use `source=url` with `package_sha256`. If
|
||||||
|
it is an Actions tarball artifact, use `source=artifact`. If it is an
|
||||||
|
unpublished source candidate, use `source=ref` with a trusted ref or SHA.
|
||||||
|
- Package acceptance tests exactly the selected package candidate. Do not apply
|
||||||
|
`openclaw update --channel beta` fallback semantics here; if `beta` is absent,
|
||||||
|
stale, older than `latest`, or points at a broken tarball, report that tag
|
||||||
|
state instead of silently testing `latest`.
|
||||||
|
|
||||||
|
Profiles:
|
||||||
|
|
||||||
|
- `smoke`: quick confidence that the tarball installs, can onboard a channel,
|
||||||
|
can run an agent turn, and basic gateway/config lanes work.
|
||||||
|
- `package`: release-package contract. Adds installer/update, doctor install
|
||||||
|
switching, bundled plugin runtime deps, plugin install/update, and package
|
||||||
|
repair lanes. This is the default native replacement for most Parallels
|
||||||
|
package/update coverage.
|
||||||
|
- `product`: package profile plus broader product surfaces: MCP channels,
|
||||||
|
cron/subagent cleanup, OpenAI web search, and OpenWebUI.
|
||||||
|
- `full`: split Docker release-path chunks with OpenWebUI.
|
||||||
|
- `custom`: exact `docker_lanes` list for a focused rerun.
|
||||||
|
|
||||||
|
Candidate sources:
|
||||||
|
|
||||||
|
- `source=npm`: `openclaw@beta`, `openclaw@latest`, or an exact release version.
|
||||||
|
- `source=ref`: pack `package_ref` using the trusted `workflow_ref` harness.
|
||||||
|
This intentionally separates old package commits from new workflow/test code.
|
||||||
|
- `source=url`: HTTPS `.tgz` plus required `package_sha256`.
|
||||||
|
- `source=artifact`: download one `.tgz` from `artifact_run_id`/`artifact_name`.
|
||||||
|
|
||||||
|
Ref model:
|
||||||
|
|
||||||
|
- `gh workflow run ... --ref <workflow-ref>` selects the workflow file revision
|
||||||
|
GitHub executes.
|
||||||
|
- `workflow_ref` is the trusted harness/script ref passed to reusable Docker
|
||||||
|
E2E.
|
||||||
|
- `package_ref` is the source ref to build when `source=ref`. It can be an
|
||||||
|
older branch/tag/SHA as long as it is reachable from an OpenClaw branch or
|
||||||
|
release tag.
|
||||||
|
|
||||||
|
Example: run latest package acceptance harness against an older trusted commit:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh workflow run package-acceptance.yml --ref main \
|
||||||
|
-f workflow_ref=main \
|
||||||
|
-f source=ref \
|
||||||
|
-f package_ref=<branch-or-sha> \
|
||||||
|
-f suite_profile=package \
|
||||||
|
-f telegram_mode=mock-openai
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `telegram_mode=mock-openai` or `telegram_mode=live-frontier` when the same
|
||||||
|
resolved `package-under-test` tarball should also run through the Telegram QA
|
||||||
|
workflow in the `qa-live-shared` environment. The standalone Telegram workflow
|
||||||
|
still accepts a published npm spec for post-publish checks, but Package
|
||||||
|
Acceptance passes the resolved artifact for `source=npm`, `ref`, `url`, and
|
||||||
|
`artifact`. Use `telegram_mode=none` only when intentionally skipping Telegram
|
||||||
|
credentialed package proof for a focused rerun.
|
||||||
|
|
||||||
|
Docker E2E images never copy repo sources as the app under test: the bare image
|
||||||
|
is a Node/Git runner, and the functional image installs the same prebuilt npm
|
||||||
|
tarball that bare lanes mount. `scripts/package-openclaw-for-docker.mjs` is the
|
||||||
|
single packer for local scripts and CI and validates the tarball inventory
|
||||||
|
before Docker consumes it. `scripts/test-docker-all.mjs --plan-json` is the
|
||||||
|
scheduler-owned CI plan for image kind, package, live image, lane, and
|
||||||
|
credential needs. Docker lane definitions live in the single scenario catalog
|
||||||
|
`scripts/lib/docker-e2e-scenarios.mjs`; planner logic lives in
|
||||||
|
`scripts/lib/docker-e2e-plan.mjs`. `scripts/docker-e2e.mjs` converts plan and
|
||||||
|
summary JSON into GitHub outputs and step summaries. Every scheduler run writes
|
||||||
|
`.artifacts/docker-tests/**/summary.json` plus `failures.json`. Read those
|
||||||
|
before rerunning. Lane entries include `command`, `rerunCommand`, status,
|
||||||
|
timing, timeout state, image kind, and log file path. The summary also includes
|
||||||
|
top-level phase timings for preflight, image build, package prep, lane pools,
|
||||||
|
and cleanup. Use `pnpm test:docker:timings <summary.json>` to rank slow lanes
|
||||||
|
and phases before deciding whether a broader rerun is justified.
|
||||||
|
|
||||||
|
Skill install proof: use `pnpm test:docker:skill-install` or targeted
|
||||||
|
`docker_lanes=skill-install` for live ClawHub skill-install validation. The
|
||||||
|
lane installs the package tarball in a bare runner, keeps
|
||||||
|
`skills.install.allowUploadedArchives=false`, resolves the current live slug
|
||||||
|
from `openclaw skills search`, installs it, and verifies `.clawhub` origin/lock
|
||||||
|
metadata. Prefer this checked-in script over inline heredoc Testbox recipes.
|
||||||
|
|
||||||
|
## Cheap Docker Reruns
|
||||||
|
|
||||||
|
First derive the smallest rerun command from artifacts:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pnpm test:docker:rerun <github-run-id>
|
||||||
|
pnpm test:docker:rerun .artifacts/docker-tests/<run>/failures.json
|
||||||
|
```
|
||||||
|
|
||||||
|
The script downloads Docker E2E artifacts for a GitHub run, reads
|
||||||
|
`summary.json`/`failures.json`, and prints a combined targeted workflow command
|
||||||
|
plus per-lane commands. Prefer the combined targeted command when several lanes
|
||||||
|
failed for the same patch:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh workflow run openclaw-live-and-e2e-checks-reusable.yml \
|
||||||
|
-f ref=<sha> \
|
||||||
|
-f include_repo_e2e=false \
|
||||||
|
-f include_release_path_suites=false \
|
||||||
|
-f include_openwebui=false \
|
||||||
|
-f docker_lanes='install-e2e bundled-channel-update-acpx' \
|
||||||
|
-f include_live_suites=false \
|
||||||
|
-f live_models_only=false
|
||||||
|
```
|
||||||
|
|
||||||
|
That path still runs the prepare job, so it creates a new tarball for `<sha>`.
|
||||||
|
If the SHA-tagged GHCR bare/functional image already exists, CI skips rebuilding
|
||||||
|
that image and only uploads the fresh package artifact before the targeted lane
|
||||||
|
job. Do not rerun the full release path unless the failed lane list
|
||||||
|
or touched surface really requires it.
|
||||||
|
|
||||||
|
## Docker Expected Timings
|
||||||
|
|
||||||
|
Treat these as ballpark. Blacksmith queue time, GHCR pull speed, provider
|
||||||
|
latency, npm cache state, and Docker daemon health can dominate.
|
||||||
|
|
||||||
|
Current local timing artifact (`.artifacts/docker-tests/lane-timings.json`) has
|
||||||
|
these rough bands:
|
||||||
|
|
||||||
|
- Tiny lanes, seconds to under 1 minute:
|
||||||
|
`agents-delete-shared-workspace` ~3s, `plugin-update` ~7s,
|
||||||
|
`config-reload` ~14s, `pi-bundle-mcp-tools` ~15s, `onboard` ~18s,
|
||||||
|
`session-runtime-context` ~20s, `gateway-network` ~34s, `qr` ~44s.
|
||||||
|
- Medium deterministic lanes, ~1-5 minutes:
|
||||||
|
`npm-onboard-channel-agent` ~96s, `openai-image-auth` ~99s,
|
||||||
|
bundled channel/update lanes usually ~90-300s when split, `openwebui` ~225s,
|
||||||
|
`mcp-channels` ~274s.
|
||||||
|
- Heavy deterministic lanes, ~6-10 minutes:
|
||||||
|
`bundled-channel-root-owned` ~429s,
|
||||||
|
`bundled-channel-setup-entry` ~420s,
|
||||||
|
`bundled-channel-load-failure` ~383s,
|
||||||
|
`cron-mcp-cleanup` ~567s.
|
||||||
|
- Live provider lanes, often ~15-20 minutes:
|
||||||
|
`live-gateway` ~958s, `live-models` ~1054s.
|
||||||
|
- Installer/release lanes:
|
||||||
|
`install-e2e` and package-update paths can vary widely with npm, provider,
|
||||||
|
and package registry behavior. Budget tens of minutes; prefer GitHub targeted
|
||||||
|
reruns over local repeats.
|
||||||
|
|
||||||
|
Default fallback lane timeout is 120 minutes. A timeout usually means debug the
|
||||||
|
lane log/artifacts first, not “run the whole thing again.”
|
||||||
|
|
||||||
|
## Failure Workflow
|
||||||
|
|
||||||
|
1. Identify exact failing job, SHA, lane, and artifact path.
|
||||||
|
2. Read `failures.json`, `summary.json`, and the failed lane log tail.
|
||||||
|
3. Use `pnpm test:docker:rerun <run-id|failures.json>` to generate targeted
|
||||||
|
GitHub rerun commands.
|
||||||
|
4. If the lane has `rerunCommand`, use that only as a local starting point.
|
||||||
|
5. For Docker release failures, dispatch targeted `docker_lanes=<failed-lane>`
|
||||||
|
on GitHub before considering local Docker.
|
||||||
|
6. Patch narrowly, then rerun the failed file/lane only.
|
||||||
|
7. Broaden to `pnpm check:changed` or CI only after the isolated proof passes.
|
||||||
|
|
||||||
|
## When To Escalate
|
||||||
|
|
||||||
|
- Public SDK/plugin contract changes: run changed gate plus relevant extension
|
||||||
|
validation.
|
||||||
|
- Build output, lazy imports, package boundaries, or published surfaces:
|
||||||
|
include `pnpm build`.
|
||||||
|
- Workflow edits: run `pnpm check:workflows`.
|
||||||
|
- Release branch or tag validation: use release docs and GitHub workflows; avoid
|
||||||
|
local Docker unless Peter explicitly asks.
|
||||||
4
.agents/skills/openclaw-testing/agents/openai.yaml
Normal file
4
.agents/skills/openclaw-testing/agents/openai.yaml
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "OpenClaw Testing"
|
||||||
|
short_description: "Choose cheap, targeted OpenClaw validation"
|
||||||
|
default_prompt: "Use $openclaw-testing to choose the cheapest safe test or CI verification path, inspect failures, and rerun only the relevant OpenClaw lane."
|
||||||
63
.agents/skills/parallels-discord-roundtrip/SKILL.md
Normal file
63
.agents/skills/parallels-discord-roundtrip/SKILL.md
Normal file
@@ -0,0 +1,63 @@
|
|||||||
|
---
|
||||||
|
name: parallels-discord-roundtrip
|
||||||
|
description: Run macOS Parallels smoke with Discord send, host verification, host reply, and guest readback proof.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Parallels Discord Roundtrip
|
||||||
|
|
||||||
|
Use when macOS Parallels smoke must prove Discord two-way delivery end to end.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Cover:
|
||||||
|
|
||||||
|
- install on fresh macOS snapshot
|
||||||
|
- onboard + gateway health
|
||||||
|
- guest `message send` to Discord
|
||||||
|
- host sees that message on Discord
|
||||||
|
- host posts a new Discord message
|
||||||
|
- guest `message read` sees that new message
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
- host env var with Discord bot token
|
||||||
|
- Discord guild ID
|
||||||
|
- Discord channel ID
|
||||||
|
- `OPENAI_API_KEY`
|
||||||
|
|
||||||
|
## Preferred run
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export OPENCLAW_PARALLELS_DISCORD_TOKEN="$(
|
||||||
|
ssh peters-mac-studio-1 'jq -r ".channels.discord.token" ~/.openclaw/openclaw.json' | tr -d '\n'
|
||||||
|
)"
|
||||||
|
|
||||||
|
pnpm test:parallels:macos \
|
||||||
|
--discord-token-env OPENCLAW_PARALLELS_DISCORD_TOKEN \
|
||||||
|
--discord-guild-id 1456350064065904867 \
|
||||||
|
--discord-channel-id 1456744319972282449 \
|
||||||
|
--json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- Snapshot target: closest to `macOS 26.3.1 fresh`.
|
||||||
|
- Snapshot resolver now prefers matching `*-poweroff*` clones when the base hint also matches. That lets the harness reuse disk-only recovery snapshots without passing a longer hint.
|
||||||
|
- If Windows/Linux snapshot restore logs show `PET_QUESTION_SNAPSHOT_STATE_INCOMPATIBLE_CPU`, drop the suspended state once, create a `*-poweroff*` replacement snapshot, and rerun. The smoke scripts now auto-start restored power-off snapshots.
|
||||||
|
- Harness configures Discord inside the guest; no checked-in token/config.
|
||||||
|
- Use the `openclaw` wrapper for guest `message send/read`; `node openclaw.mjs message ...` does not expose the lazy message subcommands the same way.
|
||||||
|
- Write `channels.discord.guilds` in one JSON object (`--strict-json`), not dotted `config set channels.discord.guilds.<snowflake>...` paths; numeric snowflakes get treated like array indexes.
|
||||||
|
- Avoid `prlctl enter` / expect for long Discord setup scripts; it line-wraps/corrupts long commands. Use `prlctl exec --current-user /bin/sh -lc ...` for the Discord config phase.
|
||||||
|
- Full 3-OS sweeps: the shared build lock is safe in parallel, but snapshot restore is still a Parallels bottleneck. Prefer serialized Windows/Linux restore-heavy reruns if the host is already under load.
|
||||||
|
- Harness cleanup deletes the temporary Discord smoke messages at exit.
|
||||||
|
- After a successful Discord roundtrip, shut down the macOS guest before handoff (`prlctl stop "macOS Tahoe"`). The macOS smoke harness should do this automatically after successful Discord proof; still stop the VM manually after ad-hoc Discord checks. Do not leave the Discord-configured VM running; it can keep reading/posting in `#maintainer` and spam Discord after the proof is complete.
|
||||||
|
- Per-phase logs: `/tmp/openclaw-parallels-smoke.*`
|
||||||
|
- Machine summary: pass `--json`
|
||||||
|
- If roundtrip flakes, inspect `fresh.discord-roundtrip.log` and `discord-last-readback.json` in the run dir first.
|
||||||
|
|
||||||
|
## Pass criteria
|
||||||
|
|
||||||
|
- fresh lane or upgrade lane requested passes
|
||||||
|
- summary reports `discord=pass` for that lane
|
||||||
|
- guest outbound nonce appears in channel history
|
||||||
|
- host inbound nonce appears in `openclaw message read` output
|
||||||
148
.agents/skills/security-triage/SKILL.md
Normal file
148
.agents/skills/security-triage/SKILL.md
Normal file
@@ -0,0 +1,148 @@
|
|||||||
|
---
|
||||||
|
name: security-triage
|
||||||
|
description: "Triage OpenClaw security advisories, drafts, and GHSA reports with shipped-tag and trust-model proof."
|
||||||
|
---
|
||||||
|
|
||||||
|
# Security Triage
|
||||||
|
|
||||||
|
Use when reviewing OpenClaw security advisories, drafts, or GHSA reports.
|
||||||
|
|
||||||
|
Goal: high-confidence maintainers' triage without over-closing real issues or shipping unnecessary regressions.
|
||||||
|
|
||||||
|
## Close Bar
|
||||||
|
|
||||||
|
Close only if one of these is true:
|
||||||
|
|
||||||
|
- duplicate of an existing advisory or fixed issue
|
||||||
|
- invalid against shipped behavior
|
||||||
|
- out of scope under `SECURITY.md`
|
||||||
|
- fixed before any affected release/tag
|
||||||
|
|
||||||
|
Do not close only because `main` is fixed. If latest shipped tag or npm release is affected, keep it open until released or published with the right status.
|
||||||
|
|
||||||
|
## Required Reads
|
||||||
|
|
||||||
|
Before answering:
|
||||||
|
|
||||||
|
1. Read `SECURITY.md`.
|
||||||
|
2. Read the GHSA body with `gh api /repos/openclaw/openclaw/security-advisories/<GHSA>`.
|
||||||
|
3. Inspect the exact implicated code paths.
|
||||||
|
4. Verify shipped state:
|
||||||
|
- `git tag --sort=-creatordate | head`
|
||||||
|
- `npm view openclaw version --userconfig "$(mktemp)"`
|
||||||
|
- `git tag --contains <fix-commit>`
|
||||||
|
- if needed: `git show <tag>:path/to/file`
|
||||||
|
5. Search for canonical overlap:
|
||||||
|
- existing published GHSAs
|
||||||
|
- older fixed bugs
|
||||||
|
- same trust-model class already covered in `SECURITY.md`
|
||||||
|
|
||||||
|
## Review Method
|
||||||
|
|
||||||
|
For each advisory, decide:
|
||||||
|
|
||||||
|
- `close`
|
||||||
|
- `keep open`
|
||||||
|
- `keep open but narrow`
|
||||||
|
|
||||||
|
Default to one advisory at a time when comments/closures are involved:
|
||||||
|
|
||||||
|
1. Review exactly one GHSA.
|
||||||
|
2. Print the GHSA URL first.
|
||||||
|
3. Summarize the decision and evidence for discussion.
|
||||||
|
4. Draft one maintainer-ready comment.
|
||||||
|
5. Copy only that one comment to the clipboard.
|
||||||
|
6. Stop and wait for Peter to post/discuss before moving to the next GHSA.
|
||||||
|
|
||||||
|
Do not batch multiple close comments unless Peter explicitly asks for a batch.
|
||||||
|
|
||||||
|
Check in this order:
|
||||||
|
|
||||||
|
1. Trust model
|
||||||
|
- Is the prerequisite already inside trusted host/local/plugin/operator state?
|
||||||
|
- Does `SECURITY.md` explicitly call this class out as out of scope or hardening-only?
|
||||||
|
2. Shipped behavior
|
||||||
|
- Is the bug present in the latest shipped tag or npm release?
|
||||||
|
- Was it fixed before release?
|
||||||
|
3. Exploit path
|
||||||
|
- Does the report show a real boundary bypass, not just prompt injection, local same-user control, or helper-level semantics?
|
||||||
|
- If data only moves between trusted workspace-memory files called out in `SECURITY.md`, do not treat "injection markers" alone as a security bug.
|
||||||
|
- In that case, frame sanitization as optional hardening only if it preserves expected memory workflows.
|
||||||
|
4. Functional tradeoff
|
||||||
|
- If a hardening change would reduce intended user functionality, call that out before proposing it.
|
||||||
|
- Prefer fixes that preserve user workflows over deny-by-default regressions unless the boundary demands it.
|
||||||
|
5. Hardening follow-up
|
||||||
|
- Even when the GHSA should close, ask whether a narrow hardening change would reduce footguns without changing the documented trust boundary.
|
||||||
|
- Separate hardening from vulnerability status. Phrase it as "not required for GHSA closure, but worth considering".
|
||||||
|
- Bring up hardening only if it is concrete, low-risk, and preserves intended maintainer/operator workflows.
|
||||||
|
- If hardening would require a product/security model change, say that explicitly and do not imply it is a required fix for closure.
|
||||||
|
|
||||||
|
## Response Format
|
||||||
|
|
||||||
|
When preparing a maintainer-ready close reply:
|
||||||
|
|
||||||
|
1. Print the GHSA URL first.
|
||||||
|
2. Then draft a detailed response the maintainer can post.
|
||||||
|
3. Include:
|
||||||
|
- exact reason for close
|
||||||
|
- exact code refs
|
||||||
|
- exact shipped tag / release facts
|
||||||
|
- fix provenance or canonical duplicate GHSA when applicable
|
||||||
|
- optional hardening note only if worthwhile and functionality-preserving
|
||||||
|
|
||||||
|
Keep tone firm, specific, non-defensive.
|
||||||
|
|
||||||
|
## Public Wording Hygiene
|
||||||
|
|
||||||
|
- Keep raw commit hashes, PR titles/numbers, and fix-mechanism summaries out of public advisory text. Use the patched release/version field only.
|
||||||
|
- Keep exact commit SHAs, PRs, and implementation notes in internal notes and verification files.
|
||||||
|
- For hardening/no-publish outcomes, do not add exploit-heavy details, "Fixed by" text, or a "Fix Commit(s)" section. Thank reporters, preserve credit, state the `SECURITY.md` boundary, and say clearly that the GHSA will close without publication.
|
||||||
|
- For published CVE/GHSA text, prefer `### Patched Versions` with the fixed release. Do not explain how the patch works unless Peter explicitly asks for that public detail.
|
||||||
|
- Keep GHSA ids out of changelog and release-note wording unless Peter explicitly asks.
|
||||||
|
|
||||||
|
## Discussion Mode
|
||||||
|
|
||||||
|
When Peter is manually posting GHSA comments, use this flow:
|
||||||
|
|
||||||
|
1. Show the URL.
|
||||||
|
2. Give a terse verdict (`close`, `keep open`, or `keep open but narrow`).
|
||||||
|
3. List the strongest evidence bullets.
|
||||||
|
4. State any optional hardening follow-up separately from the close reason.
|
||||||
|
5. Copy the proposed comment body with `pbcopy`.
|
||||||
|
6. End the reply after the one advisory. Do not continue to the next advisory until Peter says to continue.
|
||||||
|
|
||||||
|
If the GitHub API cannot post comments for private advisories, say so once and keep using clipboard/UI paste.
|
||||||
|
|
||||||
|
## Clipboard Step
|
||||||
|
|
||||||
|
After drafting the final post body for the current advisory, copy it:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pbcopy <<'EOF'
|
||||||
|
<final response>
|
||||||
|
EOF
|
||||||
|
```
|
||||||
|
|
||||||
|
Tell the user that the clipboard now contains the proposed response for that advisory.
|
||||||
|
|
||||||
|
## Useful Commands
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh api /repos/openclaw/openclaw/security-advisories/<GHSA>
|
||||||
|
gh api /repos/openclaw/openclaw/security-advisories --paginate
|
||||||
|
git tag --sort=-creatordate | head -n 20
|
||||||
|
npm view openclaw version --userconfig "$(mktemp)"
|
||||||
|
git tag --contains <commit>
|
||||||
|
git show <tag>:<path>
|
||||||
|
gh search issues --repo openclaw/openclaw --match title,body,comments -- "<terms>"
|
||||||
|
gh search prs --repo openclaw/openclaw --match title,body,comments -- "<terms>"
|
||||||
|
```
|
||||||
|
|
||||||
|
## Decision Notes
|
||||||
|
|
||||||
|
- “fixed on main, unreleased” is usually not a close.
|
||||||
|
- “needs attacker-controlled trusted local state first” is usually out of scope.
|
||||||
|
- “same-host same-user process can already read/write local state” is usually out of scope.
|
||||||
|
- “trusted workspace memory promotes/reindexes trusted workspace memory” is usually out of scope unless it crosses a documented boundary.
|
||||||
|
- “helper function behaves differently than documented config semantics” is usually invalid.
|
||||||
|
- If only the severity is wrong but the bug is real, keep it open and narrow the impact in the reply.
|
||||||
439
.agents/skills/tag-duplicate-prs-issues/SKILL.md
Normal file
439
.agents/skills/tag-duplicate-prs-issues/SKILL.md
Normal file
@@ -0,0 +1,439 @@
|
|||||||
|
---
|
||||||
|
name: tag-duplicate-prs-issues
|
||||||
|
description: Use gitcrawl to search duplicate OpenClaw PRs/issues, group related work in prtags, and sync duplicate state to GitHub.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Tag Duplicate PRs and Issues
|
||||||
|
|
||||||
|
Use this skill when a maintainer needs to decide whether a pull request or issue is a duplicate of existing work.
|
||||||
|
|
||||||
|
This skill is for maintainer triage and grouping.
|
||||||
|
It is not for reviewing the implementation quality of a PR.
|
||||||
|
|
||||||
|
## Required Setup
|
||||||
|
|
||||||
|
Do not write duplicate groups or annotations until this setup is complete.
|
||||||
|
Read-only discovery can still proceed with `gitcrawl` and live `gh`.
|
||||||
|
|
||||||
|
### Companion Skills
|
||||||
|
|
||||||
|
Use `$gitcrawl` first for local candidate discovery.
|
||||||
|
Use the `prtags` skill from the `prtags` repo at `skills/prtags/SKILL.md` when it is available.
|
||||||
|
|
||||||
|
### Install the CLIs
|
||||||
|
|
||||||
|
Install `prtags` from its latest GitHub release.
|
||||||
|
Do not rely on an old local build unless the maintainer explicitly wants to test unreleased behavior.
|
||||||
|
|
||||||
|
`prtags` CLI install path:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -fsSL https://raw.githubusercontent.com/dutifuldev/prtags/main/scripts/install-prtags.sh | bash -s -- --bin-dir "$HOME/.local/bin"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Authenticate prtags
|
||||||
|
|
||||||
|
`prtags` should be logged in with the maintainer's own GitHub account through OAuth device flow.
|
||||||
|
Do not use a shared maintainer token for interactive triage.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags auth login
|
||||||
|
prtags auth status
|
||||||
|
```
|
||||||
|
|
||||||
|
The expected outcome is that `prtags` stores the logged-in maintainer identity locally and uses that account for authenticated writes.
|
||||||
|
|
||||||
|
## Missing-Setup Rule
|
||||||
|
|
||||||
|
Do not require an up-front preflight before starting the workflow.
|
||||||
|
Proceed with the normal steps until you actually need a tool or account state.
|
||||||
|
|
||||||
|
As soon as you discover that `prtags` is missing or not logged in at the write step, stop immediately.
|
||||||
|
Do not continue in a partial write mode after that point.
|
||||||
|
|
||||||
|
If `prtags` is missing, ask the user to run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -fsSL https://raw.githubusercontent.com/dutifuldev/prtags/main/scripts/install-prtags.sh | bash -s -- --bin-dir "$HOME/.local/bin"
|
||||||
|
```
|
||||||
|
|
||||||
|
If `prtags auth status` shows that the user is not logged in, ask the user to run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags auth login
|
||||||
|
```
|
||||||
|
|
||||||
|
Resume only after the missing tool or login state has been fixed.
|
||||||
|
|
||||||
|
## Read-Path Default
|
||||||
|
|
||||||
|
For candidate discovery in this workflow, use `gitcrawl` first.
|
||||||
|
Treat it as the local history and clustering layer for related issues, duplicate attempts, and closed threads.
|
||||||
|
|
||||||
|
Use live `gh` or `gh api` for the target thread and for any candidate before making an actionable judgment.
|
||||||
|
Use live GitHub when `gitcrawl` is missing or stale for a concrete reason, such as:
|
||||||
|
|
||||||
|
- the target or candidate is not present yet
|
||||||
|
- the local data is clearly stale or incomplete for the decision you need to make
|
||||||
|
- `gitcrawl` errors, times out, or lacks the needed neighbor/search data
|
||||||
|
|
||||||
|
When you fall back to live GitHub search, note that you did so and why.
|
||||||
|
|
||||||
|
If a later `prtags` target-level write fails because its own mirror has not caught up, stop and report that the curation backend is missing the target object instead of forcing a fallback write.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
For each target PR or issue:
|
||||||
|
|
||||||
|
1. gather duplicate evidence
|
||||||
|
2. decide whether it is a real duplicate
|
||||||
|
3. create or reuse one `prtags` group for that duplicate cluster
|
||||||
|
4. save the maintainer judgment in `prtags`
|
||||||
|
5. rely on normal `prtags` group writes to drive GitHub comment sync when that integration is configured
|
||||||
|
|
||||||
|
## Tool Roles
|
||||||
|
|
||||||
|
Use the tools with these boundaries:
|
||||||
|
|
||||||
|
- `gitcrawl` is candidate generation and historical context
|
||||||
|
- use it first for local title/body search, neighbors, clusters, and closed-thread discovery
|
||||||
|
- treat every candidate as a lead until live GitHub confirms it
|
||||||
|
- `gh` is live GitHub truth
|
||||||
|
- use it for target state, body, comments, reviews, files, linked issues, and current open/closed/merged status
|
||||||
|
- use `gh search` only when `gitcrawl` is stale, missing data, or cannot express the needed query
|
||||||
|
- `prtags` is the maintainer curation layer
|
||||||
|
- use it to create or reuse one duplicate group
|
||||||
|
- use it to save the duplicate status, confidence, rationale, and group summary
|
||||||
|
- use it as the source of truth for the GitHub-facing group comment
|
||||||
|
|
||||||
|
## Working Rules
|
||||||
|
|
||||||
|
- Do not call something a duplicate only because the titles are similar.
|
||||||
|
- Do not call something a duplicate only because the same files changed.
|
||||||
|
- A duplicate cluster should be based on the same user-facing problem, the same intent, and substantially overlapping implementation or investigation context.
|
||||||
|
|
||||||
|
## One-Group Rule
|
||||||
|
|
||||||
|
Treat duplicate groups as exclusive.
|
||||||
|
A PR or issue should belong to at most one duplicate group at a time.
|
||||||
|
|
||||||
|
That means:
|
||||||
|
|
||||||
|
- before creating a new group, search for an existing group that already represents the same duplicate story
|
||||||
|
- if the target already appears to belong to a different duplicate group, stop and resolve that conflict first
|
||||||
|
- do not create a second group for the same target just because the wording is slightly different
|
||||||
|
- if two plausible existing groups overlap and you cannot safely merge the judgment, stop and ask the maintainer
|
||||||
|
|
||||||
|
This rule matters more than speed.
|
||||||
|
The skill should keep one coherent duplicate cluster per problem, not many near-duplicate clusters.
|
||||||
|
|
||||||
|
## What A Good Duplicate Group Represents
|
||||||
|
|
||||||
|
A duplicate group should describe the underlying problem and the intended fix direction.
|
||||||
|
Do not group items only because they share a keyword.
|
||||||
|
|
||||||
|
Good group shape:
|
||||||
|
|
||||||
|
- same user-facing bug or same maintainer-facing task
|
||||||
|
- same subsystem or code surface
|
||||||
|
- same intended change direction
|
||||||
|
- same likely duplicate-resolution path
|
||||||
|
|
||||||
|
Bad group shape:
|
||||||
|
|
||||||
|
- “all PRs that touch Slack”
|
||||||
|
- “all issues mentioning retry”
|
||||||
|
- “all auth-related items”
|
||||||
|
|
||||||
|
The group title should name the real problem.
|
||||||
|
The group description should summarize the intent and the code surface.
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
|
||||||
|
- `gateway: startup regression from channel status bootstrap`
|
||||||
|
- `whatsapp: QR preflight timeout handling`
|
||||||
|
- `release: cross-OS validation handoff gaps`
|
||||||
|
|
||||||
|
## Evidence Checklist
|
||||||
|
|
||||||
|
Before declaring a duplicate, gather evidence from at least two categories.
|
||||||
|
`gitcrawl` neighbors, search hits, and cluster membership count as candidate generation, not as enough proof by themselves.
|
||||||
|
|
||||||
|
For PRs:
|
||||||
|
|
||||||
|
- same or nearly same problem statement
|
||||||
|
- same changed files or overlapping file ranges
|
||||||
|
- same fix direction
|
||||||
|
- same subsystem and failure mode
|
||||||
|
- same linked issue or same user-visible symptom
|
||||||
|
|
||||||
|
For issues:
|
||||||
|
|
||||||
|
- same user-visible problem
|
||||||
|
- same reproduction story or same failure mode
|
||||||
|
- same likely fix area
|
||||||
|
- same PRs already linked or discussed
|
||||||
|
- same maintainers already steering toward the same duplicate grouping
|
||||||
|
|
||||||
|
If you only have wording similarity, that is not enough.
|
||||||
|
|
||||||
|
## Step 1: Read The Target
|
||||||
|
|
||||||
|
Start by reading the target itself.
|
||||||
|
Use live GitHub for current target state.
|
||||||
|
|
||||||
|
For a PR:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh pr view <number> --json number,title,state,mergedAt,body,closingIssuesReferences,files,comments,reviews,statusCheckRollup
|
||||||
|
```
|
||||||
|
|
||||||
|
For an issue:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh issue view <number> --json number,title,state,body,comments,closedAt
|
||||||
|
```
|
||||||
|
|
||||||
|
Record:
|
||||||
|
|
||||||
|
- target type and number
|
||||||
|
- title
|
||||||
|
- problem statement
|
||||||
|
- proposed intent
|
||||||
|
- subsystem
|
||||||
|
- whether it is open, closed, or merged
|
||||||
|
- whether there is already a likely duplicate thread mentioned by humans
|
||||||
|
|
||||||
|
## Step 2: Search Broadly With Gitcrawl
|
||||||
|
|
||||||
|
Use `gitcrawl` first because it is the local OpenClaw history and clustering source.
|
||||||
|
Do not switch to broad live GitHub search unless `gitcrawl` is missing data, stale, or failing.
|
||||||
|
|
||||||
|
Start with the target and nearby threads:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gitcrawl threads openclaw/openclaw --numbers <issue-or-pr-number> --include-closed --json
|
||||||
|
gitcrawl neighbors openclaw/openclaw --number <issue-or-pr-number> --limit 20 --json
|
||||||
|
```
|
||||||
|
|
||||||
|
Then search key phrases and subsystem terms:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gitcrawl search openclaw/openclaw --query "<key phrase from title or body>" --mode hybrid --limit 20 --json
|
||||||
|
gitcrawl search openclaw/openclaw --query "<subsystem or error phrase>" --mode hybrid --limit 20 --json
|
||||||
|
```
|
||||||
|
|
||||||
|
Inspect likely clusters:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gitcrawl cluster-detail openclaw/openclaw --id <cluster-id> --member-limit 20 --body-chars 280 --json
|
||||||
|
```
|
||||||
|
|
||||||
|
For PRs, verify likely code overlap with live file data:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh pr view <candidate-pr> --json number,title,state,mergedAt,files,body,comments,reviews
|
||||||
|
```
|
||||||
|
|
||||||
|
For issues, verify likely duplicate issue state and comments live:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh issue view <candidate-issue> --json number,title,state,body,comments,closedAt
|
||||||
|
```
|
||||||
|
|
||||||
|
## Step 3: Use Live GitHub Search For Gaps
|
||||||
|
|
||||||
|
Use targeted live GitHub search after `gitcrawl` when:
|
||||||
|
|
||||||
|
- the target is too new for the local store
|
||||||
|
- comments or reviews matter and the local store lacks them
|
||||||
|
- the exact phrase did not appear in local results but the issue/PR is current enough that GitHub should know it
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gh search prs --repo openclaw/openclaw --match title,body --limit 50 -- "<key phrase>"
|
||||||
|
gh search issues --repo openclaw/openclaw --match title,body --limit 50 -- "<key phrase>"
|
||||||
|
gh search issues --repo openclaw/openclaw --match comments --limit 50 -- "<error or maintainer phrase>"
|
||||||
|
```
|
||||||
|
|
||||||
|
## Step 4: Decide The Outcome
|
||||||
|
|
||||||
|
Choose one of these outcomes:
|
||||||
|
|
||||||
|
- `not_duplicate`
|
||||||
|
- `duplicate_needs_judgment`
|
||||||
|
- `duplicate_confirmed`
|
||||||
|
|
||||||
|
Use `duplicate_confirmed` only when the evidence is strong enough that the maintainer could safely close or retag the duplicate item.
|
||||||
|
|
||||||
|
Use `duplicate_needs_judgment` when:
|
||||||
|
|
||||||
|
- the problem looks the same but the implementation goal differs
|
||||||
|
- the code overlap is weak
|
||||||
|
- the issue wording is ambiguous
|
||||||
|
- there may be two valid duplicate group interpretations
|
||||||
|
- the target appears to intersect two existing duplicate groups
|
||||||
|
|
||||||
|
## Step 5: Reuse Or Create One prtags Group
|
||||||
|
|
||||||
|
Before creating a group, search `prtags` for an existing one.
|
||||||
|
|
||||||
|
Start with text search over groups:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags search text -R openclaw/openclaw "<problem phrase>" --types group --limit 10
|
||||||
|
prtags search similar -R openclaw/openclaw "<problem summary>" --types group --limit 10
|
||||||
|
prtags group list -R openclaw/openclaw
|
||||||
|
```
|
||||||
|
|
||||||
|
Inspect likely groups:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags group get <group-id>
|
||||||
|
prtags group get <group-id> --include-metadata
|
||||||
|
```
|
||||||
|
|
||||||
|
Reuse an existing group when:
|
||||||
|
|
||||||
|
- it represents the same problem
|
||||||
|
- it already contains clearly related members
|
||||||
|
- adding the target would keep the group coherent
|
||||||
|
|
||||||
|
Do not widen an existing group just because `gitcrawl` placed several PRs or issues near each other.
|
||||||
|
Confirm that the actual implementation path and maintainer intent still match before adding the new member.
|
||||||
|
|
||||||
|
Create a new group only when no existing group clearly fits.
|
||||||
|
|
||||||
|
Create the group with a problem-based title and an intent-based description:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags group create -R openclaw/openclaw \
|
||||||
|
--kind mixed \
|
||||||
|
--title "<problem-centered title>" \
|
||||||
|
--description "<same intent, subsystem, and duplicate-resolution path>" \
|
||||||
|
--status open
|
||||||
|
```
|
||||||
|
|
||||||
|
Then attach the target and any known duplicate members:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags group add-pr <group-id> <pr-number>
|
||||||
|
prtags group add-issue <group-id> <issue-number>
|
||||||
|
```
|
||||||
|
|
||||||
|
If a target appears to already belong to another duplicate group and you cannot safely reuse that group, stop.
|
||||||
|
Do not create a second group.
|
||||||
|
|
||||||
|
## Step 6: Ensure The Annotation Fields Exist
|
||||||
|
|
||||||
|
Use `field ensure` so the skill is idempotent.
|
||||||
|
|
||||||
|
Recommended target-level fields:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_status --scope pull_request --type enum --enum-values not_duplicate,candidate,confirmed --filterable
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_status --scope issue --type enum --enum-values not_duplicate,candidate,confirmed --filterable
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_confidence --scope pull_request --type enum --enum-values low,medium,high --filterable
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_confidence --scope issue --type enum --enum-values low,medium,high --filterable
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_rationale --scope pull_request --type text --searchable
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_rationale --scope issue --type text --searchable
|
||||||
|
```
|
||||||
|
|
||||||
|
Recommended group-level fields:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_confidence --scope group --type enum --enum-values low,medium,high --filterable
|
||||||
|
prtags field ensure -R openclaw/openclaw --name duplicate_rationale --scope group --type text --searchable
|
||||||
|
prtags field ensure -R openclaw/openclaw --name cluster_summary --scope group --type text --searchable
|
||||||
|
```
|
||||||
|
|
||||||
|
## Step 7: Save The Maintainer Judgment In prtags
|
||||||
|
|
||||||
|
For a PR:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags annotation pr set -R openclaw/openclaw <pr-number> \
|
||||||
|
duplicate_status=confirmed \
|
||||||
|
duplicate_confidence=high \
|
||||||
|
duplicate_rationale="<same problem, same fix direction, overlapping files and comments>"
|
||||||
|
```
|
||||||
|
|
||||||
|
For an issue:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags annotation issue set -R openclaw/openclaw <issue-number> \
|
||||||
|
duplicate_status=confirmed \
|
||||||
|
duplicate_confidence=high \
|
||||||
|
duplicate_rationale="<same user-visible problem and same intended fix path>"
|
||||||
|
```
|
||||||
|
|
||||||
|
For the group:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags annotation group set <group-id> \
|
||||||
|
duplicate_confidence=high \
|
||||||
|
cluster_summary="<one-sentence problem summary>" \
|
||||||
|
duplicate_rationale="<why these items belong in one duplicate cluster>"
|
||||||
|
```
|
||||||
|
|
||||||
|
When the evidence is incomplete, set `duplicate_status=candidate` and lower the confidence.
|
||||||
|
|
||||||
|
If a per-PR or per-issue annotation write fails because `prtags` cannot resolve the target, do not force a fallback write path.
|
||||||
|
Keep the group state you were able to write, report that the curation backend is still missing the target object, and defer the target-level annotation until `prtags` catches up.
|
||||||
|
|
||||||
|
## Step 8: Let prtags Sync The Group Comment
|
||||||
|
|
||||||
|
Do not tell the agent to create a GitHub comment directly.
|
||||||
|
`prtags` owns the outbound GitHub comment as a derived projection of group state.
|
||||||
|
|
||||||
|
In the normal case, do not manually trigger comment sync.
|
||||||
|
When comment sync is configured, group writes already enqueue the derived comment projection automatically.
|
||||||
|
|
||||||
|
Use manual sync only as a repair or retry path:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags group sync-comments <group-id>
|
||||||
|
```
|
||||||
|
|
||||||
|
If the maintainer needs to see which groups still need attention, use:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
prtags group list-comment-sync-targets -R openclaw/openclaw
|
||||||
|
```
|
||||||
|
|
||||||
|
The skill should treat the GitHub comment as a consequence of correct `prtags` group state.
|
||||||
|
It should not treat manual comment authoring as part of the normal duplicate workflow.
|
||||||
|
It should also not treat `sync-comments` as a required step for every duplicate decision.
|
||||||
|
|
||||||
|
## Output Format
|
||||||
|
|
||||||
|
Return a short maintainer report with these sections:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Decision: duplicate_confirmed | duplicate_needs_judgment | not_duplicate
|
||||||
|
Target: PR #<n> | Issue #<n>
|
||||||
|
Confidence: high | medium | low
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
- ...
|
||||||
|
- ...
|
||||||
|
- ...
|
||||||
|
|
||||||
|
prtags actions:
|
||||||
|
- reused group <group-id> | created group <group-id>
|
||||||
|
- added members: ...
|
||||||
|
- annotations written: ...
|
||||||
|
- comment sync: automatic if configured | manual repair triggered for <group-id>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Stop Conditions
|
||||||
|
|
||||||
|
Stop and escalate instead of forcing a duplicate decision when:
|
||||||
|
|
||||||
|
- the target appears to belong to two different duplicate groups
|
||||||
|
- the duplicate grouping is unclear
|
||||||
|
- the wording matches but the implementation goals differ
|
||||||
|
- two PRs touch the same files for different reasons
|
||||||
|
- two issues describe similar symptoms but likely different root causes
|
||||||
|
|
||||||
|
The maintainer should get one clean duplicate judgment or an explicit “needs judgment” result.
|
||||||
|
Do not blur the line.
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Tag Duplicate PRs and Issues"
|
||||||
|
short_description: "Find duplicate PRs and issues with gitcrawl, group them in prtags, and let prtags sync the GitHub comment"
|
||||||
|
default_prompt: "Use $tag-duplicate-prs-issues to decide whether an OpenClaw PR or issue is a duplicate, gather candidates with gitcrawl, verify live state with GitHub, group related items in prtags, and save the duplicate judgment."
|
||||||
79
.agents/skills/technical-documentation/SKILL.md
Normal file
79
.agents/skills/technical-documentation/SKILL.md
Normal file
@@ -0,0 +1,79 @@
|
|||||||
|
---
|
||||||
|
name: technical-documentation
|
||||||
|
description: Build and review high-quality technical docs as well as agent instruction files in your repository.
|
||||||
|
license: MIT
|
||||||
|
metadata:
|
||||||
|
source: "https://github.com/vincentkoc/dotskills"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Technical Documentation
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Produce and review technical documentation that is clear, actionable, and maintainable for both humans and agents, including contributor-governance files and agent instruction files.
|
||||||
|
|
||||||
|
## When to use
|
||||||
|
|
||||||
|
- Creating or overhauling docs in an existing product/codebase (brownfield).
|
||||||
|
- Building evergreen docs meant to stay accurate and reusable over time.
|
||||||
|
- Reviewing doc diffs for structure, clarity, and operational correctness.
|
||||||
|
- Running full-repo documentation audits that must include both governance files and product docs surfaces (`docs/`, `README*`, `.md/.mdx/.mdc`, Fern/Sphinx/Mintlify-style sources).
|
||||||
|
- Updating or reviewing AGENTS.md and/or CONTRIBUTING.md to keep agent and contributor workflows aligned with current repo practices.
|
||||||
|
- Improving repository onboarding/docs that include contribution instructions, issue templates, PR flow, and review gates.
|
||||||
|
- Designing governance documentation strategy for repos with alias instruction files (for example `CLAUDE.md`, `AGENT.md`, `.cursorrules`, `.cursor/rules/*`, `.agent/`, `.agents/`, `.pi/`) where `AGENTS.md` is treated as canonical when present and aliases should be kept as compatibility surfaces.
|
||||||
|
- Diagnosing agent-file drift where teams had to prompt iteratively to surface missing files, broken commands, or policy conflicts.
|
||||||
|
- Applying repository-specific documentation overlays, including OpenClaw page-type, docs IA, preservation, and validation rules when present.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Classify task: `build` or `review`; context: `brownfield` or `evergreen`.
|
||||||
|
2. Inventory full documentation scope early (governance + product docs): AGENTS/CONTRIBUTING/aliases plus docs directories, framework sources, and root/module READMEs.
|
||||||
|
3. Detect multilingual scope (README/docs in multiple languages) and define required parity level.
|
||||||
|
4. Read `references/agent-and-contributing.md` for agent instruction and `CONTRIBUTING.md` workflow rules (inventory, canonical/alias mapping, dual-mode balance, deliverable standards, and precedence/conflict handling).
|
||||||
|
5. Read `references/principles.md` for the governing ruleset (Matt Palmer & OpenAI).
|
||||||
|
6. For OpenClaw docs work, read `references/openclaw.md` before the build/review playbook.
|
||||||
|
7. For build tasks, follow `references/build.md`.
|
||||||
|
8. For review tasks, follow `references/review.md` and proactively detect issues without waiting for repeated prompts.
|
||||||
|
9. For complex or high-risk tasks (build or review), it is acceptable to run longer, deeper, and more exhaustive investigations when needed for confidence.
|
||||||
|
10. When available, use sub-agents for bounded parallel discovery/review work, then merge outputs into one coherent final deliverable.
|
||||||
|
11. Use `references/tooling.md` when platform/tooling choices affect recommendations.
|
||||||
|
12. Run a proactive issue sweep for both governance and docs-content surfaces, and fix high-confidence defects in the same pass unless explicitly asked for report-only mode.
|
||||||
|
13. In brownfield mode, prioritize compatibility with current docs IA, tooling, and release state.
|
||||||
|
14. In evergreen mode, prioritize timeless wording, update strategy, and durable structure.
|
||||||
|
15. Return deliverables plus validation notes, parity status, and remaining gaps.
|
||||||
|
|
||||||
|
## Sub-agent orchestration guidance
|
||||||
|
|
||||||
|
Prefer sub-agents when the repo is large or the requested change set is broad; use them by default for repo-wide, multi-framework, or high-conflict work.
|
||||||
|
|
||||||
|
- `inventory-agent` -> `agents/inventory-agent.md` (`fast` / Claude `haiku`): file/config discovery, coverage map, and missing-path checks.
|
||||||
|
- `governance-agent` -> `agents/governance-agent.md` (`thinking` / Claude `sonnet`): AGENTS/CONTRIBUTING/alias precedence, conflicts, and policy drift.
|
||||||
|
- `docs-framework-agent` -> `agents/docs-framework-agent.md` (`thinking` / Claude `sonnet`): framework config, relative path base, and file-path vs URL-path mapping checks.
|
||||||
|
- `synthesis-agent` -> `agents/synthesis-agent.md` (`long` / Claude `opus`): merge sub-agent outputs into one prioritized fix plan and unified precedence model.
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
- Doc type (tutorial, how-to, reference, explanation) and audience.
|
||||||
|
- File scope or diff scope.
|
||||||
|
- Docs framework/tooling constraints (Fern, Mintlify, Sphinx, etc.).
|
||||||
|
- Build/review mode and brownfield/evergreen intent.
|
||||||
|
- Target agent and human compatibility intent.
|
||||||
|
- Docs framework surfaces in scope (for example Fern, Sphinx, Mintlify, Markdown/MDX/MDC/RST/RSC files).
|
||||||
|
- Desired investigation depth/time budget (quick pass vs exhaustive review).
|
||||||
|
- Execution mode (`single-agent` or `sub-agent-assisted` when available).
|
||||||
|
- Remediation mode (`apply-fixes` by default, or `report-only` when requested).
|
||||||
|
- Multilingual scope: source-of-truth language, target locales, and parity expectations.
|
||||||
|
- Repository-specific overlay constraints, if any.
|
||||||
|
|
||||||
|
## Outputs
|
||||||
|
|
||||||
|
- Updated draft or review findings with clear next actions.
|
||||||
|
- Validation notes (what was checked, what remains).
|
||||||
|
- Navigation/maintenance recommendations for long-term quality.
|
||||||
|
- Governance-doc alignment summary when AGENTS/CONTRIBUTING were touched.
|
||||||
|
- Agent instruction-surface map (primary file, alias files, Codex/Claude/Cursor handling plan).
|
||||||
|
- Documentation-surface coverage map (what was reviewed under `/docs`, README hierarchy, and framework-specific source trees).
|
||||||
|
- Autodetected issue list with applied fixes (or explicit report-only findings).
|
||||||
|
- Delegation notes when sub-agents were used (scope delegated and how findings were merged).
|
||||||
|
- Multilingual parity note (in-sync, partial with rationale, or intentionally divergent).
|
||||||
|
- Repository-specific overlay notes when one was used.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
---
|
||||||
|
name: docs-framework-agent
|
||||||
|
description: Thinking-focused docs framework checker for config-relative paths and route/file mapping consistency.
|
||||||
|
model: sonnet
|
||||||
|
tools:
|
||||||
|
- Read
|
||||||
|
- Glob
|
||||||
|
- Grep
|
||||||
|
permissionMode: default
|
||||||
|
maxTurns: 10
|
||||||
|
---
|
||||||
|
|
||||||
|
You are the docs-framework sub-agent for technical documentation.
|
||||||
|
|
||||||
|
Goals:
|
||||||
|
|
||||||
|
- validate framework config-driven docs behavior
|
||||||
|
- prevent path-mapping drift between source files and published routes
|
||||||
|
|
||||||
|
Tasks:
|
||||||
|
|
||||||
|
- detect and read framework config first (Fern/Sphinx/Mintlify/custom)
|
||||||
|
- resolve paths relative to the declaring file/config
|
||||||
|
- validate both maps:
|
||||||
|
- config -> file exists
|
||||||
|
- config/nav/routing -> URL path is valid and consistent
|
||||||
|
|
||||||
|
Return:
|
||||||
|
|
||||||
|
- config files reviewed
|
||||||
|
- path assumptions made
|
||||||
|
- mismatches (`missing file`, `stale route`, `wrong base path`)
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
---
|
||||||
|
name: governance-agent
|
||||||
|
description: Thinking-focused governance reviewer for AGENTS/CONTRIBUTING/alias precedence, conflict detection, and policy drift analysis.
|
||||||
|
model: sonnet
|
||||||
|
tools:
|
||||||
|
- Read
|
||||||
|
- Glob
|
||||||
|
- Grep
|
||||||
|
permissionMode: default
|
||||||
|
maxTurns: 10
|
||||||
|
---
|
||||||
|
|
||||||
|
You are the governance sub-agent for technical documentation.
|
||||||
|
|
||||||
|
Goals:
|
||||||
|
|
||||||
|
- validate AGENTS/CONTRIBUTING/alias alignment and precedence
|
||||||
|
- identify policy drift and conflicting instructions
|
||||||
|
|
||||||
|
Tasks:
|
||||||
|
|
||||||
|
- determine canonical instruction source and alias compatibility mapping
|
||||||
|
- detect conflicts across nested scope files and tool-specific rule consumers
|
||||||
|
- validate command examples against stated governance expectations
|
||||||
|
|
||||||
|
Return:
|
||||||
|
|
||||||
|
- precedence model
|
||||||
|
- conflict list with severity
|
||||||
|
- recommended low-risk remediations
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
---
|
||||||
|
name: inventory-agent
|
||||||
|
description: Fast repo-surface discovery for technical documentation audits. Use for coverage mapping and missing-path detection before deeper review.
|
||||||
|
model: haiku
|
||||||
|
tools:
|
||||||
|
- Read
|
||||||
|
- Glob
|
||||||
|
- Grep
|
||||||
|
- LS
|
||||||
|
permissionMode: default
|
||||||
|
maxTurns: 6
|
||||||
|
---
|
||||||
|
|
||||||
|
You are the inventory sub-agent for technical documentation.
|
||||||
|
|
||||||
|
Goals:
|
||||||
|
|
||||||
|
- enumerate governance and docs-content surfaces in scope
|
||||||
|
- detect missing files, broken references, and obvious command/path failures
|
||||||
|
|
||||||
|
Tasks:
|
||||||
|
|
||||||
|
- map `AGENTS.md`/`CONTRIBUTING.md`/aliases and docs surfaces (`docs/**`, README hierarchy, `.md/.mdx/.mdc/.rst/.rsc`)
|
||||||
|
- list framework config files discovered (Fern/Sphinx/Mintlify or equivalent)
|
||||||
|
- report hard failures only, with exact file paths
|
||||||
|
|
||||||
|
Return:
|
||||||
|
|
||||||
|
- coverage map
|
||||||
|
- missing/broken path list
|
||||||
|
- unresolved blockers
|
||||||
10
.agents/skills/technical-documentation/agents/openai.yaml
Normal file
10
.agents/skills/technical-documentation/agents/openai.yaml
Normal file
@@ -0,0 +1,10 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "Technical Documentation"
|
||||||
|
short_description: "Build and review technical documentation for brownfield and evergreen systems."
|
||||||
|
icon_small: "./assets/icon.jpg"
|
||||||
|
icon_large: "./assets/icon.jpg"
|
||||||
|
brand_color: "#111827"
|
||||||
|
default_prompt: "Build or review technical documentation with a clear, maintainable, and production-ready workflow."
|
||||||
|
|
||||||
|
policy:
|
||||||
|
allow_implicit_invocation: true
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
---
|
||||||
|
name: synthesis-agent
|
||||||
|
description: Long-context synthesis agent that merges sub-agent outputs into one prioritized and deduplicated documentation action plan.
|
||||||
|
model: opus
|
||||||
|
tools:
|
||||||
|
- Read
|
||||||
|
permissionMode: default
|
||||||
|
maxTurns: 12
|
||||||
|
---
|
||||||
|
|
||||||
|
You are the synthesis sub-agent for technical documentation.
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- merge sub-agent outputs into one coherent, non-duplicated action plan
|
||||||
|
|
||||||
|
Tasks:
|
||||||
|
|
||||||
|
- prioritize blockers first, then non-blocking improvements
|
||||||
|
- normalize to one precedence model for governance decisions
|
||||||
|
- remove duplicated recommendations and contradictory fixes
|
||||||
|
- keep final output concise and execution-ready
|
||||||
|
|
||||||
|
Return:
|
||||||
|
|
||||||
|
- prioritized fix plan
|
||||||
|
- validation summary (done vs pending)
|
||||||
|
- explicit remaining gaps/blockers
|
||||||
BIN
.agents/skills/technical-documentation/assets/icon.jpg
Normal file
BIN
.agents/skills/technical-documentation/assets/icon.jpg
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 37 KiB |
@@ -0,0 +1,145 @@
|
|||||||
|
# AGENT and CONTRIBUTING Principles
|
||||||
|
|
||||||
|
This reference consolidates the core rules for agent-policy and contributor-governance docs.
|
||||||
|
|
||||||
|
You must:
|
||||||
|
|
||||||
|
1. Discover repo-level and nested instruction files with:
|
||||||
|
`rg --files -g 'AGENTS.md' -g 'CONTRIBUTING.md' -g 'CLAUDE.md' -g 'AGENT.md' -g '.cursor/rules/*' -g '.cursorrules' -g '.agent/**' -g '.agents/**' -g '.pi/**' -g 'AGENTS.*.md'`
|
||||||
|
2. Read the root and nearest-scope `AGENTS.md`/`CONTRIBUTING.md` pair before editing.
|
||||||
|
3. If alias files exist, normalize to one canonical source (`AGENTS.md` preferred when present; otherwise nearest alias), plus compatibility pointers or explicit symlink notes.
|
||||||
|
4. Document conflicting instructions and precedence decisions.
|
||||||
|
|
||||||
|
## GitHub + AGENTS baseline
|
||||||
|
|
||||||
|
Source: https://docs.github.com/en/communities/setting-up-your-project-for-healthy-contributions/setting-guidelines-for-repository-contributors
|
||||||
|
Source: https://agents.md/
|
||||||
|
Source: https://github.blog/ai-and-ml/github-copilot/how-to-write-a-great-agents-md-lessons-from-over-2500-repositories/
|
||||||
|
Source: https://cobusgreyling.substack.com/p/what-is-agentsmd
|
||||||
|
Source: https://www.infoq.com/news/2025/08/agents-md/
|
||||||
|
|
||||||
|
Use these as default operating principles:
|
||||||
|
|
||||||
|
1. Keep `CONTRIBUTING.md` discoverable and actionable (`.github`, root, or `docs`).
|
||||||
|
2. Keep agent instructions concrete: real commands, real paths, clear boundaries.
|
||||||
|
3. Use explicit behavior boundaries for agents: `Always`, `Ask first`, `Never`.
|
||||||
|
4. Keep contributor and agent rules aligned with actual repository workflows.
|
||||||
|
5. Ensure clear guidance is provided to agents on if, when and how to raise issues and pull requests.
|
||||||
|
|
||||||
|
## Canonical and alias policy
|
||||||
|
|
||||||
|
Source: https://agents.md/
|
||||||
|
Source: https://github.blog/ai-and-ml/github-copilot/how-to-write-a-great-agents-md-lessons-from-over-2500-repositories/
|
||||||
|
|
||||||
|
1. Treat `AGENTS.md` as canonical when present.
|
||||||
|
2. If `AGENTS.md` is absent, treat the nearest alias file as canonical.
|
||||||
|
3. Keep compatibility surfaces explicit: `AGENTS.md`, `AGENT.md`, `.cursorrules`, `.cursor/rules/*`, `.agent/`, `.agents/`, `.pi/`.
|
||||||
|
4. If aliases are used, document how they map back to canonical policy (or symlink when supported).
|
||||||
|
5. When repos use `.agents/` as canonical rule storage, keep `.cursor` as a compatibility symlink to `.agents` for Cursor rule auto-loading.
|
||||||
|
6. Keep policy DRY: store one shared policy core and expose it via aliases/symlinks instead of duplicating rule text.
|
||||||
|
|
||||||
|
## Context-awareness by agent platform
|
||||||
|
|
||||||
|
Source: https://github.com/vercel-labs/agent-skills/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/openai/codex/blob/main/AGENTS.md
|
||||||
|
|
||||||
|
1. For Cursor and Claude-style glob consumers, keep rule files narrow and bounded.
|
||||||
|
2. Avoid over-referencing large path sets that inflate context for glob-based agents.
|
||||||
|
3. For Codex-style workflows, prefer explicit file references and deterministic commands.
|
||||||
|
4. Keep long runbooks outside top-level policy files; link to scoped docs.
|
||||||
|
5. Ensure all agents have a happy path regardless so ensuring everything works across Codex, Claude and other coding agents.
|
||||||
|
|
||||||
|
## Symlink and compatibility operations
|
||||||
|
|
||||||
|
1. Preferred layout for multi-agent compatibility:
|
||||||
|
- canonical rule directory: `.agents/`
|
||||||
|
- Cursor compatibility path: `.cursor -> .agents` symlink
|
||||||
|
- canonical policy doc: `AGENTS.md` pointing to `.agents` paths where relevant
|
||||||
|
2. Validate symlink state before finalizing changes:
|
||||||
|
- if `.agents/` exists and `.cursor` is missing, create `.cursor` symlink to `.agents`
|
||||||
|
- if `.cursor` is a symlink to another target, fix target or document why it must differ
|
||||||
|
- if `.cursor` is a real directory/file, treat as migration conflict and ask before replacement
|
||||||
|
3. Validate rule payload through the canonical directory:
|
||||||
|
- rules: `.agents/rules/*.mdc` with valid frontmatter (`description`, `globs`, `alwaysApply` as needed)
|
||||||
|
- commands: `.agents/commands/*.md` when command routing is used
|
||||||
|
- MCP config: `.agents/mcp.json` when MCP is in scope
|
||||||
|
4. Keep Codex behavior explicit:
|
||||||
|
- `AGENTS.md` is primary for Codex repository instructions
|
||||||
|
- `.cursor` compatibility is for Cursor auto-loading and does not replace canonical AGENTS policy
|
||||||
|
5. Record applied symlink fixes and unresolved compatibility gaps in validation notes.
|
||||||
|
|
||||||
|
## Dual-mode and deliverable standards
|
||||||
|
|
||||||
|
Source: https://github.blog/ai-and-ml/github-copilot/how-to-write-a-great-agents-md-lessons-from-over-2500-repositories/
|
||||||
|
Source: https://agents.md/
|
||||||
|
Source: https://github.com/openai/codex/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/vercel-labs/agent-skills/blob/main/AGENTS.md
|
||||||
|
|
||||||
|
1. Author one shared policy core (same commands, boundaries, and precedence) for all agents.
|
||||||
|
2. For Cursor/Claude-style agents, expose that core through glob-driven and bounded files (small `AGENTS.md`/rule surface).
|
||||||
|
3. For Codex, expose that same core through explicit file references with precise scope.
|
||||||
|
4. Where styles diverge, prefer the smallest common structure that satisfies both and avoid duplicating policy text.
|
||||||
|
5. Treat AGENTS/CONTRIBUTING as first-class deliverables when in scope.
|
||||||
|
6. Preserve required structure, constraints, and examples from existing files.
|
||||||
|
7. Align wording and commands with active repository instructions.
|
||||||
|
|
||||||
|
## Proactive issue discovery and remediation
|
||||||
|
|
||||||
|
Source: https://github.blog/ai-and-ml/github-copilot/how-to-write-a-great-agents-md-lessons-from-over-2500-repositories/
|
||||||
|
Source: https://github.com/openai/codex/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/vercel-labs/agent-skills/blob/main/AGENTS.md
|
||||||
|
|
||||||
|
1. Run a conflict matrix review across AGENTS/aliases/CONTRIBUTING and related command/rule docs before finalizing.
|
||||||
|
2. Treat the following as high-priority defects: missing referenced files, non-existent setup commands, command scope mismatches, and branch/commit policy conflicts.
|
||||||
|
3. Do not stop at caveat-only notes when a low-risk fix is clear; apply the fix in the same pass.
|
||||||
|
4. If a canonical entry file is missing (for example a directory `README.md` that docs depend on), create a minimal actionable file and update references.
|
||||||
|
5. Long-running investigations are acceptable when needed to uncover cross-file drift, especially in agent-instruction ecosystems.
|
||||||
|
|
||||||
|
## Discovery
|
||||||
|
|
||||||
|
1. Agents prefer simple terminal commands so having a well defined `make *` or `npm run *` is ideal
|
||||||
|
2. Agents can discover terminal commands through shell completion so providing shell completion helps
|
||||||
|
|
||||||
|
## CONTRIBUTING size and scope control
|
||||||
|
|
||||||
|
Source: https://contributing.md/how-to-build-contributing-md/
|
||||||
|
Source: https://blog.codacy.com/best-practices-to-manage-an-open-source-project
|
||||||
|
Source: https://mozillascience.github.io/working-open-workshop/contributing/
|
||||||
|
Source: https://github.com/openclaw/openclaw/blob/main/CONTRIBUTING.md
|
||||||
|
|
||||||
|
1. Keep root `CONTRIBUTING.md` focused on setup, issue flow, PR flow, testing, and review gates.
|
||||||
|
2. Use issue/PR template links instead of embedding every process detail inline.
|
||||||
|
3. When the file grows too large, split by domain and link from root.
|
||||||
|
4. Move any large content into docs if avalible (for example Mintlify/Fern/Sphinx workflows) to avoid large contributor guide.
|
||||||
|
5. Optimize for agent/machine readability as well as humans.
|
||||||
|
|
||||||
|
## Example repos to emulate
|
||||||
|
|
||||||
|
Source: https://github.com/openclaw/openclaw/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/openclaw/openclaw/blob/main/CONTRIBUTING.md
|
||||||
|
Source: https://github.com/openclaw/openclaw/blob/main/VISION.md
|
||||||
|
Source: https://github.com/openai/codex/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/processing/p5.js/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/vercel-labs/agent-skills/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/agentsmd/agents.md/blob/main/AGENTS.md
|
||||||
|
Source: https://github.com/rails/rails/blob/main/CONTRIBUTING.md
|
||||||
|
Source: https://github.com/kubernetes/kubernetes/blob/master/CONTRIBUTING.md
|
||||||
|
Source: https://github.com/atom/atom/blob/master/CONTRIBUTING.md
|
||||||
|
Source: https://github.com/github/docs/blob/main/CONTRIBUTING.md
|
||||||
|
Source: https://github.com/facebook/react/blob/main/CONTRIBUTING.md
|
||||||
|
|
||||||
|
1. OpenClaw: strong real-world alias policy and AGENTS/CONTRIBUTING/VISION cohesion.
|
||||||
|
2. OpenAI Codex: strict command discipline and explicit scope control.
|
||||||
|
3. p5.js: explicit AI-policy guardrails in agent instructions.
|
||||||
|
4. Vercel + agentsmd spec: compact, context-efficient AGENTS patterns.
|
||||||
|
5. Rails/Kubernetes/Atom/GitHub Docs/React: contributor guidance patterns at different project scales.
|
||||||
|
|
||||||
|
## Practical merge policy
|
||||||
|
|
||||||
|
When these rules conflict:
|
||||||
|
|
||||||
|
1. Preserve contributor and reader task success first.
|
||||||
|
2. Preserve instruction clarity and unambiguous boundaries second.
|
||||||
|
3. Preserve long-term maintainability and context-efficiency third.
|
||||||
|
4. Add extra agent optimization only if it does not reduce human clarity or there is explict need.
|
||||||
|
5. Use your judgement as the expert.
|
||||||
116
.agents/skills/technical-documentation/references/build.md
Normal file
116
.agents/skills/technical-documentation/references/build.md
Normal file
@@ -0,0 +1,116 @@
|
|||||||
|
# Build Docs Playbook
|
||||||
|
|
||||||
|
Read `principles.md` first, then follow this execution flow.
|
||||||
|
|
||||||
|
## 1. Detect and align agent instruction and governance instructions
|
||||||
|
|
||||||
|
- Use `references/agent-and-contributing.md` as the source of truth for inventory, canonical/alias mapping, and precedence/conflict handling.
|
||||||
|
- Apply the symlink compatibility policy when in scope (`.agents` canonical directory with `.cursor` compatibility symlink when required by tooling).
|
||||||
|
- Long-running and extensive build investigations are acceptable when needed to resolve ambiguous or conflicting documentation sources.
|
||||||
|
- When available, use sub-agents for bounded parallel inventory/cross-check tasks and merge results into one canonical decision set.
|
||||||
|
- Capture required constraints before writing:
|
||||||
|
- nested-agent rules, command/test requirements, PR workflow, and style checks.
|
||||||
|
- Use the same command and validation expectations in proposed snippets and examples.
|
||||||
|
|
||||||
|
## 2. Inventory product documentation surfaces (not governance only)
|
||||||
|
|
||||||
|
- For repo-wide builds, include docs content surfaces in addition to AGENTS/CONTRIBUTING.
|
||||||
|
- Inventory docs files and frameworks in scope (examples): `README*.md`, `docs/**`, `**/*.md`, `**/*.mdx`, `**/*.mdc`, `**/*.rst`, `**/*.rsc`, Fern/Mintlify config, Sphinx `conf.py`.
|
||||||
|
- Build a coverage map before drafting so governance and product docs are both represented.
|
||||||
|
- If scope is ambiguous, default to broader docs discovery first, then narrow intentionally.
|
||||||
|
|
||||||
|
## 3. Framework config and path mapping rules
|
||||||
|
|
||||||
|
- Detect framework/config first (for example Fern config, Sphinx `conf.py`, Mintlify config, or equivalent).
|
||||||
|
- Resolve every referenced path relative to the file/config that declares it, not assumed repo root.
|
||||||
|
- Treat filesystem paths and published URL routes as separate mappings; do not infer one from the other without config evidence.
|
||||||
|
- Validate both layers:
|
||||||
|
- config -> file exists on disk
|
||||||
|
- config/nav/routing -> URL path is consistent and reachable
|
||||||
|
- Record path-mapping assumptions and mismatches in handoff (`missing file`, `stale route`, `wrong base path`).
|
||||||
|
|
||||||
|
## 4. Define intent and success
|
||||||
|
|
||||||
|
- Audience, prerequisites, and job-to-be-done.
|
||||||
|
- Expected reader outcome immediately after completion.
|
||||||
|
- Doc type: tutorial, how-to, reference, explanation.
|
||||||
|
- Success criteria: what must be true after publish.
|
||||||
|
|
||||||
|
## 5. Build structure before prose
|
||||||
|
|
||||||
|
- Follow the funnel: what/why, quickstart, next steps.
|
||||||
|
- Keep headings informative and scannable.
|
||||||
|
- Open each section with the takeaway sentence.
|
||||||
|
- Add decision points with concrete branch guidance.
|
||||||
|
- For OpenClaw docs work, choose a page type from `references/openclaw.md` before drafting.
|
||||||
|
- Keep task-critical OpenClaw configuration inline; link exhaustive defaults, enums, schemas, generated references, and rare debugging workflows.
|
||||||
|
|
||||||
|
## 6. Build AGENTS.md and CONTRIBUTING.md intentionally
|
||||||
|
|
||||||
|
- Keep AGENTS.md structure consistent with `agents.md` ecosystem patterns:
|
||||||
|
- include YAML frontmatter when present in repo style (`name`, `description`).
|
||||||
|
- state persona scope and explicit instruction boundaries: `Always`, `Ask first`, `Never`.
|
||||||
|
- include concrete commands and representative code examples.
|
||||||
|
- For CONTRIBUTING.md, prioritize issue triage flow, PR expectations, setup/test commands, and review gates.
|
||||||
|
- Add `Code of Conduct`, `Testing`, `Local checks`, and `PR expectations` sections when missing but required by the repo.
|
||||||
|
- If CONTRIBUTING.md is becoming too large, split by scope into linked docs (for example, framework/tool-specific setup and release workflows) and keep the root file as a concise entry point.
|
||||||
|
- Keep cross-file consistency: links from CONTRIBUTING.md to AGENTS.md (and vice versa) should be accurate and non-circular.
|
||||||
|
- If multiple AGENTS.md files exist, document the directory-level scope and avoid conflicting advice.
|
||||||
|
- If a required canonical entry file is missing (for example referenced `README.md` under a major directory), create the file in the same pass instead of adding a caveat-only note.
|
||||||
|
- For new entry files, keep them minimal and actionable: purpose, prerequisites, concrete run commands, and pointers to deeper docs.
|
||||||
|
|
||||||
|
## 7. Keep agent context tight
|
||||||
|
|
||||||
|
- Author once, expose twice:
|
||||||
|
- keep one shared policy core and avoid duplicating guidance in separate agent-specific files.
|
||||||
|
- publish that core through bounded glob-friendly files for Cursor/Claude plus explicit path references for Codex.
|
||||||
|
- For Cursor and Claude-style agents, avoid broad references. Use minimal globbing and narrow rule files that each serve one concern (for example, repo-wide setup, test rules, security checks).
|
||||||
|
- Keep AGENTS and alias files short-to-medium; move detailed runbooks to linked docs.
|
||||||
|
- For Codex, prefer explicit file references and concrete paths for exact reuse.
|
||||||
|
- Avoid adding unrelated historical or process details to avoid token/context drift during future tool reads.
|
||||||
|
|
||||||
|
## 8. Brownfield build mode
|
||||||
|
|
||||||
|
- Match existing terminology, navigation, and component patterns.
|
||||||
|
- Preserve existing IA unless there is a documented migration plan.
|
||||||
|
- For rewrites, include a migration note from old to new paths.
|
||||||
|
- Prefer smallest safe change set that improves utility.
|
||||||
|
|
||||||
|
## 9. Evergreen build mode
|
||||||
|
|
||||||
|
- Prefer stable concepts over release-tied narrative.
|
||||||
|
- Isolate volatile details under clearly marked version sections.
|
||||||
|
- Include maintenance signals: owners, refresh triggers, stale criteria.
|
||||||
|
- Include lifecycle notes: deprecation and replacement paths.
|
||||||
|
|
||||||
|
## 10. Writing constraints
|
||||||
|
|
||||||
|
- Use precise language and short, imperative instructions.
|
||||||
|
- Keep code examples copy-ready and self-contained.
|
||||||
|
- Include common failure modes and safe defaults.
|
||||||
|
- Avoid placeholder guidance that cannot be executed.
|
||||||
|
|
||||||
|
## 11. Agent and automation readiness
|
||||||
|
|
||||||
|
- Keep key facts in text (not image-only).
|
||||||
|
- Prefer structured lists/tables when choices matter.
|
||||||
|
- Add links and anchors that allow deterministic navigation.
|
||||||
|
- Document what can be checked automatically in CI.
|
||||||
|
|
||||||
|
## 12. Build validation
|
||||||
|
|
||||||
|
- Validate commands and snippets where possible.
|
||||||
|
- Verify links and references in changed sections.
|
||||||
|
- Run a reference existence sweep for every path/command you introduced.
|
||||||
|
- Verify docs-framework consistency when in scope (for example Sphinx/Fern config and referenced doc paths).
|
||||||
|
- For OpenClaw docs work, apply the validation checklist in `references/openclaw.md`.
|
||||||
|
|
||||||
|
## 13. Multilingual parity mode (when applicable)
|
||||||
|
|
||||||
|
- Pick one source-of-truth language for technical accuracy and release timing.
|
||||||
|
- Define parity target: full parity, staged parity, or intentional divergence per section.
|
||||||
|
- Keep structure aligned across locales (headings, anchors, section order) when possible.
|
||||||
|
- Preserve command/code correctness first; localize explanatory text second.
|
||||||
|
- If parity is not feasible, add a visible note with missing scope and expected sync window.
|
||||||
|
- Run a locale parity check for changed sections (added/removed steps, warnings, prerequisites).
|
||||||
|
- Record unresolved checks explicitly in handoff.
|
||||||
128
.agents/skills/technical-documentation/references/openclaw.md
Normal file
128
.agents/skills/technical-documentation/references/openclaw.md
Normal file
@@ -0,0 +1,128 @@
|
|||||||
|
# OpenClaw Documentation Overlay
|
||||||
|
|
||||||
|
Use this reference only for OpenClaw docs work. It layers OpenClaw-specific page
|
||||||
|
types, navigation, preservation, and validation rules on top of the general
|
||||||
|
technical-documentation skill.
|
||||||
|
|
||||||
|
## Reader Model
|
||||||
|
|
||||||
|
- Lead with the task the reader is trying to complete.
|
||||||
|
- Give one recommended path before alternatives.
|
||||||
|
- Keep main docs focused on the common path; move dense contracts and rare
|
||||||
|
debugging detail to linked reference or troubleshooting pages.
|
||||||
|
- Explain production risks exactly where the reader can make the mistake.
|
||||||
|
- Link concepts, guides, references, CLI pages, SDK docs, testing, and
|
||||||
|
troubleshooting so readers can continue without rereading.
|
||||||
|
|
||||||
|
## Page Types
|
||||||
|
|
||||||
|
Choose the page type before writing or reviewing:
|
||||||
|
|
||||||
|
- Overview: route readers to the right product area, integration path, or guide.
|
||||||
|
- Quickstart: get a new user to a working result with the fewest safe steps.
|
||||||
|
- Topic page: explain a major OpenClaw entity or surface end to end.
|
||||||
|
- Guide: walk through one workflow from prerequisites to production readiness.
|
||||||
|
- API/SDK/CLI reference: define every object, method, command, option, response,
|
||||||
|
error, enum, default, and version rule in scope.
|
||||||
|
- Testing guide: show sandbox setup, fixtures, simulated failures, and live-mode
|
||||||
|
differences.
|
||||||
|
- Troubleshooting guide: map observable symptoms to checks, causes, and fixes.
|
||||||
|
- Governance file: keep agent/contributor policy concrete, scoped, and aligned
|
||||||
|
with current OpenClaw repo behavior.
|
||||||
|
|
||||||
|
## Topic Pages
|
||||||
|
|
||||||
|
Use this shape for major-entity pages:
|
||||||
|
|
||||||
|
1. Title naming the entity or surface.
|
||||||
|
2. Unheaded opening that says what it is, what it owns, and what it does not own.
|
||||||
|
3. Requirements, only when setup needs accounts, versions, permissions, plugins,
|
||||||
|
operating systems, or credentials.
|
||||||
|
4. Quickstart with the recommended path and smallest reliable verification.
|
||||||
|
5. Configuration with task-critical options inline and exhaustive details linked
|
||||||
|
to reference docs.
|
||||||
|
6. Major subtopics organized by reader intent, not under a generic "Subtopics"
|
||||||
|
heading.
|
||||||
|
7. Troubleshooting with observable failures and concrete checks.
|
||||||
|
8. Related links to guides, references, commands, concepts, and adjacent topics.
|
||||||
|
|
||||||
|
## Guides
|
||||||
|
|
||||||
|
Use this shape for workflow pages:
|
||||||
|
|
||||||
|
1. Title naming the outcome, not the implementation detail.
|
||||||
|
2. Opening that states what the reader can accomplish.
|
||||||
|
3. Before you begin: accounts, keys, permissions, versions, tools, and
|
||||||
|
assumptions.
|
||||||
|
4. Choose a path, only when the reader must decide.
|
||||||
|
5. Steps with verb-led headings, commands, expected output, and checks.
|
||||||
|
6. Test with the smallest reliable proof that the workflow works.
|
||||||
|
7. Production readiness: security, retries, limits, observability, migrations,
|
||||||
|
and cleanup.
|
||||||
|
8. Troubleshooting near the workflow that causes the failures.
|
||||||
|
9. See also links to concepts, references, SDK docs, and adjacent guides.
|
||||||
|
|
||||||
|
## Docs IA And Navigation
|
||||||
|
|
||||||
|
- Read `docs/docs.json` before navigation changes.
|
||||||
|
- Keep topic pages and common workflows on the main reader path.
|
||||||
|
- Put exhaustive contracts, generated references, maintainer-only detail, and
|
||||||
|
support material under `Reference` or another clearly scoped support page.
|
||||||
|
- Keep generated `plugins/reference/*` children and redirect-only pages out of
|
||||||
|
visible navigation unless explicitly required.
|
||||||
|
- For moved pages, include a keep/drop/move/destination matrix in the handoff.
|
||||||
|
- Add "Read when" hints for docs-list routing when creating or changing pages
|
||||||
|
that participate in the docs index.
|
||||||
|
|
||||||
|
## Source-Backed Content
|
||||||
|
|
||||||
|
- CLI docs must match current flags, output, errors, and examples.
|
||||||
|
- API/SDK docs must include fields, defaults, enum values, constraints, nullable
|
||||||
|
behavior, lifecycle states, errors, and recovery guidance.
|
||||||
|
- Config docs must align exported types, schema/help output, metadata, baselines,
|
||||||
|
and current docs.
|
||||||
|
- Dependency-backed behavior must be verified from upstream docs, source, or
|
||||||
|
types before documenting defaults, timing, errors, or API behavior.
|
||||||
|
- Separate current behavior, shipped behavior, planned behavior, and maintainer
|
||||||
|
intent.
|
||||||
|
|
||||||
|
## Examples
|
||||||
|
|
||||||
|
- Prefer complete copy-pasteable commands and snippets.
|
||||||
|
- Use realistic variable names and values.
|
||||||
|
- Mark placeholders with angle-bracket names such as `<API_KEY>`.
|
||||||
|
- Show expected success output when it helps verification.
|
||||||
|
- Keep one conceptual unit per code block and use language-specific fences.
|
||||||
|
- Avoid examples that hide setup, auth, error handling, or cleanup.
|
||||||
|
- Never expose real secrets, live config, phone numbers, private videos, or
|
||||||
|
credentials.
|
||||||
|
|
||||||
|
## Preservation Reviews
|
||||||
|
|
||||||
|
For rewrites or splits:
|
||||||
|
|
||||||
|
- Identify source units before rewriting: headings, paragraphs, tables, examples,
|
||||||
|
CLI/API contracts, warnings, and troubleshooting facts.
|
||||||
|
- Map each retained unit to a destination page or section.
|
||||||
|
- Do not treat a broad "covered" row as proof for dense source material; use
|
||||||
|
line- or claim-level evidence when the source unit is dense.
|
||||||
|
- For dropped content, state whether it is obsolete, duplicated elsewhere,
|
||||||
|
unsupported, or moved to a reference/support page.
|
||||||
|
- When a docs-audit artifact is used, verify it is mapped audit data with
|
||||||
|
non-empty `mappings[]`, not only inventory or reindexed JSON.
|
||||||
|
|
||||||
|
## Validation
|
||||||
|
|
||||||
|
Choose the narrowest proof that covers the touched surface:
|
||||||
|
|
||||||
|
- `pnpm docs:list`
|
||||||
|
- `pnpm docs:check-mdx`
|
||||||
|
- `pnpm docs:check-links`
|
||||||
|
- `pnpm docs:check-i18n-glossary`
|
||||||
|
- `pnpm format:docs:check` or `pnpm lint:docs`
|
||||||
|
- `git diff --check`
|
||||||
|
- generated-doc or inventory checks when generated references, plugin catalogs,
|
||||||
|
labeler, or docs scripts changed
|
||||||
|
- behavior tests or command probes when docs claim runtime behavior
|
||||||
|
|
||||||
|
If proof is blocked, say exactly which command was not run and why.
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
# Documentation Principles
|
||||||
|
|
||||||
|
This reference consolidates the core rules used by this skill.
|
||||||
|
|
||||||
|
## Matt Palmer: 8 rules for better docs
|
||||||
|
|
||||||
|
Source: https://mattpalmer.io/posts/2025/10/8-rules-for-better-docs/
|
||||||
|
|
||||||
|
Use these as default operating principles:
|
||||||
|
|
||||||
|
1. Write for humans, optimize for agents.
|
||||||
|
2. Start with a funnel: what/why, quickstart, next steps.
|
||||||
|
3. Use Diataxis to scaffold content.
|
||||||
|
4. Write with AI, but structure for agents.
|
||||||
|
5. Offload routine docs operations to background agents.
|
||||||
|
6. Automate quality with CI.
|
||||||
|
7. Automate scaffolding and repetitive workflow tasks.
|
||||||
|
8. Make contribution easy and visible.
|
||||||
|
|
||||||
|
## OpenAI cookbook: what makes documentation good
|
||||||
|
|
||||||
|
Source: https://cookbook.openai.com/articles/what_makes_documentation_good
|
||||||
|
|
||||||
|
Key quality constraints:
|
||||||
|
|
||||||
|
- Prefer specific and accurate terminology over niche jargon.
|
||||||
|
- Keep examples self-contained and minimize dependencies.
|
||||||
|
- Prioritize high-value topics over edge-case depth.
|
||||||
|
- Do not teach unsafe patterns (for example, exposed secrets).
|
||||||
|
- Open with context that helps readers orient quickly.
|
||||||
|
- Apply empathy and override rigid rules when it clearly improves outcomes.
|
||||||
|
|
||||||
|
## Practical merge policy
|
||||||
|
|
||||||
|
When these rules conflict:
|
||||||
|
|
||||||
|
1. Preserve reader task success first.
|
||||||
|
2. Preserve structural clarity second.
|
||||||
|
3. Preserve long-term maintainability third.
|
||||||
|
4. Add agent optimization only if it does not reduce human clarity.
|
||||||
|
|
||||||
|
For agent-instructions and contributor-governance specifics (AGENTS/aliases/CONTRIBUTING), use `references/agent-and-contributing.md` as the detailed additional source of truth.
|
||||||
|
|
||||||
|
When the target repo or request is OpenClaw-specific, layer `references/openclaw.md` on top of these general rules. Otherwise ignore that repo-specific overlay.
|
||||||
|
|
||||||
|
## Execution policy for this skill
|
||||||
|
|
||||||
|
- Long-running and extensive investigations are allowed for both build and review work when needed to resolve ambiguity or cross-file drift.
|
||||||
|
- Use sub-agents when available for bounded parallel discovery, verification, or cross-source comparison.
|
||||||
|
- Keep one merged outcome: sub-agent outputs must be normalized into a single consistent recommendation/fix set.
|
||||||
|
|
||||||
|
## Multilingual parity rule
|
||||||
|
|
||||||
|
When docs exist in multiple languages, target cross-locale parity for task-critical content (steps, warnings, prerequisites, and limits). If full parity is not possible, publish explicit parity status and sync intent.
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user