Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11), free to diverge. Tree copied sans upstream .git; upstream remote added for future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19. Preserves docs/ARCHITECTURE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
216 lines
5.9 KiB
Plaintext
216 lines
5.9 KiB
Plaintext
# Calibrator
|
|
# Validates that lightweight evaluations are reliable proxies for deep evaluations
|
|
#
|
|
# Usage:
|
|
# prose run @openprose/lib/calibrator
|
|
#
|
|
# Purpose:
|
|
# Run both light and deep inspections on the same runs, compare results,
|
|
# and build confidence (or identify gaps) in light evaluations.
|
|
#
|
|
# Inputs:
|
|
# run_paths: Paths to runs to calibrate on (comma-separated or glob)
|
|
# sample_size: How many runs to sample (if more available)
|
|
#
|
|
# Outputs:
|
|
# - Agreement rate between light and deep
|
|
# - Cases where they disagree
|
|
# - Recommendations for improving light evaluation
|
|
|
|
input run_paths: "Paths to runs (comma-separated, or 'recent' for latest)"
|
|
input sample_size: "Max runs to analyze (default: 10)"
|
|
|
|
# ============================================================
|
|
# Agents
|
|
# ============================================================
|
|
|
|
agent sampler:
|
|
model: sonnet
|
|
prompt: """
|
|
You select runs for calibration analysis.
|
|
Prefer diverse runs: different programs, outcomes, sizes.
|
|
"""
|
|
|
|
agent comparator:
|
|
model: opus
|
|
prompt: """
|
|
You compare light vs deep evaluation results with nuance.
|
|
Identify agreement, disagreement, and edge cases.
|
|
"""
|
|
|
|
agent statistician:
|
|
model: sonnet
|
|
prompt: """
|
|
You compute statistics and confidence intervals.
|
|
"""
|
|
|
|
agent advisor:
|
|
model: opus
|
|
prompt: """
|
|
You recommend improvements to evaluation criteria.
|
|
"""
|
|
|
|
# ============================================================
|
|
# Phase 1: Select Runs
|
|
# ============================================================
|
|
|
|
let selected_runs = session: sampler
|
|
prompt: """
|
|
Select runs for calibration.
|
|
|
|
Input: {run_paths}
|
|
Sample size: {sample_size}
|
|
|
|
If run_paths is "recent", find recent runs in .prose/runs/
|
|
If specific paths, use those.
|
|
|
|
Select a diverse sample:
|
|
- Different programs if possible
|
|
- Mix of successful and partial/failed if available
|
|
- Different sizes (small vs large runs)
|
|
|
|
Return list of run paths.
|
|
"""
|
|
|
|
# ============================================================
|
|
# Phase 2: Run Both Inspection Depths
|
|
# ============================================================
|
|
|
|
let calibration_data = selected_runs | map:
|
|
# Run light and deep sequentially on each (can't parallel same run)
|
|
let light = session "Light inspection"
|
|
prompt: """
|
|
Run a LIGHT inspection on: {item}
|
|
|
|
Evaluate quickly:
|
|
- completion: did it finish cleanly?
|
|
- binding_integrity: do expected outputs exist?
|
|
- output_substance: do outputs have real content?
|
|
- goal_alignment: does output match program purpose?
|
|
|
|
Score each 1-10, give verdicts (pass/partial/fail).
|
|
Return JSON.
|
|
"""
|
|
|
|
let deep = session "Deep inspection"
|
|
prompt: """
|
|
Run a DEEP inspection on: {item}
|
|
|
|
Evaluate thoroughly:
|
|
- Read the full program source
|
|
- Trace execution step by step
|
|
- Check each binding's content
|
|
- Evaluate output quality in detail
|
|
- Assess fidelity (did VM follow program correctly?)
|
|
- Assess efficiency (reasonable steps for the job?)
|
|
|
|
Score each dimension 1-10, give verdicts.
|
|
Return JSON.
|
|
"""
|
|
context: light # Deep can see light's assessment
|
|
|
|
session "Package results"
|
|
prompt: """
|
|
Package the light and deep inspection results.
|
|
|
|
Run: {item}
|
|
Light: {light}
|
|
Deep: {deep}
|
|
|
|
Return:
|
|
{
|
|
"run_path": "...",
|
|
"light": { verdicts, scores },
|
|
"deep": { verdicts, scores },
|
|
"agreement": {
|
|
"vm_verdict": true/false,
|
|
"task_verdict": true/false,
|
|
"score_delta": { ... }
|
|
}
|
|
}
|
|
"""
|
|
context: { light, deep }
|
|
|
|
# ============================================================
|
|
# Phase 3: Statistical Analysis
|
|
# ============================================================
|
|
|
|
let statistics = session: statistician
|
|
prompt: """
|
|
Compute calibration statistics.
|
|
|
|
Data: {calibration_data}
|
|
|
|
Calculate:
|
|
- Overall agreement rate (how often do light and deep agree?)
|
|
- Agreement by verdict type (vm vs task)
|
|
- Score correlation (do light scores predict deep scores?)
|
|
- Disagreement patterns (when do they diverge?)
|
|
|
|
Return:
|
|
{
|
|
"sample_size": N,
|
|
"agreement_rate": { overall, vm, task },
|
|
"score_correlation": { ... },
|
|
"disagreements": [ { run, light_said, deep_said, reason } ],
|
|
"confidence": "high" | "medium" | "low"
|
|
}
|
|
"""
|
|
context: calibration_data
|
|
|
|
# ============================================================
|
|
# Phase 4: Recommendations
|
|
# ============================================================
|
|
|
|
let recommendations = session: advisor
|
|
prompt: """
|
|
Based on calibration results, recommend improvements.
|
|
|
|
Statistics: {statistics}
|
|
Raw data: {calibration_data}
|
|
|
|
If agreement is high (>90%):
|
|
- Light evaluation is reliable
|
|
- Note any edge cases to watch
|
|
|
|
If agreement is medium (70-90%):
|
|
- Identify patterns in disagreements
|
|
- Suggest criteria adjustments
|
|
|
|
If agreement is low (<70%):
|
|
- Light evaluation needs work
|
|
- Specific recommendations for improvement
|
|
|
|
Return:
|
|
{
|
|
"reliability_verdict": "reliable" | "mostly_reliable" | "needs_work",
|
|
"key_findings": [...],
|
|
"recommendations": [
|
|
{ "priority": 1, "action": "...", "rationale": "..." }
|
|
]
|
|
}
|
|
"""
|
|
context: { statistics, calibration_data }
|
|
|
|
# ============================================================
|
|
# Output
|
|
# ============================================================
|
|
|
|
output report = session "Format report"
|
|
prompt: """
|
|
Format calibration results as a report.
|
|
|
|
Statistics: {statistics}
|
|
Recommendations: {recommendations}
|
|
|
|
Include:
|
|
1. Summary: Is light evaluation reliable?
|
|
2. Agreement rates (table)
|
|
3. Disagreement cases (if any)
|
|
4. Recommendations
|
|
5. Confidence level in these results
|
|
|
|
Format as markdown.
|
|
"""
|
|
context: { statistics, recommendations, calibration_data }
|