Files
AgapHost/adolf/DESIGN-proactive-prioritization.md
alvis a5c625b9b6 adolf: bearer-authenticate the agap MCP server, fix audio config schema
openclaw.json now sends `Authorization: Bearer ${AGAP_MCP_TOKEN}` to the agap
MCP server, which requires it as of kb#180. The token is injected from
openai/.env via docker-compose.yml and only substituted here, never inlined.
It maps to agent id `adolf`, which is also what the kb#147 vault gate reads.

Fixes tools.media.audio, which had been added but never restart-validated:
the per-entry `apiKey: "not-needed"` is rejected by the schema
("tools.media.audio.models.0: Invalid input"), and an invalid config makes the
gateway refuse to start outright -- adolf crash-looped on the first restart
after the block landed. The old comment claimed the schema requires a
non-empty apiKey; it is the opposite, apiKey is not a valid per-entry key at
all. Isolated with `openclaw config validate` against the running image
(2026.6.11): {provider, model} and {provider, model, baseUrl} validate, and
adding apiKey alone reproduces the failure. baseUrl is kept -- that is the
per-entry override pointing the openai-shaped provider at the local
faster-whisper server. Provider auth follows the normal model auth order per
docs/nodes/audio.md, and faster-whisper-server has no auth to satisfy anyway.

Two lessons encoded in the comments: `enabled: false` does NOT exempt an entry
from schema validation, and a config edit is not done until a restart boots
healthy -- this sat invalid but latent because the running gateway still held
an older loaded config. The block stays enabled: false; turning STT on is
still a kb#175/#191 decision (GTX 1070 co-residency).

Also adds the proactive-prioritization and todoist-capture design notes and
the vw-mcp prototype.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 04:41:13 +00:00

15 KiB
Raw Blame History

DESIGN — Proactive action impact/cost prioritization (kb#123)

Status: draft, for review — written 2026-07-22 Owner: alvis · Written with Claude Scope: kanboard Adolf task #123. This is a design/ruleset only — no wiring into openclaw.json, no code. Claude implements it in a follow-up task.

Related, not duplicated here:

  • #122 (utility/ROI metrics via Langfuse) — that's the after-the-fact "was Adolf worth it" readout across all of Adolf's spend. This doc is the before-the-fact gate on one specific class of spend: proactive sends. #122's Langfuse data is a candidate future input to the acceptance-rate term below (§3.3), but this design does not depend on #122 landing first.
  • #124 (proactive cadence / cron schedule) — decides when Adolf looks (daily/weekly/monthly audit cadence, adapted to quota). This design decides whether a specific candidate action fires once #124 (or an ad-hoc trigger, e.g. reacting to a calendar change) has already produced one. #124 is the scheduler; this is the gate every candidate passes through regardless of what triggered it.
  • #125 (feedback loop) — this design's acceptance-probability term (§3.3) is a consumer of #125's feedback log. #125 is not built yet; §5 below specifies exactly what it needs to log, as a dependency, not an assumption.

1. Problem

Adolf can generate a proactive message (reminder, nudge, digest item) from several places: the cadence jobs in #124, a reactive trigger (calendar event changed, HA sensor fired, a task went overdue), or background reasoning noticing something. Not every candidate should be sent — some are low-value, some are redundant with something already said, some cost real Kimi quota to formulate and the window is nearly exhausted.

This design is a gate function: given a candidate proactive action, decide fire / suppress / defer, using a score computed from four inputs the task description names, plus a threshold that scales with remaining quota.

2. Where the gate sits

 [trigger: #124 cadence job | reactive event | background noticing]
                     |
                     v
          candidate proactive action
          (draft content + metadata)
                     |
                     v
        ┌─────────────────────────┐
        │   IMPACT/COST GATE      │   <-- this design
        │   (score, threshold)    │
        └─────────────────────────┘
              |            |
           fire         suppress / defer
              |            |
       send via Matrix   log decision + reason
       log outcome slot   (no send; feedback loop
       for #125           has nothing to attach to)

The gate is a pure function of the candidate + current state. It does not decide what to consider sending (#124's job) or how to learn from responses (#125's job) — only whether a given candidate clears the bar right now.

3. Scoring model

For each candidate action a, compute:

score(a) = (benefit(a) * accept_prob(a) * urgency(a)) / cost(a)

Ratio form, not a weighted sum: cost is a genuine denominator (token cost is literally what you're trading against benefit), and the three numerator terms are gates on each other, not additive alternatives — a high-benefit, low-acceptance-probability action should be suppressed even if urgency is high, not partially compensated the way a sum would allow. All four terms are normalized to comparable ranges as defined below so no single term dominates by scale alone.

3.1 Benefit — benefit(a) ∈ [0, 1]

"How important is this to the user, if accepted." Estimated by Adolf itself (the model doing the drafting) using a small fixed rubric — this is a judgment call, not a measurement, so keep the rubric coarse enough to be stable across repeated runs:

Band Value Examples
Critical 1.0 Hard deadline today/tomorrow, safety/health-adjacent, financial penalty if missed
High 0.7 Real deadline this week, blocks another person, irreversible if missed
Medium 0.4 Useful reminder, no hard deadline, low cost if ignored
Low 0.15 "Might be nice to know", trivia-adjacent, no consequence

Adolf assigns the band as part of drafting the candidate (one extra field in the same generation pass — no separate LLM call). This is inherently noisy; it is corrected over time by the acceptance-probability term (§3.3), which is grounded in actual logged outcomes rather than self-assessment.

3.2 Cost — cost(a) ∈ (0, 1], token cost normalized

Raw cost is estimable before sending: tokens_estimate(a) = prompt tokens to formulate (system + context already loaded for the turn, since it's piggy-backing on an existing generation) + estimated output tokens for the message itself. For a message that requires its own dedicated Kimi call (true incremental cost) vs. one riding along inside an already-scheduled cadence turn (near-zero marginal cost), these are very different costs — the estimate must distinguish "marginal call I wouldn't otherwise make" from "free byproduct of a call happening anyway":

cost_tokens(a) = marginal_prompt_tokens(a) + marginal_output_tokens(a)

where marginal_* is 0 (or near-0, e.g. a few output tokens) if the action rides inside a scheduled #124 audit turn that would run regardless, and the full call cost if it requires spinning up a fresh Kimi turn.

Normalize against a reference ceiling (a "typical expensive proactive send", empirically ~2-3K tokens per the #122 baseline measurements of ~32.8K for a full reply turn — a standalone proactive nudge should be far cheaper than a full conversational turn, since it's one-directional with no back-and-forth):

cost(a) = clamp(cost_tokens(a) / COST_CEILING, floor=0.05, cap=1.0)

COST_CEILING = 3000 tokens (tunable constant, revisit once #122 gives real distributions). The 0.05 floor stops a literally-free riding action from dividing by ~0 and producing a runaway score — even "free" actions carry some opportunity cost (attention, message-count against the 60/5h ceiling, not just tokens).

3.3 Acceptance probability — accept_prob(a) ∈ [0, 1]

This term has no data source yet. It depends entirely on #125 (feedback loop) being built and logging outcomes. Until then, use a flat prior:

accept_prob(a) = 0.5   # uninformative prior, pending #125

Once #125 logs (action_class, outcome) pairs (see §5's exact schema requirement — this design does not invent history, it specifies what must exist), compute a per-class empirical rate with Laplace smoothing so a class with zero or few samples doesn't overfit to noise:

accept_prob(class) = (accepted_count(class) + 1) / (total_count(class) + 2)

class is a coarse bucket, not per-message: e.g. {calendar_reminder, task_overdue, ha_anomaly, family_wiki_gap, digest_item, ...} — one row per class, not per exact message text, since exact-text history rarely repeats but the class does. A candidate's class is assigned at draft time (same pass as §3.1's benefit band).

Recency matters more than total count — a user who started dismissing calendar_reminder last week should pull that class's rate down faster than five-year-old acceptances prop it up. Use a decayed count (e.g. half-life of 30 days, or simply windowing to the trailing N=50 outcomes per class) rather than an all-time average, once enough volume exists to make decay meaningful.

3.4 Urgency — urgency(a) ∈ [0, 1]

Distinct from benefit: benefit is "how much it matters", urgency is "how soon it stops being actionable". A time-decay curve against the nearest relevant deadline (due_at) known for the candidate (calendar event start, task date_due, HA-derived risk window):

hours_to_deadline = (due_at - now) in hours

urgency(a) =
    1.0                              if hours_to_deadline <= 1
    1.0 - 0.6 * (h - 1) / 23         if 1 < h <= 24     (1.0 -> 0.4 over the day)
    0.4 * exp(-(h - 24) / 168)       if h > 24          (decays over the following week)
    0.2                              if no deadline (informational-only action)

Concretely: something due within the hour scores 1.0, something due tomorrow ~0.4-1.0 depending on how close, something a week out trails off toward the 0.2 floor for undated nudges. This is a simple monotonic decay, not a precise model — tune the constants once real cadence data exists (#124).

3.5 Putting it together

score(a) = (benefit(a) * accept_prob(a) * urgency(a)) / cost(a)

Range: numerator ∈ [0, 1], denominator ∈ [0.05, 1], so score(a) ∈ [0, 20] in the degenerate cheapest/most-urgent/most-beneficial case. In practice, typical scores cluster well below that ceiling — the threshold (§4) is calibrated empirically against observed scores, not derived analytically from the range.

4. Firing rule

fire(a)  iff  score(a) >= threshold(current_quota_state)
else          suppress(a)   # or defer(a), see below

4.1 Quota signal

Read adolf-llm:8010/usage (confirmed live shape, sampled 2026-07-22):

{
  "weekly":   {"pct": 32, "used": 32, "limit": 100, "remaining": 68, "resets": "..."},
  "window_5h":{"pct": 79, "used": 79, "limit": 100, "remaining": 21, "resets": "..."},
  "stale": false
}

Use the tighter of the two windows — whichever pct is higher is the binding constraint right now:

quota_pressure = max(weekly.pct, window_5h.pct) / 100      # ∈ [0, 1]

If stale: true (Kimi login/session broken, per the adolf-llm fallback behavior), treat as quota_pressure = 1.0 (most conservative) — an unknown quota state should suppress non-critical sends, not fire them.

4.2 Threshold as a function of quota pressure

threshold(quota_pressure) = T_BASE + (T_MAX - T_BASE) * quota_pressure^2
  • T_BASE = 0.3 — threshold when quota is abundant (pressure ~0): let most medium-benefit things through.
  • T_MAX = 3.0 — threshold when quota is nearly exhausted (pressure ~1): only near-maximal score (critical benefit, high acceptance history, urgent, cheap) still fires.
  • Squaring quota_pressure keeps the threshold flat and permissive through low-to-mid pressure (nothing changes until quota actually gets tight) and then rises steeply as the window approaches exhaustion — matching the actual failure mode (403 usage-limit) which is a cliff, not a slope.

This gives a single tunable curve with two constants, both revisitable once #122 supplies real score distributions and false-negative/positive rates.

4.3 Fire / suppress / defer

  • fire: send now.
  • suppress: below threshold and no deadline pressure — drop it. Log the decision (§5) but do not re-surface it later; if it's still relevant, the next cadence pass (#124) will regenerate it as a fresh candidate with updated urgency.
  • defer: below threshold only because of quota pressure, but urgency(a) >= 0.8 (i.e., something time-critical got starved by a quota cliff, not by low benefit). Requeue for immediate re-evaluation once quota_pressure drops (next window reset, per resets timestamp in the usage payload) rather than silently dropping it. This is the one exception to "gate is stateless" — a deferred item carries state (its own candidate record) until it either fires or its deadline passes, at which point it is logged as a missed/expired suppression, not silently lost.

5. Dependency: what #125's feedback log must contain

This design's accept_prob term (§3.3) is inert without it. #125 owns building the collection mechanism (reactions, "+/-/неактуально" replies); this design only specifies the shape the gate needs to consume, so the two tasks don't diverge on schema:

proactive_outcome {
  action_class:   string   # matches the class taxonomy in §3.3, e.g. "calendar_reminder"
  sent_at:        timestamp
  benefit_band:   float    # the benefit(a) value used at send time, for later calibration
  cost_tokens:    int      # actual cost, for calibrating COST_CEILING
  urgency_at_send: float
  outcome:        enum { accepted, dismissed, ignored, irrelevant }
  responded_at:   timestamp | null
}

ignored (no response within some window, e.g. 24h) must be distinguished from dismissed (explicit "") — an ignored item is weaker negative signal than an explicit rejection and should decay the acceptance rate less aggressively. Without this distinction the Laplace-smoothed rate in §3.3 conflates "user didn't care" with "user was just busy."

Every suppressed and deferred candidate should also be logged (not just fired ones) with outcome: not_sent — this is what lets a later audit (#122) compute false-suppression rate (was a suppressed item actually needed? only knowable in hindsight, e.g. if the same underlying deadline later caused a problem) as well as false-fire rate.

6. Worked example

Candidate: "reminder that the Seafile SSL cert renews in 3 days" (from a #124 daily cadence audit, riding along inside that scheduled call).

  • benefit: Medium band → 0.4 (annoying if missed, not critical — auto-renew likely already configured, this is a check not a fire drill).
  • cost: marginal — rides inside the already-running daily audit call, say ~150 marginal output tokens → 150/3000 = 0.05 → floored at 0.05.
  • accept_prob: no #125 data yet → flat prior 0.5.
  • urgency: hours_to_deadline = 72h → falls in the h > 24 branch: 0.4 * exp(-(72-24)/168) = 0.4 * exp(-0.286) ≈ 0.4 * 0.751 ≈ 0.30.

score = (0.4 * 0.5 * 0.30) / 0.05 = 0.06 / 0.05 = 1.2

At quota_pressure = 0 (abundant quota), threshold = 0.31.2 ≥ 0.3, fires. At quota_pressure = 1 (window nearly exhausted, matching the measured 79% 5h-window sample above, rounding up toward the cliff), threshold = 3.01.2 < 3.0, suppressed (not deferred: urgency 0.30 is well under the 0.8 defer bar) — correctly deprioritized under quota pressure in favor of anything more urgent or already proven to land well.

7. Open parameters to tune post-implementation

Everything with a concrete numeric constant above (COST_CEILING, T_BASE, T_MAX, the urgency decay constants, the defer bar) is a starting guess consistent with the measurements already on hand (#122's token baseline, the live /usage sample). None of it is load-bearing on the shape of the model — only on where the dial sits. Revisit once:

  • #125 supplies real accept_prob data (replacing the flat 0.5 prior is the single highest-value follow-up — everything else is a reasonable guess, this term is currently a placeholder).
  • #122's Langfuse integration supplies real per-action token costs to recalibrate COST_CEILING.
  • A few weeks of fire/suppress/defer logs (§5) exist to check the threshold isn't systematically over- or under-firing.

8. Acceptance check against kb#123

  • Scoring formula with each term defined, ranged, and its estimation method stated: §3.
  • Firing rule (fire only if impact/cost exceeds a threshold): §4.
  • Threshold tunable to current quota, against the real adolf-llm:8010/usage signal: §4.1-4.2.
  • Acceptance-probability term flagged as dependent on unbuilt history (#125), with the exact log schema it needs specified rather than fabricated: §3.3, §5.
  • Cross-references to sibling tasks #122, #124, #125 without duplicating their scope: header + inline.