# DESIGN — Agap Agent Platform v2: the agent algebra Status: **v2, agreed with alvis 2026-07-21** (supersedes v1 draft `2499218c`) Owner: alvis · Written with Claude One design, two axioms, one verb. Everything alvis asked for — per-model queues, quota parking, Hindsight reflect as an async task, the Claude Code loop as a task puller, semantic/tier/direct routing — falls out as a special case rather than a rule. This document is the reference; the Kanboard A2A tasks implement it. --- ## 0. Glossary - **Task plane / "the fabric"** — the task-passing substrate connecting all agents: **Kanboard** (the durable task store and queue — for humans *and* agents) + **A2A protocol semantics** (submit/status/result, Agent Cards, context-by-reference) + the **conventions** on top (claim/lease, priorities, parking, trust-class routing). Not a deployable component; the collective name, the way "the network" names cables + IP + routing. - **Card** — an agent's self-description: capabilities, tier, cost class, trust class, availability. Maps 1:1 to an A2A Agent Card. - **Backbone** — the concrete LLM an agent currently uses for reasoning. - **Context ref** — a pointer (Hindsight bank id, git ref, KB task id, file path) passed *instead of* pasted content. ## 1. Why Observed on the live stack: - **Duplicated LLM spend** — an Adolf turn costs a ~32.8K-token reply call plus a ~22.8K-token background Hindsight extraction; ~425 tokens are the conversation. - **Tool bloat** — one Adolf carries ~84 MCP tool schemas + ~26 built-in tools every turn (~22K tokens), relevant or not. - **Quota cliffs** — Kimi's flat window (~60 msgs/5h, ~300/wk measured) makes Adolf go dark with no degradation path and no way to park work. - **Hardcoded background cognition** — Hindsight reflect/consolidation call a fixed model directly: no scheduling, no priority, no quota awareness. - **No growth path** — the ambition is autonomous research agents, remote llama nodes, more GPUs, more agents. ## 2. The algebra ### Axiom 1 — everything that can receive work is an Agent An agent is `(identity, Card, Policy, State)`. The Card advertises capabilities, **tier** (model strength it offers or needs), **cost class**, **trust class** (§5), and an **availability function a(t)**. Special cases: | Agent | Persona | Memory | Card highlights | |---|---|---|---| | LLM endpoint (`kimi`, `gemma3:4b`, …) | trivial (identity) | none | tier, cost, quota-shaped a(t) | | **Adolf** | proactive auditor (SOUL.md) | Hindsight bank `adolf` | trusted; scoped core tools | | **claude-coder** (Claude Code loop) | implementer | session + repo | trusted; pulls complex coding tasks | | **Torgash** | marketplace analyst | own bank | sandboxed; marketplace tools only | | **researcher** | autonomous researcher | own bank | sandboxed/untrusted inputs; own KB project | | router | delegator | none | resolves constraints → agents | | **alvis (the human)** | — | — | trust=human; a(t)=waking hours; **inbox = KB "waiting-on-me"** | The human being an agent is not a metaphor: approval gates, escalations and decisions are ordinary tasks submitted to his inbox. The KB column he already processes *is* that inbox. ### Axiom 2 — one verb ``` submit(task, target) -> taskRef # await(taskRef) optional => sync task = (intent, context-refs, constraints, priority, deadline, provenance) target ∈ { agent-id # direct: “this backbone / this specialist” | constraint-set # tier/capability: “any large model with tools” | auto } # router decides by availability/quota/complexity ``` Context travels **by reference, never by value** — the single most important efficiency rule for inter-agent communication (A2A context-passing practice). `sync` vs `async` is not a second mechanism: sync = submit + await. ### Theorems — the old rules become consequences 1. **"Queues are per model, not per agent."** Every agent has an inbox, but queues *accumulate* only where a(t) or throughput binds — at scarce agents: model-agents and the human. Persona agents transform-and-delegate, so their inboxes stay near-empty. The v1 rule is the scarcity special case. 2. **Quota lifecycles are shapes of a(t).** always-on: a(t)=1. quota-gated (Kimi): a(t)=0 when the window is spent — the queue **parks**, nothing fails, drains on reset. cost-gated: a(t)=0 past budget. on-demand (remote llama): a(t)=0 until woken. Four lifecycles, one function. 3. **Hindsight reflect is just a submit** — `{intent: reflect, refs: bank+query, constraints: tier≥large}`, async. Same for consolidation (low priority). 4. **Backbone swap is a constraint edit.** Persona agents name constraints, not endpoints; the backbone resolves per-submit. Adolf-on-Kimi today, Adolf-on-local tomorrow — no code change. 5. **The Claude Code loop is an ordinary consumer** — an agent whose policy is "pull complex coding tasks from the fabric". It was never special. 6. **For free:** escalation = re-submit with wider constraints (gated by policy, §5); approval = submit(…, target=alvis); proactivity/cron = delayed self-submission; the researcher = a low-priority self-submitting loop. ### Granularity rule A **Task** is a durable work item with a lifecycle worth auditing. A single LLM completion inside an agent's turn is **not** a Task — it is an implementation detail, observable in Langfuse, invisible to Kanboard. This keeps the KB-literal fabric free of micro-churn by construction. ## 3. Planes ``` ┌─ Task plane (“the fabric”) ─────────────────────────────────────────┐ │ Kanboard = the queue + audit + human inboxes (KB-LITERAL: no │ │ separate store). A2A semantics; claim/lease; priorities; parking. │ └─────────────────────────────────────────────────────────────────────┘ ┌─ Agent plane ───────────────────────────────────────────────────────┐ │ Registry of Cards (persona, memory bank, tool scope, trust class, │ │ preferred tier, current backbone). Runtimes: OpenClaw (Adolf + │ │ specialists), Claude Code CLI, thin workers. │ └─────────────────────────────────────────────────────────────────────┘ ┌─ Model plane ───────────────────────────────────────────────────────┐ │ LiteLLM gateway (:4000). Auto Router v2 (2026-07-14) does the SYNC │ │ routing natively: pinned model | tier pools | complexity/semantic │ │ auto-routing (SIMPLE trusted > sandboxed > untrusted ``` - **trusted** (Adolf, claude-coder): vault access **yes**; outward actions per existing ask-first rules. - **sandboxed** (Torgash, researcher): **no vault**, no outward sends; scoped MCP allowlist; KB access **project-scoped** (researcher gets its own KB project(s)). - **untrusted** = anything ingesting the open web: its *outputs* are tainted. **Taint / prompt-injection boundary:** tainted output may be written only to the agent's own bank/notes/project. Promotion into a trusted agent's memory or into any action requires a gate (initially: a task to alvis's inbox; later possibly a reviewer-agent). **Escalation policy (initial): always-ask.** A task that fails on its tier is not silently retried on a bigger model; it becomes a decision task in alvis's inbox. Revisit once behavior is observed (debugging phase by design). **Sandboxed coding — workspace lease:** per **task**, not per agent: `workspaces///` = ephemeral gitea clone + branch; execution inside a container (no vault creds by default, network allowlist, resource caps); merge **only via PR** to gitea; autonomous agents never push to main. Reviewer = human, or later a reviewer-agent (just another persona). **Global budget governor:** near the end of a quota window, interactive agents (Adolf) outrank background ones (researcher, consolidation) — arbitration lives in the fabric (priorities + a small governor rule), not in LiteLLM. ## 6. Executor — thin KB-polling workers No Temporal/Hatchet: at homelab scale (dozens of tasks/day) a durable-execution platform would duplicate Kanboard as a second source of truth. Instead, one small worker daemon per model-queue (compose services, ~200 lines, shared lib): ``` loop: a(t) check # quota/budget/health probe; if 0 → park (sleep, re-probe) poll KB view # filtered: my queue, status=queued, by priority claim # atomic: assign-to-self + column move + lease timestamp resolve refs # fetch context by reference execute # via LiteLLM (model-agents) / agent runtime (persona) write result ref # to the shared store; never inline update status # done | failed(retry policy) | input-required(→ inbox) ``` Leases + heartbeats make dead workers safe: an expired lease returns the task to queued. Two workers on one queue never double-run a task (claim is atomic). Idempotency keys on submission prevent duplicate proactive tasks. OpenClaw cron is the proactive *submitter* (Adolf's schedule); workers are the *drainers*. ## 7. Observability — Langfuse (kept), wired for real Decision: keep **Langfuse** (already deployed; best-in-class self-hosted: traces + per-token cost + prompt management + evals, MIT). Grafana rejected for this role — generic metrics with no LLM semantics (the source of past dissatisfaction); Zabbix keeps infra monitoring. To do (it currently receives nothing): LiteLLM success/failure callbacks → Langfuse; tag every trace with `agent`, `task-id`, `queue`; per-agent cost dashboards; upgrade v2→v3. Every completion is traced here — this is where sub-Task granularity lives. ## 8. Growing the lab - **More GPUs / remote llama** → new model-agent Cards with `on-demand` a(t) (health probe, wake hook, graceful absence). Routing skips absent nodes. - **More specialists** → new Cards + scoped tools + own banks. The fabric and A2A don't change. - **Autonomous research agents** → low-priority loops on always-on local queues, escalating (via always-ask, initially) for large-model synthesis; own KB project; tainted outputs until promoted. ## 9. Migration order 1. Registries: model Cards + agent Cards (schema + populate). 2. First thin worker end-to-end on an always-on local queue. 3. Quota-gated worker (Kimi park/resume). Claim/lease semantics. 4. A2A protocol surface (JSON-RPC + Agent Cards) over the fabric. 5. Cutovers: Hindsight reflect → fabric; consolidation → fabric (low prio); claude-coder + Adolf declared as registry agents (Adolf's tools shrink to scoped core). 6. Trust enforcement: capability grants (virtual keys + MCP allowlists), taint gate, budget governor, langfuse wiring. 7. Scale: Torgash, researcher (own KB project), on-demand nodes. ## 10. Decision log (2026-07-21, alvis) 1. Single completions are not Tasks (langfuse-only) → no micro-churn. 2. **KB-literal**: Kanboard is the queue, humans included; no separate store. 3. Trust classes as §5; vault = trusted only. 4. Escalation = always-ask initially. 5. Researcher: KB access allowed, own project(s), scope-limited. 6. Real A2A protocol now (JSON-RPC + Agent Cards). 7. Proactive schedules: OpenClaw cron → fabric. 8. Langfuse kept as the observability layer; Grafana rejected; Zabbix = infra. 9. Executor = thin KB-polling workers; no Hatchet/Temporal at this scale. 10. Sync routing = LiteLLM Auto Router v2; fabric owns async/parking. 11. Sandbox = per-task workspace lease + container + PR-only merges.