Unified redesign: two axioms (everything that can receive work is an Agent with a Card and availability function; one verb submit(task, target)) from which the v1 rules follow as theorems — per-model queues (scarcity), quota parking (a(t)=0), reflect-as-task, backbone swap as constraint edit, the Claude Code loop as an ordinary consumer, human-as-agent (KB waiting column = his inbox). Adds what v1 lacked: trust classes (vault=trusted only), prompt-injection taint boundary, always-ask escalation, global budget governor, per-task workspace-lease sandboxing with PR-only merges, KB-literal fabric decision, LiteLLM Auto Router v2 for sync routing, thin KB-polling workers as the executor, Langfuse kept as observability, real A2A protocol adoption. Includes decision log from alvis. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014t8Qg9gi7H7HtT8MncoXAB
14 KiB
DESIGN — Agap Agent Platform v2: the agent algebra
Status: v2, agreed with alvis 2026-07-21 (supersedes v1 draft 2499218c)
Owner: alvis · Written with Claude
One design, two axioms, one verb. Everything alvis asked for — per-model queues, quota parking, Hindsight reflect as an async task, the Claude Code loop as a task puller, semantic/tier/direct routing — falls out as a special case rather than a rule. This document is the reference; the Kanboard A2A tasks implement it.
0. Glossary
- Task plane / "the fabric" — the task-passing substrate connecting all agents: Kanboard (the durable task store and queue — for humans and agents) + A2A protocol semantics (submit/status/result, Agent Cards, context-by-reference) + the conventions on top (claim/lease, priorities, parking, trust-class routing). Not a deployable component; the collective name, the way "the network" names cables + IP + routing.
- Card — an agent's self-description: capabilities, tier, cost class, trust class, availability. Maps 1:1 to an A2A Agent Card.
- Backbone — the concrete LLM an agent currently uses for reasoning.
- Context ref — a pointer (Hindsight bank id, git ref, KB task id, file path) passed instead of pasted content.
1. Why
Observed on the live stack:
- Duplicated LLM spend — an Adolf turn costs a ~32.8K-token reply call plus a ~22.8K-token background Hindsight extraction; ~425 tokens are the conversation.
- Tool bloat — one Adolf carries ~84 MCP tool schemas + ~26 built-in tools every turn (~22K tokens), relevant or not.
- Quota cliffs — Kimi's flat window (~60 msgs/5h, ~300/wk measured) makes Adolf go dark with no degradation path and no way to park work.
- Hardcoded background cognition — Hindsight reflect/consolidation call a fixed model directly: no scheduling, no priority, no quota awareness.
- No growth path — the ambition is autonomous research agents, remote llama nodes, more GPUs, more agents.
2. The algebra
Axiom 1 — everything that can receive work is an Agent
An agent is (identity, Card, Policy, State). The Card advertises capabilities,
tier (model strength it offers or needs), cost class, trust class
(§5), and an availability function a(t). Special cases:
| Agent | Persona | Memory | Card highlights |
|---|---|---|---|
LLM endpoint (kimi, gemma3:4b, …) |
trivial (identity) | none | tier, cost, quota-shaped a(t) |
| Adolf | proactive auditor (SOUL.md) | Hindsight bank adolf |
trusted; scoped core tools |
| claude-coder (Claude Code loop) | implementer | session + repo | trusted; pulls complex coding tasks |
| Torgash | marketplace analyst | own bank | sandboxed; marketplace tools only |
| researcher | autonomous researcher | own bank | sandboxed/untrusted inputs; own KB project |
| router | delegator | none | resolves constraints → agents |
| alvis (the human) | — | — | trust=human; a(t)=waking hours; inbox = KB "waiting-on-me" |
The human being an agent is not a metaphor: approval gates, escalations and decisions are ordinary tasks submitted to his inbox. The KB column he already processes is that inbox.
Axiom 2 — one verb
submit(task, target) -> taskRef # await(taskRef) optional => sync
task = (intent, context-refs, constraints, priority, deadline, provenance)
target ∈ { agent-id # direct: “this backbone / this specialist”
| constraint-set # tier/capability: “any large model with tools”
| auto } # router decides by availability/quota/complexity
Context travels by reference, never by value — the single most important
efficiency rule for inter-agent communication (A2A context-passing practice).
sync vs async is not a second mechanism: sync = submit + await.
Theorems — the old rules become consequences
- "Queues are per model, not per agent." Every agent has an inbox, but queues accumulate only where a(t) or throughput binds — at scarce agents: model-agents and the human. Persona agents transform-and-delegate, so their inboxes stay near-empty. The v1 rule is the scarcity special case.
- Quota lifecycles are shapes of a(t). always-on: a(t)=1. quota-gated (Kimi): a(t)=0 when the window is spent — the queue parks, nothing fails, drains on reset. cost-gated: a(t)=0 past budget. on-demand (remote llama): a(t)=0 until woken. Four lifecycles, one function.
- Hindsight reflect is just a submit —
{intent: reflect, refs: bank+query, constraints: tier≥large}, async. Same for consolidation (low priority). - Backbone swap is a constraint edit. Persona agents name constraints, not endpoints; the backbone resolves per-submit. Adolf-on-Kimi today, Adolf-on-local tomorrow — no code change.
- The Claude Code loop is an ordinary consumer — an agent whose policy is "pull complex coding tasks from the fabric". It was never special.
- For free: escalation = re-submit with wider constraints (gated by policy, §5); approval = submit(…, target=alvis); proactivity/cron = delayed self-submission; the researcher = a low-priority self-submitting loop.
Granularity rule
A Task is a durable work item with a lifecycle worth auditing. A single LLM completion inside an agent's turn is not a Task — it is an implementation detail, observable in Langfuse, invisible to Kanboard. This keeps the KB-literal fabric free of micro-churn by construction.
3. Planes
┌─ Task plane (“the fabric”) ─────────────────────────────────────────┐
│ Kanboard = the queue + audit + human inboxes (KB-LITERAL: no │
│ separate store). A2A semantics; claim/lease; priorities; parking. │
└─────────────────────────────────────────────────────────────────────┘
┌─ Agent plane ───────────────────────────────────────────────────────┐
│ Registry of Cards (persona, memory bank, tool scope, trust class, │
│ preferred tier, current backbone). Runtimes: OpenClaw (Adolf + │
│ specialists), Claude Code CLI, thin workers. │
└─────────────────────────────────────────────────────────────────────┘
┌─ Model plane ───────────────────────────────────────────────────────┐
│ LiteLLM gateway (:4000). Auto Router v2 (2026-07-14) does the SYNC │
│ routing natively: pinned model | tier pools | complexity/semantic │
│ auto-routing (SIMPLE<MEDIUM<COMPLEX<REASONING), plus virtual keys, │
│ budgets, 429-fallback. alvis's three routing modes map 1:1: │
│ specific backbone → pinned model_name │
│ “tier” routing → tier pool │
│ automatic router → auto_router/complexity_router │
│ The fabric owns everything LiteLLM cannot: ASYNC queueing, parking │
│ across quota windows, leases, cross-agent budget arbitration. │
└─────────────────────────────────────────────────────────────────────┘
┌─ Context stores ────────────────────────────────────────────────────┐
│ Hindsight banks (per-agent memory) · gitea (code, docs, this file) │
│ · KB task bodies · files. Tasks point here; payloads never inline. │
└─────────────────────────────────────────────────────────────────────┘
4. A2A: the protocol, adopted now
The algebra maps 1:1 onto A2A v1.0 (Jan 2026), which is why we implement the real protocol immediately rather than "patterns first":
| Algebra | A2A |
|---|---|
| Card | Agent Card (/.well-known/agent.json) |
| submit / await | message/send (sync-ish) / tasks/get (async) |
| task lifecycle | submitted → working → input-required → completed/failed/canceled |
| notify | push notifications |
Implementation: JSON-RPC 2.0 over HTTP on the LAN; each runtime (OpenClaw, Claude loop, workers) exposes/consumes A2A; Kanboard remains the durable state behind the endpoints. Scalability/extensibility later (remote nodes, third-party agents) then needs zero redesign.
5. Trust & sandboxing
Trust classes (on every Card):
human > trusted > sandboxed > untrusted
- trusted (Adolf, claude-coder): vault access yes; outward actions per existing ask-first rules.
- sandboxed (Torgash, researcher): no vault, no outward sends; scoped MCP allowlist; KB access project-scoped (researcher gets its own KB project(s)).
- untrusted = anything ingesting the open web: its outputs are tainted.
Taint / prompt-injection boundary: tainted output may be written only to the agent's own bank/notes/project. Promotion into a trusted agent's memory or into any action requires a gate (initially: a task to alvis's inbox; later possibly a reviewer-agent).
Escalation policy (initial): always-ask. A task that fails on its tier is not silently retried on a bigger model; it becomes a decision task in alvis's inbox. Revisit once behavior is observed (debugging phase by design).
Sandboxed coding — workspace lease: per task, not per agent:
workspaces/<agent>/<task-id>/ = ephemeral gitea clone + branch; execution
inside a container (no vault creds by default, network allowlist, resource
caps); merge only via PR to gitea; autonomous agents never push to main.
Reviewer = human, or later a reviewer-agent (just another persona).
Global budget governor: near the end of a quota window, interactive agents (Adolf) outrank background ones (researcher, consolidation) — arbitration lives in the fabric (priorities + a small governor rule), not in LiteLLM.
6. Executor — thin KB-polling workers
No Temporal/Hatchet: at homelab scale (dozens of tasks/day) a durable-execution platform would duplicate Kanboard as a second source of truth. Instead, one small worker daemon per model-queue (compose services, ~200 lines, shared lib):
loop:
a(t) check # quota/budget/health probe; if 0 → park (sleep, re-probe)
poll KB view # filtered: my queue, status=queued, by priority
claim # atomic: assign-to-self + column move + lease timestamp
resolve refs # fetch context by reference
execute # via LiteLLM (model-agents) / agent runtime (persona)
write result ref # to the shared store; never inline
update status # done | failed(retry policy) | input-required(→ inbox)
Leases + heartbeats make dead workers safe: an expired lease returns the task to queued. Two workers on one queue never double-run a task (claim is atomic). Idempotency keys on submission prevent duplicate proactive tasks. OpenClaw cron is the proactive submitter (Adolf's schedule); workers are the drainers.
7. Observability — Langfuse (kept), wired for real
Decision: keep Langfuse (already deployed; best-in-class self-hosted:
traces + per-token cost + prompt management + evals, MIT). Grafana rejected for
this role — generic metrics with no LLM semantics (the source of past
dissatisfaction); Zabbix keeps infra monitoring. To do (it currently receives
nothing): LiteLLM success/failure callbacks → Langfuse; tag every trace with
agent, task-id, queue; per-agent cost dashboards; upgrade v2→v3. Every
completion is traced here — this is where sub-Task granularity lives.
8. Growing the lab
- More GPUs / remote llama → new model-agent Cards with
on-demanda(t) (health probe, wake hook, graceful absence). Routing skips absent nodes. - More specialists → new Cards + scoped tools + own banks. The fabric and A2A don't change.
- Autonomous research agents → low-priority loops on always-on local queues, escalating (via always-ask, initially) for large-model synthesis; own KB project; tainted outputs until promoted.
9. Migration order
- Registries: model Cards + agent Cards (schema + populate).
- First thin worker end-to-end on an always-on local queue.
- Quota-gated worker (Kimi park/resume). Claim/lease semantics.
- A2A protocol surface (JSON-RPC + Agent Cards) over the fabric.
- Cutovers: Hindsight reflect → fabric; consolidation → fabric (low prio); claude-coder + Adolf declared as registry agents (Adolf's tools shrink to scoped core).
- Trust enforcement: capability grants (virtual keys + MCP allowlists), taint gate, budget governor, langfuse wiring.
- Scale: Torgash, researcher (own KB project), on-demand nodes.
10. Decision log (2026-07-21, alvis)
- Single completions are not Tasks (langfuse-only) → no micro-churn.
- KB-literal: Kanboard is the queue, humans included; no separate store.
- Trust classes as §5; vault = trusted only.
- Escalation = always-ask initially.
- Researcher: KB access allowed, own project(s), scope-limited.
- Real A2A protocol now (JSON-RPC + Agent Cards).
- Proactive schedules: OpenClaw cron → fabric.
- Langfuse kept as the observability layer; Grafana rejected; Zabbix = infra.
- Executor = thin KB-polling workers; no Hatchet/Temporal at this scale.
- Sync routing = LiteLLM Auto Router v2; fabric owns async/parking.
- Sandbox = per-task workspace lease + container + PR-only merges.