Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
alvis: a missed cron is an operational fault, not a task-lifecycle event, and re-firing it would make the keeper a trigger. Removed cron catch-up from fabric-keeper duties entirely. - Keeper is now strictly janitorial: neither assigns NOR triggers work (lease sweeps, deadlines, dead-letter, digests, metrics only). - Schedules that must not be missed are monitored in Zabbix like any other infra fault. - NEW §6d: faults are incidents, handled PER PROBLEM CLASS by a future ops-agent (watches Zabbix, applies known mitigation or files a task with the incident as context) — never a blanket catch-up policy, since a missed briefing / backup / consolidation want different responses. - Decision log entry 22. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014t8Qg9gi7H7HtT8MncoXAB
412 lines
23 KiB
Markdown
412 lines
23 KiB
Markdown
# DESIGN — Agap Agent Platform v2.1: the agent algebra
|
||
|
||
Status: **v2.1, agreed with alvis 2026-07-21** (v2 `7be30c71` + hardening review)
|
||
Owner: alvis · Written with Claude
|
||
|
||
One design, two axioms, one verb. Everything alvis asked for — per-model queues,
|
||
quota parking, Hindsight reflect as an async task, the Claude Code loop as a task
|
||
puller, semantic/tier/direct routing — falls out as a special case rather than a
|
||
rule. This document is the reference; the Kanboard A2A tasks implement it.
|
||
|
||
---
|
||
|
||
## 0. Glossary
|
||
|
||
- **Task plane / "the fabric"** — the task-passing substrate connecting all
|
||
agents: **Kanboard** (the durable task store and queue — for humans *and*
|
||
agents) + **A2A protocol semantics** (submit/status/result, Agent Cards,
|
||
context-by-reference) + the **conventions** on top (claim/lease, priorities,
|
||
parking, trust-class routing). Not a deployable component; the collective name,
|
||
the way "the network" names cables + IP + routing.
|
||
- **Card** — an agent's self-description: capabilities, tier, cost class, trust
|
||
class, availability. Maps 1:1 to an A2A Agent Card.
|
||
- **Backbone** — the concrete LLM an agent currently uses for reasoning.
|
||
- **Context ref** — a pointer (Hindsight bank id, git ref, KB task id, file
|
||
path) passed *instead of* pasted content.
|
||
- **fabric-keeper** — the janitor daemon owning time semantics (§6b): lease
|
||
sweeps, deadlines, dead-letter, inbox digests. It never assigns work and
|
||
never triggers work.
|
||
|
||
## 1. Why
|
||
|
||
Observed on the live stack:
|
||
|
||
- **Duplicated LLM spend** — an Adolf turn costs a ~32.8K-token reply call plus a
|
||
~22.8K-token background Hindsight extraction; ~425 tokens are the conversation.
|
||
- **Tool bloat** — one Adolf carries ~84 MCP tool schemas + ~26 built-in tools
|
||
every turn (~22K tokens), relevant or not.
|
||
- **Quota cliffs** — Kimi's flat window (~60 msgs/5h, ~300/wk measured) makes
|
||
Adolf go dark with no degradation path and no way to park work.
|
||
- **Hardcoded background cognition** — Hindsight reflect/consolidation call a
|
||
fixed model directly: no scheduling, no priority, no quota awareness.
|
||
- **No growth path** — the ambition is autonomous research agents, remote llama
|
||
nodes, more GPUs, more agents.
|
||
|
||
## 2. The algebra
|
||
|
||
### Axiom 1 — everything that can receive work is an Agent
|
||
|
||
An agent is `(identity, Card, Policy, State)`. The Card advertises capabilities,
|
||
**tier** (model strength it offers or needs), **cost class**, **trust class**
|
||
(§5), and an **availability function a(t)**. Special cases:
|
||
|
||
| Agent | Persona | Memory | Card highlights |
|
||
|---|---|---|---|
|
||
| LLM endpoint (`kimi`, `gemma3:4b`, …) | trivial (identity) | none | tier, cost, quota-shaped a(t) |
|
||
| **Adolf** | proactive auditor (SOUL.md) | Hindsight bank `adolf` | trusted; scoped core tools |
|
||
| **claude-coder** (Claude Code loop) | implementer | session + repo | trusted; pulls complex coding tasks |
|
||
| **Torgash** | marketplace analyst | own bank | sandboxed; marketplace tools only |
|
||
| **researcher** | autonomous researcher | own bank | sandboxed/untrusted inputs; own KB project |
|
||
| router | delegator | none | resolves constraints → agents |
|
||
| **alvis (the human)** | — | — | trust=human; a(t)=waking hours; **inbox = KB "waiting-on-me"** |
|
||
|
||
The human being an agent is not a metaphor: approval gates, escalations and
|
||
decisions are ordinary tasks submitted to his inbox. The KB column he already
|
||
processes *is* that inbox.
|
||
|
||
### Axiom 2 — one verb
|
||
|
||
```
|
||
submit(task, target) -> taskRef # await(taskRef) optional => sync
|
||
task = (intent, context-refs, constraints, priority, deadline, provenance)
|
||
target ∈ { agent-id # direct: “this backbone / this specialist”
|
||
| constraint-set # tier/capability: “any large model with tools”
|
||
| auto } # router decides by availability/quota/complexity
|
||
```
|
||
|
||
Context travels **by reference, never by value** — the single most important
|
||
efficiency rule for inter-agent communication (A2A context-passing practice).
|
||
`sync` vs `async` is not a second mechanism: sync = submit + await.
|
||
|
||
**Transport rule.** Sync and async share the algebra but not the transport:
|
||
**sync goes direct** — an A2A `message/send` RPC straight to the target agent's
|
||
endpoint, journaled to KB afterwards; **async/durable goes through KB** and is
|
||
drained by polling workers. KB polling must never sit on a sync path — a sync
|
||
call may not inherit poll-interval latency.
|
||
|
||
### Completion vs verification
|
||
|
||
KB convention (native semantics, no new machinery): the **Done column =
|
||
unverified completion** — the worker/agent finished and self-reported. **Closing
|
||
the task = verified completion.** The producer never closes its own task; the
|
||
submitter, a human, or (later) a reviewer-agent closes after checking the
|
||
task's acceptance criteria. Lifecycle: … → done (unverified) → closed
|
||
(verified). For code, the PR review is the verification; closing follows merge.
|
||
|
||
### Theorems — the old rules become consequences
|
||
|
||
1. **"Queues are per model, not per agent."** Every agent has an inbox, but
|
||
queues *accumulate* only where a(t) or throughput binds — at scarce agents:
|
||
model-agents and the human. Persona agents transform-and-delegate, so their
|
||
inboxes stay near-empty. The v1 rule is the scarcity special case.
|
||
2. **Quota lifecycles are shapes of a(t).** always-on: a(t)=1. quota-gated
|
||
(Kimi): a(t)=0 when the window is spent — the queue **parks**, nothing fails,
|
||
drains on reset. cost-gated: a(t)=0 past budget. on-demand (remote llama):
|
||
a(t)=0 until woken. Four lifecycles, one function.
|
||
3. **Hindsight reflect is just a submit** — `{intent: reflect, refs: bank+query,
|
||
constraints: tier≥large}`, async. Same for consolidation (low priority).
|
||
4. **Backbone swap is a constraint edit.** Persona agents name constraints, not
|
||
endpoints; the backbone resolves per-submit. Adolf-on-Kimi today,
|
||
Adolf-on-local tomorrow — no code change.
|
||
5. **The Claude Code loop is an ordinary consumer** — an agent whose policy is
|
||
"pull complex coding tasks from the fabric". It was never special.
|
||
6. **For free:** escalation = re-submit with wider constraints (gated by policy,
|
||
§5); approval = submit(…, target=alvis); proactivity/cron = delayed
|
||
self-submission; the researcher = a low-priority self-submitting loop.
|
||
|
||
### Granularity rule
|
||
|
||
A **Task** is a durable work item with a lifecycle worth auditing. A single LLM
|
||
completion inside an agent's turn is **not** a Task — it is an implementation
|
||
detail, observable in Langfuse, invisible to Kanboard. This keeps the KB-literal
|
||
fabric free of micro-churn by construction.
|
||
|
||
## 3. Planes
|
||
|
||
```
|
||
┌─ Task plane (“the fabric”) ─────────────────────────────────────────┐
|
||
│ Kanboard = the queue + audit + human inboxes (KB-LITERAL: no │
|
||
│ separate store). A2A semantics; claim/lease; priorities; parking. │
|
||
└─────────────────────────────────────────────────────────────────────┘
|
||
┌─ Agent plane ───────────────────────────────────────────────────────┐
|
||
│ Registry of Cards (persona, memory bank, tool scope, trust class, │
|
||
│ preferred tier, current backbone). Runtimes: OpenClaw (Adolf + │
|
||
│ specialists), Claude Code CLI, thin workers. │
|
||
└─────────────────────────────────────────────────────────────────────┘
|
||
┌─ Model plane ───────────────────────────────────────────────────────┐
|
||
│ LiteLLM gateway (:4000). Auto Router v2 (2026-07-14) does the SYNC │
|
||
│ routing natively: pinned model | tier pools | complexity/semantic │
|
||
│ auto-routing (SIMPLE<MEDIUM<COMPLEX<REASONING), plus virtual keys, │
|
||
│ budgets, 429-fallback. alvis's three routing modes map 1:1: │
|
||
│ specific backbone → pinned model_name │
|
||
│ “tier” routing → tier pool │
|
||
│ automatic router → auto_router/complexity_router │
|
||
│ The fabric owns everything LiteLLM cannot: ASYNC queueing, parking │
|
||
│ across quota windows, leases, cross-agent budget arbitration. │
|
||
└─────────────────────────────────────────────────────────────────────┘
|
||
┌─ Context stores ────────────────────────────────────────────────────┐
|
||
│ Hindsight banks (per-agent memory) · gitea (code, docs, this file) │
|
||
│ · KB task bodies · files. Tasks point here; payloads never inline. │
|
||
└─────────────────────────────────────────────────────────────────────┘
|
||
```
|
||
|
||
### GPU residency — a local model's a(t) is not 1
|
||
|
||
Local "free" models contend for VRAM with interactive components (measured on
|
||
the 8 GB GTX 1070: bge-m3 + gemma3:4b + tei-reranker ≈ 6.2 GB; loading anything
|
||
bigger evicts the reranker and silently regresses recall latency). So a local
|
||
model's availability is **a(t) = f(VRAM headroom)**, and the model registry
|
||
carries a **residency policy**: a never-evict set (embedder, reranker —
|
||
interactive-critical), allowed co-residency groups, and a pre-load check every
|
||
worker must pass before pulling a model onto a GPU. With more GPUs this becomes
|
||
a placement problem — same policy, more slots.
|
||
|
||
### Personas and Cards are code
|
||
|
||
SOUL.md files and agent Cards live **in git** and are deployed to runtimes —
|
||
never edited live in volumes. "Who changed Adolf's soul" must be a `git log`
|
||
answer. (The current SOUL.md in the adolf-state volume is migration debt.)
|
||
|
||
### Eval gate on backbone/routing changes
|
||
|
||
Backbone swap being "one constraint edit" is quality-blind. Each agent keeps a
|
||
**golden set** (10–20 canonical exchanges); any backbone or routing change is
|
||
shadow-replayed against it and compared (Langfuse datasets/evals) before taking
|
||
effect. The algebra's flexibility must not become a silent-degradation machine.
|
||
|
||
## 4. A2A: the protocol, adopted now
|
||
|
||
The algebra maps 1:1 onto A2A v1.0 (Jan 2026), which is why we implement the
|
||
real protocol immediately rather than "patterns first":
|
||
|
||
| Algebra | A2A |
|
||
|---|---|
|
||
| Card | Agent Card (`/.well-known/agent.json`) |
|
||
| submit / await | `message/send` (sync-ish) / `tasks/get` (async) |
|
||
| task lifecycle | submitted → working → input-required → completed/failed/canceled |
|
||
| notify | push notifications |
|
||
|
||
Implementation: JSON-RPC 2.0 over HTTP on the LAN; each runtime (OpenClaw,
|
||
Claude loop, workers) exposes/consumes A2A; Kanboard remains the durable state
|
||
behind the endpoints. Scalability/extensibility later (remote nodes, third-party
|
||
agents) then needs zero redesign.
|
||
|
||
**Auth is mandatory on every A2A surface.** The LAN is **not trusted** — the
|
||
xray/3x-ui VPN terminates other people's peers on it. No unauthenticated
|
||
JSON-RPC listener, ever: shared tokens minimum, mTLS preferred.
|
||
|
||
## 5. Trust & sandboxing
|
||
|
||
**Trust classes** (on every Card):
|
||
|
||
```
|
||
human > trusted > sandboxed > untrusted
|
||
```
|
||
|
||
- **trusted** (Adolf, claude-coder): vault access **yes**; outward actions per
|
||
existing ask-first rules.
|
||
- **sandboxed** (Torgash, researcher): **no vault**, no outward sends; scoped
|
||
MCP allowlist; KB access **project-scoped** (researcher gets its own KB
|
||
project(s)).
|
||
- **untrusted** = anything ingesting the open web: its *outputs* are tainted.
|
||
|
||
**Taint / prompt-injection boundary:** tainted output may be written only to the
|
||
agent's own bank/notes/project. Promotion into a trusted agent's memory or into
|
||
any action requires a gate (initially: a task to alvis's inbox; later possibly a
|
||
reviewer-agent).
|
||
|
||
**Escalation policy (initial): always-ask.** A task that fails on its tier is
|
||
not silently retried on a bigger model; it becomes a decision task in alvis's
|
||
inbox. Revisit once behavior is observed (debugging phase by design).
|
||
|
||
**Sandboxed coding — workspace lease:** per **task**, not per agent:
|
||
`workspaces/<agent>/<task-id>/` = ephemeral gitea clone + branch; execution
|
||
inside a container (no vault creds by default, network allowlist, resource
|
||
caps); merge **only via PR** to gitea; autonomous agents never push to main.
|
||
Reviewer = human, or later a reviewer-agent (just another persona).
|
||
|
||
**Global budget governor:** near the end of a quota window, interactive agents
|
||
(Adolf) outrank background ones (researcher, consolidation) — arbitration lives
|
||
in the fabric (priorities + a small governor rule), not in LiteLLM.
|
||
|
||
**Fabric hygiene (runaway protection):** agents submit tasks that cause agents
|
||
to submit tasks — idempotency keys stop duplicates, not generative loops. So:
|
||
per-agent **task-creation quotas**; an **ancestry depth cap** on provenance
|
||
chains; cycle detection at submit; and a **dead-letter** state for poison tasks
|
||
after max-retries — never an infinite retry loop through paid quota.
|
||
|
||
## 5b. Humans (plural) and memory partitioning
|
||
|
||
There is more than one human already (alvis and elizaveta are both on Adolf's
|
||
Matrix allowlist) and there will be more. Every human is an agent with
|
||
trust=human, their own inbox, and — critically — **their own privacy domain**.
|
||
|
||
**Memory partitioning (hard rules):**
|
||
|
||
- **Per-human private banks**: `adolf-alvis`, `adolf-elizaveta`, … Everything
|
||
learned in conversation with human H goes to H's private bank by default.
|
||
**Content from one human's conversations must never surface to another
|
||
human.** This is a correctness property, not a preference.
|
||
- **One shared household bank** for facts that are explicitly household-wide
|
||
(addresses, devices, routines, shared plans). Trusted agents may write;
|
||
**promotion from a private bank happens only by that human's explicit action
|
||
or approval task** — never automatically.
|
||
- **Recall is interlocutor-scoped**: when Adolf talks to H it recalls from H's
|
||
private bank + the shared bank, nothing else. The recall/retain hooks select
|
||
the bank by interlocutor identity.
|
||
- Sandboxed agents (Torgash, researcher) read at most the shared bank; never
|
||
any private bank. This is the cross-human face of the memory matrix.
|
||
- The current single `adolf` bank is migration debt: split into
|
||
`adolf-alvis` + shared.
|
||
|
||
**Human inbox design:** notifications are priority-routed — gate/urgent tasks
|
||
ping the human via Matrix (Adolf initiates them; cf. proactive-messaging work),
|
||
everything else lands in a daily digest from the fabric-keeper. Ignored gate
|
||
tasks park and re-remind; **they never default-approve**. Vacation mode: a
|
||
human's a(t)=0 parks their inbox like any other scarce queue — gated flows
|
||
wait; predefined degraded defaults apply where explicitly configured.
|
||
|
||
## 6. Executor — thin KB-polling workers
|
||
|
||
No Temporal/Hatchet: at homelab scale (dozens of tasks/day) a durable-execution
|
||
platform would duplicate Kanboard as a second source of truth. Instead, one
|
||
small worker daemon per model-queue (compose services, ~200 lines, shared lib):
|
||
|
||
```
|
||
loop:
|
||
a(t) check # quota/budget/health probe; if 0 → park (sleep, re-probe)
|
||
poll KB view # filtered: my queue, status=queued, by priority
|
||
claim # atomic: assign-to-self + column move + lease timestamp
|
||
resolve refs # fetch context by reference
|
||
execute # via LiteLLM (model-agents) / agent runtime (persona)
|
||
write result ref # to the shared store; never inline
|
||
update status # done | failed(retry policy) | input-required(→ inbox)
|
||
```
|
||
|
||
Leases + heartbeats make dead workers safe: an expired lease returns the task to
|
||
queued. Two workers on one queue never double-run a task (claim is atomic).
|
||
Idempotency keys on submission prevent duplicate proactive tasks. OpenClaw cron
|
||
is the proactive *submitter* (Adolf's schedule); workers are the *drainers*.
|
||
|
||
### 6b. Who is "the scheduler"? — decomposed, plus one janitor
|
||
|
||
There is deliberately **no central dispatcher**. Scheduling decomposes into
|
||
four concerns, each with its own owner:
|
||
|
||
| Concern | Question | Owner |
|
||
|---|---|---|
|
||
| Triggering | when do tasks appear? | OpenClaw cron, agents' delayed self-submissions, humans |
|
||
| Dispatch | which task runs next? | each queue's worker (claim by priority under its a(t)) |
|
||
| Admission | may it run now? | LiteLLM budgets/rate + the budget governor |
|
||
| **Time semantics** | expired leases, deadlines, stuck tasks? | **the fabric-keeper** |
|
||
| Failure/anomaly response | something went wrong — now what? | Zabbix → the ops-agent (§6d) |
|
||
|
||
The **fabric-keeper** is one tiny always-on daemon that neither assigns nor
|
||
triggers work. It sweeps expired leases back to queued, enforces task deadlines
|
||
(escalating to the responsible inbox), moves poison tasks to dead-letter, emits
|
||
the human daily digest, and exports queue depths/ages to Zabbix.
|
||
|
||
Its defining property is being **off the critical path**: kill it and work still
|
||
flows (workers keep pulling and running) — only hygiene degrades. That is what
|
||
separates a janitor from a scheduler, which in a push system would stop
|
||
everything.
|
||
|
||
**Deliberately NOT a keeper duty: missed crons.** A cron window missed because
|
||
the host was down is not a task-lifecycle event — it's an *operational fault*,
|
||
and re-firing it would make the keeper a trigger. Instead: schedules that must
|
||
not be missed are **monitored in Zabbix** (a missed run raises a warning like
|
||
any other infra fault), and the response is handled per-incident, not by a
|
||
blanket catch-up policy (§6d).
|
||
|
||
### 6d. Faults are incidents, not keeper chores
|
||
|
||
Anything that "went wrong" — a missed critical cron, a service down, a queue
|
||
backing up, a stale backup — surfaces as a **Zabbix problem**. Zabbix is the
|
||
single place operational faults are detected.
|
||
|
||
Response is **per-incident, not global**: each Zabbix problem class gets its own
|
||
mitigation path, expressed as a task. Later this is automated by a dedicated
|
||
**ops-agent** — a worker that watches Zabbix problems and, per problem class,
|
||
either applies a known mitigation or files a task (to the right agent, or to a
|
||
human inbox) with the incident as context. One blanket "catch-up policy" would
|
||
be exactly the wrong abstraction: a missed briefing, a missed backup and a
|
||
missed consolidation want completely different responses.
|
||
|
||
### 6c. Kanboard is tier-0 now
|
||
|
||
Promoting KB to the fabric's backbone promotes its ops class: **backups on par
|
||
with the vault**, Zabbix monitoring of the service and API, and a defined
|
||
**degraded mode** — if KB is down, Adolf still answers Matrix chat (no fabric
|
||
operations, no task memory), workers park, nothing crashes or data-loses.
|
||
|
||
## 7. Observability — Langfuse (kept), wired for real
|
||
|
||
Decision: keep **Langfuse** (already deployed; best-in-class self-hosted:
|
||
traces + per-token cost + prompt management + evals, MIT). Grafana rejected for
|
||
this role — generic metrics with no LLM semantics (the source of past
|
||
dissatisfaction); Zabbix keeps infra monitoring. To do (it currently receives
|
||
nothing): LiteLLM success/failure callbacks → Langfuse; tag every trace with
|
||
`agent`, `task-id`, `queue`; per-agent cost dashboards; upgrade v2→v3. Every
|
||
completion is traced here — this is where sub-Task granularity lives.
|
||
|
||
## 8. Growing the lab
|
||
|
||
- **More GPUs / remote llama** → new model-agent Cards with `on-demand` a(t)
|
||
(health probe, wake hook, graceful absence). Routing skips absent nodes.
|
||
- **More specialists** → new Cards + scoped tools + own banks. The fabric and
|
||
A2A don't change.
|
||
- **Autonomous research agents** → low-priority loops on always-on local queues,
|
||
escalating (via always-ask, initially) for large-model synthesis; own KB
|
||
project; tainted outputs until promoted.
|
||
|
||
## 9. Migration order
|
||
|
||
1. Registries: model Cards + agent Cards (schema + populate).
|
||
2. First thin worker end-to-end on an always-on local queue.
|
||
3. Quota-gated worker (Kimi park/resume). Claim/lease semantics.
|
||
4. A2A protocol surface (JSON-RPC + Agent Cards) over the fabric.
|
||
5. Cutovers: Hindsight reflect → fabric; consolidation → fabric (low prio);
|
||
claude-coder + Adolf declared as registry agents (Adolf's tools shrink to
|
||
scoped core).
|
||
6. Trust enforcement: capability grants (virtual keys + MCP allowlists), taint
|
||
gate, budget governor, langfuse wiring.
|
||
7. Scale: Torgash, researcher (own KB project), on-demand nodes.
|
||
|
||
## 10. Decision log (2026-07-21, alvis)
|
||
|
||
1. Single completions are not Tasks (langfuse-only) → no micro-churn.
|
||
2. **KB-literal**: Kanboard is the queue, humans included; no separate store.
|
||
3. Trust classes as §5; vault = trusted only.
|
||
4. Escalation = always-ask initially.
|
||
5. Researcher: KB access allowed, own project(s), scope-limited.
|
||
6. Real A2A protocol now (JSON-RPC + Agent Cards).
|
||
7. Proactive schedules: OpenClaw cron → fabric.
|
||
8. Langfuse kept as the observability layer; Grafana rejected; Zabbix = infra.
|
||
9. Executor = thin KB-polling workers; no Hatchet/Temporal at this scale.
|
||
10. Sync routing = LiteLLM Auto Router v2; fabric owns async/parking.
|
||
11. Sandbox = per-task workspace lease + container + PR-only merges.
|
||
|
||
Added in v2.1 (hardening review, same day):
|
||
|
||
12. **Multi-human**: per-human private banks + shared household bank;
|
||
interlocutor-scoped recall; cross-human leakage forbidden (hard rule);
|
||
promotion to shared only by the owning human's action/approval.
|
||
13. A2A auth mandatory everywhere — the LAN is untrusted (VPN peers).
|
||
14. KB = tier-0 infrastructure (backup, monitoring, degraded mode).
|
||
15. Done column = unverified completion; closed task = verified; the producer
|
||
never closes its own task.
|
||
16. Scheduler = decomposed (cron/self-submission triggers; workers dispatch;
|
||
LiteLLM+governor admit); the **fabric-keeper** janitor owns time semantics
|
||
(leases, deadlines, dead-letter, digests) and neither assigns nor triggers.
|
||
17. Personas/Cards live in git, deployed — never edited live.
|
||
18. Backbone/routing changes gated by golden-set shadow eval (Langfuse).
|
||
19. Fabric hygiene: creation quotas, ancestry depth cap, cycle detection,
|
||
dead-letter for poison tasks.
|
||
20. GPU residency policy: local a(t)=f(VRAM); never-evict set (embedder,
|
||
reranker); pre-load checks.
|
||
21. Transport: sync = direct A2A RPC (journaled); async = KB polling; polling
|
||
never on a sync path.
|
||
22. **Missed crons are not a keeper duty**: schedules that must not be missed
|
||
are monitored in **Zabbix**; faults are incidents handled **per problem
|
||
class** by a future **ops-agent** (§6d), never by a blanket catch-up policy.
|