Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
8.3 KiB
Adolf — Solution Architecture
Status: Proposed · Date: 2026-07-05 · Owner: alvis
Adolf is a self-hosted personal AI assistant on Agap, reachable first over Matrix
(@bot:mtx.alogins.net). It is built on OpenClaw (the multi-channel assistant shell), driven by
the Kimi Code CLI on the existing flat kimi login subscription (no per-token cost), with
Cognee as long-term memory.
1. Goals & decisions
- Base = OpenClaw (
github.com/openclaw/openclaw, Node.js). Provides the multi-channel inbox, session/agent routing, SOUL.md persona, tool/MCP registry, and later WebChat + voice. - Brain = Kimi Code CLI subscription, hybrid. The CLI owns the agent loop, but OpenClaw's tools/MCP are not bypassed — they are exposed to the CLI as a shared MCP layer.
- Memory = Cognee, "auto + tool". Auto-ingest every exchange and inject retrieved memory each
turn, and expose explicit
memory_search/memory_addMCP tools. - Cognee's backend LLM also runs on the Kimi subscription — via its own separate wrapper
(
cognee-llm), because its usage (stateless, structured, batch, no media/streaming) is unlike OpenClaw's. Embeddings stay on LiteLLM (embedder= nomic-embed). - The existing
kimi-agentservice is left untouched (still used for OpenWebUI experiments); we build new, separate wrappers. - First surface = Matrix chat only (voice/pipecat, WebChat, OpenWebUI deferred).
Why the hybrid works (de-risked)
OpenClaw runs its own agent loop and expects a raw LLM backend (provider api: openai-completions
or anthropic-messages, configured with baseUrl/apiKey). The Kimi CLI is itself an agent, so
naively wrapping it would bypass OpenClaw's tools. Confirmed escape hatch: Kimi Code CLI supports
MCP (~/.kimi/mcp.json, headless --mcp-config-file <file>, kimi mcp add/list/remove/auth) plus
image/video input, session resume (-r), headless -p, and --output-format stream-json.
We point the CLI at the same MCP servers OpenClaw uses → the CLI can call OpenClaw/cognee tools while
OpenClaw stays the channel/session/persona/memory layer.
2. Topology & data flow
Matrix ⇄ OpenClaw (Gateway: inbox, session routing, SOUL.md "Adolf")
│ OpenAI /v1/chat/completions (streaming, images) model = "adolf"
▼
adolf-llm (NEW wrapper container, :8010)
• maps OpenClaw session → Kimi session (1:1 resume)
• cognee.search() → inject memories into prompt (auto-retrieve)
• spawn: kimi -r <sess> -p <prompt> --output-format stream-json
--mcp-config-file <shared-mcp.json>
• stream stream-json → SSE deltas (real streaming)
• persist inbound images to session dir, reference by path (media)
• after turn: cognee.add(user+assistant) (auto-ingest)
│ ▲ MCP
▼ │
kimi CLI (agent) ── MCP ──► shared MCP servers:
• cognee-mcp (memory_search / add / cognify)
• openclaw-tools bridge (OpenClaw integrations)
cognee ──► cognee-llm (SEPARATE wrapper, :8011)
• stateless one-shot: kimi -p (fresh disposable session per call)
• NO resume, NO media, NO MCP, NON-streaming, structured/JSON out
• high concurrency for batch cognify
──► embeddings: LiteLLM `embedder` (nomic-embed), NOT the CLI
Reused infra (all in agap_git/openai/docker-compose.yml): Qdrant :6333 (cognee vectors),
LiteLLM :4000 (embeddings + optional observability + fallback LLM), Langfuse, Postgres.
3. Components
3.1 OpenClaw ("Adolf") — the shell
Node 24 (≥22.19). Cloned into this repo. Configuration:
- Model provider (custom):
{ baseUrl: "http://adolf-llm:8010/v1", apiKey: "${ADOLF_KEY}", api: "openai-completions", models: [{ id: "adolf", input: ["text","image"] }] }; default agent modeladolf. - SOUL.md persona "Adolf" (Russian-friendly, concise).
- Matrix channel with bot creds from Vaultwarden (
MATRIX_ADOLF_PASSWORD/MATRIX_ADOLF_TOKEN); route room/peer → the Adolf agent. - MCP registry = cognee-mcp + the openclaw-tools bridge; the same server list feeds
shared-mcp.json(single source of truth).
3.2 adolf-llm — conversational wrapper (:8010)
New agap_git/openai/adolf-llm/ (Dockerfile + server.js), own CLI-home volume + conversations
workspace. OpenAI-compatible GET /v1/models + POST /v1/chat/completions (model adolf).
Reuses proven bits from agap_git/openai/kimi-agent/server.js (session map w/ prune, convTurns/
textOf, runKimi spawn + stream-json parse, resume). Adds:
- Real streaming — parse stream-json assistant deltas → OpenAI
chat.completion.chunkSSE. - Media — persist
image_url/base64 parts to the session dir, reference file paths in the prompt. - Session correspondence — key the map on OpenClaw's session id (not a history hash) → 1:1
kimi -rresume → Moonshot context-cache hits → token savings. Fallback: improved history-hash. - Shared MCP —
--mcp-config-file shared-mcp.jsonso the CLI can call OpenClaw/cognee tools. - Cognee hooks — pre-turn
cognee.search()inject; post-turn asynccognee.add().
3.3 cognee-llm — stateless wrapper (:8011)
New agap_git/openai/cognee-llm/, own disposable workspace + CLI-home volume (same subscription).
Shares the runKimi/stream-json core but with the opposite policy: stateless one-shot
(kimi -p, fresh temp dir per call, no resume), non-streaming, no media, no MCP,
structured/JSON-oriented, low temperature, bounded parallelism for batch cognify. Cognee's
LLM_API_BASE → http://cognee-llm:8011/v1.
⚠️ Risk: cognify issues many structured calls; the agentic CLI adds per-call latency and may not reliably emit clean JSON. Fallback: route cognee's LLM to a LiteLLM model (
judge/local qwen) while keeping Kimi for the assistant. Decide after a batch cognify test.
3.4 Cognee memory service
cognee + cognee-mcp containers. Backends: Qdrant (vectors), graph store = Kuzu (embedded)
or Neo4j (decision), LLM via cognee-llm, embeddings via LiteLLM embedder. Persist under
/mnt/ssd/dbs/cognee/. cognee-mcp exposes memory_search/memory_add/cognify into
shared-mcp.json. Memories scoped per OpenClaw user/session.
3.5 Shared MCP layer ("smart tool integration")
shared-mcp.json is the contract both consumers (OpenClaw + Kimi CLI) use. Cognee-mcp is
off-the-shelf. The openclaw-tools bridge exposes the OpenClaw integrations Adolf needs
(send-to-channel, cron/reminders, browser/nodes) as MCP tools. Start minimal, grow.
4. Open gates to verify early
- Kimi CLI home/MCP path & package — current container uses
@moonshot-ai/kimi-code+/root/.kimi-code; upstream is nowMoonshotAI/kimi-cli+~/.kimi/mcp.json. Confirm package, home dir,--mcp-config-file, and thatkimi loginsubscription auth persists in the new containers' CLI-home volumes. - OpenClaw → provider session identity — does OpenClaw forward a stable session id to a custom
provider (header /
user/metadata)? Decides robust mapping vs. hash fallback. - Headless image input — confirm the CLI accepts images in
-pruns via file-path references. - Graph store — Kuzu embedded vs Neo4j service for cognee.
- Cognee-llm suitability — batch cognify latency/JSON reliability via the CLI vs LiteLLM fallback.
5. End-to-end verification
- Matrix message → OpenClaw →
adolf-llm→kimi→ streamed reply token-by-token in Matrix. - Send an image on Matrix → CLI reads it (describe-the-image test).
- Follow-up → same Kimi session resumed (
sessions.jsonmaps to one session; reduced prompt tokens vs. cold turn). - Memory — a stated fact is recalled in a later session via
memory_search; cognee graph/Qdrant populated (auto-ingest); cognify's LLM served bycognee-llm(or LiteLLM fallback). - MCP tool — a tool call works mid-turn (e.g. set a reminder via the bridge).
- The old
kimi-agentservice is unchanged and still serves OpenWebUI.