Files
adolf/docs/SPIKE-FINDINGS.md
alvis f214eb2fae
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
docs: add P0 spike findings for the 5 open architecture gates
Resolves ARCHITECTURE.md §4: confirms kimi-code (not kimi-cli) is
already the current upstream package with working -r/-p/stream-json
and portable CLI-home auth; documents OpenClaw's chat_id metadata
block as the session-identity mechanism; verifies headless image
input via the built-in ReadMediaFile tool; recommends Kuzu over
Neo4j for Cognee's graph store; and recommends LiteLLM over the
Kimi CLI for cognee's batch-cognify LLM backend based on measured
per-call latency.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
2026-07-05 09:43:54 +00:00

185 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Adolf — P0 De-Risking Spike Findings
**Date:** 2026-07-05 · **Status:** Resolved — see verdicts below · **Scope:** ARCHITECTURE.md §4 gates 1-5
Method: live inspection of the running `kimi-agent` container/image (Dockerfile, `server.js`,
`docker exec` into it), throwaway containers built from the same image with a **copy** of its
`/root/.kimi-code` volume (created, tested, and fully destroyed — the live `kimi-agent` service was
never stopped, restarted, or modified), decompiled/grepped `@moonshot-ai/kimi-code` bundle source,
and direct reads of the vendored OpenClaw source already in this repo (`src/auto-reply/reply/*`,
`docs/concepts/model-providers.md`) plus upstream web docs/search.
---
## Gate 1 — Kimi CLI package, home dir, MCP path, flags, auth persistence
**Verdict: architecture assumption was backwards. No change needed — current setup is already correct.**
- The running `kimi-agent` container uses `@moonshot-ai/kimi-code@0.22.3` (binary `kimi`), home dir
`/root/.kimi-code`. This **is** the current upstream package — `MoonshotAI/kimi-code` is the
active/maintained CLI; `MoonshotAI/kimi-cli` (home `~/.kimi/`) is the **legacy** predecessor being
phased out. Confirmed both externally (web search: "Kimi CLI is evolving into Kimi Code CLI... the
kimi-cli project will be gradually wound down") and internally — the shipped binary's own `kimi
migrate` command copies data **from** `~/.kimi/` (legacy) **into** `~/.kimi-code/` (current), and its
bundle contains the string `"Old data kept at ~/.kimi/ — kimi-cli still works."` So `~/.kimi/mcp.json`
is the *old* path, not the upstream target.
- **MCP config**: there is no `kimi mcp` subcommand and no `--mcp-config-file` flag in this version
(`kimi --help` confirmed; grepped the bundle for the flag string — absent). MCP servers are
configured via a **project-root `.mcp.json`** (Claude-Code-compatible schema, auto-discovered by
walking up from cwd) or the interactive `/mcp-config` TUI command. Verified empirically: placing
`.mcp.json` in a headless run's working directory does not error and `kimi doctor` validates config
files cleanly from that directory. **Action for `adolf-llm`/`cognee-llm`: drop `shared-mcp.json` as
`.mcp.json` into each session's working directory** rather than relying on a CLI flag that doesn't exist.
- **Resume flags**: `-r`/`--resume` are **hidden, undocumented aliases for `-S`/`--session`** (confirmed
by the CLI's own `session.resume_hint` meta line: `"command":"kimi -r session_..."`, and empirically
by resuming a session with `-r <id>` and getting correct context recall). `-p`, `--output-format
stream-json` both exist and work exactly as `kimi-agent/server.js` already uses them.
- **Auth persistence in a fresh container/volume — confirmed YES.** Copied the live `kimi-agent`'s
`/root/.kimi-code` into a brand-new named volume, mounted it on a throwaway container (different
container, same image), and made a real API call (`kimi -p "..." --output-format stream-json`) —
it authenticated and returned a real completion, and session resume (`-r <id>`) worked correctly
across separate `docker exec` invocations. **Subscription auth is fully portable via the CLI-home
volume** — `adolf-llm`/`cognee-llm` can each get their own volume seeded from (or sharing) the same
OAuth credential files.
- Cleanup: throwaway container, throwaway volume, and the host-side credential copy under `/tmp` were
all deleted after testing. The live `kimi-agent` container was untouched throughout (never
stopped/restarted/execed with anything destructive) and is still `Up` on its original uptime.
---
## Gate 2 — OpenClaw → provider session identity
**Verdict: no header/`user`-field wiring for custom providers — but OpenClaw gives us something
better: a structured, parseable session-identity JSON block embedded directly in every request's
prompt content. Use it as the primary session key; history-hash is not needed as a fallback for the
common case.**
Read directly from the vendored OpenClaw source in this repo
(`src/auto-reply/reply/inbound-meta.ts`) plus `docs/concepts/model-providers.md`:
- **Confirmed no hidden attribution headers or dynamic `user` field for custom `openai-completions`
baseUrl providers.** `docs/concepts/model-providers.md:680`: *"Proxy-style OpenAI-compatible routes
also skip native OpenAI-only request shaping: ... and no hidden OpenClaw attribution headers."*
Attribution headers (`originator`, `version`, `User-Agent`) are attached **only** on verified native
hosts (`api.openai.com`, `chatgpt.com/backend-api`) — `adolf-llm` will not receive them.
- **The actual mechanism**: `buildInboundMetaSystemPrompt()` injects a stable
`schema: "openclaw.inbound_meta.v2"` JSON block (`account_id`, `channel`, `provider`, `surface`,
`chat_type`) into the **system prompt** of every turn — deliberately excluding anything that
changes per-turn, to preserve provider-side prompt-cache prefix stability.
- **The stable per-conversation key we need — `chat_id`** — is emitted by
`buildInboundUserContextPrefix()` into the **user-role** message content instead, specifically as a
`"Conversation info (untrusted metadata):"` JSON code block containing
`chat_id: ctx.OriginatingTo` (plus `message_id`, `sender`, `timestamp`, etc.). `OriginatingTo` is a
channel-agnostic routing id (e.g. `"whatsapp:+1555..."`, `"telegram:chat-1"`; for us it will be the
Matrix room/peer id) and is confirmed present for direct-message channels that aren't `webchat`
(Matrix qualifies) via `shouldIncludeConversationInfo = !isDirect || (channel && channel !== "webchat")`.
This is populated by the core reply pipeline (`get-reply-run.*`) for **every** provider, native or
custom — it's provider-agnostic.
- **Action for `adolf-llm`**: on each incoming request, scan the latest user message content for the
`"Conversation info (untrusted metadata):"` JSON block, parse `chat_id`, and use it directly as the
session-map key → 1:1 Kimi session resume. This is strictly more robust than a history hash (it's
stable even across edited/truncated history) and needs no OpenClaw-side config changes. Keep the
history-hash approach only as a defensive fallback if the block is ever absent (e.g. webchat surface,
or a future OpenClaw version relocates the field — grep on the label string, don't assume a fixed
line offset).
- Escape hatch confirmed but not needed here: `agents.defaults.models["provider/model"].params.extra_body`
can merge static extra JSON into the outbound body for vendor-specific fields, but it's a config-time
value, not a per-request dynamic session id, so it doesn't help with mapping.
---
## Gate 3 — Headless image input via file-path reference
**Verdict: confirmed working end-to-end.** No CLI flag is needed — the agent has a built-in
`ReadMediaFile` tool it invokes autonomously.
Empirical test in a throwaway authed container: prompted `kimi -p "Look at ./test2.png and tell me its
exact pixel dimensions and dominant color."` (a real 64×64 solid-color PNG placed in the cwd). Observed
in the `stream-json` output:
1. `{"role":"assistant","tool_calls":[{"function":{"name":"ReadMediaFile","arguments":"{\"path\":\"./test2.png\"}"}}]}`
2. Tool result: reads the file, reports `Original dimensions: 64x64 pixels`, and returns an
`image_url` content part with a base64 `data:image/png;base64,...` payload wired into the model context.
3. Final assistant answer correctly identified the exact color used (`rgb(200, 30, 30)`).
(A first attempt with a degenerate 1×1 pixel PNG failed at the provider with `400 ... failed to decode
image: invalid or unsupported image format` — that's the upstream vision API rejecting a
near-empty test fixture, not a CLI/tool-path limitation; the 64×64 real PNG round-tripped cleanly.)
**Action for `adolf-llm`**: persist inbound Matrix images to the session's working directory and
reference them by relative path in the prompt text (e.g. "See attached image: `./img_3.jpg`") — the CLI
will autonomously call `ReadMediaFile` on it. No special flag or placeholder syntax
(`[image #N ...]`, which is TUI-paste-only) is required for headless mode.
---
## Gate 4 — Graph store: Kuzu vs Neo4j for Cognee
**Verdict: Kuzu (embedded). Use it; do not stand up Neo4j.**
- **Kuzu is Cognee's own default graph backend** (`GRAPH_DATABASE_PROVIDER=kuzu`), file-based, runs
embedded in-process with no network setup, no extra container, no auth surface, and stores its data
as plain files (fits directly under `/mnt/ssd/dbs/cognee/`, consistent with the rest of Adolf's
storage layout).
- Cognee's own guidance: Kuzu is recommended for local/single-user use; it uses file-based locking and
is explicitly **not** meant for multi-agent/concurrent-process access — that's when Cognee's docs
point to Neo4j instead.
- Adolf is explicitly single-user, single-agent, home-server scale (§1: "auto + tool" memory, one
Matrix bot). Neo4j would add: a whole extra JVM-based service (~0.5-1GB+ RAM), its own
backup/upgrade/security surface, and a network-exposed port — for zero benefit at this scale, since
we don't need Cypher-browser visualization or concurrent multi-writer access.
- **Recommendation**: Kuzu embedded, no separate container. Revisit only if/when a second concurrent
agent needs to write to the same graph, or ad-hoc Cypher-browser graph exploration becomes a real
requirement — neither applies to the current design.
---
## Gate 5 — `cognee-llm` suitability: agentic Kimi CLI vs LiteLLM fallback for batch cognify
**Verdict: default cognee's LLM backend to LiteLLM (fallback plan in ARCHITECTURE.md §3.3); keep
`cognee-llm`/Kimi CLI as an optional low-volume path only. JSON reliability is fine; per-call latency
and subscription/concurrency risk are the real blockers for batch use.**
Empirical test in a throwaway authed container (same one used for gate 1, cleaned up after):
| Call | Wall time | JSON cleanliness |
|---|---|---|
| Trivial prompt (`{"ok":true}`) — process floor | **~5.1 s** | clean |
| Structured entity/relationship extraction, inline text (no tool call) | **~21.9 s** | clean, exact schema match |
| Same extraction via file reference (adds a `Read` tool round-trip) | **~24.0 s** | clean, exact schema match |
Findings:
- **JSON reliability is good** — in both extraction runs the final assistant `content` was strict,
schema-conformant JSON with no prose/markdown fences, when explicitly instructed. This de-risks the
"won't emit clean JSON" half of the concern.
- **Latency is the real problem.** ~5 s is the CLI's fixed per-invocation floor (Node process cold
start, config/credential load, provider round-trip) even for a near-empty completion; a realistic
extraction call runs ~20-25 s. Cognify issues one such call per chunk/entity-extraction step — for a
batch of even a few dozen chunks, serialized wall time reaches many minutes, and true parallelism is
unverified: the Kimi subscription is a single-seat, interactive-coding-oriented plan, and hammering it
with concurrent batch CLI invocations risks rate-limiting or provider-side throttling that this spike
did not (and should not) test at scale — that's a live-account risk, not something to probe casually.
- **Every call is also agentic** (tool-call round trips are possible/likely even for "just extract
JSON" prompts, as seen in the file-reference test), adding non-determinism to latency and requiring
the wrapper to reliably strip tool-call/meta lines from `stream-json` output before parsing the final
JSON payload — extra parsing complexity for zero benefit in a task that doesn't need agentic tool use.
- **Recommendation**: point Cognee's `LLM_API_BASE` at a LiteLLM-routed model (`judge`/local qwen, per
ARCHITECTURE.md §3.3's own stated fallback) as the default for cognify. This is materially faster (no
process-spawn floor, no agent loop, no tool-call risk) and avoids risking the shared subscription
under batch load. Keep the `cognee-llm` Kimi-CLI wrapper buildable/available for low-volume or
experimental use, but do not make it the default path.
---
## Summary table
| Gate | Verdict |
|---|---|
| 1. Kimi CLI package/home/MCP/auth | `@moonshot-ai/kimi-code` + `/root/.kimi-code` **is already current upstream** (kimi-cli/`~/.kimi/` is legacy). No `--mcp-config-file`/`kimi mcp`; use project-root `.mcp.json` instead. `-r`/`-p`/`--output-format stream-json` all confirmed. Subscription auth **persists** across a fresh container given a copied CLI-home volume. |
| 2. OpenClaw session identity | No header/`user`-field wiring for custom providers. Use the `chat_id` (`ctx.OriginatingTo`) embedded in every turn's user-content "Conversation info" JSON block as the session key — parse it from prompt content, not headers. |
| 3. Headless image input | Confirmed working: CLI autonomously calls a built-in `ReadMediaFile` tool on a referenced file path in `-p` mode; no flag/placeholder syntax needed. |
| 4. Graph store | **Kuzu (embedded)** — Cognee's own default, matches single-user/no-extra-service goal. Neo4j not justified at this scale. |
| 5. cognee-llm suitability | JSON output is clean, but ~5-24 s per call plus unverified batch-concurrency limits on a single-seat subscription make it unsuitable as the default. **Default to LiteLLM fallback** for cognify; keep Kimi-CLI path optional. |
All gates are resolved with either direct empirical verification (1, 3, 5) or authoritative source
reading (2, plus corroborating docs for 4). No gate required an unresolved forbidden action.