Vendor OpenClaw source as Adolf fork baseline
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Some checks failed
ClawSweeper Dispatch / dispatch (push) Has been cancelled
CodeQL / Security High (actions) (push) Has been cancelled
CodeQL / Security High (channel-runtime-boundary) (push) Has been cancelled
CodeQL / Security High (core-auth-secrets) (push) Has been cancelled
CodeQL / Security High (mcp-process-tool-boundary) (push) Has been cancelled
CodeQL / Security High (network-ssrf-boundary) (push) Has been cancelled
CodeQL / Security High (plugin-trust-boundary) (push) Has been cancelled
CodeQL / Security High (process-exec-boundary) (push) Has been cancelled
Docs Sync Publish Repo / sync-publish-repo (push) Has been cancelled
Docs / docs (push) Has been cancelled
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Has been cancelled
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Has been cancelled
Workflow Sanity / no-tabs (push) Has been cancelled
Workflow Sanity / actionlint (push) Has been cancelled
Workflow Sanity / generated-doc-baselines (push) Has been cancelled
CI / runner-admission (push) Has been cancelled
CI / preflight (push) Has been cancelled
CI / security-fast (push) Has been cancelled
CI / pnpm-store-warmup (push) Has been cancelled
CI / build-artifacts (push) Has been cancelled
CI / native-i18n (push) Has been cancelled
CI / ${{ matrix.check_name }} (push) Has been cancelled
CI / ${{ matrix.checkName }} (push) Has been cancelled
CI / checks-node-compat-node22 (push) Has been cancelled
CI / check-bundled-channel-config-metadata (push) Has been cancelled
CI / check-dependencies (push) Has been cancelled
CI / check-guards (push) Has been cancelled
CI / check-lint (push) Has been cancelled
CI / check-prod-types (push) Has been cancelled
CI / check-shrinkwrap (push) Has been cancelled
CI / check-test-types (push) Has been cancelled
CI / check-additional-boundaries-a (push) Has been cancelled
CI / check-additional-boundaries-bcd (push) Has been cancelled
CI / check-additional-extension-bundled (push) Has been cancelled
CI / check-additional-extension-channels (push) Has been cancelled
CI / check-additional-extension-package-boundary (push) Has been cancelled
CI / check-additional-runtime-topology-architecture (push) Has been cancelled
CI / check-session-accessor-boundary (push) Has been cancelled
CI / check-session-transcript-reader-boundary (push) Has been cancelled
CI / check-docs (push) Has been cancelled
CI / skills-python (push) Has been cancelled
CI / macos-swift (push) Has been cancelled
CI / ios-build (push) Has been cancelled
CI / ci-timings-summary (push) Has been cancelled
Native App Locale Refresh / Refresh native fa (push) Has been cancelled
Native App Locale Refresh / Refresh native fr (push) Has been cancelled
Native App Locale Refresh / Refresh native hi (push) Has been cancelled
Native App Locale Refresh / Refresh native id (push) Has been cancelled
Native App Locale Refresh / Refresh native it (push) Has been cancelled
Native App Locale Refresh / Refresh native ja-JP (push) Has been cancelled
Control UI Locale Refresh / plan (push) Has been cancelled
Control UI Locale Refresh / Refresh ${{ matrix.locale }} (push) Has been cancelled
Control UI Locale Refresh / Commit control UI locale refresh (push) Has been cancelled
Live Media Runner Image / Build live media runner image (push) Has been cancelled
Native App Locale Refresh / Refresh native ar (push) Has been cancelled
Native App Locale Refresh / Refresh native de (push) Has been cancelled
Native App Locale Refresh / Refresh native es (push) Has been cancelled
Native App Locale Refresh / Refresh native ko (push) Has been cancelled
Native App Locale Refresh / Refresh native nl (push) Has been cancelled
Native App Locale Refresh / Refresh native pl (push) Has been cancelled
Native App Locale Refresh / Refresh native pt-BR (push) Has been cancelled
Native App Locale Refresh / Refresh native ru (push) Has been cancelled
Native App Locale Refresh / Refresh native sv (push) Has been cancelled
Native App Locale Refresh / Refresh native th (push) Has been cancelled
Native App Locale Refresh / Refresh native tr (push) Has been cancelled
Native App Locale Refresh / Refresh native uk (push) Has been cancelled
Native App Locale Refresh / Refresh native vi (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-CN (push) Has been cancelled
Native App Locale Refresh / Refresh native zh-TW (push) Has been cancelled
Native App Locale Refresh / Commit native locale refresh (push) Has been cancelled
Plugin Init Scaffold Validation / Validate provider scaffold (push) Has been cancelled
Plugin NPM Release / preview_plugins_npm (push) Has been cancelled
Plugin NPM Release / Validate release publish approval (push) Has been cancelled
Plugin NPM Release / preview_plugin_pack (push) Has been cancelled
Plugin NPM Release / publish_plugins_npm (push) Has been cancelled
Sandbox Common Smoke / sandbox-common-smoke (push) Has been cancelled
Website Installer Sync / static (push) Has been cancelled
Website Installer Sync / linux-docker (push) Has been cancelled
Website Installer Sync / macos-installer (push) Has been cancelled
Website Installer Sync / windows-installer (push) Has been cancelled
Website Installer Sync / sync-website (push) Has been cancelled
Adolf is a fork/vendored clone of github.com/openclaw/openclaw (v2026.6.11), free to diverge. Tree copied sans upstream .git; upstream remote added for future syncs. Node pinned to 24 (.nvmrc); engines already require >=22.19. Preserves docs/ARCHITECTURE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LeqyaxJF2nbRXJtae2kNB2
This commit is contained in:
158
docs/providers/alibaba.md
Normal file
158
docs/providers/alibaba.md
Normal file
@@ -0,0 +1,158 @@
|
||||
---
|
||||
summary: "Alibaba Model Studio Wan video generation in OpenClaw"
|
||||
title: "Alibaba Model Studio"
|
||||
read_when:
|
||||
- You want to use Alibaba Wan video generation in OpenClaw
|
||||
- You need Model Studio or DashScope API key setup for video generation
|
||||
---
|
||||
|
||||
The bundled `alibaba` plugin registers a video-generation provider for Wan models on Alibaba Model Studio (the international name for DashScope). It is enabled by default; only an API key is needed.
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | ------------------------------------------------------------------------------- |
|
||||
| Provider id | `alibaba` |
|
||||
| Plugin | bundled, `enabledByDefault: true` |
|
||||
| Auth env vars | `MODELSTUDIO_API_KEY` → `DASHSCOPE_API_KEY` → `QWEN_API_KEY` (first match wins) |
|
||||
| Onboarding flag | `--auth-choice alibaba-model-studio-api-key` |
|
||||
| Direct CLI flag | `--alibaba-model-studio-api-key <key>` |
|
||||
| Default model | `alibaba/wan2.6-t2v` |
|
||||
| Default base URL | `https://dashscope-intl.aliyuncs.com` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set an API key">
|
||||
Store the key against the `alibaba` provider through onboarding:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice alibaba-model-studio-api-key
|
||||
```
|
||||
|
||||
Or pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --alibaba-model-studio-api-key <your-key>
|
||||
```
|
||||
|
||||
Or export one of the accepted env vars before starting the Gateway:
|
||||
|
||||
```bash
|
||||
export MODELSTUDIO_API_KEY=sk-...
|
||||
# or DASHSCOPE_API_KEY=...
|
||||
# or QWEN_API_KEY=...
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Set a default video model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "alibaba/wan2.6-t2v",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the provider is configured">
|
||||
```bash
|
||||
openclaw models list --provider alibaba
|
||||
```
|
||||
|
||||
The list includes all five bundled Wan models. If `MODELSTUDIO_API_KEY` cannot be resolved, `openclaw models status --json` reports the missing credential under `auth.unusableProfiles`.
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Note>
|
||||
The Alibaba plugin and the [Qwen plugin](/providers/qwen) both authenticate against DashScope and accept overlapping env vars. Use `alibaba/...` model ids for the dedicated Wan video surface; use `qwen/...` ids for Qwen chat, embedding, or media-understanding.
|
||||
</Note>
|
||||
|
||||
## Built-in Wan models
|
||||
|
||||
| Model ref | Mode |
|
||||
| -------------------------- | ------------------------- |
|
||||
| `alibaba/wan2.6-t2v` | Text-to-video (default) |
|
||||
| `alibaba/wan2.6-i2v` | Image-to-video |
|
||||
| `alibaba/wan2.6-r2v` | Reference-to-video |
|
||||
| `alibaba/wan2.6-r2v-flash` | Reference-to-video (fast) |
|
||||
| `alibaba/wan2.7-r2v` | Reference-to-video |
|
||||
|
||||
## Capabilities and limits
|
||||
|
||||
All three modes share the same per-request video count and duration cap; only the input shape differs.
|
||||
|
||||
| Mode | Max output videos | Max input images | Max input videos | Max duration | Supported controls |
|
||||
| ------------------ | ----------------- | ---------------- | ---------------- | ------------ | --------------------------------------------------------- |
|
||||
| Text-to-video | 1 | n/a | n/a | 10 s | `size`, `aspectRatio`, `resolution`, `audio`, `watermark` |
|
||||
| Image-to-video | 1 | 1 | n/a | 10 s | `size`, `aspectRatio`, `resolution`, `audio`, `watermark` |
|
||||
| Reference-to-video | 1 | n/a | 4 | 10 s | `size`, `aspectRatio`, `resolution`, `audio`, `watermark` |
|
||||
|
||||
A request that omits `durationSeconds` gets DashScope's accepted default of **5 seconds**. Set `durationSeconds` explicitly on the [video generation tool](/tools/video-generation) to extend up to 10 s.
|
||||
|
||||
<Warning>
|
||||
Reference image and video inputs must be remote `http(s)` URLs; DashScope's reference modes reject local file paths. Upload to object storage first, or use the [media tool](/tools/media-overview) flow that already produces a public URL.
|
||||
</Warning>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Override the DashScope base URL">
|
||||
The provider defaults to the international DashScope endpoint. To target the China-region endpoint:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
alibaba: {
|
||||
baseUrl: "https://dashscope.aliyuncs.com",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
The provider strips trailing slashes before constructing AIGC task URLs.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Auth env priority">
|
||||
OpenClaw resolves the Alibaba API key from environment variables in this order, taking the first non-empty value:
|
||||
|
||||
1. `MODELSTUDIO_API_KEY`
|
||||
2. `DASHSCOPE_API_KEY`
|
||||
3. `QWEN_API_KEY`
|
||||
|
||||
Configured `auth.profiles` entries (set via `openclaw models auth login`) override env-var resolution. See [Auth profiles in the models FAQ](/help/faq-models#auth-profiles-what-they-are-and-how-to-manage-them) for profile rotation, cooldown, and override mechanics.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Relationship to the Qwen plugin">
|
||||
Both bundled plugins talk to DashScope and accept overlapping API keys. Use:
|
||||
|
||||
- `alibaba/wan*.*` ids for the dedicated Wan video provider documented on this page.
|
||||
- `qwen/*` ids for Qwen chat, embedding, and media understanding (see [Qwen](/providers/qwen)).
|
||||
|
||||
Setting `MODELSTUDIO_API_KEY` once authenticates both plugins, since the auth env var list intentionally overlaps; onboarding each plugin separately is not required.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Qwen" href="/providers/qwen" icon="microchip">
|
||||
Qwen chat, embedding, and media-understanding setup on the same DashScope auth.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/config-agents#agent-defaults" icon="gear">
|
||||
Agent defaults and model configuration.
|
||||
</Card>
|
||||
<Card title="Models FAQ" href="/help/faq-models" icon="circle-question">
|
||||
Auth profiles, switching models, and resolving "no profile" errors.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
476
docs/providers/anthropic.md
Normal file
476
docs/providers/anthropic.md
Normal file
@@ -0,0 +1,476 @@
|
||||
---
|
||||
summary: "Use Anthropic Claude via API keys or Claude CLI in OpenClaw"
|
||||
read_when:
|
||||
- You want to use Anthropic models in OpenClaw
|
||||
title: "Anthropic"
|
||||
---
|
||||
|
||||
Anthropic builds the **Claude** model family. OpenClaw supports two auth routes:
|
||||
|
||||
- **API key** - direct Anthropic API access with usage-based billing (`anthropic/*` models)
|
||||
- **Claude CLI** - reuse an existing Claude Code login on the same host
|
||||
|
||||
<Warning>
|
||||
OpenClaw's Claude CLI backend runs the installed Claude Code CLI in
|
||||
non-interactive print mode (`claude -p`). Anthropic's current Claude Code docs
|
||||
describe that mode as Agent SDK/programmatic usage. Anthropic's June 15, 2026
|
||||
support update paused the announced separate Agent SDK billing change: Claude
|
||||
Agent SDK, `claude -p`, and third-party app usage still draw from a signed-in
|
||||
subscription's usage limits, and the previously announced monthly Agent SDK
|
||||
credit is not available while Anthropic revises that plan.
|
||||
|
||||
Interactive Claude Code still draws from the signed-in Claude plan's limits.
|
||||
API key auth is direct pay-as-you-go billing and does not depend on that plan.
|
||||
For long-lived gateway hosts, shared automation, and predictable production
|
||||
spend, use an Anthropic API key.
|
||||
|
||||
Anthropic's current support articles can change this behavior without an
|
||||
OpenClaw release:
|
||||
|
||||
- [Claude Code CLI reference](https://code.claude.com/docs/en/cli-usage)
|
||||
- [Use the Claude Agent SDK with your Claude plan](https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan)
|
||||
- [Use Claude Code with your Pro or Max plan](https://support.claude.com/en/articles/11145838-use-claude-code-with-your-pro-or-max-plan)
|
||||
- [Use Claude Code with your Team or Enterprise plan](https://support.claude.com/en/articles/11845131-using-claude-code-with-your-team-or-enterprise-plan)
|
||||
- [Manage Claude Code costs](https://code.claude.com/docs/en/costs)
|
||||
|
||||
</Warning>
|
||||
|
||||
## Getting started
|
||||
|
||||
<Tabs>
|
||||
<Tab title="API key">
|
||||
**Best for:** standard API access and usage-based billing.
|
||||
|
||||
<Steps>
|
||||
<Step title="Get your API key">
|
||||
Create an API key in the [Anthropic Console](https://console.anthropic.com/).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard
|
||||
# choose: Anthropic API key
|
||||
```
|
||||
|
||||
Or pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --anthropic-api-key "$ANTHROPIC_API_KEY"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider anthropic
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
### Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { ANTHROPIC_API_KEY: "example-anthropic-key-not-real" },
|
||||
agents: { defaults: { model: { primary: "anthropic/claude-opus-4-8" } } },
|
||||
}
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Claude CLI">
|
||||
**Best for:** reusing an existing Claude CLI login without a separate API key.
|
||||
|
||||
<Steps>
|
||||
<Step title="Ensure Claude CLI is installed and logged in">
|
||||
Verify with:
|
||||
|
||||
```bash
|
||||
claude --version
|
||||
```
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard
|
||||
# choose: Claude CLI
|
||||
```
|
||||
|
||||
OpenClaw detects and reuses the existing Claude CLI credentials.
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider anthropic
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Note>
|
||||
Setup and runtime details for the Claude CLI backend are in [CLI Backends](/gateway/cli-backends).
|
||||
</Note>
|
||||
|
||||
<Warning>
|
||||
Claude CLI reuse expects the OpenClaw process to run on the same host as the
|
||||
Claude CLI login. Docker installs can persist a container home and log in to
|
||||
Claude Code there; see
|
||||
[Claude CLI backend in Docker](/install/docker#claude-cli-backend-in-docker).
|
||||
Other container installs such as [Podman](/install/podman) do not mount host
|
||||
`~/.claude` into setup or runtime; use an Anthropic API key there, or choose
|
||||
a provider with OpenClaw-managed OAuth such as
|
||||
[OpenAI Codex](/providers/openai).
|
||||
</Warning>
|
||||
|
||||
### Config example
|
||||
|
||||
Prefer the canonical Anthropic model ref plus a CLI runtime override:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "anthropic/claude-opus-4-8" },
|
||||
models: {
|
||||
"anthropic/claude-opus-4-8": {
|
||||
agentRuntime: { id: "claude-cli" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Legacy `claude-cli/claude-opus-4-7` model refs still work for
|
||||
compatibility, but new config should keep provider/model selection as
|
||||
`anthropic/*` and put the execution backend in provider/model runtime policy.
|
||||
|
||||
### Billing and `claude -p`
|
||||
|
||||
OpenClaw uses Claude Code's non-interactive `claude -p` path for Claude CLI
|
||||
runs. Anthropic currently treats that path as Agent SDK/programmatic usage:
|
||||
|
||||
- Anthropic's June 15, 2026 support update paused the previously announced
|
||||
separate Agent SDK credit plan.
|
||||
- Subscription-plan Claude Agent SDK, `claude -p`, and third-party app usage
|
||||
still draw from the signed-in subscription's usage limits.
|
||||
- The previously announced monthly Agent SDK credit is not available while
|
||||
Anthropic revises that plan.
|
||||
- Console/API-key logins use pay-as-you-go API billing and do not receive
|
||||
the subscription Agent SDK credit.
|
||||
|
||||
See Anthropic's [Agent SDK plan
|
||||
article](https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan)
|
||||
for the pause notice, and the Claude Code plan articles for
|
||||
[Pro/Max](https://support.claude.com/en/articles/11145838-use-claude-code-with-your-pro-or-max-plan)
|
||||
and
|
||||
[Team/Enterprise](https://support.claude.com/en/articles/11845131-use-claude-code-with-your-team-or-enterprise-plan)
|
||||
subscription behavior.
|
||||
|
||||
Anthropic can change Claude Code billing and rate-limit behavior without an
|
||||
OpenClaw release. Check `claude auth status`, `/status`, and
|
||||
Anthropic's linked docs when billing predictability matters.
|
||||
|
||||
<Tip>
|
||||
For shared production automation, use an Anthropic API key instead of
|
||||
Claude CLI. OpenClaw also supports subscription-style options from
|
||||
[OpenAI Codex](/providers/openai), [Qwen Cloud](/providers/qwen),
|
||||
[MiniMax](/providers/minimax), and [Z.AI / GLM](/providers/zai).
|
||||
</Tip>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Thinking defaults (Claude Fable 5, 4.8, and 4.6)
|
||||
|
||||
`anthropic/claude-fable-5` always uses adaptive thinking and defaults to `high`
|
||||
effort. Anthropic does not allow thinking to be disabled for this model, so
|
||||
`/think off` and `/think minimal` map to `low` effort instead. OpenClaw also
|
||||
omits custom temperature values for Fable 5 requests, since Anthropic rejects
|
||||
a temperature override on any thinking-enabled request.
|
||||
|
||||
Claude Opus 4.8 keeps thinking off by default in OpenClaw. When you explicitly
|
||||
enable adaptive thinking with `/think high|xhigh|max`, OpenClaw sends
|
||||
Anthropic's Opus 4.8 effort values; Claude 4.6 models (Opus 4.6 and Sonnet 4.6)
|
||||
default to `adaptive`.
|
||||
|
||||
Override per-message with `/think:<level>` or in model params:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"anthropic/claude-opus-4-8": {
|
||||
params: { thinking: "high" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
Related Anthropic docs:
|
||||
- [Adaptive thinking](https://platform.claude.com/docs/en/build-with-claude/adaptive-thinking)
|
||||
- [Extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking)
|
||||
|
||||
</Note>
|
||||
|
||||
## Safety refusal fallback (Claude Fable 5)
|
||||
|
||||
<Warning>
|
||||
Using Claude Fable 5 means also using Claude Opus 4.8. Fable 5 ships with
|
||||
safety classifiers that can decline a request, and Anthropic's sanctioned
|
||||
recovery is to have `claude-opus-4-8` serve that turn. OpenClaw opts into this
|
||||
automatically for direct API-key requests, so some Fable turns are answered
|
||||
and billed as Claude Opus 4.8. If your policy or budget cannot accept
|
||||
Opus-served turns, do not select `anthropic/claude-fable-5`.
|
||||
</Warning>
|
||||
|
||||
### Why this exists
|
||||
|
||||
Fable 5 classifiers return `stop_reason: "refusal"` on requests in restricted
|
||||
domains, and they also false-positive on benign-adjacent work (security
|
||||
tooling, life sciences, or even asking the model to reproduce its raw
|
||||
reasoning). Without a fallback, the turn dies with an error even though
|
||||
another Claude model would happily serve it - Anthropic's own refusal message
|
||||
tells API integrators to configure a fallback model.
|
||||
|
||||
### How it works
|
||||
|
||||
1. For every direct API-key request to `anthropic/claude-fable-5`, OpenClaw
|
||||
sends Anthropic's server-side fallback opt-in: the
|
||||
`server-side-fallback-2026-06-01` beta header plus
|
||||
`fallbacks: [{"model": "claude-opus-4-8"}]`. Claude Opus 4.8 is the only
|
||||
fallback target Anthropic permits for Fable 5.
|
||||
2. Only a safety-classifier decline triggers the fallback. Rate limits,
|
||||
overloads, and server errors behave exactly as before and go through
|
||||
OpenClaw's normal [model failover](/concepts/model-failover).
|
||||
3. The rescue happens inside the same call. A decline before any output is
|
||||
invisible apart from latency; the whole answer comes from Opus 4.8. On a
|
||||
mid-stream decline the partial text is kept as the prefix the fallback
|
||||
model continues from, while the declined model's reasoning and tool calls
|
||||
are discarded per Anthropic's replay rules (they must not be echoed back or
|
||||
executed).
|
||||
4. If Claude Opus 4.8 declines as well, the turn surfaces the refusal as an
|
||||
error, exactly like before this feature.
|
||||
|
||||
The fallback happens at the Anthropic API level, so `claude-opus-4-8` does not
|
||||
need to be in your configured model list or fallback chain - a Fable-capable
|
||||
API key can always serve Opus.
|
||||
|
||||
### Observability and billing
|
||||
|
||||
- A fallback-served turn records a `provider_fallback` diagnostic on the
|
||||
assistant message naming `fromModel` and `toModel`, and the message's
|
||||
`responseModel` reports `claude-opus-4-8`.
|
||||
- Anthropic bills per attempt: a decline before output is free, and the rescue
|
||||
bills at Claude Opus 4.8 rates (currently half of Fable 5 rates). OpenClaw's
|
||||
per-turn cost estimate prices fallback-served turns at Opus rates to match.
|
||||
- A mid-stream decline additionally bills the already-streamed Fable partial
|
||||
on Anthropic's side; that portion is reported in the API's per-attempt
|
||||
usage but not folded into OpenClaw's per-turn estimate.
|
||||
|
||||
### Scope
|
||||
|
||||
Applies to `anthropic/claude-fable-5` with API-key auth against
|
||||
`api.anthropic.com`. OAuth (Claude CLI subscription reuse), proxy base URLs,
|
||||
Bedrock, Vertex, and Foundry requests are unchanged and still surface
|
||||
refusals as errors there.
|
||||
|
||||
Verified live: a benign prompt asking Fable 5 to reproduce its raw chain of
|
||||
thought is declined with `category: "reasoning_extraction"` when sent without
|
||||
fallbacks, and the same prompt through OpenClaw returns a normal Opus-served
|
||||
answer with the `provider_fallback` diagnostic attached.
|
||||
|
||||
See Anthropic's [refusals and fallback
|
||||
guide](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback)
|
||||
for the underlying behavior.
|
||||
|
||||
## Prompt caching
|
||||
|
||||
OpenClaw supports Anthropic's prompt caching feature for API-key auth.
|
||||
|
||||
| Value | Cache duration | Description |
|
||||
| ------------------- | -------------- | -------------------------------------- |
|
||||
| `"short"` (default) | 5 minutes | Applied automatically for API-key auth |
|
||||
| `"long"` | 1 hour | Extended cache |
|
||||
| `"none"` | No caching | Disable prompt caching |
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"anthropic/claude-opus-4-6": {
|
||||
params: { cacheRetention: "long" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Per-agent cache overrides">
|
||||
Use model-level params as your baseline, then override specific agents via `agents.list[].params`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "anthropic/claude-opus-4-6" },
|
||||
models: {
|
||||
"anthropic/claude-opus-4-6": {
|
||||
params: { cacheRetention: "long" },
|
||||
},
|
||||
},
|
||||
},
|
||||
list: [
|
||||
{ id: "research", default: true },
|
||||
{ id: "alerts", params: { cacheRetention: "none" } },
|
||||
],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Config merge order:
|
||||
|
||||
1. `agents.defaults.models["provider/model"].params`
|
||||
2. `agents.list[].params` (matching `id`, overrides by key)
|
||||
|
||||
This lets one agent keep a long-lived cache while another agent on the same model disables caching for bursty/low-reuse traffic.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Bedrock Claude notes">
|
||||
- Anthropic Claude models on Bedrock (`amazon-bedrock/*anthropic.claude*`) accept `cacheRetention` pass-through when configured.
|
||||
- Non-Anthropic Bedrock models are forced to `cacheRetention: "none"` at runtime.
|
||||
- API-key smart defaults also seed `cacheRetention: "short"` for Claude-on-Bedrock refs when no explicit value is set.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Fast mode">
|
||||
OpenClaw's shared `/fast` toggle sets Anthropic's `service_tier` field for direct API-key traffic to `api.anthropic.com`.
|
||||
|
||||
| Command | Maps to |
|
||||
|---------|---------|
|
||||
| `/fast on` | `service_tier: "auto"` |
|
||||
| `/fast off` | `service_tier: "standard_only"` |
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"anthropic/claude-sonnet-4-6": {
|
||||
params: { fastMode: true },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
- Only applies to direct `api.anthropic.com` requests made with an API key. OAuth/subscription-token requests and proxy routes never get a `service_tier` field.
|
||||
- Explicit `serviceTier` or `service_tier` params override `/fast` when both are set.
|
||||
- On accounts without Priority Tier capacity, `service_tier: "auto"` may resolve to `standard`.
|
||||
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Media understanding (image and PDF)">
|
||||
The bundled Anthropic plugin registers image and PDF understanding. OpenClaw
|
||||
auto-resolves media capabilities from the configured Anthropic auth; no
|
||||
additional config is needed.
|
||||
|
||||
| Property | Value |
|
||||
| --------------- | --------------------- |
|
||||
| Default model | `claude-opus-4-8` |
|
||||
| Supported input | Images, PDF documents |
|
||||
|
||||
When an image or PDF is attached to a conversation, OpenClaw automatically
|
||||
routes it through the Anthropic media understanding provider.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="1M context window">
|
||||
Anthropic's 1M context window is GA on Claude 4.x models with adaptive
|
||||
thinking: Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6. OpenClaw sizes those
|
||||
models at 1,048,576 tokens automatically, no `params.context1m` needed:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"anthropic/claude-opus-4-6": {},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Older configs can keep `params.context1m: true`; it is a harmless no-op for
|
||||
these models and OpenClaw no longer sends the retired
|
||||
`context-1m-2025-08-07` beta header regardless. Older `anthropicBeta` config
|
||||
entries with that value are dropped during request header resolution, and
|
||||
unsupported older Claude models stay on their normal context window.
|
||||
|
||||
`params.context1m: true` behaves the same way for the Claude CLI backend
|
||||
(`claude-cli/*`): eligible GA-capable Opus and Sonnet models already get the
|
||||
1M window automatically, so the param is optional there too.
|
||||
|
||||
<Warning>
|
||||
Requires long-context access on your Anthropic credential. OAuth/subscription token auth keeps its required Anthropic beta headers, but OpenClaw strips the retired 1M beta header if it remains in older config.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Claude Opus 4.8 1M context">
|
||||
`anthropic/claude-opus-4-8` and its `claude-cli` variant have a 1M context
|
||||
window by default; no `params.context1m: true` needed.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="401 errors / token suddenly invalid">
|
||||
Anthropic token auth expires and can be revoked. For new setups, use an Anthropic API key instead.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title='No API key found for provider "anthropic"'>
|
||||
Anthropic auth is **per agent**; new agents do not inherit the main agent's keys. Re-run onboarding for that agent (or configure an API key on the gateway host), then verify with `openclaw models status`.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title='No credentials found for profile "anthropic:default"'>
|
||||
Run `openclaw models status` to see which auth profile is active. Re-run onboarding, or configure an API key for that profile path.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="No available auth profile (all in cooldown)">
|
||||
Check `openclaw models status --json` for `auth.unusableProfiles`. Anthropic rate-limit cooldowns can be model-scoped, so a sibling Anthropic model may still be usable. Add another Anthropic profile or wait for cooldown.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Note>
|
||||
More help: [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq).
|
||||
</Note>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="CLI backends" href="/gateway/cli-backends" icon="terminal">
|
||||
Claude CLI backend setup and runtime details.
|
||||
</Card>
|
||||
<Card title="Prompt caching" href="/reference/prompt-caching" icon="database">
|
||||
How prompt caching works across providers.
|
||||
</Card>
|
||||
<Card title="OAuth and auth" href="/gateway/authentication" icon="key">
|
||||
Auth details and credential reuse rules.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
149
docs/providers/arcee.md
Normal file
149
docs/providers/arcee.md
Normal file
@@ -0,0 +1,149 @@
|
||||
---
|
||||
summary: "Arcee AI setup (auth + model selection)"
|
||||
title: "Arcee AI"
|
||||
read_when:
|
||||
- You want to use Arcee AI with OpenClaw
|
||||
- You need the API key env var or CLI auth choice
|
||||
---
|
||||
|
||||
[Arcee AI](https://arcee.ai) provides the Trinity family of mixture-of-experts models through an OpenAI-compatible API. All Trinity models are Apache 2.0 licensed. Arcee is an official OpenClaw plugin, not bundled with core, so it needs an install step before onboarding.
|
||||
|
||||
Access Arcee models directly through the Arcee platform or through [OpenRouter](/providers/openrouter).
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ------------------------------------------------------------------------------------- |
|
||||
| Provider | `arcee` |
|
||||
| Auth | `ARCEEAI_API_KEY` (direct) or `OPENROUTER_API_KEY` (via OpenRouter) |
|
||||
| API | OpenAI-compatible |
|
||||
| Base URL | `https://api.arcee.ai/api/v1` (direct) or `https://openrouter.ai/api/v1` (OpenRouter) |
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/arcee-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Direct (Arcee platform)">
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Create an API key at [Arcee AI](https://chat.arcee.ai/).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice arceeai-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "arcee/trinity-large-thinking" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
|
||||
<Tab title="Via OpenRouter">
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Create an API key at [OpenRouter](https://openrouter.ai/keys).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice arceeai-openrouter
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "arcee/trinity-large-thinking" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
The same model refs work for both direct and OpenRouter setups.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Non-interactive setup
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Direct (Arcee platform)">
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice arceeai-api-key \
|
||||
--arceeai-api-key "$ARCEEAI_API_KEY"
|
||||
```
|
||||
</Tab>
|
||||
|
||||
<Tab title="Via OpenRouter">
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice arceeai-openrouter \
|
||||
--openrouter-api-key "$OPENROUTER_API_KEY"
|
||||
```
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
| Model ref | Name | Input | Context | Max output | Cost (in/out per 1M) | Tools | Notes |
|
||||
| ------------------------------ | ---------------------- | ----- | ------- | ---------- | -------------------- | ----- | ----------------------------------------- |
|
||||
| `arcee/trinity-large-thinking` | Trinity Large Thinking | text | 256K | 80K | $0.25 / $0.90 | No | Default model; extended thinking |
|
||||
| `arcee/trinity-large-preview` | Trinity Large Preview | text | 128K | 16K | $0.25 / $1.00 | Yes | General-purpose; 400B params, 13B active |
|
||||
| `arcee/trinity-mini` | Trinity Mini 26B | text | 128K | 80K | $0.045 / $0.15 | Yes | Fast and cost-efficient; function calling |
|
||||
|
||||
<Tip>
|
||||
The onboarding preset sets `arcee/trinity-large-thinking` as the default model.
|
||||
</Tip>
|
||||
|
||||
## Supported features
|
||||
|
||||
| Feature | Supported |
|
||||
| --------------------------------------------- | -------------------------------------------- |
|
||||
| Streaming | Yes |
|
||||
| Tool use / function calling | Yes (Trinity Mini, Trinity Large Preview) |
|
||||
| Structured output (JSON mode and JSON schema) | Yes |
|
||||
| Extended thinking | Yes (Trinity Large Thinking; tools disabled) |
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Environment note">
|
||||
If the Gateway runs as a daemon (launchd/systemd), make sure `ARCEEAI_API_KEY`
|
||||
(or `OPENROUTER_API_KEY`) is available to that process, for example in
|
||||
`~/.openclaw/.env` or via `env.shellEnv`.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="OpenRouter routing">
|
||||
When using Arcee models via OpenRouter, the same `arcee/*` model refs apply.
|
||||
OpenClaw routes transparently based on your auth choice. See the
|
||||
[OpenRouter provider docs](/providers/openrouter) for OpenRouter-specific
|
||||
configuration details.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="OpenRouter" href="/providers/openrouter" icon="shuffle">
|
||||
Access Arcee models and many others through a single API key.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
128
docs/providers/azure-speech.md
Normal file
128
docs/providers/azure-speech.md
Normal file
@@ -0,0 +1,128 @@
|
||||
---
|
||||
summary: "Azure AI Speech text-to-speech for OpenClaw replies"
|
||||
read_when:
|
||||
- You want Azure Speech synthesis for outbound replies
|
||||
- You need native Ogg Opus voice-note output from Azure Speech
|
||||
title: "Azure Speech"
|
||||
---
|
||||
|
||||
Azure Speech is a bundled Azure AI Speech text-to-speech provider. OpenClaw
|
||||
calls the Azure Speech REST API directly with SSML, synthesizing MP3 for
|
||||
standard replies, native Ogg/Opus for voice notes, and 8 kHz mulaw for
|
||||
telephony channels such as Voice Call. The request sends the provider-owned
|
||||
output format through the `X-Microsoft-OutputFormat` header.
|
||||
|
||||
| Detail | Value |
|
||||
| ----------------------- | -------------------------------------------------------------------------------------------------------------- |
|
||||
| Provider ID | `azure-speech` (alias: `azure`) |
|
||||
| Website | [Azure AI Speech](https://azure.microsoft.com/products/ai-services/ai-speech) |
|
||||
| Docs | [Speech REST text-to-speech](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech) |
|
||||
| Auth | `AZURE_SPEECH_KEY` plus `AZURE_SPEECH_REGION` |
|
||||
| Default voice | `en-US-JennyNeural` |
|
||||
| Default file output | `audio-24khz-48kbitrate-mono-mp3` |
|
||||
| Default voice-note file | `ogg-24khz-16bit-mono-opus` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Create an Azure Speech resource">
|
||||
In the Azure portal, create a Speech resource. Copy **KEY 1** from
|
||||
Resource Management > Keys and Endpoint, and copy the resource location
|
||||
such as `eastus`.
|
||||
|
||||
```
|
||||
AZURE_SPEECH_KEY=<speech-resource-key>
|
||||
AZURE_SPEECH_REGION=eastus
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Select Azure Speech in messages.tts">
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "azure-speech",
|
||||
providers: {
|
||||
"azure-speech": {
|
||||
voice: "en-US-JennyNeural",
|
||||
lang: "en-US",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Send a message">
|
||||
Send a reply through any connected channel. OpenClaw synthesizes the audio
|
||||
with Azure Speech and delivers MP3 for standard audio, or Ogg/Opus when
|
||||
the channel expects a voice note.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Configuration options
|
||||
|
||||
All options live under `messages.tts.providers["azure-speech"]`.
|
||||
|
||||
| Option | Description |
|
||||
| ----------------------- | ----------------------------------------------------------------------------------------------------- |
|
||||
| `apiKey` | Azure Speech resource key. Falls back to `AZURE_SPEECH_KEY`, `AZURE_SPEECH_API_KEY`, or `SPEECH_KEY`. |
|
||||
| `region` | Azure Speech resource region. Falls back to `AZURE_SPEECH_REGION` or `SPEECH_REGION`. |
|
||||
| `endpoint` | Optional Azure Speech endpoint override. Falls back to `AZURE_SPEECH_ENDPOINT`. |
|
||||
| `baseUrl` | Optional Azure Speech base URL override. |
|
||||
| `voice` | Azure voice ShortName (default `en-US-JennyNeural`). Legacy alias: `voiceId`. |
|
||||
| `lang` | SSML language code (default `en-US`). |
|
||||
| `outputFormat` | Audio-file output format (default `audio-24khz-48kbitrate-mono-mp3`). |
|
||||
| `voiceNoteOutputFormat` | Voice-note output format (default `ogg-24khz-16bit-mono-opus`). |
|
||||
| `timeoutMs` | Request timeout override in milliseconds. Falls back to the global `messages.tts.timeoutMs`. |
|
||||
|
||||
The provider is considered configured once `apiKey` is set plus one of
|
||||
`region`, `endpoint`, or `baseUrl`. Env vars are only checked as a fallback
|
||||
for config keys left unset.
|
||||
|
||||
## Notes
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Authentication">
|
||||
Azure Speech uses a Speech resource key, not an Azure OpenAI key. The key
|
||||
is sent as `Ocp-Apim-Subscription-Key`; OpenClaw derives
|
||||
`https://<region>.tts.speech.microsoft.com` from `region` unless you
|
||||
provide `endpoint` or `baseUrl`.
|
||||
</Accordion>
|
||||
<Accordion title="Voice names">
|
||||
Use the Azure Speech voice `ShortName` value, for example
|
||||
`en-US-JennyNeural`. The bundled provider can list voices through the
|
||||
same Speech resource and filters out voices marked deprecated, retired,
|
||||
or disabled.
|
||||
</Accordion>
|
||||
<Accordion title="Audio outputs">
|
||||
Azure accepts output formats such as `audio-24khz-48kbitrate-mono-mp3`,
|
||||
`ogg-24khz-16bit-mono-opus`, and `riff-24khz-16bit-mono-pcm`. OpenClaw
|
||||
requests Ogg/Opus for `voice-note` targets so channels can send native
|
||||
voice bubbles without an extra MP3 conversion, and forces
|
||||
`raw-8khz-8bit-mono-mulaw` for telephony targets.
|
||||
</Accordion>
|
||||
<Accordion title="Alias">
|
||||
`azure` is accepted as a provider alias for existing config, but new
|
||||
config should use `azure-speech` to avoid confusion with Azure OpenAI
|
||||
model providers.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Text-to-speech" href="/tools/tts" icon="waveform-lines">
|
||||
TTS overview, providers, and `messages.tts` config.
|
||||
</Card>
|
||||
<Card title="Configuration" href="/gateway/configuration" icon="gear">
|
||||
Full config reference including `messages.tts` settings.
|
||||
</Card>
|
||||
<Card title="Providers" href="/providers" icon="grid">
|
||||
All bundled OpenClaw providers.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
Common issues and debugging steps.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
208
docs/providers/bedrock-mantle.md
Normal file
208
docs/providers/bedrock-mantle.md
Normal file
@@ -0,0 +1,208 @@
|
||||
---
|
||||
summary: "Use Amazon Bedrock Mantle (OpenAI-compatible) models with OpenClaw"
|
||||
read_when:
|
||||
- You want to use Bedrock Mantle hosted OSS models with OpenClaw
|
||||
- You need the Mantle OpenAI-compatible endpoint for GPT-OSS, Qwen, Kimi, or GLM
|
||||
title: "Amazon Bedrock Mantle"
|
||||
---
|
||||
|
||||
OpenClaw includes a bundled **Amazon Bedrock Mantle** provider that connects to
|
||||
the Mantle OpenAI-compatible endpoint. Mantle hosts open-source and
|
||||
third-party models (GPT-OSS, Qwen, Kimi, GLM, and similar) through a standard
|
||||
`/v1/chat/completions` surface backed by Bedrock infrastructure. Mantle also
|
||||
exposes two Anthropic Claude models through an Anthropic Messages route.
|
||||
|
||||
| Property | Value |
|
||||
| -------------- | ---------------------------------------------------------------------------------------------- |
|
||||
| Provider ID | `amazon-bedrock-mantle` |
|
||||
| API | `openai-completions` for discovered OSS models, `anthropic-messages` for the two Claude models |
|
||||
| Auth | Explicit `AWS_BEARER_TOKEN_BEDROCK` or IAM credential-chain bearer-token generation |
|
||||
| Default region | `us-east-1` (override with `AWS_REGION` or `AWS_DEFAULT_REGION`) |
|
||||
|
||||
## Getting started
|
||||
|
||||
Choose your preferred auth method and follow the setup steps.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Explicit bearer token">
|
||||
**Best for:** environments where you already have a Mantle bearer token.
|
||||
|
||||
<Steps>
|
||||
<Step title="Set the bearer token on the gateway host">
|
||||
```bash
|
||||
export AWS_BEARER_TOKEN_BEDROCK="..."
|
||||
```
|
||||
|
||||
Optionally set a region (defaults to `us-east-1`):
|
||||
|
||||
```bash
|
||||
export AWS_REGION="us-west-2"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are discovered">
|
||||
```bash
|
||||
openclaw models list
|
||||
```
|
||||
|
||||
Discovered models appear under the `amazon-bedrock-mantle` provider. No
|
||||
additional config is required unless you want to override defaults.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="IAM credentials">
|
||||
**Best for:** using AWS SDK-compatible credentials (shared config, SSO, web identity, instance or task roles).
|
||||
|
||||
<Steps>
|
||||
<Step title="Configure AWS credentials on the gateway host">
|
||||
Any AWS SDK-compatible auth source works:
|
||||
|
||||
```bash
|
||||
export AWS_PROFILE="default"
|
||||
export AWS_REGION="us-west-2"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are discovered">
|
||||
```bash
|
||||
openclaw models list
|
||||
```
|
||||
|
||||
OpenClaw generates a Mantle bearer token from the credential chain automatically.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
When `AWS_BEARER_TOKEN_BEDROCK` is not set, OpenClaw mints the bearer token for you from the AWS default credential chain, including shared credentials/config profiles, SSO, web identity, and instance or task roles.
|
||||
</Tip>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Automatic model discovery
|
||||
|
||||
When `AWS_BEARER_TOKEN_BEDROCK` is set, OpenClaw uses it directly. Otherwise,
|
||||
OpenClaw attempts to generate a Mantle bearer token from the AWS default
|
||||
credential chain. It then discovers available Mantle models by querying the
|
||||
region's `/v1/models` endpoint.
|
||||
|
||||
| Behavior | Detail |
|
||||
| ----------------- | ------------------------------------------------------------------------------------ |
|
||||
| Discovery cache | Results cached for 1 hour per region; a fetch failure returns the last cached result |
|
||||
| IAM token refresh | Every 2 hours, cached per region |
|
||||
|
||||
To keep the Mantle plugin enabled but suppress automatic discovery and IAM
|
||||
bearer-token generation, disable the plugin-owned discovery toggle:
|
||||
|
||||
```bash
|
||||
openclaw config set plugins.entries.amazon-bedrock-mantle.config.discovery.enabled false
|
||||
```
|
||||
|
||||
<Note>
|
||||
The bearer token is the same `AWS_BEARER_TOKEN_BEDROCK` used by the standard [Amazon Bedrock](/providers/bedrock) provider.
|
||||
</Note>
|
||||
|
||||
### Supported regions
|
||||
|
||||
`us-east-1`, `us-east-2`, `us-west-2`, `ap-northeast-1`,
|
||||
`ap-south-1`, `ap-southeast-3`, `eu-central-1`, `eu-west-1`, `eu-west-2`,
|
||||
`eu-south-1`, `eu-north-1`, `sa-east-1`.
|
||||
|
||||
## Manual configuration
|
||||
|
||||
If you prefer explicit config instead of auto-discovery:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
"amazon-bedrock-mantle": {
|
||||
baseUrl: "https://bedrock-mantle.us-east-1.api.aws/v1",
|
||||
api: "openai-completions",
|
||||
auth: "api-key",
|
||||
apiKey: "env:AWS_BEARER_TOKEN_BEDROCK",
|
||||
models: [
|
||||
{
|
||||
id: "gpt-oss-120b",
|
||||
name: "GPT-OSS 120B",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 32000,
|
||||
maxTokens: 4096,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Reasoning support">
|
||||
Reasoning support is inferred from model IDs containing patterns like
|
||||
`thinking`, `reasoner`, `reasoning`, `deepseek.r`, `gpt-oss-120b`, or
|
||||
`gpt-oss-safeguard-120b`. OpenClaw sets `reasoning: true` automatically for
|
||||
matching models during discovery.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Endpoint unavailability">
|
||||
If the Mantle endpoint is unavailable, returns no models, or bearer-token
|
||||
resolution fails, discovery returns an empty result and the implicit
|
||||
provider is skipped. OpenClaw does not error; other configured providers
|
||||
continue to work normally.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Claude Opus 4.7 and Claude Mythos Preview via the Anthropic Messages route">
|
||||
OpenClaw always appends two Claude models to the Mantle catalog after
|
||||
successful discovery, regardless of what `/v1/models` returns:
|
||||
`amazon-bedrock-mantle/anthropic.claude-opus-4-7` (Claude Opus 4.7) and
|
||||
`amazon-bedrock-mantle/anthropic.claude-mythos-preview` (Claude Mythos
|
||||
Preview). Both use the `anthropic-messages` API surface and stream through
|
||||
the same bearer-authenticated Anthropic-compatible endpoint
|
||||
(`<mantle-base>/anthropic`), so the AWS bearer token is not treated like an
|
||||
Anthropic API key.
|
||||
|
||||
Claude Mythos Preview always requests reasoning, defaulting to `high`
|
||||
effort when no `/think` level is set (mapped from `xhigh`/`max` down to
|
||||
`high`, and `minimal` up to `low`). Opus 4.7 on Mantle streams without
|
||||
model-provided reasoning, and OpenClaw omits its `temperature` parameter
|
||||
since Opus 4.7 does not accept sampling overrides on this route; Mythos
|
||||
Preview accepts a `temperature` override normally.
|
||||
|
||||
These two models are not configurable through `models.providers["amazon-bedrock-mantle"].models`
|
||||
entries — they are always added by discovery when it succeeds, and are
|
||||
only removed by disabling discovery entirely.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Relationship to Amazon Bedrock provider">
|
||||
Bedrock Mantle is a separate provider from the standard
|
||||
[Amazon Bedrock](/providers/bedrock) provider. Mantle uses an
|
||||
OpenAI-compatible `/v1` surface for its OSS catalog, while the standard
|
||||
Bedrock provider uses the native Bedrock Converse API.
|
||||
|
||||
Both providers share the same `AWS_BEARER_TOKEN_BEDROCK` credential when
|
||||
present.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Amazon Bedrock" href="/providers/bedrock" icon="cloud">
|
||||
Native Bedrock provider for Anthropic Claude, Titan, and other models.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="OAuth and auth" href="/gateway/authentication" icon="key">
|
||||
Auth details and credential reuse rules.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
Common issues and how to resolve them.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
452
docs/providers/bedrock.md
Normal file
452
docs/providers/bedrock.md
Normal file
@@ -0,0 +1,452 @@
|
||||
---
|
||||
summary: "Use Amazon Bedrock (Converse API) models with OpenClaw"
|
||||
read_when:
|
||||
- You want to use Amazon Bedrock models with OpenClaw
|
||||
- You need AWS credential/region setup for model calls
|
||||
title: "Amazon Bedrock"
|
||||
---
|
||||
|
||||
OpenClaw can use **Amazon Bedrock** models via its **Bedrock Converse**
|
||||
streaming provider. Bedrock auth uses the **AWS SDK default credential chain**,
|
||||
not an API key.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ----------------------------------------------------------- |
|
||||
| Provider | `amazon-bedrock` |
|
||||
| API | `bedrock-converse-stream` |
|
||||
| Auth | AWS credentials (env vars, shared config, or instance role) |
|
||||
| Region | `AWS_REGION` or `AWS_DEFAULT_REGION` (default: `us-east-1`) |
|
||||
|
||||
## Getting started
|
||||
|
||||
Choose your preferred auth method and follow the setup steps.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Access keys / env vars">
|
||||
**Best for:** developer machines, CI, or hosts where you manage AWS credentials directly.
|
||||
|
||||
<Steps>
|
||||
<Step title="Set AWS credentials on the gateway host">
|
||||
```bash
|
||||
export AWS_ACCESS_KEY_ID="EXAMPLE_AWS_ACCESS_KEY_ID"
|
||||
export AWS_SECRET_ACCESS_KEY="..."
|
||||
export AWS_REGION="us-east-1"
|
||||
# Optional:
|
||||
export AWS_SESSION_TOKEN="..."
|
||||
export AWS_PROFILE="your-profile"
|
||||
# Optional (Bedrock API key/bearer token):
|
||||
export AWS_BEARER_TOKEN_BEDROCK="..."
|
||||
```
|
||||
</Step>
|
||||
<Step title="Add a Bedrock provider and model to your config">
|
||||
No `apiKey` is required. Configure the provider with `auth: "aws-sdk"`:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
"amazon-bedrock": {
|
||||
baseUrl: "https://bedrock-runtime.us-east-1.amazonaws.com",
|
||||
api: "bedrock-converse-stream",
|
||||
auth: "aws-sdk",
|
||||
models: [
|
||||
{
|
||||
id: "us.anthropic.claude-opus-4-6-v1:0",
|
||||
name: "Claude Opus 4.6 (Bedrock)",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 200000,
|
||||
maxTokens: 8192,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "amazon-bedrock/us.anthropic.claude-opus-4-6-v1:0" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
With env-marker auth (`AWS_ACCESS_KEY_ID`, `AWS_PROFILE`, or `AWS_BEARER_TOKEN_BEDROCK`), OpenClaw auto-enables the implicit Bedrock provider for model discovery without extra config.
|
||||
</Tip>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="EC2 instance roles (IMDS)">
|
||||
**Best for:** EC2 instances with an IAM role attached, using the instance metadata service for authentication.
|
||||
|
||||
<Steps>
|
||||
<Step title="Enable discovery explicitly">
|
||||
When using IMDS, OpenClaw cannot detect AWS auth from env markers alone, so you must opt in:
|
||||
|
||||
```bash
|
||||
openclaw config set plugins.entries.amazon-bedrock.config.discovery.enabled true
|
||||
openclaw config set plugins.entries.amazon-bedrock.config.discovery.region us-east-1
|
||||
```
|
||||
</Step>
|
||||
<Step title="Optionally add an env marker for auto mode">
|
||||
If you also want the env-marker auto-detection path to work (for example, for `openclaw status` surfaces):
|
||||
|
||||
```bash
|
||||
export AWS_PROFILE=default
|
||||
export AWS_REGION=us-east-1
|
||||
```
|
||||
|
||||
You do **not** need a fake API key.
|
||||
</Step>
|
||||
<Step title="Verify models are discovered">
|
||||
```bash
|
||||
openclaw models list
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Warning>
|
||||
The IAM role attached to your EC2 instance must have the following permissions:
|
||||
|
||||
- `bedrock:InvokeModel`
|
||||
- `bedrock:InvokeModelWithResponseStream`
|
||||
- `bedrock:ListFoundationModels` (for automatic discovery)
|
||||
- `bedrock:ListInferenceProfiles` (for inference profile discovery)
|
||||
|
||||
Or attach the managed policy `AmazonBedrockFullAccess`.
|
||||
</Warning>
|
||||
|
||||
<Note>
|
||||
You only need `AWS_PROFILE=default` if you specifically want an env marker for auto mode or status surfaces. The actual Bedrock runtime auth path uses the AWS SDK default chain, so IMDS instance-role auth works even without env markers.
|
||||
</Note>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Automatic model discovery
|
||||
|
||||
OpenClaw can automatically discover Bedrock models that support **streaming**
|
||||
and **text output**. Discovery uses `bedrock:ListFoundationModels` and
|
||||
`bedrock:ListInferenceProfiles`, and results are cached (default: 1 hour).
|
||||
|
||||
How the implicit provider is enabled:
|
||||
|
||||
- If `plugins.entries.amazon-bedrock.config.discovery.enabled` is `true`,
|
||||
OpenClaw will try discovery even when no AWS env marker is present.
|
||||
- If `plugins.entries.amazon-bedrock.config.discovery.enabled` is unset,
|
||||
OpenClaw only auto-adds the
|
||||
implicit Bedrock provider when it sees one of these AWS auth markers:
|
||||
`AWS_BEARER_TOKEN_BEDROCK`, `AWS_ACCESS_KEY_ID` +
|
||||
`AWS_SECRET_ACCESS_KEY`, or `AWS_PROFILE`.
|
||||
- The actual Bedrock runtime auth path still uses the AWS SDK default chain, so
|
||||
shared config, SSO, and IMDS instance-role auth can work even when discovery
|
||||
needed `enabled: true` to opt in.
|
||||
|
||||
<Note>
|
||||
For explicit `models.providers["amazon-bedrock"]` entries, OpenClaw can still resolve Bedrock env-marker auth early from AWS env markers such as `AWS_BEARER_TOKEN_BEDROCK` without forcing full runtime auth loading. The actual model-call auth path still uses the AWS SDK default chain.
|
||||
</Note>
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Discovery config options">
|
||||
Config options live under `plugins.entries.amazon-bedrock.config.discovery`:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"amazon-bedrock": {
|
||||
config: {
|
||||
discovery: {
|
||||
enabled: true,
|
||||
region: "us-east-1",
|
||||
providerFilter: ["anthropic", "amazon"],
|
||||
refreshInterval: 3600,
|
||||
defaultContextWindow: 32000,
|
||||
defaultMaxTokens: 4096,
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
| Option | Default | Description |
|
||||
| ------ | ------- | ----------- |
|
||||
| `enabled` | auto | In auto mode, OpenClaw only enables the implicit Bedrock provider when it sees a supported AWS env marker. Set `true` to force discovery. |
|
||||
| `region` | `AWS_REGION` / `AWS_DEFAULT_REGION` / `us-east-1` | AWS region used for discovery API calls. |
|
||||
| `providerFilter` | (all) | Matches Bedrock provider names (for example `anthropic`, `amazon`). |
|
||||
| `refreshInterval` | `3600` | Cache duration in seconds. Set to `0` to disable caching. |
|
||||
| `defaultContextWindow` | `32000` | Context window used for discovered models with no known token limits (override if you know your model limits). |
|
||||
| `defaultMaxTokens` | `4096` | Max output tokens used for discovered models with no known token limits (override if you know your model limits). |
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Context window and max-token limits">
|
||||
The Bedrock `ListFoundationModels` and `GetFoundationModel` APIs return no
|
||||
token-limit metadata, only model ID, name, modalities, and lifecycle
|
||||
status. OpenClaw ships a lookup table of known context windows and output
|
||||
limits for popular Bedrock models (Claude, Nova, Llama, Mistral, DeepSeek,
|
||||
and others) so session management, compaction thresholds, and
|
||||
context-overflow detection work correctly for those models.
|
||||
|
||||
Discovered models not in the table fall back to `defaultContextWindow`
|
||||
and `defaultMaxTokens`. If a model you use is missing accurate limits,
|
||||
override it with an explicit
|
||||
`models.providers["amazon-bedrock"].models` entry.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Quick setup (AWS path)
|
||||
|
||||
This walkthrough creates an IAM role, attaches Bedrock permissions, associates
|
||||
the instance profile, and enables OpenClaw discovery on the EC2 host.
|
||||
|
||||
```bash
|
||||
# 1. Create IAM role and instance profile
|
||||
aws iam create-role --role-name EC2-Bedrock-Access \
|
||||
--assume-role-policy-document '{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [{
|
||||
"Effect": "Allow",
|
||||
"Principal": {"Service": "ec2.amazonaws.com"},
|
||||
"Action": "sts:AssumeRole"
|
||||
}]
|
||||
}'
|
||||
|
||||
aws iam attach-role-policy --role-name EC2-Bedrock-Access \
|
||||
--policy-arn arn:aws:iam::aws:policy/AmazonBedrockFullAccess
|
||||
|
||||
aws iam create-instance-profile --instance-profile-name EC2-Bedrock-Access
|
||||
aws iam add-role-to-instance-profile \
|
||||
--instance-profile-name EC2-Bedrock-Access \
|
||||
--role-name EC2-Bedrock-Access
|
||||
|
||||
# 2. Attach to your EC2 instance
|
||||
aws ec2 associate-iam-instance-profile \
|
||||
--instance-id i-xxxxx \
|
||||
--iam-instance-profile Name=EC2-Bedrock-Access
|
||||
|
||||
# 3. On the EC2 instance, enable discovery explicitly
|
||||
openclaw config set plugins.entries.amazon-bedrock.config.discovery.enabled true
|
||||
openclaw config set plugins.entries.amazon-bedrock.config.discovery.region us-east-1
|
||||
|
||||
# 4. Optional: add an env marker if you want auto mode without explicit enable
|
||||
echo 'export AWS_PROFILE=default' >> ~/.bashrc
|
||||
echo 'export AWS_REGION=us-east-1' >> ~/.bashrc
|
||||
source ~/.bashrc
|
||||
|
||||
# 5. Verify models are discovered
|
||||
openclaw models list
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Inference profiles">
|
||||
OpenClaw discovers **regional and global inference profiles** alongside
|
||||
foundation models. When a profile maps to a known foundation model, the
|
||||
profile inherits that model's capabilities (context window, max tokens,
|
||||
reasoning, vision) and the correct Bedrock request region is injected
|
||||
automatically. This means cross-region Claude profiles work without manual
|
||||
provider overrides. Global cross-region profiles (`global.*`) are listed
|
||||
first in `openclaw models list` since they generally offer better capacity
|
||||
and automatic failover.
|
||||
|
||||
Inference profile IDs look like `us.anthropic.claude-opus-4-6-v1:0` (regional)
|
||||
or `anthropic.claude-opus-4-6-v1:0` (global). If the backing model is already
|
||||
in the discovery results, the profile inherits its full capability set;
|
||||
otherwise safe defaults apply.
|
||||
|
||||
No extra configuration is needed. As long as discovery is enabled and the IAM
|
||||
principal has `bedrock:ListInferenceProfiles`, profiles appear alongside
|
||||
foundation models in `openclaw models list`.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Service tier">
|
||||
Some Bedrock models support a `service_tier` parameter to optimize for cost
|
||||
or latency. The following tiers are available:
|
||||
|
||||
| Tier | Description |
|
||||
|------|-------------|
|
||||
| `default` | Standard Bedrock tier |
|
||||
| `flex` | Discounted processing for workloads that can tolerate longer latency |
|
||||
| `priority` | Prioritized processing for latency-sensitive workloads |
|
||||
| `reserved` | Reserved capacity for steady-state workloads |
|
||||
|
||||
Set `serviceTier` (or `service_tier`) via `agents.defaults.params` for
|
||||
Bedrock model requests, or per-model in
|
||||
`agents.defaults.models["<model-key>"].params`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
params: {
|
||||
serviceTier: "flex", // applies to all models
|
||||
},
|
||||
models: {
|
||||
"amazon-bedrock/mistral.mistral-large-3-675b-instruct": {
|
||||
params: {
|
||||
serviceTier: "priority", // per-model override
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Valid values are `default`, `flex`, `priority`, and `reserved`. Claude
|
||||
Fable 5 only supports the `default` tier; OpenClaw warns and ignores
|
||||
`flex`, `priority`, or `reserved` requested for that model. For other
|
||||
models, not every model supports every tier -- an unsupported tier
|
||||
returns a Bedrock validation error, and the error message can be
|
||||
misleading (for example "The provided model identifier is invalid"
|
||||
rather than naming the tier as the problem). If you see this error, check
|
||||
whether the model supports the requested tier.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Claude Opus 4.7 and 4.8 temperature">
|
||||
Bedrock rejects the `temperature` parameter for Claude Opus 4.7 and Opus
|
||||
4.8. OpenClaw omits `temperature` automatically for any matching Bedrock
|
||||
ref, including foundation model ids, named inference profiles, application
|
||||
inference profiles whose underlying model resolves to Opus 4.7/4.8 via
|
||||
`bedrock:GetInferenceProfile`, and dotted `opus-4.7`/`opus-4.8` variants
|
||||
with optional region prefixes (`us.`, `eu.`, `ap.`, `apac.`, `au.`, `jp.`,
|
||||
`global.`). No config knob is required, and the omission applies to both
|
||||
the request options object and the `inferenceConfig` payload field.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Claude Fable 5">
|
||||
Use `amazon-bedrock/anthropic.claude-fable-5` in `us-east-1`, or the
|
||||
regional inference ids such as `us.anthropic.claude-fable-5`.
|
||||
OpenClaw applies Fable's 1M context window, 128K output limit, always-on
|
||||
adaptive thinking, and supported effort mapping. `/think off` and
|
||||
`/think minimal` map to `low`; temperature and forced tool choice controls
|
||||
are omitted, matching the Opus 4.7/4.8 route. Streaming output is held
|
||||
until Bedrock returns a terminal status so mid-stream refusals do not
|
||||
expose partial text.
|
||||
|
||||
AWS requires an explicit `provider_data_share` data-retention opt-in before
|
||||
Fable is available. Prompts and completions are shared with Anthropic and
|
||||
retained for up to 30 days for trust and safety. Review and configure
|
||||
[Bedrock data retention](https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html)
|
||||
before enabling the model.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Guardrails">
|
||||
You can apply [Amazon Bedrock Guardrails](https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails.html)
|
||||
to all Bedrock model invocations by adding a `guardrail` object to the
|
||||
`amazon-bedrock` plugin config. Guardrails let you enforce content filtering,
|
||||
topic denial, word filters, sensitive information filters, and contextual
|
||||
grounding checks.
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"amazon-bedrock": {
|
||||
config: {
|
||||
guardrail: {
|
||||
guardrailIdentifier: "abc123", // guardrail ID or full ARN
|
||||
guardrailVersion: "1", // version number or "DRAFT"
|
||||
streamProcessingMode: "sync", // optional: "sync" or "async"
|
||||
trace: "enabled", // optional: "enabled", "disabled", or "enabled_full"
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`guardrailIdentifier` and `guardrailVersion` are required.
|
||||
|
||||
| Option | Description |
|
||||
| ------ | ----------- |
|
||||
| `guardrailIdentifier` | Guardrail ID (e.g. `abc123`) or full ARN (e.g. `arn:aws:bedrock:us-east-1:123456789012:guardrail/abc123`). |
|
||||
| `guardrailVersion` | Published version number, or `"DRAFT"` for the working draft. |
|
||||
| `streamProcessingMode` | `"sync"` or `"async"` for guardrail evaluation during streaming. If omitted, Bedrock uses its default. |
|
||||
| `trace` | `"enabled"` or `"enabled_full"` for debugging; omit or set `"disabled"` for production. |
|
||||
|
||||
<Warning>
|
||||
The IAM principal used by the gateway must have the `bedrock:ApplyGuardrail` permission in addition to the standard invoke permissions.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Embeddings for memory search">
|
||||
Bedrock can also serve as the embedding provider for
|
||||
[memory search](/concepts/memory-search). This is configured separately from the
|
||||
inference provider -- set `agents.defaults.memorySearch.provider` to `"bedrock"`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
memorySearch: {
|
||||
provider: "bedrock",
|
||||
model: "amazon.titan-embed-text-v2:0", // default
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Bedrock embeddings use the same AWS SDK credential chain as inference (instance
|
||||
roles, SSO, access keys, shared config, and web identity). No API key is
|
||||
needed.
|
||||
|
||||
Supported embedding models include Amazon Titan Embed (v1, v2), Amazon Nova
|
||||
Embed, Cohere Embed (v3, v4), and TwelveLabs Marengo. See
|
||||
[Memory configuration reference -- Bedrock](/reference/memory-config#bedrock-embedding-config)
|
||||
for the full model list and dimension options.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Notes and caveats">
|
||||
- Bedrock requires **model access** enabled in your AWS account/region.
|
||||
- Automatic discovery needs the `bedrock:ListFoundationModels` and
|
||||
`bedrock:ListInferenceProfiles` permissions.
|
||||
- If you rely on auto mode, set one of the supported AWS auth env markers on the
|
||||
gateway host. If you prefer IMDS/shared-config auth without env markers, set
|
||||
`plugins.entries.amazon-bedrock.config.discovery.enabled: true`.
|
||||
- OpenClaw surfaces the credential source in this order: `AWS_BEARER_TOKEN_BEDROCK`,
|
||||
then `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, then `AWS_PROFILE`, then the
|
||||
default AWS SDK chain.
|
||||
- Reasoning support depends on the model; check the Bedrock model card for
|
||||
current capabilities.
|
||||
- If you prefer a managed key flow, you can also place an OpenAI-compatible
|
||||
proxy in front of Bedrock and configure it as an OpenAI provider instead.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Memory search" href="/concepts/memory-search" icon="magnifying-glass">
|
||||
Bedrock embeddings for memory search configuration.
|
||||
</Card>
|
||||
<Card title="Memory config reference" href="/reference/memory-config#bedrock-embedding-config" icon="database">
|
||||
Full Bedrock embedding model list and dimension options.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
General troubleshooting and FAQ.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
137
docs/providers/cerebras.md
Normal file
137
docs/providers/cerebras.md
Normal file
@@ -0,0 +1,137 @@
|
||||
---
|
||||
summary: "Cerebras setup (auth + model selection)"
|
||||
title: "Cerebras"
|
||||
read_when:
|
||||
- You want to use Cerebras with OpenClaw
|
||||
- You need the Cerebras API key env var or CLI auth choice
|
||||
---
|
||||
|
||||
[Cerebras](https://www.cerebras.ai) provides high-speed OpenAI-compatible inference on custom inference hardware. The plugin ships a static four-model catalog (no live discovery).
|
||||
|
||||
| Property | Value |
|
||||
| --------------- | --------------------------------------------------------- |
|
||||
| Provider id | `cerebras` |
|
||||
| Plugin | official external package (`@openclaw/cerebras-provider`) |
|
||||
| Auth env var | `CEREBRAS_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice cerebras-api-key` |
|
||||
| Direct CLI flag | `--cerebras-api-key <key>` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://api.cerebras.ai/v1` |
|
||||
| Default model | `cerebras/zai-glm-4.7` |
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/cerebras-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Create an API key in the [Cerebras Cloud Console](https://cloud.cerebras.ai).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
<CodeGroup>
|
||||
|
||||
```bash Onboarding
|
||||
openclaw onboard --auth-choice cerebras-api-key
|
||||
```
|
||||
|
||||
```bash Direct flag
|
||||
openclaw onboard --non-interactive \
|
||||
--auth-choice cerebras-api-key \
|
||||
--cerebras-api-key "$CEREBRAS_API_KEY"
|
||||
```
|
||||
|
||||
```bash Env only
|
||||
export CEREBRAS_API_KEY=csk-...
|
||||
```
|
||||
|
||||
</CodeGroup>
|
||||
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider cerebras
|
||||
```
|
||||
|
||||
Lists all four static models. If `CEREBRAS_API_KEY` is unresolved, `openclaw models status --json` reports the missing credential under `auth.unusableProfiles`.
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Non-interactive setup
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice cerebras-api-key \
|
||||
--cerebras-api-key "$CEREBRAS_API_KEY"
|
||||
```
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
All four models share a 128k context window and 8,192 max output tokens.
|
||||
|
||||
| Model ref | Name | Reasoning | Notes |
|
||||
| ----------------------------------------- | -------------------- | --------- | -------------------------------------- |
|
||||
| `cerebras/zai-glm-4.7` | Z.ai GLM 4.7 | yes | Default model; preview reasoning model |
|
||||
| `cerebras/gpt-oss-120b` | GPT OSS 120B | yes | Production reasoning model |
|
||||
| `cerebras/qwen-3-235b-a22b-instruct-2507` | Qwen 3 235B Instruct | no | Preview non-reasoning model |
|
||||
| `cerebras/llama3.1-8b` | Llama 3.1 8B | no | Production speed-focused model |
|
||||
|
||||
<Warning>
|
||||
Cerebras marks `zai-glm-4.7` and `qwen-3-235b-a22b-instruct-2507` as preview models, and `llama3.1-8b` plus `qwen-3-235b-a22b-instruct-2507` are documented for deprecation on May 27, 2026. Check Cerebras' [supported-models page](https://inference-docs.cerebras.ai/models/overview) before relying on them for production workloads.
|
||||
</Warning>
|
||||
|
||||
## Manual config
|
||||
|
||||
Most setups only need the API key. Use explicit `models.providers.cerebras` config to override model metadata or run in `mode: "merge"` against the static catalog:
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { CEREBRAS_API_KEY: "csk-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "cerebras/zai-glm-4.7" },
|
||||
},
|
||||
},
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
cerebras: {
|
||||
baseUrl: "https://api.cerebras.ai/v1",
|
||||
apiKey: "${CEREBRAS_API_KEY}",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
{ id: "zai-glm-4.7", name: "Z.ai GLM 4.7" },
|
||||
{ id: "gpt-oss-120b", name: "GPT OSS 120B" },
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
If the Gateway runs as a daemon (launchd, systemd, Docker), make sure `CEREBRAS_API_KEY` is available to that process — for example in `~/.openclaw/.env` or through `env.shellEnv`. A key exported only in an interactive shell will not help a managed service unless the env is imported separately.
|
||||
</Note>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Thinking modes" href="/tools/thinking" icon="brain">
|
||||
Reasoning effort levels for the two reasoning-capable Cerebras models.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/config-agents#agent-defaults" icon="gear">
|
||||
Agent defaults and model configuration.
|
||||
</Card>
|
||||
<Card title="Models FAQ" href="/help/faq-models" icon="circle-question">
|
||||
Auth profiles, switching models, and resolving "no profile" errors.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
156
docs/providers/chutes.md
Normal file
156
docs/providers/chutes.md
Normal file
@@ -0,0 +1,156 @@
|
||||
---
|
||||
summary: "Chutes setup (OAuth or API key, model discovery, aliases)"
|
||||
title: "Chutes"
|
||||
read_when:
|
||||
- You want to use Chutes with OpenClaw
|
||||
- You need the OAuth or API key setup path
|
||||
- You want the default model, aliases, or discovery behavior
|
||||
---
|
||||
|
||||
[Chutes](https://chutes.ai) exposes open-source model catalogs through an
|
||||
OpenAI-compatible API. OpenClaw supports both browser OAuth and API-key auth.
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | ------------------------------------------------------- |
|
||||
| Provider | `chutes` |
|
||||
| Plugin | official external package (`@openclaw/chutes-provider`) |
|
||||
| API | OpenAI-compatible |
|
||||
| Base URL | `https://llm.chutes.ai/v1` |
|
||||
| Auth | OAuth or API key (see below) |
|
||||
| Runtime env vars | `CHUTES_API_KEY`, `CHUTES_OAUTH_TOKEN` |
|
||||
|
||||
`CHUTES_OAUTH_TOKEN` supplies an already-obtained OAuth access token directly
|
||||
(for example in CI), bypassing the interactive browser flow below.
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/chutes-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
Both paths set the default model to `chutes/zai-org/GLM-4.7-TEE` and register
|
||||
the Chutes catalog.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="OAuth">
|
||||
<Steps>
|
||||
<Step title="Run the OAuth onboarding flow">
|
||||
```bash
|
||||
openclaw onboard --auth-choice chutes
|
||||
```
|
||||
OpenClaw launches the browser flow locally, or shows a URL + redirect-paste
|
||||
flow on remote/headless hosts. OAuth tokens auto-refresh through OpenClaw auth
|
||||
profiles.
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
<Tab title="API key">
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Create a key at
|
||||
[chutes.ai/settings/api-keys](https://chutes.ai/settings/api-keys).
|
||||
</Step>
|
||||
<Step title="Run the API key onboarding flow">
|
||||
```bash
|
||||
openclaw onboard --auth-choice chutes-api-key
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Discovery behavior
|
||||
|
||||
When Chutes auth is available, OpenClaw queries `GET /v1/models` with that
|
||||
credential and uses the discovered models, cached for 5 minutes per
|
||||
credential. On an expired/unauthorized key (HTTP 401), OpenClaw retries once
|
||||
without credentials. If discovery still returns no rows, fails, or returns any
|
||||
other non-2xx status, it falls back to the bundled static catalog (both API-key
|
||||
and OAuth discovery use this same path). If discovery fails at startup, the
|
||||
static catalog is used automatically.
|
||||
|
||||
## Default aliases
|
||||
|
||||
OpenClaw registers three convenience aliases for the Chutes catalog:
|
||||
|
||||
| Alias | Target model |
|
||||
| --------------- | ----------------------------------------------------- |
|
||||
| `chutes-fast` | `chutes/zai-org/GLM-4.7-FP8` |
|
||||
| `chutes-pro` | `chutes/deepseek-ai/DeepSeek-V3.2-TEE` |
|
||||
| `chutes-vision` | `chutes/chutesai/Mistral-Small-3.2-24B-Instruct-2506` |
|
||||
|
||||
## Built-in starter catalog
|
||||
|
||||
The bundled fallback catalog has 47 models. A representative sample of current refs:
|
||||
|
||||
| Model ref |
|
||||
| ----------------------------------------------------- |
|
||||
| `chutes/zai-org/GLM-4.7-TEE` |
|
||||
| `chutes/zai-org/GLM-5-TEE` |
|
||||
| `chutes/deepseek-ai/DeepSeek-V3.2-TEE` |
|
||||
| `chutes/deepseek-ai/DeepSeek-R1-0528-TEE` |
|
||||
| `chutes/moonshotai/Kimi-K2.5-TEE` |
|
||||
| `chutes/chutesai/Mistral-Small-3.2-24B-Instruct-2506` |
|
||||
| `chutes/Qwen/Qwen3-Coder-Next-TEE` |
|
||||
| `chutes/openai/gpt-oss-120b-TEE` |
|
||||
|
||||
Run `openclaw models list --all --provider chutes` for the full list.
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "chutes/zai-org/GLM-4.7-TEE" },
|
||||
models: {
|
||||
"chutes/zai-org/GLM-4.7-TEE": { alias: "Chutes GLM 4.7" },
|
||||
"chutes/deepseek-ai/DeepSeek-V3.2-TEE": { alias: "Chutes DeepSeek V3.2" },
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="OAuth overrides">
|
||||
Customize the OAuth flow with optional environment variables:
|
||||
|
||||
| Variable | Purpose |
|
||||
| -------- | ------- |
|
||||
| `CHUTES_CLIENT_ID` | OAuth client id (prompted if unset) |
|
||||
| `CHUTES_CLIENT_SECRET` | OAuth client secret |
|
||||
| `CHUTES_OAUTH_REDIRECT_URI` | Redirect URI (default `http://127.0.0.1:1456/oauth-callback`) |
|
||||
| `CHUTES_OAUTH_SCOPES` | Space-separated scopes (default `openid profile chutes:invoke`) |
|
||||
|
||||
See the [Chutes OAuth docs](https://chutes.ai/docs/sign-in-with-chutes/overview)
|
||||
for redirect-app requirements and help.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Notes">
|
||||
- Chutes models are registered as `chutes/<model-id>`.
|
||||
- Chutes does not report token usage while streaming (`supportsUsageInStreaming: false`); usage totals still show once the stream completes.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Provider rules, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema including provider settings.
|
||||
</Card>
|
||||
<Card title="Chutes" href="https://chutes.ai" icon="arrow-up-right-from-square">
|
||||
Chutes dashboard and API docs.
|
||||
</Card>
|
||||
<Card title="Chutes API keys" href="https://chutes.ai/settings/api-keys" icon="key">
|
||||
Create and manage Chutes API keys.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
201
docs/providers/claude-max-api-proxy.md
Normal file
201
docs/providers/claude-max-api-proxy.md
Normal file
@@ -0,0 +1,201 @@
|
||||
---
|
||||
summary: "Community proxy to expose Claude subscription credentials as an OpenAI-compatible endpoint"
|
||||
read_when:
|
||||
- You want to use Claude Max subscription with OpenAI-compatible tools
|
||||
- You want a local API server that wraps Claude Code CLI
|
||||
- You want to evaluate subscription-based vs API-key-based Anthropic access
|
||||
title: "Claude Max API proxy"
|
||||
---
|
||||
|
||||
**claude-max-api-proxy** is a community npm package (not an OpenClaw plugin) that
|
||||
exposes a Claude Max/Pro subscription as an OpenAI-compatible API endpoint, so
|
||||
you can point any OpenAI-compatible tool at your subscription instead of an
|
||||
Anthropic API key.
|
||||
|
||||
<Warning>
|
||||
Technical compatibility only, not an officially sanctioned path. Anthropic has
|
||||
blocked some subscription usage outside Claude Code in the past; verify
|
||||
Anthropic's current billing rules before relying on this.
|
||||
|
||||
Anthropic's Claude Code docs describe `claude -p` as Agent SDK/programmatic
|
||||
usage. As of Anthropic's June 15, 2026 support update, Claude Agent SDK,
|
||||
`claude -p`, and third-party app usage draw from the signed-in subscription's
|
||||
usage limits (the previously announced separate Agent SDK credit plan is
|
||||
paused). See Anthropic's [Agent SDK plan
|
||||
article](https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan),
|
||||
the [Pro/Max](https://support.claude.com/en/articles/11145838-use-claude-code-with-your-pro-or-max-plan)
|
||||
and [Team/Enterprise](https://support.claude.com/en/articles/11845131-use-claude-code-with-your-team-or-enterprise-plan)
|
||||
plan articles, and [Anthropic provider](/providers/anthropic) for OpenClaw's
|
||||
own Claude CLI billing notes.
|
||||
</Warning>
|
||||
|
||||
## Why use this
|
||||
|
||||
| Approach | Cost route | Best for |
|
||||
| ------------------------- | ----------------------------------------------- | ------------------------------------------ |
|
||||
| Anthropic API key | Pay per token through Claude Console | Production apps, shared automation, volume |
|
||||
| Claude subscription proxy | Claude Code / `claude -p` plan and credit rules | Personal experiments with compatible tools |
|
||||
|
||||
This proxy lets a Claude Max or Pro subscription work with OpenAI-compatible
|
||||
tools. It is not an unlimited flat-rate path — it inherits Claude Code's usage
|
||||
limits. API keys remain the clearer billing path for production use.
|
||||
|
||||
## How it works
|
||||
|
||||
```text
|
||||
Your App -> claude-max-api-proxy -> Claude Code CLI / claude -p -> Anthropic
|
||||
(OpenAI format) (converts format) (uses your login)
|
||||
```
|
||||
|
||||
The proxy spawns the Claude Code CLI as a subprocess per request, converts
|
||||
OpenAI-format chat requests to CLI prompts, and streams (or returns) the
|
||||
response back in OpenAI format.
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the proxy">
|
||||
Requires Node.js 20+ and an authenticated Claude Code CLI.
|
||||
|
||||
```bash
|
||||
npm install -g claude-max-api-proxy
|
||||
|
||||
# Verify Claude CLI is authenticated
|
||||
claude --version
|
||||
claude auth login # if not already authenticated
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Start the server">
|
||||
```bash
|
||||
claude-max-api
|
||||
# Server runs at http://localhost:3456
|
||||
```
|
||||
</Step>
|
||||
<Step title="Test the proxy">
|
||||
```bash
|
||||
curl http://localhost:3456/health
|
||||
curl http://localhost:3456/v1/models
|
||||
|
||||
curl http://localhost:3456/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "claude-opus-4",
|
||||
"messages": [{"role": "user", "content": "Hello!"}]
|
||||
}'
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Configure OpenClaw">
|
||||
Point OpenClaw at the proxy as a custom OpenAI-compatible endpoint:
|
||||
|
||||
```json5
|
||||
{
|
||||
env: {
|
||||
OPENAI_API_KEY: "not-needed",
|
||||
OPENAI_BASE_URL: "http://localhost:3456/v1",
|
||||
},
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "openai/claude-opus-4" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Note>
|
||||
The model ids below are the proxy's own catalog, not OpenClaw's Anthropic
|
||||
model refs. Each id maps to a Claude Code CLI model alias (`opus`, `sonnet`,
|
||||
`haiku`), so the underlying model shifts whenever Anthropic updates that
|
||||
alias in the CLI. Check the proxy's current README before relying on a
|
||||
specific mapping.
|
||||
</Note>
|
||||
|
||||
| Model ID | CLI alias | Current mapping |
|
||||
| ----------------- | --------- | --------------- |
|
||||
| `claude-opus-4` | `opus` | Claude Opus 4.5 |
|
||||
| `claude-sonnet-4` | `sonnet` | Claude Sonnet 4 |
|
||||
| `claude-haiku-4` | `haiku` | Claude Haiku 4 |
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Proxy-style OpenAI-compatible notes">
|
||||
This uses OpenClaw's generic custom `/v1` OpenAI-compatible route, the same
|
||||
path as any other self-hosted OpenAI-compatible backend:
|
||||
|
||||
- Native OpenAI-only request shaping does not apply.
|
||||
- `/fast` and `service_tier` only apply to direct `api.anthropic.com`
|
||||
traffic; proxy routes leave `service_tier` untouched (see
|
||||
[Anthropic provider fast mode](/providers/anthropic#advanced-configuration)).
|
||||
- No Responses `store`, prompt-cache hints, or OpenAI reasoning-compat
|
||||
payload shaping.
|
||||
- OpenClaw's OpenAI/Codex attribution headers (`originator`, `version`,
|
||||
`User-Agent`) are only sent on native `api.openai.com` OAuth traffic, not
|
||||
on custom `OPENAI_BASE_URL` targets like this proxy.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Auto-start on macOS with LaunchAgent">
|
||||
```bash
|
||||
cat > ~/Library/LaunchAgents/com.claude-max-api.plist << 'EOF'
|
||||
<?xml version="1.0" encoding="UTF-8"?>
|
||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||
<plist version="1.0">
|
||||
<dict>
|
||||
<key>Label</key>
|
||||
<string>com.claude-max-api</string>
|
||||
<key>RunAtLoad</key>
|
||||
<true/>
|
||||
<key>KeepAlive</key>
|
||||
<true/>
|
||||
<key>ProgramArguments</key>
|
||||
<array>
|
||||
<string>/usr/local/bin/node</string>
|
||||
<string>/usr/local/lib/node_modules/claude-max-api-proxy/dist/server/standalone.js</string>
|
||||
</array>
|
||||
<key>EnvironmentVariables</key>
|
||||
<dict>
|
||||
<key>PATH</key>
|
||||
<string>/usr/local/bin:/opt/homebrew/bin:~/.local/bin:/usr/bin:/bin</string>
|
||||
</dict>
|
||||
</dict>
|
||||
</plist>
|
||||
EOF
|
||||
|
||||
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.claude-max-api.plist
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Notes
|
||||
|
||||
- Inherits Claude Code's `claude -p` billing, usage-credit, and rate-limit behavior.
|
||||
- Binds to `127.0.0.1` only; does not send data to any third-party server beyond the CLI's own call to Anthropic.
|
||||
- Streaming responses are supported.
|
||||
- Auth failures are not checked at startup and only surface once a chat request actually runs; if the CLI is unauthenticated, expect the first request to fail rather than the server to refuse to start.
|
||||
|
||||
<Note>
|
||||
For native Anthropic integration with Claude CLI or API keys, see [Anthropic provider](/providers/anthropic). For OpenAI/Codex subscriptions, see [OpenAI provider](/providers/openai).
|
||||
</Note>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Anthropic provider" href="/providers/anthropic" icon="bolt">
|
||||
Native OpenClaw integration with Claude CLI or API keys.
|
||||
</Card>
|
||||
<Card title="OpenAI provider" href="/providers/openai" icon="robot">
|
||||
For OpenAI/Codex subscriptions.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Overview of all providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration" href="/gateway/configuration" icon="gear">
|
||||
Full config reference.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
156
docs/providers/clawrouter.md
Normal file
156
docs/providers/clawrouter.md
Normal file
@@ -0,0 +1,156 @@
|
||||
---
|
||||
summary: "Route credential-scoped models through ClawRouter and show managed quotas"
|
||||
title: "ClawRouter"
|
||||
read_when:
|
||||
- You want one managed key for multiple model providers
|
||||
- You need ClawRouter model discovery or quota reporting in OpenClaw
|
||||
---
|
||||
|
||||
ClawRouter gives OpenClaw one policy-scoped key for multiple upstream model
|
||||
providers. The bundled `clawrouter` plugin discovers only the models allowed
|
||||
for that key, routes each model through its declared protocol, and reports
|
||||
the key's budget and aggregate usage on OpenClaw usage surfaces.
|
||||
|
||||
Upstream credentials and provider-specific forwarding stay in ClawRouter, so
|
||||
you never install or authenticate each upstream provider plugin on the
|
||||
OpenClaw host. The plugin ships bundled with OpenClaw (`enabledByDefault: true`);
|
||||
you only need an issued ClawRouter credential.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | ---------------------------------------- |
|
||||
| Provider | `clawrouter` |
|
||||
| Plugin | bundled (included in OpenClaw) |
|
||||
| Auth | `CLAWROUTER_API_KEY` |
|
||||
| Default URL | `https://clawrouter.openclaw.ai` |
|
||||
| Model catalog | Credential-scoped via `/v1/catalog` |
|
||||
| Quotas | Monthly budget and usage via `/v1/usage` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get a scoped credential">
|
||||
Ask your ClawRouter administrator for a credential whose policy includes
|
||||
the providers, models, and monthly budget you should use. Credentials are
|
||||
revealed once when issued.
|
||||
</Step>
|
||||
<Step title="Configure OpenClaw">
|
||||
```bash
|
||||
export CLAWROUTER_API_KEY="..."
|
||||
openclaw onboard --auth-choice clawrouter-api-key
|
||||
openclaw plugins enable clawrouter
|
||||
```
|
||||
|
||||
`clawrouter` is bundled and enabled by default. If your configuration sets
|
||||
`plugins.allow`, add `clawrouter` to that list before enabling it. For a
|
||||
custom deployment, set `models.providers.clawrouter.baseUrl` to the
|
||||
ClawRouter origin; the default is `https://clawrouter.openclaw.ai`.
|
||||
|
||||
</Step>
|
||||
<Step title="List granted models">
|
||||
```bash
|
||||
openclaw models list --all --provider clawrouter
|
||||
```
|
||||
|
||||
Use the returned model refs exactly as shown. They retain the upstream
|
||||
namespace, such as `clawrouter/openai/gpt-5.5`,
|
||||
`clawrouter/anthropic/claude-sonnet-4-6`, or
|
||||
`clawrouter/google/gemini-3.5-flash`. If `agents.defaults.models` is an
|
||||
allowlist in your configuration, add each selected ClawRouter ref to it.
|
||||
|
||||
</Step>
|
||||
<Step title="Select a model">
|
||||
```bash
|
||||
openclaw models set clawrouter/<provider>/<model>
|
||||
```
|
||||
|
||||
You can also select a returned model for one run with
|
||||
`openclaw agent --model clawrouter/<provider>/<model> --message "..."`.
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Model discovery
|
||||
|
||||
`GET /v1/catalog` returns `{ providers: [...] }`, where each provider entry
|
||||
lists its own `models[]` (with upstream id, capabilities, and pricing) and its
|
||||
supported request routes. OpenClaw does not ship a second, fixed list of
|
||||
ClawRouter models. A catalog model is advertised as an OpenClaw model when:
|
||||
|
||||
- the credential's policy grants its provider;
|
||||
- the catalog model advertises a supported LLM capability (`llm.responses`,
|
||||
`llm.chat`, `llm.messages`, or `llm.stream` with a matching streaming
|
||||
route); and
|
||||
- the provider exposes a matching route for one of the transports below.
|
||||
|
||||
Adding a model to a supported ClawRouter provider needs no OpenClaw release:
|
||||
the next catalog refresh (cached 60 seconds per credential scope) discovers
|
||||
it. A model that needs a new wire protocol requires plugin support first.
|
||||
|
||||
## Protocol and provider plugins
|
||||
|
||||
ClawRouter owns upstream credentials; its catalog tells OpenClaw which
|
||||
transport to use, so you never install every upstream company's auth plugin.
|
||||
|
||||
| Catalog capability / route | OpenClaw transport |
|
||||
| -------------------------------------------------------- | ---------------------- |
|
||||
| `llm.responses` (OpenAI-compatible provider) | `openai-responses` |
|
||||
| `llm.chat` (OpenAI-compatible provider) | `openai-completions` |
|
||||
| `llm.messages` + `anthropic.messages` route | `anthropic-messages` |
|
||||
| `llm.stream` + streaming `google.generate_content` route | `google-generative-ai` |
|
||||
|
||||
The plugin also applies the matching replay and tool-schema policies for those
|
||||
families (OpenAI/DeepSeek/Gemini tool-schema compat; native Anthropic and
|
||||
Google Gemini replay policies). A catalog provider exposing only an
|
||||
unsupported request format is intentionally not advertised as an OpenClaw
|
||||
text model. Normalize those providers to one of the supported contracts in
|
||||
ClawRouter rather than sending an incompatible payload.
|
||||
|
||||
## Quotas and usage
|
||||
|
||||
ClawRouter's `/v1/usage` response feeds the normal OpenClaw provider-usage
|
||||
surfaces: request, token, and spend totals, plus a monthly budget window when
|
||||
the key has a limit. Unmetered keys still show aggregate usage without a
|
||||
percentage window.
|
||||
|
||||
Quota lookup uses the same scoped key as model discovery. A failed quota
|
||||
lookup does not block model execution.
|
||||
|
||||
Check the live snapshot with:
|
||||
|
||||
```bash
|
||||
openclaw status --usage
|
||||
openclaw models status
|
||||
```
|
||||
|
||||
The same provider snapshot is available to `/status` in chat and OpenClaw's
|
||||
usage UI. The budget is policy-wide, so requests made by another client using
|
||||
the same ClawRouter policy can change the remaining percentage.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Check |
|
||||
| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| No ClawRouter models | Confirm the plugin is enabled and allowed by `plugins.allow`, then check that the credential is active and grants at least one ready provider. |
|
||||
| A configured ClawRouter model is missing | Inspect its `/v1/catalog` capability and route support. Unsupported transport contracts are intentionally filtered. |
|
||||
| `Unknown model: clawrouter/...` | Add the exact catalog ref to `agents.defaults.models` when that configuration map is being used as an allowlist. |
|
||||
| `401` or `403` from catalog or usage | Reissue or re-scope the ClawRouter credential; OpenClaw does not fall back to upstream provider keys. |
|
||||
| Model call fails after discovery | Check the provider connection and upstream health in ClawRouter, then retry after its readiness state recovers. |
|
||||
| Usage has totals but no percentage | The policy is unmetered; add a monthly budget in ClawRouter to expose a percentage window. |
|
||||
|
||||
## Security behavior
|
||||
|
||||
- Catalog discovery is scoped to the configured proxy key and cached per credential scope (agent dir, workspace dir, auth profile id, and base URL).
|
||||
- The proxy key is attached only at request dispatch; it is not stored in model metadata.
|
||||
- Native Anthropic and Gemini model ids are rewritten to their upstream ids only at dispatch.
|
||||
- Unsupported or ungranted catalog rows fail closed and are not selectable.
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Provider configuration and model selection.
|
||||
</Card>
|
||||
<Card title="Usage tracking" href="/concepts/usage-tracking" icon="chart-line">
|
||||
OpenClaw usage and status surfaces.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
129
docs/providers/cloudflare-ai-gateway.md
Normal file
129
docs/providers/cloudflare-ai-gateway.md
Normal file
@@ -0,0 +1,129 @@
|
||||
---
|
||||
summary: "Cloudflare AI Gateway setup (auth + model selection)"
|
||||
title: "Cloudflare AI gateway"
|
||||
read_when:
|
||||
- You want to use Cloudflare AI Gateway with OpenClaw
|
||||
- You need the account ID, gateway ID, or API key env var
|
||||
---
|
||||
|
||||
[Cloudflare AI Gateway](https://developers.cloudflare.com/ai-gateway/) sits in front of provider APIs and adds analytics, caching, and controls. For Anthropic, OpenClaw uses the Anthropic Messages API through your Gateway endpoint.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | ---------------------------------------------------------------------------------------- |
|
||||
| Provider | `cloudflare-ai-gateway` |
|
||||
| Plugin | official external package (`@openclaw/cloudflare-ai-gateway-provider`) |
|
||||
| Base URL | `https://gateway.ai.cloudflare.com/v1/<account_id>/<gateway_id>/anthropic` |
|
||||
| Default model | `cloudflare-ai-gateway/claude-sonnet-4-6` |
|
||||
| API key | `CLOUDFLARE_AI_GATEWAY_API_KEY` (your provider API key for requests through the Gateway) |
|
||||
|
||||
<Note>
|
||||
For Anthropic models routed through Cloudflare AI Gateway, use your **Anthropic API key** as the provider key.
|
||||
</Note>
|
||||
|
||||
When thinking is enabled for Anthropic Messages models, OpenClaw strips trailing
|
||||
assistant prefill turns before sending the payload through Cloudflare AI Gateway.
|
||||
Anthropic rejects response prefilling with extended thinking, while ordinary
|
||||
non-thinking prefill remains available.
|
||||
|
||||
## Install plugin
|
||||
|
||||
Install the official plugin, then restart Gateway:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/cloudflare-ai-gateway-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set the provider API key and Gateway details">
|
||||
Run onboarding and choose the Cloudflare AI Gateway auth option:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice cloudflare-ai-gateway-api-key
|
||||
```
|
||||
|
||||
This prompts for your account ID, gateway ID, and API key.
|
||||
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
Add the model to your OpenClaw config:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "cloudflare-ai-gateway/claude-sonnet-4-6" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider cloudflare-ai-gateway
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Non-interactive example
|
||||
|
||||
For scripted or CI setups, pass all values on the command line:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice cloudflare-ai-gateway-api-key \
|
||||
--cloudflare-ai-gateway-account-id "your-account-id" \
|
||||
--cloudflare-ai-gateway-gateway-id "your-gateway-id" \
|
||||
--cloudflare-ai-gateway-api-key "$CLOUDFLARE_AI_GATEWAY_API_KEY"
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Authenticated gateways">
|
||||
If you enabled Gateway authentication in Cloudflare, add the `cf-aig-authorization` header. This is **in addition to** your provider API key.
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
"cloudflare-ai-gateway": {
|
||||
headers: {
|
||||
"cf-aig-authorization": "Bearer <cloudflare-ai-gateway-token>",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Tip>
|
||||
The `cf-aig-authorization` header authenticates with the Cloudflare Gateway itself, while the provider API key (for example, your Anthropic key) authenticates with the upstream provider.
|
||||
</Tip>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment note">
|
||||
If the Gateway runs as a daemon (launchd/systemd), make sure `CLOUDFLARE_AI_GATEWAY_API_KEY` is available to that process.
|
||||
|
||||
<Warning>
|
||||
A key exported only in an interactive shell will not help a launchd/systemd daemon unless that environment is imported there as well. Set the key in `~/.openclaw/.env` or via `env.shellEnv` to ensure the gateway process can read it.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
General troubleshooting and FAQ.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
71
docs/providers/cohere.md
Normal file
71
docs/providers/cohere.md
Normal file
@@ -0,0 +1,71 @@
|
||||
---
|
||||
summary: "Cohere setup (auth + model selection)"
|
||||
title: "Cohere"
|
||||
read_when:
|
||||
- You want to use Cohere with OpenClaw
|
||||
- You need the Cohere API key env var or CLI auth choice
|
||||
---
|
||||
|
||||
[Cohere](https://cohere.com) provides OpenAI-compatible inference through its Compatibility API. OpenClaw bundles the Cohere provider during its externalization transition and also publishes it as an official external plugin.
|
||||
|
||||
| Property | Value |
|
||||
| --------------- | ---------------------------------------------------- |
|
||||
| Provider id | `cohere` |
|
||||
| Plugin | bundled during transition; official external package |
|
||||
| Auth env var | `COHERE_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice cohere-api-key` |
|
||||
| Direct CLI flag | `--cohere-api-key <key>` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://api.cohere.ai/compatibility/v1` |
|
||||
| Default model | `cohere/command-a-03-2025` |
|
||||
| Context window | 256,000 tokens |
|
||||
|
||||
## Get started
|
||||
|
||||
1. Cohere ships with current OpenClaw packages. If it is missing, install the external package and restart the Gateway:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/cohere-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
2. Create a Cohere API key.
|
||||
3. Run onboarding:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--auth-choice cohere-api-key \
|
||||
--cohere-api-key "$COHERE_API_KEY"
|
||||
```
|
||||
|
||||
4. Confirm the catalog is available:
|
||||
|
||||
```bash
|
||||
openclaw models list --provider cohere
|
||||
```
|
||||
|
||||
Onboarding only sets Cohere as the primary model when no primary model is already configured.
|
||||
|
||||
## Environment-only setup
|
||||
|
||||
Make `COHERE_API_KEY` available to the Gateway process, then select the Cohere model:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "cohere/command-a-03-2025" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
If the Gateway runs as a daemon or in Docker, set `COHERE_API_KEY` for that service. Exporting it only in an interactive shell does not make it available to an already-running Gateway.
|
||||
</Note>
|
||||
|
||||
## Related
|
||||
|
||||
- [Model providers](/concepts/model-providers)
|
||||
- [Models CLI](/cli/models)
|
||||
- [Provider directory](/providers/index)
|
||||
364
docs/providers/comfy.md
Normal file
364
docs/providers/comfy.md
Normal file
@@ -0,0 +1,364 @@
|
||||
---
|
||||
summary: "ComfyUI workflow image, video, and music generation setup in OpenClaw"
|
||||
title: "ComfyUI"
|
||||
read_when:
|
||||
- You want to use local ComfyUI workflows with OpenClaw
|
||||
- You want to use Comfy Cloud with image, video, or music workflows
|
||||
- You need the bundled comfy plugin config keys
|
||||
---
|
||||
|
||||
OpenClaw ships a bundled `comfy` plugin for workflow-driven ComfyUI runs. The
|
||||
plugin is entirely workflow-driven: OpenClaw does not map generic `size`,
|
||||
`aspectRatio`, `resolution`, `durationSeconds`, or TTS-style controls onto
|
||||
your graph.
|
||||
|
||||
| Property | Detail |
|
||||
| ------------ | -------------------------------------------------------------------------------- |
|
||||
| Provider | `comfy` |
|
||||
| Model | `comfy/workflow` |
|
||||
| Shared tools | `image_generate`, `video_generate`, `music_generate` |
|
||||
| Auth | None for local ComfyUI; `COMFY_API_KEY` or `COMFY_CLOUD_API_KEY` for Comfy Cloud |
|
||||
| API | ComfyUI `/prompt` / `/history` / `/view`; Comfy Cloud `/api/*` |
|
||||
|
||||
## What it supports
|
||||
|
||||
- Image generation and editing from a workflow JSON (edit takes 1 uploaded reference image)
|
||||
- Video generation from a workflow JSON, text-to-video or image-to-video (1 reference image)
|
||||
- Music/audio generation through the shared `music_generate` tool, with an optional 1 reference image
|
||||
- Output download from a configured node, or from all matching output nodes when none is configured
|
||||
|
||||
## Getting started
|
||||
|
||||
Choose between running ComfyUI on your own machine or using Comfy Cloud.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Local">
|
||||
**Best for:** running your own ComfyUI instance on your machine or LAN.
|
||||
|
||||
<Steps>
|
||||
<Step title="Start ComfyUI locally">
|
||||
Make sure your local ComfyUI instance is running (defaults to `http://127.0.0.1:8188`).
|
||||
</Step>
|
||||
<Step title="Prepare your workflow JSON">
|
||||
Export or create a ComfyUI workflow JSON file. Note the node IDs for the prompt input node and the output node you want OpenClaw to read from.
|
||||
</Step>
|
||||
<Step title="Configure the provider">
|
||||
Set `mode: "local"` and point at your workflow file. Minimal image example:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
comfy: {
|
||||
config: {
|
||||
mode: "local",
|
||||
baseUrl: "http://127.0.0.1:8188",
|
||||
image: {
|
||||
workflowPath: "./workflows/flux-api.json",
|
||||
promptNodeId: "6",
|
||||
outputNodeId: "9",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set the default model">
|
||||
Point OpenClaw at the `comfy/workflow` model for the capability you configured:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "comfy/workflow",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify">
|
||||
```bash
|
||||
openclaw models list --provider comfy
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Comfy Cloud">
|
||||
**Best for:** running workflows on Comfy Cloud without managing local GPU resources.
|
||||
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Sign up at [comfy.org](https://comfy.org) and generate an API key from your account dashboard.
|
||||
</Step>
|
||||
<Step title="Set the API key">
|
||||
Provide your key through any of these methods:
|
||||
|
||||
```bash
|
||||
# Onboarding flag
|
||||
openclaw onboard --comfy-api-key "your-key"
|
||||
|
||||
# Environment variable (preferred for daemons)
|
||||
export COMFY_API_KEY="your-key"
|
||||
|
||||
# Alternative environment variable
|
||||
export COMFY_CLOUD_API_KEY="your-key"
|
||||
|
||||
# Or inline in config
|
||||
openclaw config set plugins.entries.comfy.config.apiKey "your-key"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Prepare your workflow JSON">
|
||||
Export or create a ComfyUI workflow JSON file. Note the node IDs for the prompt input node and the output node.
|
||||
</Step>
|
||||
<Step title="Configure the provider">
|
||||
Set `mode: "cloud"` and point at your workflow file:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
comfy: {
|
||||
config: {
|
||||
mode: "cloud",
|
||||
image: {
|
||||
workflowPath: "./workflows/flux-api.json",
|
||||
promptNodeId: "6",
|
||||
outputNodeId: "9",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Tip>
|
||||
Cloud mode defaults `baseUrl` to `https://cloud.comfy.org`. Set `baseUrl` only for a custom cloud endpoint.
|
||||
</Tip>
|
||||
</Step>
|
||||
<Step title="Set the default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "comfy/workflow",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify">
|
||||
```bash
|
||||
openclaw models list --provider comfy
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Configuration
|
||||
|
||||
Comfy supports shared top-level connection settings plus per-capability workflow sections (`image`, `video`, `music`):
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
comfy: {
|
||||
config: {
|
||||
mode: "local",
|
||||
baseUrl: "http://127.0.0.1:8188",
|
||||
image: {
|
||||
workflowPath: "./workflows/flux-api.json",
|
||||
promptNodeId: "6",
|
||||
outputNodeId: "9",
|
||||
},
|
||||
video: {
|
||||
workflowPath: "./workflows/video-api.json",
|
||||
promptNodeId: "12",
|
||||
outputNodeId: "21",
|
||||
},
|
||||
music: {
|
||||
workflowPath: "./workflows/music-api.json",
|
||||
promptNodeId: "3",
|
||||
outputNodeId: "18",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### Shared keys
|
||||
|
||||
| Key | Type | Description |
|
||||
| --------------------- | ---------------------- | ------------------------------------------------------------------------------------- |
|
||||
| `mode` | `"local"` or `"cloud"` | Connection mode. Defaults to `"local"`. |
|
||||
| `baseUrl` | string | Defaults to `http://127.0.0.1:8188` for local or `https://cloud.comfy.org` for cloud. |
|
||||
| `apiKey` | string | Optional inline key, alternative to `COMFY_API_KEY` / `COMFY_CLOUD_API_KEY` env vars. |
|
||||
| `allowPrivateNetwork` | boolean | Allow a private/LAN `baseUrl` in cloud mode. |
|
||||
|
||||
### Per-capability keys
|
||||
|
||||
These keys apply inside the `image`, `video`, or `music` sections:
|
||||
|
||||
| Key | Required | Default | Description |
|
||||
| ---------------------------- | -------- | -------- | ---------------------------------------------------------------------------- |
|
||||
| `workflow` or `workflowPath` | Yes | -- | Inline workflow JSON, or path to the ComfyUI workflow JSON file. |
|
||||
| `promptNodeId` | Yes | -- | Node ID that receives the text prompt. |
|
||||
| `promptInputName` | No | `"text"` | Input name on the prompt node. |
|
||||
| `outputNodeId` | No | -- | Node ID to read output from. If omitted, all matching output nodes are used. |
|
||||
| `pollIntervalMs` | No | `1500` | Polling interval in milliseconds for job completion. |
|
||||
| `timeoutMs` | No | `300000` | Timeout in milliseconds for the workflow run. |
|
||||
|
||||
The `image` and `video` sections also support a reference-image input node:
|
||||
|
||||
| Key | Required | Default | Description |
|
||||
| --------------------- | ------------------------------------ | --------- | --------------------------------------------------- |
|
||||
| `inputImageNodeId` | Yes (when passing a reference image) | -- | Node ID that receives the uploaded reference image. |
|
||||
| `inputImageInputName` | No | `"image"` | Input name on the image node. |
|
||||
|
||||
`apiKey` accepts either a literal string or a [secret reference](/gateway/configuration-reference#secrets) object.
|
||||
|
||||
## Workflow details
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Image workflows">
|
||||
Set the default image model to `comfy/workflow`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "comfy/workflow",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
**Reference-image editing example:**
|
||||
|
||||
To enable image editing with an uploaded reference image, add `inputImageNodeId` to your image config:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
comfy: {
|
||||
config: {
|
||||
image: {
|
||||
workflowPath: "./workflows/edit-api.json",
|
||||
promptNodeId: "6",
|
||||
inputImageNodeId: "7",
|
||||
inputImageInputName: "image",
|
||||
outputNodeId: "9",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Video workflows">
|
||||
Set the default video model to `comfy/workflow`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "comfy/workflow",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Comfy video workflows support text-to-video and image-to-video through the configured graph.
|
||||
|
||||
<Note>
|
||||
OpenClaw does not pass input videos into Comfy workflows. Only text prompts and single reference images are supported as inputs.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Music workflows">
|
||||
The bundled plugin registers a music-generation provider for workflow-defined audio or music outputs, surfaced through the shared `music_generate` tool. It accepts an optional reference image (up to 1):
|
||||
|
||||
```text
|
||||
/tool music_generate prompt="Warm ambient synth loop with soft tape texture"
|
||||
```
|
||||
|
||||
Use the `music` config section to point at your audio workflow JSON and output node.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Backward compatibility">
|
||||
Existing top-level image config (without the nested `image` section) still works:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
comfy: {
|
||||
config: {
|
||||
workflowPath: "./workflows/flux-api.json",
|
||||
promptNodeId: "6",
|
||||
outputNodeId: "9",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw treats that legacy shape as the image workflow config. You do not need to migrate immediately, but the nested `image` / `video` / `music` sections are recommended for new setups. If you only use image generation, the legacy flat config and the new nested `image` section are functionally equivalent.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Live tests">
|
||||
Opt-in live coverage exists for the bundled plugin:
|
||||
|
||||
```bash
|
||||
OPENCLAW_LIVE_TEST=1 COMFY_LIVE_TEST=1 pnpm test:live -- extensions/comfy/comfy.live.test.ts
|
||||
```
|
||||
|
||||
The live test skips individual image, video, or music cases unless the matching Comfy workflow section is configured.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Image Generation" href="/tools/image-generation" icon="image">
|
||||
Image generation tool configuration and usage.
|
||||
</Card>
|
||||
<Card title="Video Generation" href="/tools/video-generation" icon="video">
|
||||
Video generation tool configuration and usage.
|
||||
</Card>
|
||||
<Card title="Music Generation" href="/tools/music-generation" icon="music">
|
||||
Music and audio generation tool setup.
|
||||
</Card>
|
||||
<Card title="Provider Directory" href="/providers/index" icon="layers">
|
||||
Overview of all providers and model refs.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/config-agents#agent-defaults" icon="gear">
|
||||
Full config reference including agent defaults.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
182
docs/providers/deepgram.md
Normal file
182
docs/providers/deepgram.md
Normal file
@@ -0,0 +1,182 @@
|
||||
---
|
||||
summary: "Deepgram transcription for inbound voice notes"
|
||||
read_when:
|
||||
- You want Deepgram speech-to-text for audio attachments
|
||||
- You want Deepgram streaming transcription for Voice Call
|
||||
- You need a quick Deepgram config example
|
||||
title: "Deepgram"
|
||||
---
|
||||
|
||||
Deepgram is a speech-to-text API. OpenClaw uses it for inbound audio/voice-note
|
||||
transcription through `tools.media.audio` and for Voice Call streaming STT
|
||||
through `plugins.entries.voice-call.config.streaming`.
|
||||
|
||||
Batch transcription uploads the complete audio file to Deepgram and injects
|
||||
the transcript into the reply pipeline (`{{Transcript}}` + `[Audio]` block).
|
||||
Voice Call streaming forwards live G.711 u-law frames over Deepgram's
|
||||
WebSocket `listen` endpoint and emits partial/final transcripts as Deepgram
|
||||
returns them.
|
||||
|
||||
| Detail | Value |
|
||||
| ------------- | ---------------------------------------------------------- |
|
||||
| Website | [deepgram.com](https://deepgram.com) |
|
||||
| Docs | [developers.deepgram.com](https://developers.deepgram.com) |
|
||||
| Auth | `DEEPGRAM_API_KEY` |
|
||||
| Default model | `nova-3` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set your API key">
|
||||
```bash
|
||||
DEEPGRAM_API_KEY=dg_...
|
||||
```
|
||||
</Step>
|
||||
<Step title="Enable the audio provider">
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
enabled: true,
|
||||
models: [{ provider: "deepgram", model: "nova-3" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Send a voice note">
|
||||
Send an audio message through any connected channel. OpenClaw transcribes it
|
||||
via Deepgram and injects the transcript into the reply pipeline.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Configuration options
|
||||
|
||||
| Option | Path | Description |
|
||||
| ---------- | ------------------------------------- | ------------------------------------- |
|
||||
| `model` | `tools.media.audio.models[].model` | Deepgram model id (default: `nova-3`) |
|
||||
| `language` | `tools.media.audio.models[].language` | Language hint (optional) |
|
||||
|
||||
`providerOptions.deepgram` merges extra query params directly into the
|
||||
Deepgram `/listen` request, so any Deepgram-supported param name works
|
||||
(for example `detect_language`, `punctuate`, `smart_format`):
|
||||
|
||||
<Tabs>
|
||||
<Tab title="With language hint">
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
enabled: true,
|
||||
models: [{ provider: "deepgram", model: "nova-3", language: "en" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Tab>
|
||||
<Tab title="With Deepgram options">
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
enabled: true,
|
||||
providerOptions: {
|
||||
deepgram: {
|
||||
detect_language: true,
|
||||
punctuate: true,
|
||||
smart_format: true,
|
||||
},
|
||||
},
|
||||
models: [{ provider: "deepgram", model: "nova-3" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Voice Call streaming STT
|
||||
|
||||
The bundled `deepgram` plugin also registers a realtime transcription provider
|
||||
for the Voice Call plugin.
|
||||
|
||||
| Setting | Config path | Default |
|
||||
| --------------- | ----------------------------------------------------------------------- | -------------------------------- |
|
||||
| API key | `plugins.entries.voice-call.config.streaming.providers.deepgram.apiKey` | Falls back to `DEEPGRAM_API_KEY` |
|
||||
| Model | `...deepgram.model` | `nova-3` |
|
||||
| Language | `...deepgram.language` | (unset) |
|
||||
| Encoding | `...deepgram.encoding` | `mulaw` |
|
||||
| Sample rate | `...deepgram.sampleRate` | `8000` |
|
||||
| Endpointing | `...deepgram.endpointingMs` | `800` |
|
||||
| Interim results | `...deepgram.interimResults` | `true` |
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"voice-call": {
|
||||
config: {
|
||||
streaming: {
|
||||
enabled: true,
|
||||
provider: "deepgram",
|
||||
providers: {
|
||||
deepgram: {
|
||||
apiKey: "${DEEPGRAM_API_KEY}",
|
||||
model: "nova-3",
|
||||
endpointingMs: 800,
|
||||
language: "en-US",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
Voice Call receives telephony audio as 8 kHz G.711 u-law. The Deepgram
|
||||
streaming provider defaults to `encoding: "mulaw"` and `sampleRate: 8000`, so
|
||||
Twilio media frames can be forwarded directly.
|
||||
</Note>
|
||||
|
||||
## Notes
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Authentication">
|
||||
Authentication follows the standard provider auth order. `DEEPGRAM_API_KEY` is
|
||||
the simplest path.
|
||||
</Accordion>
|
||||
<Accordion title="Proxy and custom endpoints">
|
||||
Override endpoints or headers with `tools.media.audio.baseUrl` and
|
||||
`tools.media.audio.headers` when using a proxy.
|
||||
</Accordion>
|
||||
<Accordion title="Output behavior">
|
||||
Output follows the same audio rules as other providers (size caps, timeouts,
|
||||
transcript injection).
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Media tools" href="/tools/media-overview" icon="photo-film">
|
||||
Audio, image, and video processing pipeline overview.
|
||||
</Card>
|
||||
<Card title="Configuration" href="/gateway/configuration" icon="gear">
|
||||
Full config reference including media tool settings.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
Common issues and debugging steps.
|
||||
</Card>
|
||||
<Card title="FAQ" href="/help/faq" icon="circle-question">
|
||||
Frequently asked questions about OpenClaw setup.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
100
docs/providers/deepinfra.md
Normal file
100
docs/providers/deepinfra.md
Normal file
@@ -0,0 +1,100 @@
|
||||
---
|
||||
summary: "Use DeepInfra's unified API to access the most popular open source and frontier models in OpenClaw"
|
||||
read_when:
|
||||
- You want a single API key for the top open source LLMs
|
||||
- You want to run models via DeepInfra's API in OpenClaw
|
||||
title: "DeepInfra"
|
||||
---
|
||||
|
||||
DeepInfra routes requests to popular open source and frontier models behind a
|
||||
single OpenAI-compatible endpoint and API key. Most OpenAI SDKs work against
|
||||
it by switching the base URL.
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/deepinfra-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Get an API key
|
||||
|
||||
1. Sign in at [deepinfra.com](https://deepinfra.com/)
|
||||
2. Go to Dashboard / Keys and generate a key, or use the auto-created one
|
||||
|
||||
## CLI setup
|
||||
|
||||
```bash
|
||||
openclaw onboard --deepinfra-api-key <key>
|
||||
```
|
||||
|
||||
Or set the environment variable:
|
||||
|
||||
```bash
|
||||
export DEEPINFRA_API_KEY="<your-deepinfra-api-key>" # pragma: allowlist secret
|
||||
```
|
||||
|
||||
## Config snippet
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { DEEPINFRA_API_KEY: "<your-deepinfra-api-key>" }, // pragma: allowlist secret
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "deepinfra/deepseek-ai/DeepSeek-V4-Flash" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Supported surfaces
|
||||
|
||||
Chat, image generation, and video generation refresh their model catalogs
|
||||
live from `https://api.deepinfra.com/v1/openai/models?sort_by=openclaw&filter=with_meta`
|
||||
once `DEEPINFRA_API_KEY` is configured. Other surfaces use the static
|
||||
defaults below until they move onto the same live catalog.
|
||||
|
||||
| Surface | Default model | OpenClaw config/tool |
|
||||
| ------------------------ | ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------- |
|
||||
| Chat / model provider | first chat-tagged entry from live catalog (static fallback `deepseek-ai/DeepSeek-V4-Flash`) | `agents.defaults.model` |
|
||||
| Image generation/editing | first `image-gen`-tagged entry from live catalog (static fallback `black-forest-labs/FLUX-1-schnell`) | `image_generate`, `agents.defaults.imageGenerationModel` |
|
||||
| Media understanding | `moonshotai/Kimi-K2.5` for images | inbound image understanding |
|
||||
| Speech-to-text | `openai/whisper-large-v3-turbo` | inbound audio transcription |
|
||||
| Text-to-speech | `hexgrad/Kokoro-82M` | `messages.tts.provider: "deepinfra"` |
|
||||
| Video generation | static fallback `Pixverse/Pixverse-T2V` (no live video-gen rows from DeepInfra today) | `video_generate`, `agents.defaults.videoGenerationModel` |
|
||||
| Memory embeddings | `BAAI/bge-m3` | `agents.defaults.memorySearch.provider: "deepinfra"` |
|
||||
|
||||
DeepInfra also exposes reranking, classification, object-detection, and other
|
||||
native model types. OpenClaw has no provider contract for those categories
|
||||
yet, so this plugin does not register them.
|
||||
|
||||
## Available models
|
||||
|
||||
OpenClaw discovers DeepInfra models dynamically once a key is configured. Use
|
||||
`/models deepinfra` or `openclaw models list --provider deepinfra` to see the
|
||||
current list.
|
||||
|
||||
Any model on [deepinfra.com](https://deepinfra.com/) works with the
|
||||
`deepinfra/` prefix:
|
||||
|
||||
```text
|
||||
deepinfra/deepseek-ai/DeepSeek-V4-Flash
|
||||
deepinfra/deepseek-ai/DeepSeek-V3.2
|
||||
deepinfra/MiniMaxAI/MiniMax-M2.5
|
||||
deepinfra/moonshotai/Kimi-K2.5
|
||||
deepinfra/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B
|
||||
deepinfra/zai-org/GLM-5.1
|
||||
...and many more
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
- Model refs are `deepinfra/<provider>/<model>` (for example `deepinfra/Qwen/Qwen3-Max`).
|
||||
- Default chat model: `deepinfra/deepseek-ai/DeepSeek-V4-Flash`
|
||||
- Base URL: `https://api.deepinfra.com/v1/openai`
|
||||
- Native video generation uses `https://api.deepinfra.com/v1/inference/<model>`.
|
||||
|
||||
## Related
|
||||
|
||||
- [Model providers](/concepts/model-providers)
|
||||
- [All providers](/providers/index)
|
||||
147
docs/providers/deepseek.md
Normal file
147
docs/providers/deepseek.md
Normal file
@@ -0,0 +1,147 @@
|
||||
---
|
||||
summary: "DeepSeek setup (auth + model selection)"
|
||||
title: "DeepSeek"
|
||||
read_when:
|
||||
- You want to use DeepSeek with OpenClaw
|
||||
- You need the API key env var or CLI auth choice
|
||||
---
|
||||
|
||||
[DeepSeek](https://www.deepseek.com) provides powerful AI models with an OpenAI-compatible API.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | -------------------------- |
|
||||
| Provider | `deepseek` |
|
||||
| Auth | `DEEPSEEK_API_KEY` |
|
||||
| API | OpenAI-compatible |
|
||||
| Base URL | `https://api.deepseek.com` |
|
||||
|
||||
## Install plugin
|
||||
|
||||
Install the official plugin, then restart Gateway:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/deepseek-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get your API key">
|
||||
Create an API key at [platform.deepseek.com](https://platform.deepseek.com/api_keys).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice deepseek-api-key
|
||||
```
|
||||
|
||||
Prompts for your API key and sets `deepseek/deepseek-v4-flash` as the default model.
|
||||
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider deepseek
|
||||
```
|
||||
|
||||
To inspect the plugin's static catalog without a running Gateway:
|
||||
|
||||
```bash
|
||||
openclaw models list --all --provider deepseek
|
||||
```
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Non-interactive setup">
|
||||
For scripted or headless installations, pass all flags directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice deepseek-api-key \
|
||||
--deepseek-api-key "$DEEPSEEK_API_KEY" \
|
||||
--skip-health \
|
||||
--accept-risk
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Warning>
|
||||
If Gateway runs as a daemon (launchd/systemd), make sure `DEEPSEEK_API_KEY` is
|
||||
available to that process (for example, in `~/.openclaw/.env` or via
|
||||
`env.shellEnv`).
|
||||
</Warning>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
| Model ref | Name | Input | Context | Max output | Notes |
|
||||
| ---------------------------- | ----------------- | ----- | --------- | ---------- | ------------------------------------------ |
|
||||
| `deepseek/deepseek-v4-flash` | DeepSeek V4 Flash | text | 1,000,000 | 384,000 | Default model; V4 thinking-capable surface |
|
||||
| `deepseek/deepseek-v4-pro` | DeepSeek V4 Pro | text | 1,000,000 | 384,000 | V4 thinking-capable surface |
|
||||
| `deepseek/deepseek-chat` | DeepSeek Chat | text | 131,072 | 8,192 | DeepSeek V3.2 non-thinking surface |
|
||||
| `deepseek/deepseek-reasoner` | DeepSeek Reasoner | text | 131,072 | 65,536 | Reasoning-enabled V3.2 surface |
|
||||
|
||||
<Tip>
|
||||
V4 models support DeepSeek's `thinking` control. OpenClaw also replays
|
||||
DeepSeek `reasoning_content` on follow-up turns so thinking sessions with tool
|
||||
calls can continue.
|
||||
Use `/think xhigh` or `/think max` with DeepSeek V4 models to request DeepSeek's
|
||||
maximum `reasoning_effort`; both map to `"max"`.
|
||||
</Tip>
|
||||
|
||||
## Thinking and tools
|
||||
|
||||
DeepSeek V4 thinking sessions require replayed assistant messages from a
|
||||
thinking-enabled turn to include `reasoning_content` on follow-up requests.
|
||||
OpenClaw's DeepSeek plugin backfills that field automatically, so normal
|
||||
multi-turn tool use works on `deepseek/deepseek-v4-flash` and
|
||||
`deepseek/deepseek-v4-pro` even when history came from another
|
||||
OpenAI-compatible provider (no native `reasoning_content`) or from a plain
|
||||
assistant message. No `/new` required after switching providers mid-session.
|
||||
|
||||
When thinking is disabled (including the UI **None** selection), OpenClaw
|
||||
sends `thinking: { type: "disabled" }` and strips replayed `reasoning_content`
|
||||
from outgoing history, keeping the session on the non-thinking DeepSeek path.
|
||||
|
||||
Use `deepseek/deepseek-v4-flash` for the default fast path. Use
|
||||
`deepseek/deepseek-v4-pro` for the stronger model when you can accept higher
|
||||
cost or latency.
|
||||
|
||||
## Live testing
|
||||
|
||||
To run only the DeepSeek V4 direct-model checks from the modern model live suite:
|
||||
|
||||
```bash
|
||||
OPENCLAW_LIVE_PROVIDERS=deepseek \
|
||||
OPENCLAW_LIVE_MODELS="deepseek/deepseek-v4-flash,deepseek/deepseek-v4-pro" \
|
||||
pnpm test:live src/agents/models.profiles.live.test.ts
|
||||
```
|
||||
|
||||
Verifies both V4 models complete and that thinking/tool follow-up turns
|
||||
preserve the replay payload DeepSeek requires.
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { DEEPSEEK_API_KEY: "sk-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "deepseek/deepseek-v4-flash" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config reference for agents, models, and providers.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
311
docs/providers/ds4.md
Normal file
311
docs/providers/ds4.md
Normal file
@@ -0,0 +1,311 @@
|
||||
---
|
||||
summary: "Run OpenClaw through ds4, a local DeepSeek V4 Flash OpenAI-compatible server"
|
||||
read_when:
|
||||
- You want to run OpenClaw against antirez/ds4
|
||||
- You want a local DeepSeek V4 Flash backend with tool calls
|
||||
- You need the OpenClaw config for ds4-server
|
||||
title: "ds4"
|
||||
---
|
||||
|
||||
[ds4](https://github.com/antirez/ds4) serves DeepSeek V4 Flash from a local
|
||||
Metal backend with an OpenAI-compatible `/v1` API. OpenClaw connects to ds4
|
||||
through the generic `openai-completions` provider family.
|
||||
|
||||
ds4 is not a bundled OpenClaw provider plugin. Configure it under
|
||||
`models.providers.ds4`, then select `ds4/deepseek-v4-flash`.
|
||||
|
||||
| Property | Value |
|
||||
| ----------- | --------------------------------------------------------- |
|
||||
| Provider id | `ds4` |
|
||||
| Plugin | none (config-only) |
|
||||
| API | OpenAI-compatible Chat Completions (`openai-completions`) |
|
||||
| Base URL | `http://127.0.0.1:18000/v1` (suggested) |
|
||||
| Model id | `deepseek-v4-flash` |
|
||||
| Tool calls | OpenAI-style `tools` / `tool_calls` |
|
||||
| Reasoning | DeepSeek-style `thinking` and `reasoning_effort` |
|
||||
|
||||
## Requirements
|
||||
|
||||
- macOS with Metal support.
|
||||
- A working ds4 checkout with `ds4-server` and the DeepSeek V4 Flash GGUF file.
|
||||
- Enough memory for the context you choose; larger `--ctx` values allocate more
|
||||
KV memory at server startup.
|
||||
|
||||
<Warning>
|
||||
OpenClaw agent turns include tool schemas and workspace context. A tiny context
|
||||
such as `--ctx 4096` can pass direct curl tests but fail full agent runs with
|
||||
`500 prompt exceeds context`. Use at least `--ctx 32768` for agent and tool
|
||||
smoke tests. Use `--ctx 393216` only with enough memory and to enable ds4
|
||||
Think Max.
|
||||
</Warning>
|
||||
|
||||
## Quickstart
|
||||
|
||||
<Steps>
|
||||
<Step title="Start ds4-server">
|
||||
Replace `<DS4_DIR>` with your ds4 checkout path.
|
||||
|
||||
```bash
|
||||
<DS4_DIR>/ds4-server \
|
||||
--model <DS4_DIR>/ds4flash.gguf \
|
||||
--host 127.0.0.1 \
|
||||
--port 18000 \
|
||||
--ctx 32768 \
|
||||
--tokens 128
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the OpenAI-compatible endpoint">
|
||||
```bash
|
||||
curl http://127.0.0.1:18000/v1/models
|
||||
```
|
||||
|
||||
The response should include `deepseek-v4-flash`.
|
||||
|
||||
</Step>
|
||||
<Step title="Add the OpenClaw provider config">
|
||||
Add the config from [Full config](#full-config), then run a one-shot model
|
||||
check:
|
||||
|
||||
```bash
|
||||
openclaw infer model run \
|
||||
--local \
|
||||
--model ds4/deepseek-v4-flash \
|
||||
--thinking off \
|
||||
--prompt "Reply with exactly: openclaw-ds4-ok" \
|
||||
--json
|
||||
```
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Full config
|
||||
|
||||
Use this config when ds4 is already running on `127.0.0.1:18000`.
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "ds4/deepseek-v4-flash" },
|
||||
models: {
|
||||
"ds4/deepseek-v4-flash": {
|
||||
alias: "DS4 local",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
ds4: {
|
||||
baseUrl: "http://127.0.0.1:18000/v1",
|
||||
apiKey: "ds4-local",
|
||||
api: "openai-completions",
|
||||
timeoutSeconds: 300,
|
||||
models: [
|
||||
{
|
||||
id: "deepseek-v4-flash",
|
||||
name: "DeepSeek V4 Flash (ds4)",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 32768,
|
||||
maxTokens: 128,
|
||||
compat: {
|
||||
supportsUsageInStreaming: true,
|
||||
supportsReasoningEffort: true,
|
||||
maxTokensField: "max_tokens",
|
||||
supportsStrictMode: false,
|
||||
thinkingFormat: "deepseek",
|
||||
supportedReasoningEfforts: ["low", "medium", "high", "xhigh"],
|
||||
},
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Keep `contextWindow` aligned with `ds4-server --ctx`. Keep `maxTokens` aligned
|
||||
with `--tokens` unless you intentionally want OpenClaw to request less output
|
||||
than the server default.
|
||||
|
||||
## On-demand startup
|
||||
|
||||
OpenClaw can start ds4 only when a `ds4/...` model is selected. Add
|
||||
`localService` to the same provider entry:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
ds4: {
|
||||
baseUrl: "http://127.0.0.1:18000/v1",
|
||||
apiKey: "ds4-local",
|
||||
api: "openai-completions",
|
||||
timeoutSeconds: 300,
|
||||
localService: {
|
||||
command: "<DS4_DIR>/ds4-server",
|
||||
args: [
|
||||
"--model",
|
||||
"<DS4_DIR>/ds4flash.gguf",
|
||||
"--host",
|
||||
"127.0.0.1",
|
||||
"--port",
|
||||
"18000",
|
||||
"--ctx",
|
||||
"32768",
|
||||
"--tokens",
|
||||
"128",
|
||||
],
|
||||
cwd: "<DS4_DIR>",
|
||||
healthUrl: "http://127.0.0.1:18000/v1/models",
|
||||
readyTimeoutMs: 300000,
|
||||
idleStopMs: 0,
|
||||
},
|
||||
models: [
|
||||
{
|
||||
id: "deepseek-v4-flash",
|
||||
name: "DeepSeek V4 Flash (ds4)",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 32768,
|
||||
maxTokens: 128,
|
||||
compat: {
|
||||
supportsUsageInStreaming: true,
|
||||
supportsReasoningEffort: true,
|
||||
maxTokensField: "max_tokens",
|
||||
supportsStrictMode: false,
|
||||
thinkingFormat: "deepseek",
|
||||
supportedReasoningEfforts: ["low", "medium", "high", "xhigh"],
|
||||
},
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`command` must be an absolute executable path. Shell lookup and `~` expansion
|
||||
are not used. See [Local model services](/gateway/local-model-services) for
|
||||
every `localService` field.
|
||||
|
||||
## Think Max
|
||||
|
||||
ds4 applies Think Max only when both are true:
|
||||
|
||||
- `ds4-server` starts with `--ctx 393216` or higher.
|
||||
- The request uses `reasoning_effort: "max"` (or the equivalent ds4 effort field).
|
||||
|
||||
If you run that large context, update both the server flags and OpenClaw model
|
||||
metadata:
|
||||
|
||||
```json5
|
||||
{
|
||||
contextWindow: 393216,
|
||||
maxTokens: 384000,
|
||||
compat: {
|
||||
supportsUsageInStreaming: true,
|
||||
supportsReasoningEffort: true,
|
||||
maxTokensField: "max_tokens",
|
||||
supportsStrictMode: false,
|
||||
thinkingFormat: "deepseek",
|
||||
supportedReasoningEfforts: ["low", "medium", "high", "xhigh", "max"],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Test
|
||||
|
||||
Direct HTTP check, bypassing OpenClaw:
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:18000/v1/chat/completions \
|
||||
-H 'content-type: application/json' \
|
||||
-d '{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"Reply with exactly: ds4-ok"}],"max_tokens":16,"stream":false,"thinking":{"type":"disabled"}}'
|
||||
```
|
||||
|
||||
OpenClaw model routing (same as the Quickstart check):
|
||||
|
||||
```bash
|
||||
openclaw infer model run \
|
||||
--local \
|
||||
--model ds4/deepseek-v4-flash \
|
||||
--thinking off \
|
||||
--prompt "Reply with exactly: openclaw-ds4-ok" \
|
||||
--json
|
||||
```
|
||||
|
||||
Full agent and tool-call smoke test, with context of at least 32768:
|
||||
|
||||
```bash
|
||||
openclaw agent \
|
||||
--local \
|
||||
--session-id ds4-tool-smoke \
|
||||
--model ds4/deepseek-v4-flash \
|
||||
--thinking off \
|
||||
--message "Use the shell command pwd once, then reply exactly: tool-ok <output>" \
|
||||
--json \
|
||||
--timeout 240
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- `executionTrace.winnerProvider` is `ds4`
|
||||
- `executionTrace.winnerModel` is `deepseek-v4-flash`
|
||||
- `toolSummary.calls` is at least `1`
|
||||
- `finalAssistantVisibleText` starts with `tool-ok`
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="curl /v1/models cannot connect">
|
||||
ds4 is not running or not bound to the host/port in `baseUrl`. Start
|
||||
`ds4-server`, then retry:
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:18000/v1/models
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="500 prompt exceeds context">
|
||||
The configured `--ctx` is too small for the OpenClaw turn. Raise
|
||||
`ds4-server --ctx`, then update `models.providers.ds4.models[].contextWindow`
|
||||
to match. Full agent turns with tools need substantially more context than a
|
||||
direct one-message curl request.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Think Max does not activate">
|
||||
ds4 only uses Think Max when `--ctx` is at least `393216` and the request
|
||||
asks for `reasoning_effort: "max"`. Smaller contexts fall back to high
|
||||
reasoning.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="The first request is slow">
|
||||
ds4 has a cold Metal residency and model warmup phase. Set
|
||||
`localService.readyTimeoutMs: 300000` when OpenClaw starts the server on
|
||||
demand.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Local model services" href="/gateway/local-model-services" icon="play">
|
||||
Start local model servers on demand before model requests.
|
||||
</Card>
|
||||
<Card title="Local models" href="/gateway/local-models" icon="server">
|
||||
Choose and operate local model backends.
|
||||
</Card>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Configure provider refs, auth, and failover.
|
||||
</Card>
|
||||
<Card title="DeepSeek" href="/providers/deepseek" icon="brain">
|
||||
Native DeepSeek provider behavior and thinking controls.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
131
docs/providers/elevenlabs.md
Normal file
131
docs/providers/elevenlabs.md
Normal file
@@ -0,0 +1,131 @@
|
||||
---
|
||||
summary: "Use ElevenLabs speech, Scribe STT, and realtime transcription with OpenClaw"
|
||||
read_when:
|
||||
- You want ElevenLabs text-to-speech in OpenClaw
|
||||
- You want ElevenLabs Scribe speech-to-text for audio attachments
|
||||
- You want ElevenLabs realtime transcription for Voice Call or Google Meet
|
||||
title: "ElevenLabs"
|
||||
---
|
||||
|
||||
OpenClaw uses ElevenLabs for text-to-speech, batch speech-to-text with Scribe
|
||||
v2, and streaming STT with Scribe v2 Realtime. The plugin is bundled and
|
||||
enabled by default; no `plugins install` step is needed.
|
||||
|
||||
| Capability | OpenClaw surface | Default |
|
||||
| ------------------------ | -------------------------------------------------------------------- | ------------------------ |
|
||||
| Text-to-speech | `messages.tts` / `talk` | `eleven_multilingual_v2` |
|
||||
| Batch speech-to-text | `tools.media.audio` | `scribe_v2` |
|
||||
| Streaming speech-to-text | Voice Call streaming or Google Meet `realtime.transcriptionProvider` | `scribe_v2_realtime` |
|
||||
|
||||
## Authentication
|
||||
|
||||
Set `ELEVENLABS_API_KEY` in the environment. `XI_API_KEY` is also accepted for
|
||||
compatibility with existing ElevenLabs tooling.
|
||||
|
||||
```bash
|
||||
export ELEVENLABS_API_KEY="..."
|
||||
```
|
||||
|
||||
## Text-to-speech
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
providers: {
|
||||
elevenlabs: {
|
||||
apiKey: "${ELEVENLABS_API_KEY}",
|
||||
voiceId: "pMsXgVXv3BLzUgSXRplE",
|
||||
modelId: "eleven_multilingual_v2",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Set `modelId` to `eleven_v3` to use ElevenLabs v3 TTS. OpenClaw keeps
|
||||
`eleven_multilingual_v2` as the default for existing installs.
|
||||
|
||||
Discord voice channels use ElevenLabs' streaming TTS endpoint when ElevenLabs
|
||||
is the selected `voice.tts`/`messages.tts` provider: playback starts from the
|
||||
returned audio stream instead of waiting for OpenClaw to download the whole
|
||||
audio file first. `latencyTier` maps to ElevenLabs' `optimize_streaming_latency`
|
||||
query parameter for models that accept it; OpenClaw omits that parameter for
|
||||
`eleven_v3`, which rejects it.
|
||||
|
||||
## Speech-to-text
|
||||
|
||||
Use Scribe v2 for inbound audio attachments and short recorded voice segments:
|
||||
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
enabled: true,
|
||||
models: [{ provider: "elevenlabs", model: "scribe_v2" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw sends multipart audio to ElevenLabs `/v1/speech-to-text` with
|
||||
`model_id: "scribe_v2"`. Language hints map to `language_code` when present.
|
||||
|
||||
## Streaming STT
|
||||
|
||||
The bundled `elevenlabs` plugin registers Scribe v2 Realtime for Voice Call and
|
||||
Google Meet agent-mode streaming transcription.
|
||||
|
||||
| Setting | Config path | Default |
|
||||
| --------------- | ------------------------------------------------------------------------- | ------------------------------------------------- |
|
||||
| API key | `plugins.entries.voice-call.config.streaming.providers.elevenlabs.apiKey` | Falls back to `ELEVENLABS_API_KEY` / `XI_API_KEY` |
|
||||
| Model | `...elevenlabs.modelId` | `scribe_v2_realtime` |
|
||||
| Audio format | `...elevenlabs.audioFormat` | `ulaw_8000` |
|
||||
| Sample rate | `...elevenlabs.sampleRate` | `8000` |
|
||||
| Commit strategy | `...elevenlabs.commitStrategy` | `vad` |
|
||||
| Language | `...elevenlabs.languageCode` | (unset) |
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"voice-call": {
|
||||
config: {
|
||||
streaming: {
|
||||
enabled: true,
|
||||
provider: "elevenlabs",
|
||||
providers: {
|
||||
elevenlabs: {
|
||||
apiKey: "${ELEVENLABS_API_KEY}",
|
||||
audioFormat: "ulaw_8000",
|
||||
commitStrategy: "vad",
|
||||
languageCode: "en",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
Voice Call receives Twilio media as 8 kHz G.711 u-law. The ElevenLabs realtime
|
||||
provider defaults to `ulaw_8000`, so telephony frames can be forwarded without
|
||||
transcoding.
|
||||
</Note>
|
||||
|
||||
For Google Meet agent mode, set
|
||||
`plugins.entries.google-meet.config.realtime.transcriptionProvider` to
|
||||
`"elevenlabs"` and configure the same provider block under
|
||||
`plugins.entries.google-meet.config.realtime.providers.elevenlabs`.
|
||||
|
||||
## Related
|
||||
|
||||
- [Text-to-speech](/tools/tts)
|
||||
- [Google Meet](/plugins/google-meet)
|
||||
- [Model selection](/concepts/model-providers)
|
||||
267
docs/providers/fal.md
Normal file
267
docs/providers/fal.md
Normal file
@@ -0,0 +1,267 @@
|
||||
---
|
||||
summary: "fal image, video, and music generation setup in OpenClaw"
|
||||
title: "Fal"
|
||||
read_when:
|
||||
- You want to use fal image generation in OpenClaw
|
||||
- You need the FAL_KEY auth flow
|
||||
- You want fal defaults for image_generate, video_generate, or music_generate
|
||||
---
|
||||
|
||||
OpenClaw ships a bundled `fal` provider for hosted image, video, and music
|
||||
generation.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ------------------------------------------------------------------------------- |
|
||||
| Provider | `fal` |
|
||||
| Auth | `FAL_KEY` (canonical; `FAL_API_KEY` also works as a fallback) |
|
||||
| API | fal model endpoints (`https://fal.run`; video jobs use `https://queue.fal.run`) |
|
||||
| Base URL | Override with `models.providers.fal.baseUrl` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set the API key">
|
||||
```bash
|
||||
openclaw onboard --auth-choice fal-api-key
|
||||
```
|
||||
|
||||
Non-interactive setups can pass `--fal-api-key <key>` or export `FAL_KEY`.
|
||||
Onboarding also sets `fal/fal-ai/flux/dev` as the default image model when
|
||||
none is configured.
|
||||
|
||||
</Step>
|
||||
<Step title="Set a default image model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "fal/fal-ai/flux/dev",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Image generation
|
||||
|
||||
The bundled `fal` image-generation provider defaults to
|
||||
`fal/fal-ai/flux/dev`.
|
||||
|
||||
| Capability | Value |
|
||||
| -------------- | ------------------------------------------------------------------ |
|
||||
| Max images | 4 per request; Krea 2: 1 per request |
|
||||
| Size overrides | `1024x1024`, `1024x1536`, `1536x1024`, `1024x1792`, `1792x1024` |
|
||||
| Aspect ratio | Supported everywhere except Flux image-to-image |
|
||||
| Resolution | `1K`, `2K`, `4K` (per-model limits below) |
|
||||
| Output format | `png` (default) or `jpeg`; Krea 2 rejects `outputFormat` overrides |
|
||||
|
||||
Edit requests (reference images via the shared `image` / `images` parameters)
|
||||
route to a per-model edit endpoint with per-model reference limits:
|
||||
|
||||
| Model family | Model ref after `fal/` | Edit endpoint | Max reference images |
|
||||
| ------------------------- | -------------------------------------- | ----------------- | -------------------- |
|
||||
| Flux and other fal models | `fal-ai/flux/dev` (default) | `/image-to-image` | 1 |
|
||||
| GPT Image | `openai/gpt-image-*` | `/edit` | 10 |
|
||||
| Grok Imagine | `xai/grok-imagine-image` | `/edit` | 3 |
|
||||
| Nano Banana (legacy) | `fal-ai/nano-banana` | `/edit` | 3 |
|
||||
| Nano Banana 2 | `fal-ai/nano-banana-*` | `/edit` | 14 |
|
||||
| Nano Banana 2 Lite | `google/nano-banana-2-lite` | `/edit` | 14 |
|
||||
| Krea 2 | `krea/v2/{medium,large}/text-to-image` | none (style refs) | 10 style references |
|
||||
|
||||
<Warning>
|
||||
Flux image-to-image requests do **not** support `aspectRatio` overrides. GPT
|
||||
Image and Nano Banana 2 edit requests use fal's `/edit` endpoint and accept
|
||||
aspect-ratio hints. Nano Banana 2 also accepts extra-native wide/tall ratios
|
||||
such as `4:1`, `1:4`, `8:1`, and `1:8`; Krea 2 validates its own smaller
|
||||
aspect-ratio subset. Grok Imagine has its own ratio list (including `2:1`,
|
||||
`20:9`, `19.5:9`, and their inverses) and only accepts `1K`/`2K` resolutions;
|
||||
legacy Nano Banana and Nano Banana 2 Lite reject `resolution` overrides.
|
||||
</Warning>
|
||||
|
||||
Krea 2 models use fal's native Krea payload schema. OpenClaw sends
|
||||
`aspect_ratio`, `creativity`, and `image_style_references` instead of the
|
||||
generic `image_size` / edit-endpoint payload used by Flux. The model refs are:
|
||||
|
||||
- `fal/krea/v2/medium/text-to-image`
|
||||
- `fal/krea/v2/large/text-to-image`
|
||||
|
||||
Use Medium for faster expressive illustration, anime, painting, and artistic
|
||||
styles. Use Large for slower photoreal, raw texture, film grain, and detailed
|
||||
looks. Krea defaults to `fal.creativity: "medium"`; supported values are
|
||||
`raw`, `low`, `medium`, and `high`.
|
||||
|
||||
Krea 2 exposes aspect ratio, not `image_size`, in fal's request schema. Prefer
|
||||
`aspectRatio`; OpenClaw maps `size` to the closest supported Krea aspect ratio
|
||||
and rejects `resolution` for Krea rather than dropping it.
|
||||
|
||||
Use `outputFormat: "png"` when you want PNG output from fal models that expose
|
||||
`output_format`. fal does not declare an explicit transparent-background
|
||||
control in OpenClaw, so `background: "transparent"` is reported as an ignored
|
||||
override for fal models.
|
||||
Krea 2 endpoints do not expose an `output_format` request field through fal, so
|
||||
OpenClaw rejects `outputFormat` overrides for Krea requests.
|
||||
|
||||
To use Krea 2 Medium:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "fal/krea/v2/medium/text-to-image",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Video generation
|
||||
|
||||
The bundled `fal` video-generation provider defaults to
|
||||
`fal/fal-ai/minimax/video-01-live`.
|
||||
|
||||
| Capability | Value |
|
||||
| ---------- | ------------------------------------------------------------------ |
|
||||
| Modes | Text-to-video, single-image reference, Seedance reference-to-video |
|
||||
| Runtime | Queue-backed submit/status/result flow for long-running jobs |
|
||||
| Timeout | 20 minutes per job by default; status polled every 5 seconds |
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Available video models">
|
||||
**MiniMax (default):**
|
||||
|
||||
- `fal/fal-ai/minimax/video-01-live`
|
||||
|
||||
**HeyGen video-agent:**
|
||||
|
||||
- `fal/fal-ai/heygen/v2/video-agent`
|
||||
|
||||
**Kling and Wan:**
|
||||
|
||||
- `fal/fal-ai/kling-video/v2.1/master/text-to-video`
|
||||
- `fal/fal-ai/wan/v2.2-a14b/text-to-video`
|
||||
- `fal/fal-ai/wan/v2.2-a14b/image-to-video`
|
||||
|
||||
**Seedance 2.0:**
|
||||
|
||||
- `fal/bytedance/seedance-2.0/fast/text-to-video`
|
||||
- `fal/bytedance/seedance-2.0/fast/image-to-video`
|
||||
- `fal/bytedance/seedance-2.0/fast/reference-to-video`
|
||||
- `fal/bytedance/seedance-2.0/text-to-video`
|
||||
- `fal/bytedance/seedance-2.0/image-to-video`
|
||||
- `fal/bytedance/seedance-2.0/reference-to-video`
|
||||
|
||||
MiniMax Live and HeyGen requests send only the prompt plus an optional
|
||||
single reference image; other overrides are not forwarded. Seedance models
|
||||
accept `aspectRatio`, `size`, `resolution`, durations of 4-15 seconds, and
|
||||
an audio toggle.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Seedance 2.0 config example">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "fal/bytedance/seedance-2.0/fast/text-to-video",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Seedance 2.0 reference-to-video config example">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "fal/bytedance/seedance-2.0/fast/reference-to-video",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Reference-to-video accepts up to 9 images, 3 videos, and 3 audio references
|
||||
through the shared `video_generate` `images`, `videos`, and `audioRefs`
|
||||
parameters, with at most 12 total reference files. Audio references require
|
||||
at least one image or video reference in the same request.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="HeyGen video-agent config example">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "fal/fal-ai/heygen/v2/video-agent",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Music generation
|
||||
|
||||
The bundled `fal` plugin also registers a music-generation provider for the
|
||||
shared `music_generate` tool.
|
||||
|
||||
| Capability | Value |
|
||||
| ------------- | ------------------------------------------------------------------------------------------------------------------------ |
|
||||
| Default model | `fal/fal-ai/minimax-music/v2.6` |
|
||||
| Models | `fal-ai/minimax-music/v2.6` (mp3), `fal-ai/ace-step/prompt-to-audio` (wav), `fal-ai/stable-audio-25/text-to-audio` (wav) |
|
||||
| Max duration | 240 seconds |
|
||||
| Runtime | Synchronous request plus generated audio download |
|
||||
|
||||
Use fal as the default music provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
musicGenerationModel: {
|
||||
primary: "fal/fal-ai/minimax-music/v2.6",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`fal-ai/minimax-music/v2.6` supports explicit lyrics and instrumental mode,
|
||||
but not both in the same request. ACE-Step and Stable Audio are
|
||||
prompt-to-audio endpoints; choose them with the `model` override when you want
|
||||
those model families. ACE-Step rejects explicit lyrics; Stable Audio rejects
|
||||
both lyrics and instrumental mode.
|
||||
|
||||
<Tip>
|
||||
The tables and accordions above cover the model families the bundled fal
|
||||
provider special-cases. Other fal image endpoint ids can still be selected as
|
||||
the image model; they are treated like Flux (generic `image_size` payload, one
|
||||
reference image via `/image-to-image`).
|
||||
</Tip>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Image generation" href="/tools/image-generation" icon="image">
|
||||
Shared image tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Music generation" href="/tools/music-generation" icon="music">
|
||||
Shared music tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/config-agents#agent-defaults" icon="gear">
|
||||
Agent defaults including image, video, and music model selection.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
149
docs/providers/fireworks.md
Normal file
149
docs/providers/fireworks.md
Normal file
@@ -0,0 +1,149 @@
|
||||
---
|
||||
summary: "Fireworks setup (auth + model selection)"
|
||||
title: "Fireworks"
|
||||
read_when:
|
||||
- You want to use Fireworks with OpenClaw
|
||||
- You need the Fireworks API key env var or default model id
|
||||
- You are debugging Kimi thinking-off behavior on Fireworks
|
||||
---
|
||||
|
||||
[Fireworks](https://fireworks.ai) exposes open-weight and routed models through an OpenAI-compatible API. Install the official Fireworks provider plugin to use two pre-cataloged Kimi models and any Fireworks model or router id at runtime.
|
||||
|
||||
| Property | Value |
|
||||
| --------------- | ------------------------------------------------------ |
|
||||
| Provider id | `fireworks` (alias: `fireworks-ai`) |
|
||||
| Package | `@openclaw/fireworks-provider` |
|
||||
| Auth env var | `FIREWORKS_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice fireworks-api-key` |
|
||||
| Direct CLI flag | `--fireworks-api-key <key>` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://api.fireworks.ai/inference/v1` |
|
||||
| Default model | `fireworks/accounts/fireworks/routers/kimi-k2p5-turbo` |
|
||||
| Default alias | `Kimi K2.5 Turbo` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the plugin">
|
||||
```bash
|
||||
openclaw plugins install @openclaw/fireworks-provider
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set the Fireworks API key">
|
||||
<CodeGroup>
|
||||
|
||||
```bash Onboarding
|
||||
openclaw onboard --auth-choice fireworks-api-key
|
||||
```
|
||||
|
||||
```bash Direct flag
|
||||
openclaw onboard --non-interactive \
|
||||
--auth-choice fireworks-api-key \
|
||||
--fireworks-api-key "$FIREWORKS_API_KEY"
|
||||
```
|
||||
|
||||
```bash Env only
|
||||
export FIREWORKS_API_KEY=fw-...
|
||||
```
|
||||
|
||||
</CodeGroup>
|
||||
|
||||
Onboarding stores the key against the `fireworks` provider in your auth profiles and sets the **Fire Pass** Kimi K2.5 Turbo router as the default model.
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider fireworks
|
||||
```
|
||||
|
||||
The list should include `Kimi K2.6` and `Kimi K2.5 Turbo (Fire Pass)`. If `FIREWORKS_API_KEY` is unresolved, `openclaw models status --json` reports the missing credential under `auth.unusableProfiles`.
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Non-interactive setup
|
||||
|
||||
For scripted or CI installs, pass everything on the command line:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice fireworks-api-key \
|
||||
--fireworks-api-key "$FIREWORKS_API_KEY" \
|
||||
--skip-health \
|
||||
--accept-risk
|
||||
```
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
| Model ref | Name | Input | Context | Max output | Thinking |
|
||||
| ------------------------------------------------------ | --------------------------- | ------------ | ------- | ---------- | -------------------- |
|
||||
| `fireworks/accounts/fireworks/models/kimi-k2p6` | Kimi K2.6 | text + image | 262,144 | 262,144 | Forced off |
|
||||
| `fireworks/accounts/fireworks/routers/kimi-k2p5-turbo` | Kimi K2.5 Turbo (Fire Pass) | text + image | 256,000 | 256,000 | Forced off (default) |
|
||||
|
||||
<Note>
|
||||
OpenClaw pins all Fireworks Kimi models to `thinking: off` because Kimi on Fireworks can leak chain-of-thought into the visible reply unless the request explicitly disables thinking. Routing the same model through [Moonshot](/providers/moonshot) directly preserves Kimi reasoning output. See [thinking modes](/tools/thinking) for switching between providers.
|
||||
</Note>
|
||||
|
||||
## Custom Fireworks model ids
|
||||
|
||||
OpenClaw accepts any Fireworks model or router id at runtime. Use the exact id shown by Fireworks and prefix it with `fireworks/`. Dynamic resolution clones the Fire Pass template (text + image input, OpenAI-compatible API, default cost zero) and disables thinking automatically when the id matches the Kimi pattern. GLM dynamic ids are marked text-only unless you configure a custom model entry with image input.
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: {
|
||||
primary: "fireworks/accounts/fireworks/models/<your-model-id>",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="How model id prefixing works">
|
||||
Every Fireworks model ref in OpenClaw starts with `fireworks/` followed by the exact id or router path from the Fireworks platform. For example:
|
||||
|
||||
- Router model: `fireworks/accounts/fireworks/routers/kimi-k2p5-turbo`
|
||||
- Direct model: `fireworks/accounts/fireworks/models/<model-name>`
|
||||
|
||||
OpenClaw strips the `fireworks/` prefix when constructing the API request and sends the remaining path to the Fireworks endpoint as the OpenAI-compatible `model` field.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Why thinking is forced off for Kimi">
|
||||
Fireworks serves Kimi without a separate reasoning channel, so chain-of-thought can surface in the visible `content` stream. On every Fireworks Kimi request OpenClaw sends `thinking: { type: "disabled" }` and strips `reasoning`, `reasoning_effort`, and `reasoningEffort` from the payload (`extensions/fireworks/stream.ts`). The provider policy (`extensions/fireworks/thinking-policy.ts`) advertises only the `off` thinking level for Kimi model ids, so manual `/think` switches and provider-policy surfaces stay aligned with the runtime contract.
|
||||
|
||||
To use Kimi reasoning end-to-end, configure the [Moonshot provider](/providers/moonshot) and route the same model through it.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment availability for the daemon">
|
||||
If the Gateway runs as a managed service (launchd, systemd, Docker), the Fireworks key must be visible to that process — not just to your interactive shell.
|
||||
|
||||
<Warning>
|
||||
A key exported only in an interactive shell will not help a launchd or systemd daemon unless that environment is imported there too. Set the key in `~/.openclaw/.env` or via `env.shellEnv` to make it readable from the gateway process.
|
||||
</Warning>
|
||||
|
||||
OpenClaw loads `~/.openclaw/.env` when it loads config, so keys stored there reach managed gateway services on every platform. Restart the gateway (or re-run `openclaw doctor --fix`) after rotating the key.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Thinking modes" href="/tools/thinking" icon="brain">
|
||||
`/think` levels, provider policies, and routing reasoning-capable models.
|
||||
</Card>
|
||||
<Card title="Moonshot" href="/providers/moonshot" icon="moon">
|
||||
Run Kimi with native thinking output through Moonshot's own API.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
General troubleshooting and FAQ.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
264
docs/providers/github-copilot.md
Normal file
264
docs/providers/github-copilot.md
Normal file
@@ -0,0 +1,264 @@
|
||||
---
|
||||
summary: "Sign in to GitHub Copilot from OpenClaw using the device flow or non-interactive token import"
|
||||
read_when:
|
||||
- You want to use GitHub Copilot as a model provider
|
||||
- You need the `openclaw models auth login-github-copilot` flow
|
||||
- You are choosing between the built-in Copilot provider, Copilot SDK harness, and Copilot Proxy
|
||||
title: "GitHub Copilot"
|
||||
---
|
||||
|
||||
GitHub Copilot is GitHub's AI coding assistant. It provides access to Copilot
|
||||
models for your GitHub account and plan. OpenClaw can use Copilot as a model
|
||||
provider or agent runtime in three different ways.
|
||||
|
||||
## Three ways to use Copilot in OpenClaw
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Built-in provider (github-copilot)">
|
||||
Use the native device-login flow to obtain a GitHub token, then exchange it for
|
||||
Copilot API tokens when OpenClaw runs. This is the **default** and simplest path
|
||||
because it does not require VS Code.
|
||||
|
||||
<Steps>
|
||||
<Step title="Run the login command">
|
||||
```bash
|
||||
openclaw models auth login-github-copilot
|
||||
```
|
||||
|
||||
You will be prompted to visit a URL and enter a one-time code. Keep the
|
||||
terminal open until it completes.
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```bash
|
||||
openclaw models set github-copilot/claude-opus-4.7
|
||||
```
|
||||
|
||||
Or in config:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: { model: { primary: "github-copilot/claude-opus-4.7" } },
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Copilot SDK harness plugin (copilot)">
|
||||
Install the external `@openclaw/copilot` plugin when you want GitHub's
|
||||
Copilot CLI and SDK to own the low-level agent loop for selected
|
||||
`github-copilot/*` models.
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/copilot
|
||||
```
|
||||
|
||||
Then opt a model or provider into the runtime:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: "github-copilot/gpt-5.5",
|
||||
models: {
|
||||
"github-copilot/gpt-5.5": {
|
||||
agentRuntime: { id: "copilot" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Choose this when you want native Copilot CLI sessions, SDK-managed thread
|
||||
state, and Copilot-owned compaction for those agent turns. Without the
|
||||
explicit `agentRuntime` opt-in, `github-copilot/*` models keep using the
|
||||
built-in provider. See [Copilot SDK harness](/plugins/copilot) for the full
|
||||
runtime contract.
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Copilot Proxy plugin (copilot-proxy)">
|
||||
Use the **Copilot Proxy** VS Code extension as a local bridge. OpenClaw talks to
|
||||
the proxy's `/v1` endpoint (default `http://localhost:3000/v1`) and uses the
|
||||
model list you configure.
|
||||
|
||||
The `copilot-proxy` plugin ships with OpenClaw and is enabled by default.
|
||||
Configure the base URL and model ids with:
|
||||
|
||||
```bash
|
||||
openclaw models auth login --provider copilot-proxy --set-default
|
||||
```
|
||||
|
||||
<Note>
|
||||
Choose this when you already run Copilot Proxy in VS Code or need to route
|
||||
through it. The VS Code extension must stay running.
|
||||
</Note>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Optional flags
|
||||
|
||||
| Command | Flag | Description |
|
||||
| ---------------------------------------------------------------------- | --------------- | ---------------------------------------------------- |
|
||||
| `openclaw models auth login-github-copilot` | `--yes` | Overwrite an existing auth profile without prompting |
|
||||
| `openclaw models auth login --provider github-copilot --method device` | `--set-default` | Also apply the provider's recommended default model |
|
||||
|
||||
```bash
|
||||
# Skip the re-login confirmation
|
||||
openclaw models auth login-github-copilot --yes
|
||||
|
||||
# Login and set the default model in one step
|
||||
openclaw models auth login --provider github-copilot --method device --set-default
|
||||
```
|
||||
|
||||
## Non-interactive onboarding
|
||||
|
||||
The device-login flow requires an interactive TTY. For headless setup, import
|
||||
an existing GitHub OAuth access token with `openclaw onboard --non-interactive`:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive --accept-risk \
|
||||
--auth-choice github-copilot \
|
||||
--github-copilot-token "$COPILOT_GITHUB_TOKEN" \
|
||||
--skip-channels --skip-health
|
||||
```
|
||||
|
||||
You can also omit `--auth-choice`; passing `--github-copilot-token` infers the
|
||||
GitHub Copilot provider auth choice. If the flag is omitted, onboarding falls
|
||||
back to `COPILOT_GITHUB_TOKEN`, `GH_TOKEN`, then `GITHUB_TOKEN`. Use
|
||||
`--secret-input-mode ref` with `COPILOT_GITHUB_TOKEN` set to store an env-backed
|
||||
`tokenRef` instead of plaintext in `auth-profiles.json`.
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Interactive TTY required">
|
||||
The device-login flow requires an interactive TTY. Run it directly in a
|
||||
terminal, not in a non-interactive script or CI pipeline.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Model availability depends on your plan">
|
||||
Copilot model availability depends on your GitHub plan. If a model is
|
||||
rejected, try another ID (for example `github-copilot/gpt-5.5`). See
|
||||
GitHub's [supported models per Copilot plan](https://docs.github.com/en/copilot/reference/ai-models/supported-models#supported-ai-models-per-copilot-plan)
|
||||
for the current model list.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Live catalog refresh from the Copilot API">
|
||||
Once the device-login (or env-var) auth path has resolved a GitHub token,
|
||||
OpenClaw refreshes the model catalog on demand from `${baseUrl}/models`
|
||||
(the same endpoint VS Code Copilot uses) so the runtime tracks
|
||||
per-account entitlement and accurate context windows without manifest
|
||||
churn. Newly published Copilot models become visible without an OpenClaw
|
||||
upgrade, and context windows reflect the real per-model limits
|
||||
(e.g. 400k for the gpt-5.x series, 1M for the internal
|
||||
`claude-opus-*-1m` variants).
|
||||
|
||||
The bundled static catalog stays as the visible fallback when discovery
|
||||
is disabled, the user has no GitHub auth profile, the token-exchange
|
||||
fails, or the `/models` HTTPS call errors. To opt out and rely entirely
|
||||
on the static manifest catalog (offline / air-gapped scenarios):
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"github-copilot": {
|
||||
config: { discovery: { enabled: false } },
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Transport selection">
|
||||
Claude model IDs use the Anthropic Messages transport automatically.
|
||||
Gemini models use the OpenAI Chat Completions transport; GPT and o-series
|
||||
models keep the OpenAI Responses transport. OpenClaw selects the correct
|
||||
transport based on the model ref.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Request compatibility">
|
||||
OpenClaw sends Copilot IDE-style request headers on Copilot transports
|
||||
(VS Code editor/plugin versions and the `vscode-chat` integration id),
|
||||
marks tool-result follow-up turns as agent-initiated, and sets the Copilot
|
||||
vision header when a turn carries image input.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment variable resolution order">
|
||||
OpenClaw resolves Copilot auth from environment variables in the following
|
||||
priority order:
|
||||
|
||||
| Priority | Variable | Notes |
|
||||
| -------- | --------------------- | -------------------------------- |
|
||||
| 1 | `COPILOT_GITHUB_TOKEN` | Highest priority, Copilot-specific |
|
||||
| 2 | `GH_TOKEN` | GitHub CLI token (fallback) |
|
||||
| 3 | `GITHUB_TOKEN` | Standard GitHub token (lowest) |
|
||||
|
||||
When multiple variables are set, OpenClaw uses the highest-priority one.
|
||||
The device-login flow (`openclaw models auth login-github-copilot`) stores
|
||||
its token in the auth profile store and takes precedence over all environment
|
||||
variables.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Token storage">
|
||||
The login stores a GitHub token in the auth profile store (profile id
|
||||
`github-copilot:github`) and exchanges it for a short-lived Copilot API
|
||||
token when OpenClaw runs. You do not need to manage the token manually.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Memory search embeddings
|
||||
|
||||
GitHub Copilot can also serve as an embedding provider for
|
||||
[memory search](/concepts/memory-search). If you have a Copilot subscription and
|
||||
have logged in, OpenClaw can use it for embeddings without a separate API key.
|
||||
|
||||
### Config
|
||||
|
||||
Set `memorySearch.provider` explicitly to use GitHub Copilot embeddings. If a
|
||||
GitHub token is available, OpenClaw discovers available embedding models from
|
||||
the Copilot API and picks the best one automatically.
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
memorySearch: {
|
||||
provider: "github-copilot",
|
||||
// Optional: override the auto-discovered model
|
||||
model: "text-embedding-3-small",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### How it works
|
||||
|
||||
1. OpenClaw resolves your GitHub token (from env vars or auth profile).
|
||||
2. Exchanges it for a short-lived Copilot API token.
|
||||
3. Queries the Copilot `/models` endpoint to discover available embedding models.
|
||||
4. Picks the best model (preference order: `text-embedding-3-small`,
|
||||
`text-embedding-3-large`, `text-embedding-ada-002`).
|
||||
5. Sends embedding requests to the Copilot `/embeddings` endpoint.
|
||||
|
||||
Model availability depends on your GitHub plan. If no embedding models are
|
||||
available, OpenClaw skips Copilot and tries the next provider.
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="OAuth and auth" href="/gateway/authentication" icon="key">
|
||||
Auth details and credential reuse rules.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
99
docs/providers/gmi.md
Normal file
99
docs/providers/gmi.md
Normal file
@@ -0,0 +1,99 @@
|
||||
---
|
||||
summary: "Use GMI Cloud's OpenAI-compatible API with OpenClaw"
|
||||
read_when:
|
||||
- You want to run OpenClaw with GMI Cloud models
|
||||
- You need the GMI provider id, key, or endpoint
|
||||
title: "GMI Cloud"
|
||||
---
|
||||
|
||||
GMI Cloud is a hosted inference platform for frontier and open-weight models
|
||||
behind an OpenAI-compatible API. In OpenClaw it is an official external provider
|
||||
plugin: install it once, store credentials through normal model auth, and use
|
||||
model refs like `gmi/google/gemini-3.1-flash-lite`.
|
||||
|
||||
Use GMI when you want one API key for several hosted model families, including
|
||||
Anthropic, DeepSeek, Google, Moonshot, OpenAI, and Z.AI routes exposed by GMI's
|
||||
catalog. It works as a secondary provider for model fallback, for comparing
|
||||
hosted routes across vendors, or when GMI has a model available before your
|
||||
primary provider does. OpenClaw owns the provider id, auth profile, aliases,
|
||||
model catalog seed, and base URL; GMI owns live model availability, billing,
|
||||
rate limits, and any provider-side routing policy.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | ---------------------------------------- |
|
||||
| Provider id | `gmi` (aliases: `gmi-cloud`, `gmicloud`) |
|
||||
| Package | `@openclaw/gmi-provider` |
|
||||
| Auth env var | `GMI_API_KEY` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://api.gmi-serving.com/v1` |
|
||||
| Default model | `gmi/google/gemini-3.1-flash-lite` |
|
||||
|
||||
## Setup
|
||||
|
||||
Install the plugin, restart the gateway, then create an API key in GMI Cloud
|
||||
(`https://www.gmicloud.ai/`):
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/gmi-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
Then run:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice gmi-api-key
|
||||
```
|
||||
|
||||
Non-interactive setups can pass `--gmi-api-key <key>`, or set:
|
||||
|
||||
```bash
|
||||
export GMI_API_KEY="<your-gmi-api-key>" # pragma: allowlist secret
|
||||
```
|
||||
|
||||
## When to choose GMI
|
||||
|
||||
- You want a hosted OpenAI-compatible endpoint rather than a local model server.
|
||||
- You want to try several commercial and open-weight model families through one
|
||||
provider account.
|
||||
- You want a fallback provider with different upstream routing from DeepInfra,
|
||||
OpenRouter, Together, or the direct vendor APIs.
|
||||
- You need GMI-specific model ids, pricing, or account controls.
|
||||
|
||||
Choose the direct vendor provider instead when you need vendor-native features
|
||||
that GMI does not expose through its OpenAI-compatible route. Choose a local
|
||||
provider such as LM Studio, Ollama, SGLang, or vLLM when data locality or local
|
||||
GPU control matters more than hosted convenience.
|
||||
|
||||
## Models
|
||||
|
||||
The plugin catalog seeds commonly available GMI Cloud route ids:
|
||||
|
||||
| Model ref | Input | Context | Max output |
|
||||
| ---------------------------------- | ------------ | --------- | ---------- |
|
||||
| `gmi/anthropic/claude-sonnet-4.6` | text + image | 200,000 | 64,000 |
|
||||
| `gmi/deepseek-ai/DeepSeek-V3.2` | text | 163,840 | 65,536 |
|
||||
| `gmi/google/gemini-3.1-flash-lite` | text + image | 1,048,576 | 65,536 |
|
||||
| `gmi/moonshotai/Kimi-K2.5` | text + image | 262,144 | 65,536 |
|
||||
| `gmi/openai/gpt-5.4` | text + image | 400,000 | 128,000 |
|
||||
| `gmi/zai-org/GLM-5.1-FP8` | text | 202,752 | 65,536 |
|
||||
|
||||
The catalog is a seed, not a promise that every account can call every model at
|
||||
all times. List what the configured provider reports in your environment:
|
||||
|
||||
```bash
|
||||
openclaw models list --provider gmi
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- `401` or `403`: check that `GMI_API_KEY` is set for the process running
|
||||
OpenClaw, or re-run onboarding to store the key in the provider auth profile.
|
||||
- Unknown model errors: confirm the model exists in your GMI account and use the
|
||||
full `gmi/<route-id>` ref shown by `openclaw models list --provider gmi`.
|
||||
- Intermittent provider errors: try a different GMI route or configure GMI as a
|
||||
fallback rather than the only primary model provider.
|
||||
|
||||
## Related
|
||||
|
||||
- [Model providers](/concepts/model-providers)
|
||||
- [All providers](/providers/index)
|
||||
480
docs/providers/google.md
Normal file
480
docs/providers/google.md
Normal file
@@ -0,0 +1,480 @@
|
||||
---
|
||||
summary: "Google Gemini setup (API key + OAuth, image generation, media understanding, TTS, web search)"
|
||||
title: "Google (Gemini)"
|
||||
read_when:
|
||||
- You want to use Google Gemini models with OpenClaw
|
||||
- You need the API key or OAuth auth flow
|
||||
---
|
||||
|
||||
The Google plugin provides access to Gemini models through Google AI Studio, plus image generation, media understanding (image/audio/video), text-to-speech, and web search via Gemini Grounding.
|
||||
|
||||
- Provider: `google`
|
||||
- Auth: `GEMINI_API_KEY` or `GOOGLE_API_KEY`
|
||||
- API: Google Gemini API
|
||||
- Runtime option: `agentRuntime.id: "google-gemini-cli"` reuses Gemini CLI OAuth while keeping model refs canonical as `google/*`.
|
||||
|
||||
## Getting started
|
||||
|
||||
Choose your preferred auth method and follow the setup steps.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="API key">
|
||||
**Best for:** standard Gemini API access through Google AI Studio.
|
||||
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice gemini-api-key
|
||||
```
|
||||
|
||||
Or pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice gemini-api-key \
|
||||
--gemini-api-key "$GEMINI_API_KEY"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "google/gemini-3.1-pro-preview" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider google
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
`GEMINI_API_KEY` and `GOOGLE_API_KEY` are both accepted. Use whichever you already have configured.
|
||||
</Tip>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Gemini CLI (OAuth)">
|
||||
**Best for:** reusing an existing Gemini CLI login via PKCE OAuth instead of a separate API key.
|
||||
|
||||
<Warning>
|
||||
The `google-gemini-cli` provider is an unofficial integration. Some users
|
||||
report account restrictions when using OAuth this way. Use at your own risk.
|
||||
</Warning>
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the Gemini CLI">
|
||||
The local `gemini` command must be available on `PATH`.
|
||||
|
||||
```bash
|
||||
# Homebrew
|
||||
brew install gemini-cli
|
||||
|
||||
# or npm
|
||||
npm install -g @google/gemini-cli
|
||||
```
|
||||
|
||||
OpenClaw supports both Homebrew installs and global npm installs, including
|
||||
common Windows/npm layouts.
|
||||
</Step>
|
||||
<Step title="Log in via OAuth">
|
||||
```bash
|
||||
openclaw models auth login --provider google-gemini-cli --set-default
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider google
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
- Default model: `google/gemini-3.1-pro-preview`
|
||||
- Runtime: `google-gemini-cli`
|
||||
- Alias: `gemini-cli`
|
||||
|
||||
Gemini 3.1 Pro's Gemini API model id is `gemini-3.1-pro-preview`. OpenClaw accepts the shorter `google/gemini-3.1-pro` as a convenience alias and normalizes it before provider calls.
|
||||
|
||||
**Environment variables:**
|
||||
|
||||
- `OPENCLAW_GEMINI_OAUTH_CLIENT_ID` / `GEMINI_CLI_OAUTH_CLIENT_ID`
|
||||
- `OPENCLAW_GEMINI_OAUTH_CLIENT_SECRET` / `GEMINI_CLI_OAUTH_CLIENT_SECRET`
|
||||
|
||||
<Note>
|
||||
If Gemini CLI OAuth requests fail after login, set `GOOGLE_CLOUD_PROJECT` or
|
||||
`GOOGLE_CLOUD_PROJECT_ID` on the gateway host and retry.
|
||||
</Note>
|
||||
|
||||
<Note>
|
||||
If login fails before the browser flow starts, make sure the local `gemini`
|
||||
command is installed and on `PATH`.
|
||||
</Note>
|
||||
|
||||
`google-gemini-cli/*` model refs are legacy compatibility aliases. New
|
||||
configs should use `google/*` model refs plus the `google-gemini-cli`
|
||||
runtime when they want local Gemini CLI execution.
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
<Note>
|
||||
`google/gemini-3-pro-preview` was retired on 2026-03-09; use `google/gemini-3.1-pro-preview` instead. Re-running Gemini API key setup (`openclaw onboard --auth-choice gemini-api-key` or `openclaw models auth login --provider google`) rewrites a stale configured default to the current model.
|
||||
</Note>
|
||||
|
||||
## Capabilities
|
||||
|
||||
| Capability | Supported |
|
||||
| ---------------------- | ----------------------------- |
|
||||
| Chat completions | Yes |
|
||||
| Image generation | Yes |
|
||||
| Music generation | Yes |
|
||||
| Text-to-speech | Yes |
|
||||
| Realtime voice | Yes (Google Live API) |
|
||||
| Image understanding | Yes |
|
||||
| Audio transcription | Yes |
|
||||
| Video understanding | Yes |
|
||||
| Web search (Grounding) | Yes |
|
||||
| Thinking/reasoning | Yes (Gemini 2.5+ / Gemini 3+) |
|
||||
| Gemma 4 models | Yes |
|
||||
|
||||
## Web search
|
||||
|
||||
The bundled `gemini` web-search provider uses Gemini Google Search grounding.
|
||||
Configure a dedicated search key under `plugins.entries.google.config.webSearch`,
|
||||
or let it reuse `models.providers.google.apiKey` after `GEMINI_API_KEY`:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
google: {
|
||||
config: {
|
||||
webSearch: {
|
||||
apiKey: "AIza...", // optional if GEMINI_API_KEY or models.providers.google.apiKey is set
|
||||
baseUrl: "https://generativelanguage.googleapis.com/v1beta", // falls back to models.providers.google.baseUrl
|
||||
model: "gemini-2.5-flash",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Credential precedence is dedicated `webSearch.apiKey`, then `GEMINI_API_KEY`,
|
||||
then `models.providers.google.apiKey`. `webSearch.baseUrl` is optional and
|
||||
exists for operator proxies or compatible Gemini API endpoints; when omitted,
|
||||
Gemini web search reuses `models.providers.google.baseUrl`. See
|
||||
[Gemini search](/tools/gemini-search) for the provider-specific tool behavior.
|
||||
|
||||
<Tip>
|
||||
Gemini 3 models use `thinkingLevel` rather than `thinkingBudget`. OpenClaw maps
|
||||
Gemini 3, Gemini 3.1, and `gemini-*-latest` alias reasoning controls to
|
||||
`thinkingLevel` so default/low-latency runs do not send disabled
|
||||
`thinkingBudget` values.
|
||||
|
||||
`/think adaptive` keeps Google's dynamic thinking semantics instead of choosing
|
||||
a fixed OpenClaw level. Gemini 3 and Gemini 3.1 omit a fixed `thinkingLevel` so
|
||||
Google can choose the level; Gemini 2.5 sends Google's dynamic sentinel
|
||||
`thinkingBudget: -1`.
|
||||
|
||||
Gemma 4 models (for example `gemma-4-26b-a4b-it`) support thinking mode. OpenClaw
|
||||
rewrites `thinkingBudget` to a supported Google `thinkingLevel` for Gemma 4.
|
||||
Setting thinking to `off` preserves thinking disabled instead of mapping to
|
||||
`MINIMAL`.
|
||||
|
||||
Gemini 2.5 Pro only works in thinking mode and rejects an explicit
|
||||
`thinkingBudget: 0`; OpenClaw strips that value for Gemini 2.5 Pro requests
|
||||
instead of sending it.
|
||||
</Tip>
|
||||
|
||||
## Image generation
|
||||
|
||||
The bundled `google` image-generation provider defaults to
|
||||
`google/gemini-3.1-flash-image-preview`.
|
||||
|
||||
- Also supports `google/gemini-3-pro-image-preview`
|
||||
- Generate: up to 4 images per request
|
||||
- Edit mode: enabled, up to 5 input images
|
||||
- Geometry controls: `size`, `aspectRatio`, and `resolution`
|
||||
|
||||
To use Google as the default image provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "google/gemini-3.1-flash-image-preview",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
See [Image Generation](/tools/image-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
## Video generation
|
||||
|
||||
The bundled `google` plugin also registers video generation through the shared
|
||||
`video_generate` tool.
|
||||
|
||||
- Default video model: `google/veo-3.1-fast-generate-preview`
|
||||
- Modes: text-to-video, image-to-video, and single-video reference flows
|
||||
- Supports `aspectRatio` (`16:9`, `9:16`) and `resolution` (`720P`, `1080P`); audio output is not supported by Veo today
|
||||
- Supported durations: **4, 6, or 8 seconds** (other values snap to the nearest allowed value)
|
||||
|
||||
To use Google as the default video provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "google/veo-3.1-fast-generate-preview",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
See [Video Generation](/tools/video-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
## Music generation
|
||||
|
||||
The bundled `google` plugin also registers music generation through the shared
|
||||
`music_generate` tool.
|
||||
|
||||
- Default music model: `google/lyria-3-clip-preview`
|
||||
- Also supports `google/lyria-3-pro-preview`
|
||||
- Prompt controls: `lyrics` and `instrumental`
|
||||
- Output format: `mp3` by default, plus `wav` on `google/lyria-3-pro-preview`
|
||||
- Reference inputs: up to 10 images
|
||||
- Session-backed runs detach through the shared task/status flow, including `action: "status"`
|
||||
|
||||
To use Google as the default music provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
musicGenerationModel: {
|
||||
primary: "google/lyria-3-clip-preview",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
See [Music Generation](/tools/music-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
## Text-to-speech
|
||||
|
||||
The bundled `google` speech provider uses the Gemini API TTS path with
|
||||
`gemini-3.1-flash-tts-preview`.
|
||||
|
||||
- Default voice: `Kore`
|
||||
- Auth: `messages.tts.providers.google.apiKey`, `models.providers.google.apiKey`, `GEMINI_API_KEY`, or `GOOGLE_API_KEY`
|
||||
- Output: WAV for regular TTS attachments, Opus for voice-note targets, PCM for Talk/telephony
|
||||
- Voice-note output: Google PCM is wrapped as WAV and transcoded to 48 kHz Opus with `ffmpeg`
|
||||
|
||||
Google's batch Gemini TTS path returns generated audio in the completed
|
||||
`generateContent` response. For lowest-latency spoken conversations, use the
|
||||
Google realtime voice provider backed by the Gemini Live API instead of batch
|
||||
TTS.
|
||||
|
||||
To use Google as the default TTS provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "google",
|
||||
providers: {
|
||||
google: {
|
||||
model: "gemini-3.1-flash-tts-preview",
|
||||
speakerVoice: "Kore",
|
||||
audioProfile: "Speak professionally with a calm tone.",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Gemini API TTS uses natural-language prompting for style control. Set
|
||||
`audioProfile` to prepend a reusable style prompt before the spoken text. Set
|
||||
`speakerName` when your prompt text refers to a named speaker.
|
||||
|
||||
Gemini API TTS also accepts expressive square-bracket audio tags in the text,
|
||||
such as `[whispers]` or `[laughs]`. To keep tags out of the visible chat reply
|
||||
while sending them to TTS, put them inside a `[[tts:text]]...[[/tts:text]]`
|
||||
block:
|
||||
|
||||
```text
|
||||
Here is the clean reply text.
|
||||
|
||||
[[tts:text]][whispers] Here is the spoken version.[[/tts:text]]
|
||||
```
|
||||
|
||||
<Note>
|
||||
A Google Cloud Console API key restricted to the Gemini API is valid for this
|
||||
provider. This is not the separate Cloud Text-to-Speech API path.
|
||||
</Note>
|
||||
|
||||
## Realtime voice
|
||||
|
||||
The bundled `google` plugin registers a realtime voice provider backed by the
|
||||
Gemini Live API for backend audio bridges such as Voice Call and Google Meet.
|
||||
|
||||
| Setting | Config path | Default |
|
||||
| --------------------- | ------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
|
||||
| Model | `plugins.entries.voice-call.config.realtime.providers.google.model` | `gemini-2.5-flash-native-audio-preview-12-2025` |
|
||||
| Voice | `...google.voice` | `Kore` |
|
||||
| Temperature | `...google.temperature` | (unset) |
|
||||
| VAD start sensitivity | `...google.startSensitivity` | (unset) |
|
||||
| VAD end sensitivity | `...google.endSensitivity` | (unset) |
|
||||
| Silence duration | `...google.silenceDurationMs` | (unset) |
|
||||
| Activity handling | `...google.activityHandling` | Google default, `start-of-activity-interrupts` |
|
||||
| Turn coverage | `...google.turnCoverage` | Google default, `only-activity` |
|
||||
| Disable auto VAD | `...google.automaticActivityDetectionDisabled` | `false` |
|
||||
| Session resumption | `...google.sessionResumption` | `true` |
|
||||
| Context compression | `...google.contextWindowCompression` | `true` |
|
||||
| API key | `...google.apiKey` | Falls back to `models.providers.google.apiKey`, `GEMINI_API_KEY`, or `GOOGLE_API_KEY` |
|
||||
|
||||
Example Voice Call realtime config:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"voice-call": {
|
||||
enabled: true,
|
||||
config: {
|
||||
realtime: {
|
||||
enabled: true,
|
||||
provider: "google",
|
||||
providers: {
|
||||
google: {
|
||||
model: "gemini-2.5-flash-native-audio-preview-12-2025",
|
||||
speakerVoice: "Kore",
|
||||
activityHandling: "start-of-activity-interrupts",
|
||||
turnCoverage: "only-activity",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
Google Live API uses bidirectional audio and function calling over a WebSocket.
|
||||
OpenClaw adapts telephony/Meet bridge audio to Gemini's PCM Live API stream and
|
||||
keeps tool calls on the shared realtime voice contract. Leave `temperature`
|
||||
unset unless you need sampling changes; OpenClaw omits non-positive values
|
||||
because Google Live can return transcripts without audio for `temperature: 0`.
|
||||
Gemini API transcription is enabled without `languageCodes`; the current Google
|
||||
SDK rejects language-code hints on this API path.
|
||||
</Note>
|
||||
|
||||
<Note>
|
||||
Control UI Talk supports Google Live browser sessions with constrained one-use
|
||||
tokens. Backend-only realtime voice providers can also run through the generic
|
||||
Gateway relay transport, which keeps provider credentials on the Gateway.
|
||||
</Note>
|
||||
|
||||
For maintainer live verification, run
|
||||
`OPENAI_API_KEY=... GEMINI_API_KEY=... node --import tsx scripts/dev/realtime-talk-live-smoke.ts`.
|
||||
The smoke also covers OpenAI backend/WebRTC paths; the Google leg mints the same
|
||||
constrained Live API token shape used by Control UI Talk, opens the browser
|
||||
WebSocket endpoint, sends the initial setup payload, and waits for
|
||||
`setupComplete`.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Direct Gemini cache reuse">
|
||||
For direct Gemini API runs (`api: "google-generative-ai"`), OpenClaw
|
||||
passes a configured `cachedContent` handle through to Gemini requests.
|
||||
|
||||
- Configure per-model or global params with either
|
||||
`cachedContent` or legacy `cached_content`
|
||||
- Params from a more specific scope (model-level over global) always win.
|
||||
Within the same scope, if both keys are set, `cached_content` wins.
|
||||
Use only one key per scope to avoid surprises.
|
||||
- Example value: `cachedContents/prebuilt-context`
|
||||
- Gemini cache-hit usage is normalized into OpenClaw `cacheRead` from
|
||||
upstream `cachedContentTokenCount`
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"google/gemini-2.5-pro": {
|
||||
params: {
|
||||
cachedContent: "cachedContents/prebuilt-context",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Gemini CLI usage notes">
|
||||
When using the `google-gemini-cli` OAuth provider, OpenClaw uses Gemini
|
||||
CLI `stream-json` output by default and normalizes usage from the final
|
||||
`stats` payload. Legacy `--output-format json` overrides still use the
|
||||
JSON parser.
|
||||
|
||||
- Streamed reply text comes from assistant `message` events.
|
||||
- For legacy JSON output, reply text comes from the CLI JSON `response` field.
|
||||
- Usage falls back to `stats` when the CLI leaves `usage` empty.
|
||||
- `stats.cached` is normalized into OpenClaw `cacheRead`.
|
||||
- If `stats.input` is missing, OpenClaw derives input tokens from
|
||||
`stats.input_tokens - stats.cached`.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment and daemon setup">
|
||||
If the Gateway runs as a daemon (launchd/systemd), make sure `GEMINI_API_KEY`
|
||||
is available to that process (for example, in `~/.openclaw/.env` or via
|
||||
`env.shellEnv`).
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Image generation" href="/tools/image-generation" icon="image">
|
||||
Shared image tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Music generation" href="/tools/music-generation" icon="music">
|
||||
Shared music tool parameters and provider selection.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
128
docs/providers/gradium.md
Normal file
128
docs/providers/gradium.md
Normal file
@@ -0,0 +1,128 @@
|
||||
---
|
||||
summary: "Use Gradium text-to-speech in OpenClaw"
|
||||
read_when:
|
||||
- You want Gradium for text-to-speech
|
||||
- You need Gradium API key, voice, or directive token configuration
|
||||
title: "Gradium"
|
||||
---
|
||||
|
||||
[Gradium](https://gradium.ai) is a text-to-speech provider for OpenClaw. It renders standard audio replies (WAV), voice-note-compatible Opus output, and 8 kHz u-law audio for telephony surfaces.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | ------------------------------------ |
|
||||
| Provider id | `gradium` |
|
||||
| Auth | `GRADIUM_API_KEY` or config `apiKey` |
|
||||
| Base URL | `https://api.gradium.ai` (default) |
|
||||
| Default voice | `Emma` (`YTpq7expH9539ERJ`) |
|
||||
|
||||
## Install plugin
|
||||
|
||||
Gradium is an official external plugin. Install it, then restart Gateway:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/gradium-speech
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Setup
|
||||
|
||||
Create a Gradium API key, then expose it with an env var or the config key. Config takes precedence over the env var.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Env var">
|
||||
```bash
|
||||
export GRADIUM_API_KEY="gsk_..."
|
||||
```
|
||||
</Tab>
|
||||
|
||||
<Tab title="Config key">
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "gradium",
|
||||
providers: {
|
||||
gradium: {
|
||||
apiKey: "${GRADIUM_API_KEY}",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Config
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "gradium",
|
||||
providers: {
|
||||
gradium: {
|
||||
speakerVoiceId: "YTpq7expH9539ERJ",
|
||||
// apiKey: "${GRADIUM_API_KEY}",
|
||||
// baseUrl: "https://api.gradium.ai",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
| Key | Type | Description |
|
||||
| ----------------------------------------------- | ------ | --------------------------------------------------------------------------------- |
|
||||
| `messages.tts.providers.gradium.apiKey` | string | Resolved API key. Supports `${ENV}` and secret refs. |
|
||||
| `messages.tts.providers.gradium.baseUrl` | string | API origin override. Trailing slashes stripped. Default `https://api.gradium.ai`. |
|
||||
| `messages.tts.providers.gradium.speakerVoiceId` | string | Default voice id used when no directive override is present. |
|
||||
|
||||
Output format is chosen automatically by target surface (see [Output](#output)) and is not configurable in `openclaw.json`.
|
||||
|
||||
## Voices
|
||||
|
||||
| Name | Voice ID |
|
||||
| ------------------ | ------------------ |
|
||||
| Arthur | `3jUdJyOi9pgbxBTK` |
|
||||
| Christina | `2H4HY2CBNyJHBCrP` |
|
||||
| Emma **(default)** | `YTpq7expH9539ERJ` |
|
||||
| John | `KWJiFWu2O9nMPYcR` |
|
||||
| Kent | `LFZvm12tW_z0xfGo` |
|
||||
| Sydney | `jtEKaLYNn6iif5PR` |
|
||||
| Tiffany | `Eu9iL_CYe8N-Gkx_` |
|
||||
|
||||
### Per-message voice override
|
||||
|
||||
When the active speech policy allows voice overrides, switch voices inline with a directive token (any of these are equivalent, all take a provider-native voice id):
|
||||
|
||||
```text
|
||||
/voice:LFZvm12tW_z0xfGo
|
||||
/voice_id:LFZvm12tW_z0xfGo
|
||||
/voiceid:LFZvm12tW_z0xfGo
|
||||
/gradium_voice:LFZvm12tW_z0xfGo
|
||||
/gradiumvoice:LFZvm12tW_z0xfGo
|
||||
```
|
||||
|
||||
If the speech policy disables voice overrides, the directive is consumed but ignored.
|
||||
|
||||
## Output
|
||||
|
||||
Output format is selected by target surface; the provider does not synthesize other formats.
|
||||
|
||||
| Target | Format | File ext | Sample rate | Voice-compatible flag |
|
||||
| -------------- | ----------- | -------- | ----------- | --------------------- |
|
||||
| Standard audio | `wav` | `.wav` | provider | no |
|
||||
| Voice note | `opus` | `.opus` | provider | yes |
|
||||
| Telephony | `ulaw_8000` | n/a | 8 kHz | n/a |
|
||||
|
||||
## Auto-select order
|
||||
|
||||
Among configured TTS providers, Gradium's auto-select order is `30`. See [Text-to-Speech](/tools/tts) for how OpenClaw picks the active provider when `messages.tts.provider` is not pinned.
|
||||
|
||||
## Related
|
||||
|
||||
- [Text-to-Speech](/tools/tts)
|
||||
- [Media Overview](/tools/media-overview)
|
||||
166
docs/providers/groq.md
Normal file
166
docs/providers/groq.md
Normal file
@@ -0,0 +1,166 @@
|
||||
---
|
||||
summary: "Groq setup (auth + model selection + Whisper transcription)"
|
||||
title: "Groq"
|
||||
read_when:
|
||||
- You want to use Groq with OpenClaw
|
||||
- You need the API key env var or CLI auth choice
|
||||
- You are configuring Whisper audio transcription on Groq
|
||||
---
|
||||
|
||||
[Groq](https://groq.com) provides ultra-fast inference on open-weight models (Llama, Gemma, Kimi, Qwen, GPT OSS, and more) using custom LPU hardware. The Groq plugin registers both an OpenAI-compatible chat provider and an audio media-understanding provider.
|
||||
|
||||
| Property | Value |
|
||||
| ---------------------- | ---------------------------------------- |
|
||||
| Provider id | `groq` |
|
||||
| Plugin | official external package |
|
||||
| Auth env var | `GROQ_API_KEY` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://api.groq.com/openai/v1` |
|
||||
| Audio transcription | `whisper-large-v3-turbo` (default) |
|
||||
| Suggested chat default | `groq/llama-3.3-70b-versatile` |
|
||||
|
||||
## Install plugin
|
||||
|
||||
Install the official plugin, then restart Gateway:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/groq-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Create an API key at [console.groq.com/keys](https://console.groq.com/keys).
|
||||
</Step>
|
||||
<Step title="Set the API key">
|
||||
```bash
|
||||
export GROQ_API_KEY=gsk_...
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "groq/llama-3.3-70b-versatile" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the catalog is reachable">
|
||||
```bash
|
||||
openclaw models list --provider groq
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
### Config file example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { GROQ_API_KEY: "gsk_..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "groq/llama-3.3-70b-versatile" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
OpenClaw ships a manifest-backed Groq catalog with both reasoning and non-reasoning entries. Run `openclaw models list --provider groq` to see the static rows for your installed version, or check [console.groq.com/docs/models](https://console.groq.com/docs/models) for Groq's authoritative list.
|
||||
|
||||
| Model ref | Name | Reasoning | Input | Context |
|
||||
| ------------------------------------------------ | ----------------------- | --------- | ------------ | ------- |
|
||||
| `groq/llama-3.3-70b-versatile` | Llama 3.3 70B Versatile | no | text | 131,072 |
|
||||
| `groq/llama-3.1-8b-instant` | Llama 3.1 8B Instant | no | text | 131,072 |
|
||||
| `groq/meta-llama/llama-4-scout-17b-16e-instruct` | Llama 4 Scout 17B | no | text + image | 131,072 |
|
||||
| `groq/openai/gpt-oss-120b` | GPT OSS 120B | yes | text | 131,072 |
|
||||
| `groq/openai/gpt-oss-20b` | GPT OSS 20B | yes | text | 131,072 |
|
||||
| `groq/openai/gpt-oss-safeguard-20b` | Safety GPT OSS 20B | yes | text | 131,072 |
|
||||
| `groq/qwen/qwen3-32b` | Qwen3 32B | yes | text | 131,072 |
|
||||
| `groq/groq/compound` | Compound | yes | text | 131,072 |
|
||||
| `groq/groq/compound-mini` | Compound Mini | yes | text | 131,072 |
|
||||
|
||||
<Tip>
|
||||
The catalog evolves with each OpenClaw release. `openclaw models list --provider groq` shows the rows known to your installed version; cross-check with [console.groq.com/docs/models](https://console.groq.com/docs/models) for newly-added or deprecated models.
|
||||
</Tip>
|
||||
|
||||
## Reasoning models
|
||||
|
||||
Groq reasoning models (`reasoning: true` in the table above) map OpenClaw's shared `/think` levels onto `reasoning_effort` values of `low`, `medium`, or `high`. `/think off` or `/think none` omits `reasoning_effort` from the request rather than sending a disabled value.
|
||||
|
||||
See [Thinking modes](/tools/thinking) for the shared `/think` levels and how OpenClaw translates them per provider.
|
||||
|
||||
## Audio transcription
|
||||
|
||||
Groq's plugin also registers an **audio media-understanding provider** so voice messages can be transcribed through the shared `tools.media.audio` surface.
|
||||
|
||||
| Property | Value |
|
||||
| ------------------ | ----------------------------------------- |
|
||||
| Shared config path | `tools.media.audio` |
|
||||
| Default base URL | `https://api.groq.com/openai/v1` |
|
||||
| Default model | `whisper-large-v3-turbo` |
|
||||
| Auto priority | 20 |
|
||||
| API endpoint | OpenAI-compatible `/audio/transcriptions` |
|
||||
|
||||
To make Groq the default audio backend:
|
||||
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
models: [{ provider: "groq" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Environment availability for the daemon">
|
||||
If the Gateway runs as a managed service (launchd, systemd, Docker), `GROQ_API_KEY` must be visible to that process — not just to your interactive shell.
|
||||
|
||||
<Warning>
|
||||
A key exported only in an interactive shell will not help a launchd or systemd daemon unless that environment is imported there too. Set the key in `~/.openclaw/.env` or via `env.shellEnv` to make it readable from the gateway process.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Custom Groq model ids">
|
||||
OpenClaw accepts any Groq model id at runtime. Use the exact id shown by Groq and prefix it with `groq/`. The static catalog covers the common cases; uncatalogued ids fall through to the default OpenAI-compatible template.
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "groq/<your-model-id>" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Thinking modes" href="/tools/thinking" icon="brain">
|
||||
Reasoning effort levels and provider-policy interaction.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema including provider and audio settings.
|
||||
</Card>
|
||||
<Card title="Groq Console" href="https://console.groq.com" icon="arrow-up-right-from-square">
|
||||
Groq dashboard, API docs, and pricing.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
214
docs/providers/huggingface.md
Normal file
214
docs/providers/huggingface.md
Normal file
@@ -0,0 +1,214 @@
|
||||
---
|
||||
summary: "Hugging Face Inference setup (auth + model selection)"
|
||||
read_when:
|
||||
- You want to use Hugging Face Inference with OpenClaw
|
||||
- You need the HF token env var or CLI auth choice
|
||||
title: "Hugging Face (inference)"
|
||||
---
|
||||
|
||||
[Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers) exposes an OpenAI-compatible chat completions router in front of many hosted models (DeepSeek, Llama, and more) under one token. OpenClaw talks to the **chat completions endpoint only**; for text-to-image, embeddings, or speech use the [HF inference clients](https://huggingface.co/docs/api-inference/quicktour) directly.
|
||||
|
||||
| Property | Value |
|
||||
| ------------ | --------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Provider id | `huggingface` |
|
||||
| Plugin | bundled (enabled by default, no install step) |
|
||||
| Auth env var | `HUGGINGFACE_HUB_TOKEN` or `HF_TOKEN` (fine-grained token) |
|
||||
| API | OpenAI-compatible (`https://router.huggingface.co/v1`) |
|
||||
| Billing | Single HF token; [pricing](https://huggingface.co/docs/inference-providers/pricing) follows provider rates with a free tier |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Create a fine-grained token">
|
||||
Go to [Hugging Face Settings Tokens](https://huggingface.co/settings/tokens/new?ownUserPermissions=inference.serverless.write&tokenType=fineGrained) and create a new fine-grained token.
|
||||
|
||||
<Warning>
|
||||
The token must have the **Make calls to Inference Providers** permission enabled or API requests will be rejected.
|
||||
</Warning>
|
||||
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
Choose **Hugging Face** in the provider dropdown, then enter your API key when prompted:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice huggingface-api-key
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Select a default model">
|
||||
In the **Default Hugging Face model** dropdown, pick a model. The list loads from the Inference API when your token is valid; otherwise OpenClaw shows the built-in catalog below. Your choice is saved as `agents.defaults.model.primary`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "huggingface/deepseek-ai/DeepSeek-R1" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider huggingface
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
### Non-interactive setup
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice huggingface-api-key \
|
||||
--huggingface-api-key "$HF_TOKEN"
|
||||
```
|
||||
|
||||
Sets `huggingface/deepseek-ai/DeepSeek-R1` as the default model.
|
||||
|
||||
## Model IDs
|
||||
|
||||
Model refs use the form `huggingface/<org>/<model>` (Hub-style IDs). OpenClaw's built-in catalog:
|
||||
|
||||
| Model | Ref (prefix with `huggingface/`) |
|
||||
| ---------------------------- | ----------------------------------------- |
|
||||
| DeepSeek R1 | `deepseek-ai/DeepSeek-R1` |
|
||||
| DeepSeek V3.1 | `deepseek-ai/DeepSeek-V3.1` |
|
||||
| GPT-OSS 120B | `openai/gpt-oss-120b` |
|
||||
| Llama 3.3 70B Instruct Turbo | `meta-llama/Llama-3.3-70B-Instruct-Turbo` |
|
||||
|
||||
<Tip>
|
||||
When your token is valid, OpenClaw also discovers any other model from **GET** `https://router.huggingface.co/v1/models` at onboarding time and Gateway startup, so your catalog can include far more than the four models above. You can append `:fastest` or `:cheapest` to any model id; HF's router routes to the matching inference provider. Set your default provider order in [Inference Provider settings](https://hf.co/settings/inference-providers).
|
||||
</Tip>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Model discovery and onboarding dropdown">
|
||||
OpenClaw discovers models with:
|
||||
|
||||
```bash
|
||||
GET https://router.huggingface.co/v1/models
|
||||
Authorization: Bearer $HUGGINGFACE_HUB_TOKEN # or $HF_TOKEN
|
||||
```
|
||||
|
||||
The response is OpenAI-style: `{ "object": "list", "data": [ { "id": "Qwen/Qwen3-8B", "owned_by": "Qwen", ... }, ... ] }`.
|
||||
|
||||
With a configured key (onboarding, `HUGGINGFACE_HUB_TOKEN`, or `HF_TOKEN`), the **Default Hugging Face model** dropdown during interactive setup is populated from this endpoint. Gateway startup repeats the same call to refresh the catalog. Discovered models are merged with the built-in catalog above (used for metadata like context window and cost when an id matches). If the request fails, returns no data, or no key is set, OpenClaw falls back to the built-in catalog only.
|
||||
|
||||
Disable discovery without removing the provider:
|
||||
|
||||
```bash
|
||||
openclaw config set plugins.entries.huggingface.config.discovery.enabled false
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Model names, aliases, and policy suffixes">
|
||||
- **Name from API:** discovered models use the API's `name`, `title`, or `display_name` when present; otherwise OpenClaw derives a name from the model id (e.g. `deepseek-ai/DeepSeek-R1` becomes "DeepSeek R1").
|
||||
- **Override display name:** set a custom label per model in config:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"huggingface/deepseek-ai/DeepSeek-R1": { alias: "DeepSeek R1 (fast)" },
|
||||
"huggingface/deepseek-ai/DeepSeek-R1:cheapest": { alias: "DeepSeek R1 (cheap)" },
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
- **Policy suffixes:** `:fastest` and `:cheapest` are HF router conventions, not something OpenClaw rewrites: the suffix is sent verbatim as part of the model id and HF's router picks the matching inference provider. Add each variant as its own entry under `models.providers.huggingface.models` (or in `model.primary`) if you want a distinct alias per suffix.
|
||||
- **Config merge:** existing entries in `models.providers.huggingface.models` (e.g. in `models.json`) are kept on config merge, so any custom `name`, `alias`, or model options you set there persist across restarts.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment and daemon setup">
|
||||
If the Gateway runs as a daemon (launchd/systemd), make sure `HUGGINGFACE_HUB_TOKEN` or `HF_TOKEN` is available to that process (for example, in `~/.openclaw/.env` or via `env.shellEnv`).
|
||||
|
||||
<Note>
|
||||
OpenClaw accepts both `HUGGINGFACE_HUB_TOKEN` and `HF_TOKEN`. If both are set, `HUGGINGFACE_HUB_TOKEN` takes precedence.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Config: DeepSeek R1 with fallback">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: {
|
||||
primary: "huggingface/deepseek-ai/DeepSeek-R1",
|
||||
fallbacks: ["huggingface/openai/gpt-oss-120b"],
|
||||
},
|
||||
models: {
|
||||
"huggingface/deepseek-ai/DeepSeek-R1": { alias: "DeepSeek R1" },
|
||||
"huggingface/openai/gpt-oss-120b": { alias: "GPT-OSS 120B" },
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Config: DeepSeek with cheapest and fastest variants">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "huggingface/deepseek-ai/DeepSeek-R1" },
|
||||
models: {
|
||||
"huggingface/deepseek-ai/DeepSeek-R1": { alias: "DeepSeek R1" },
|
||||
"huggingface/deepseek-ai/DeepSeek-R1:cheapest": { alias: "DeepSeek R1 (cheapest)" },
|
||||
"huggingface/deepseek-ai/DeepSeek-R1:fastest": { alias: "DeepSeek R1 (fastest)" },
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Config: DeepSeek + Llama + GPT-OSS with aliases">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: {
|
||||
primary: "huggingface/deepseek-ai/DeepSeek-V3.1",
|
||||
fallbacks: [
|
||||
"huggingface/meta-llama/Llama-3.3-70B-Instruct-Turbo",
|
||||
"huggingface/openai/gpt-oss-120b",
|
||||
],
|
||||
},
|
||||
models: {
|
||||
"huggingface/deepseek-ai/DeepSeek-V3.1": { alias: "DeepSeek V3.1" },
|
||||
"huggingface/meta-llama/Llama-3.3-70B-Instruct-Turbo": { alias: "Llama 3.3 70B Turbo" },
|
||||
"huggingface/openai/gpt-oss-120b": { alias: "GPT-OSS 120B" },
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Overview of all providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/models" icon="brain">
|
||||
How to choose and configure models.
|
||||
</Card>
|
||||
<Card title="Inference Providers docs" href="https://huggingface.co/docs/inference-providers" icon="book">
|
||||
Official Hugging Face Inference Providers documentation.
|
||||
</Card>
|
||||
<Card title="Configuration" href="/gateway/configuration" icon="gear">
|
||||
Full config reference.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
107
docs/providers/index.md
Normal file
107
docs/providers/index.md
Normal file
@@ -0,0 +1,107 @@
|
||||
---
|
||||
summary: "Model providers (LLMs) supported by OpenClaw"
|
||||
read_when:
|
||||
- You want to choose a model provider
|
||||
- You need a quick overview of supported LLM backends
|
||||
title: "Provider directory"
|
||||
---
|
||||
|
||||
OpenClaw can use many LLM providers. Pick a provider, authenticate, then set the
|
||||
default model as `provider/model`.
|
||||
|
||||
Looking for chat channel docs (WhatsApp/Telegram/Discord/Slack/Mattermost (plugin)/etc.)? See [Channels](/channels).
|
||||
|
||||
## Quick start
|
||||
|
||||
1. Authenticate with the provider (usually via `openclaw onboard`).
|
||||
2. Set the default model:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: { defaults: { model: { primary: "anthropic/claude-opus-4-6" } } },
|
||||
}
|
||||
```
|
||||
|
||||
## Provider docs
|
||||
|
||||
- [Alibaba Model Studio](/providers/alibaba)
|
||||
- [Amazon Bedrock](/providers/bedrock)
|
||||
- [Amazon Bedrock Mantle](/providers/bedrock-mantle)
|
||||
- [Anthropic (API + Claude CLI)](/providers/anthropic)
|
||||
- [Arcee AI (Trinity models)](/providers/arcee)
|
||||
- [Azure Speech](/providers/azure-speech)
|
||||
- [BytePlus (International)](/concepts/model-providers#byteplus-international)
|
||||
- [Cerebras](/providers/cerebras)
|
||||
- [Chutes](/providers/chutes)
|
||||
- [ClawRouter (managed multi-provider routing)](/providers/clawrouter)
|
||||
- [Cloudflare AI Gateway](/providers/cloudflare-ai-gateway)
|
||||
- [Cohere](/providers/cohere)
|
||||
- [ComfyUI](/providers/comfy)
|
||||
- [DeepSeek](/providers/deepseek)
|
||||
- [ds4 (local DeepSeek V4)](/providers/ds4)
|
||||
- [ElevenLabs](/providers/elevenlabs)
|
||||
- [fal](/providers/fal)
|
||||
- [Fireworks](/providers/fireworks)
|
||||
- [GitHub Copilot](/providers/github-copilot)
|
||||
- [GMI Cloud](/providers/gmi)
|
||||
- [Google (Gemini)](/providers/google)
|
||||
- [Gradium](/providers/gradium)
|
||||
- [Groq (LPU inference)](/providers/groq)
|
||||
- [Hugging Face (Inference)](/providers/huggingface)
|
||||
- [inferrs (local models)](/providers/inferrs)
|
||||
- [Kilocode](/providers/kilocode)
|
||||
- [LiteLLM (unified gateway)](/providers/litellm)
|
||||
- [LM Studio (local models)](/providers/lmstudio)
|
||||
- [MiniMax](/providers/minimax)
|
||||
- [Mistral](/providers/mistral)
|
||||
- [Moonshot AI (Kimi + Kimi Coding)](/providers/moonshot)
|
||||
- [NovitaAI](/providers/novita)
|
||||
- [NVIDIA](/providers/nvidia)
|
||||
- [Ollama (cloud + local models)](/providers/ollama)
|
||||
- [Ollama Cloud](/providers/ollama-cloud)
|
||||
- [OpenAI (API + Codex)](/providers/openai)
|
||||
- [OpenCode](/providers/opencode)
|
||||
- [OpenCode Go](/providers/opencode-go)
|
||||
- [OpenRouter](/providers/openrouter)
|
||||
- [Perplexity (web search)](/providers/perplexity-provider)
|
||||
- [Qianfan](/providers/qianfan)
|
||||
- [Qwen Cloud](/providers/qwen)
|
||||
- [Qwen OAuth / Portal](/providers/qwen-oauth)
|
||||
- [Runway](/providers/runway)
|
||||
- [SenseAudio](/providers/senseaudio)
|
||||
- [SGLang (local models)](/providers/sglang)
|
||||
- [StepFun](/providers/stepfun)
|
||||
- [Synthetic](/providers/synthetic)
|
||||
- [Tencent Cloud (TokenHub)](/providers/tencent)
|
||||
- [Together AI](/providers/together)
|
||||
- [Venice (Venice AI, privacy-focused)](/providers/venice)
|
||||
- [Vercel AI Gateway](/providers/vercel-ai-gateway)
|
||||
- [vLLM (local models)](/providers/vllm)
|
||||
- [Volcengine (Doubao)](/providers/volcengine)
|
||||
- [Vydra](/providers/vydra)
|
||||
- [xAI](/providers/xai)
|
||||
- [Xiaomi](/providers/xiaomi)
|
||||
- [Z.AI (GLM)](/providers/zai)
|
||||
|
||||
## Shared overview pages
|
||||
|
||||
- [Additional provider variants](/providers/models#additional-provider-variants) - Anthropic Vertex, Copilot Proxy, and Gemini CLI OAuth
|
||||
- [Image Generation](/tools/image-generation) - Shared `image_generate` tool, provider selection, and failover
|
||||
- [Music Generation](/tools/music-generation) - Shared `music_generate` tool, provider selection, and failover
|
||||
- [Video Generation](/tools/video-generation) - Shared `video_generate` tool, provider selection, and failover
|
||||
|
||||
## Transcription providers
|
||||
|
||||
- [Deepgram (audio transcription)](/providers/deepgram)
|
||||
- [ElevenLabs](/providers/elevenlabs#speech-to-text)
|
||||
- [Mistral](/providers/mistral#audio-transcription-voxtral)
|
||||
- [OpenAI](/providers/openai)
|
||||
- [SenseAudio](/providers/senseaudio)
|
||||
- [xAI](/providers/xai)
|
||||
|
||||
## Community tools
|
||||
|
||||
- [Claude Max API Proxy](/providers/claude-max-api-proxy) - Community proxy for Claude subscription credentials (verify Anthropic policy/terms before use)
|
||||
|
||||
For the full provider catalog (xAI, Groq, Mistral, etc.) and advanced configuration,
|
||||
see [Model providers](/concepts/model-providers).
|
||||
236
docs/providers/inferrs.md
Normal file
236
docs/providers/inferrs.md
Normal file
@@ -0,0 +1,236 @@
|
||||
---
|
||||
summary: "Run OpenClaw through inferrs (OpenAI-compatible local server)"
|
||||
read_when:
|
||||
- You want to run OpenClaw against a local inferrs server
|
||||
- You are serving Gemma or another model through inferrs
|
||||
- You need the exact OpenClaw compat flags for inferrs
|
||||
title: "Inferrs"
|
||||
---
|
||||
|
||||
[inferrs](https://github.com/ericcurtin/inferrs) serves local models behind an OpenAI-compatible `/v1` API. OpenClaw talks to it through the generic `openai-completions` adapter.
|
||||
|
||||
| Property | Value |
|
||||
| ------------------ | -------------------------------------------------------------------- |
|
||||
| Provider id | `inferrs` (custom; configure under `models.providers.inferrs`) |
|
||||
| Plugin | none — not a bundled OpenClaw provider plugin |
|
||||
| Auth env var | none required; any value works if your inferrs server has no auth |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Suggested base URL | `http://127.0.0.1:8080/v1` (or wherever your inferrs server listens) |
|
||||
|
||||
<Note>
|
||||
`inferrs` is a custom self-hosted OpenAI-compatible backend, not a dedicated OpenClaw provider plugin: you configure it under `models.providers.inferrs` instead of picking an onboarding auth choice. For a bundled plugin with auto-discovery, see [SGLang](/providers/sglang) or [vLLM](/providers/vllm).
|
||||
</Note>
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Start inferrs with a model">
|
||||
```bash
|
||||
inferrs serve google/gemma-4-E2B-it \
|
||||
--host 127.0.0.1 \
|
||||
--port 8080 \
|
||||
--device metal
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the server is reachable">
|
||||
```bash
|
||||
curl http://127.0.0.1:8080/health
|
||||
curl http://127.0.0.1:8080/v1/models
|
||||
```
|
||||
</Step>
|
||||
<Step title="Add an OpenClaw provider entry">
|
||||
Add an explicit provider entry and point your default model at it. See the config example below.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Full config example
|
||||
|
||||
Gemma 4 on a local `inferrs` server:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "inferrs/google/gemma-4-E2B-it" },
|
||||
models: {
|
||||
"inferrs/google/gemma-4-E2B-it": {
|
||||
alias: "Gemma 4 (inferrs)",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
inferrs: {
|
||||
baseUrl: "http://127.0.0.1:8080/v1",
|
||||
apiKey: "inferrs-local",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
{
|
||||
id: "google/gemma-4-E2B-it",
|
||||
name: "Gemma 4 E2B (inferrs)",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 131072,
|
||||
maxTokens: 4096,
|
||||
compat: {
|
||||
requiresStringContent: true,
|
||||
},
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## On-demand startup
|
||||
|
||||
OpenClaw can start `inferrs` itself only when an `inferrs/...` model is selected. Add `localService` to the same provider entry:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
inferrs: {
|
||||
baseUrl: "http://127.0.0.1:8080/v1",
|
||||
apiKey: "inferrs-local",
|
||||
api: "openai-completions",
|
||||
timeoutSeconds: 300,
|
||||
localService: {
|
||||
command: "/opt/homebrew/bin/inferrs",
|
||||
args: [
|
||||
"serve",
|
||||
"google/gemma-4-E2B-it",
|
||||
"--host",
|
||||
"127.0.0.1",
|
||||
"--port",
|
||||
"8080",
|
||||
"--device",
|
||||
"metal",
|
||||
],
|
||||
healthUrl: "http://127.0.0.1:8080/v1/models",
|
||||
readyTimeoutMs: 180000,
|
||||
idleStopMs: 0,
|
||||
},
|
||||
models: [
|
||||
{
|
||||
id: "google/gemma-4-E2B-it",
|
||||
name: "Gemma 4 E2B (inferrs)",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 131072,
|
||||
maxTokens: 4096,
|
||||
compat: {
|
||||
requiresStringContent: true,
|
||||
},
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`command` must be an absolute path. Run `which inferrs` on the Gateway host and use that path. Full field reference: [Local model services](/gateway/local-model-services).
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Why requiresStringContent matters">
|
||||
Some `inferrs` Chat Completions routes accept only string `messages[].content`, not structured content-part arrays.
|
||||
|
||||
<Warning>
|
||||
If OpenClaw runs fail with:
|
||||
|
||||
```text
|
||||
messages[1].content: invalid type: sequence, expected a string
|
||||
```
|
||||
|
||||
set `compat.requiresStringContent: true` in the model entry. OpenClaw then flattens pure text content parts into plain strings before sending the request.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Gemma and tool-schema caveat">
|
||||
Some `inferrs` + Gemma combinations accept small direct `/v1/chat/completions` requests but fail on full OpenClaw agent-runtime turns. Try disabling the tool schema surface first:
|
||||
|
||||
```json5
|
||||
compat: {
|
||||
requiresStringContent: true,
|
||||
supportsTools: false
|
||||
}
|
||||
```
|
||||
|
||||
That reduces prompt pressure on stricter local backends. If tiny direct requests still work but normal OpenClaw agent turns keep crashing inside `inferrs`, treat it as an upstream model/server limitation rather than an OpenClaw transport issue.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Manual smoke test">
|
||||
Test both layers once configured:
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:8080/v1/chat/completions \
|
||||
-H 'content-type: application/json' \
|
||||
-d '{"model":"google/gemma-4-E2B-it","messages":[{"role":"user","content":"What is 2 + 2?"}],"stream":false}'
|
||||
```
|
||||
|
||||
```bash
|
||||
openclaw infer model run \
|
||||
--model inferrs/google/gemma-4-E2B-it \
|
||||
--prompt "What is 2 + 2? Reply with one short sentence." \
|
||||
--json
|
||||
```
|
||||
|
||||
If the first command works but the second fails, see Troubleshooting below.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Proxy-style behavior">
|
||||
Because `inferrs` uses the generic `openai-completions` adapter (not `openai-responses`), native-OpenAI-only request shaping never applies: no `service_tier`, no Responses `store`, no prompt-cache hints, and no OpenAI reasoning-compat payload shaping get sent.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="curl /v1/models fails">
|
||||
`inferrs` is not running, not reachable, or not bound to the host/port you configured. Confirm the server is started and listening on that address.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="messages[].content expected a string">
|
||||
Set `compat.requiresStringContent: true` in the model entry (see above).
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Direct /v1/chat/completions calls pass but openclaw infer model run fails">
|
||||
Set `compat.supportsTools: false` to disable the tool schema surface (see the Gemma caveat above).
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="inferrs still crashes on larger agent turns">
|
||||
If schema errors are gone but `inferrs` still crashes on larger agent turns, treat it as an upstream `inferrs` or model limitation. Reduce prompt pressure or switch backend/model.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Tip>
|
||||
For general help, see [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq).
|
||||
</Tip>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Local models" href="/gateway/local-models" icon="server">
|
||||
Running OpenClaw against local model servers.
|
||||
</Card>
|
||||
<Card title="Local model services" href="/gateway/local-model-services" icon="play">
|
||||
Starting local model servers on demand for configured providers.
|
||||
</Card>
|
||||
<Card title="Gateway troubleshooting" href="/gateway/troubleshooting#local-openai-compatible-backend-passes-direct-probes-but-agent-runs-fail" icon="wrench">
|
||||
Debugging local OpenAI-compatible backends that pass probes but fail agent runs.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Overview of all providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
109
docs/providers/inworld.md
Normal file
109
docs/providers/inworld.md
Normal file
@@ -0,0 +1,109 @@
|
||||
---
|
||||
summary: "Inworld streaming text-to-speech for OpenClaw replies"
|
||||
read_when:
|
||||
- You want Inworld speech synthesis for outbound replies
|
||||
- You need PCM telephony or OGG_OPUS voice-note output from Inworld
|
||||
title: "Inworld"
|
||||
---
|
||||
|
||||
Inworld is a streaming text-to-speech (TTS) provider. In OpenClaw it synthesizes outbound reply audio (MP3 by default, OGG_OPUS for voice notes) and raw PCM audio for telephony channels such as Voice Call.
|
||||
|
||||
OpenClaw posts to Inworld's streaming TTS endpoint, concatenates the returned base64 audio chunks into a single buffer, and hands the result to the standard reply-audio pipeline.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | --------------------------------------------------------------- |
|
||||
| Provider id | `inworld` |
|
||||
| Plugin | official external package (`@openclaw/inworld-speech`) |
|
||||
| Contract | `speechProviders` (TTS only) |
|
||||
| Auth env var | `INWORLD_API_KEY` (HTTP Basic, Base64 dashboard credential) |
|
||||
| Base URL | `https://api.inworld.ai` |
|
||||
| Default voice | `Sarah` |
|
||||
| Default model | `inworld-tts-1.5-max` |
|
||||
| Output | MP3 (default), OGG_OPUS (voice notes), PCM 22050 Hz (telephony) |
|
||||
| Website | [inworld.ai](https://inworld.ai) |
|
||||
| Docs | [docs.inworld.ai/tts/tts](https://docs.inworld.ai/tts/tts) |
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/inworld-speech
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set your API key">
|
||||
Copy the credential from your Inworld dashboard (Workspace > API Keys) and set it as an env var. The value is sent verbatim as the HTTP Basic credential, so do not Base64-encode it again or convert it to a bearer token.
|
||||
|
||||
```bash
|
||||
INWORLD_API_KEY=<base64-credential-from-dashboard>
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Select Inworld in messages.tts">
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "inworld",
|
||||
providers: {
|
||||
inworld: {
|
||||
voiceId: "Sarah",
|
||||
modelId: "inworld-tts-1.5-max",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Send a message">
|
||||
Send a reply through any connected channel. OpenClaw synthesizes the audio with Inworld and delivers it as MP3 (or OGG_OPUS when the channel expects a voice note).
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Configuration options
|
||||
|
||||
| Option | Path | Description |
|
||||
| ------------- | -------------------------------------------- | ------------------------------------------------------------------- |
|
||||
| `apiKey` | `messages.tts.providers.inworld.apiKey` | Base64 dashboard credential. Falls back to `INWORLD_API_KEY`. |
|
||||
| `baseUrl` | `messages.tts.providers.inworld.baseUrl` | Override Inworld API base URL (default `https://api.inworld.ai`). |
|
||||
| `voiceId` | `messages.tts.providers.inworld.voiceId` | Voice identifier (default `Sarah`). Legacy alias: `speakerVoiceId`. |
|
||||
| `modelId` | `messages.tts.providers.inworld.modelId` | TTS model id (default `inworld-tts-1.5-max`). |
|
||||
| `temperature` | `messages.tts.providers.inworld.temperature` | Sampling temperature, `0` (exclusive) to `2` (optional). |
|
||||
|
||||
## Notes
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Authentication">
|
||||
Inworld uses HTTP Basic auth with a single Base64-encoded credential string. Copy it verbatim from the Inworld dashboard. The provider sends it as `Authorization: Basic <apiKey>` without any further encoding, so do not Base64-encode it yourself and do not pass a bearer-style token. See [TTS auth notes](/tools/tts#inworld-primary) for the same callout.
|
||||
</Accordion>
|
||||
<Accordion title="Models">
|
||||
Supported model ids: `inworld-tts-1.5-max` (default), `inworld-tts-1.5-mini`, `inworld-tts-1-max`, `inworld-tts-1`.
|
||||
</Accordion>
|
||||
<Accordion title="Audio outputs">
|
||||
Replies use MP3 by default. When the channel target is `voice-note`, OpenClaw asks Inworld for `OGG_OPUS` so the audio plays as a native voice bubble. Telephony synthesis uses raw `PCM` at 22050 Hz to feed the telephony bridge.
|
||||
</Accordion>
|
||||
<Accordion title="Custom endpoints">
|
||||
Override the API host with `messages.tts.providers.inworld.baseUrl`. Trailing slashes are stripped before requests are sent.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Text-to-speech" href="/tools/tts" icon="waveform-lines">
|
||||
TTS overview, providers, and `messages.tts` config.
|
||||
</Card>
|
||||
<Card title="Configuration" href="/gateway/configuration" icon="gear">
|
||||
Full config reference including `messages.tts` settings.
|
||||
</Card>
|
||||
<Card title="Providers" href="/providers" icon="grid">
|
||||
All supported OpenClaw providers.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
Common issues and debugging steps.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
121
docs/providers/kilocode.md
Normal file
121
docs/providers/kilocode.md
Normal file
@@ -0,0 +1,121 @@
|
||||
---
|
||||
summary: "Use Kilo Gateway's unified API to access many models in OpenClaw"
|
||||
title: "Kilo Gateway"
|
||||
read_when:
|
||||
- You want a single API key for many LLMs
|
||||
- You want to run models via Kilo Gateway in OpenClaw
|
||||
---
|
||||
|
||||
Kilo Gateway routes requests to many models behind a single OpenAI-compatible endpoint and API key.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ---------------------------------- |
|
||||
| Provider | `kilocode` |
|
||||
| Auth | `KILOCODE_API_KEY` |
|
||||
| API | OpenAI-compatible |
|
||||
| Base URL | `https://api.kilo.ai/api/gateway/` |
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/kilocode-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Setup
|
||||
|
||||
<Steps>
|
||||
<Step title="Create an account">
|
||||
Go to [app.kilo.ai](https://app.kilo.ai), sign in or create an account, then generate an API key.
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice kilocode-api-key
|
||||
```
|
||||
|
||||
Or set the environment variable directly:
|
||||
|
||||
```bash
|
||||
export KILOCODE_API_KEY="<your-kilocode-api-key>" # pragma: allowlist secret
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider kilocode
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Default model and catalog
|
||||
|
||||
The default model is `kilocode/kilo/auto`, a provider-owned smart-routing model. OpenClaw does not
|
||||
publish a task-to-upstream-model mapping for it; routing behind `kilo/auto` is owned by Kilo Gateway.
|
||||
|
||||
At startup OpenClaw queries `GET https://api.kilo.ai/api/gateway/models` and merges discovered models
|
||||
ahead of a static fallback catalog. The static fallback contains only `kilocode/kilo/auto` (`Kilo Auto`,
|
||||
`input: ["text", "image"]`, `reasoning: true`, `contextWindow: 1000000`, `maxTokens: 128000`).
|
||||
|
||||
Any model on the gateway is addressable as `kilocode/<upstream-id>` (for example
|
||||
`kilocode/anthropic/claude-sonnet-4`, `kilocode/openai/gpt-5.5`). Run `/models kilocode` or
|
||||
`openclaw models list --provider kilocode` to see the full discovered list.
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { KILOCODE_API_KEY: "<your-kilocode-api-key>" }, // pragma: allowlist secret
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "kilocode/kilo/auto" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Behavior notes
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Transport and compatibility">
|
||||
Kilo Gateway is OpenRouter-compatible, so it uses the proxy-style OpenAI-compatible request
|
||||
path rather than native OpenAI request shaping (no `store`, no OpenAI reasoning-effort payload).
|
||||
|
||||
- Gemini-backed Kilo refs stay on the proxy-Gemini path: OpenClaw sanitizes Gemini thought
|
||||
signatures there but does not enable native Gemini replay validation or bootstrap rewrites.
|
||||
- Requests use a Bearer token built from your API key.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Stream wrapper and reasoning">
|
||||
The Kilo stream wrapper adds an `X-KILOCODE-FEATURE` request header (default `openclaw`,
|
||||
override with the `KILOCODE_FEATURE` env var) and normalizes reasoning-effort payloads for
|
||||
models that support it.
|
||||
|
||||
<Warning>
|
||||
`kilocode/kilo/auto` and `x-ai/*` refs skip reasoning-effort injection. Use a concrete model
|
||||
ref such as `kilocode/anthropic/claude-sonnet-4` if you need reasoning support.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Troubleshooting">
|
||||
- If model discovery fails at startup, OpenClaw falls back to the static catalog containing `kilocode/kilo/auto`.
|
||||
- Confirm your API key is valid and that your Kilo account has the desired models enabled.
|
||||
- When Gateway runs as a daemon, ensure `KILOCODE_API_KEY` is available to that process (for example in `~/.openclaw/.env` or via `env.shellEnv`).
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full OpenClaw configuration reference.
|
||||
</Card>
|
||||
<Card title="Kilo Gateway" href="https://app.kilo.ai" icon="arrow-up-right-from-square">
|
||||
Kilo Gateway dashboard, API keys, and account management.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
204
docs/providers/litellm.md
Normal file
204
docs/providers/litellm.md
Normal file
@@ -0,0 +1,204 @@
|
||||
---
|
||||
summary: "Run OpenClaw through LiteLLM Proxy for unified model access and cost tracking"
|
||||
title: "LiteLLM"
|
||||
read_when:
|
||||
- You want to route OpenClaw through a LiteLLM proxy
|
||||
- You need cost tracking, logging, or model routing through LiteLLM
|
||||
---
|
||||
|
||||
[LiteLLM](https://litellm.ai) is an open-source LLM gateway with a unified API to 100+ model
|
||||
providers. Route OpenClaw through LiteLLM for centralized cost tracking, logging, virtual keys with
|
||||
spend limits, and backend failover without changing OpenClaw config.
|
||||
|
||||
## Quick start
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Onboarding (recommended)">
|
||||
```bash
|
||||
openclaw onboard --auth-choice litellm-api-key
|
||||
```
|
||||
|
||||
For non-interactive setup against a remote proxy, pass the proxy URL explicitly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive --accept-risk --auth-choice litellm-api-key \
|
||||
--litellm-api-key "$LITELLM_API_KEY" --custom-base-url "https://litellm.example/v1"
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Manual setup">
|
||||
<Steps>
|
||||
<Step title="Start LiteLLM Proxy">
|
||||
```bash
|
||||
pip install 'litellm[proxy]'
|
||||
litellm --model claude-opus-4-6
|
||||
```
|
||||
</Step>
|
||||
<Step title="Point OpenClaw to LiteLLM">
|
||||
```bash
|
||||
export LITELLM_API_KEY="your-litellm-key"
|
||||
openclaw
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Configuration
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
litellm: {
|
||||
baseUrl: "http://localhost:4000",
|
||||
apiKey: "${LITELLM_API_KEY}",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
{
|
||||
id: "claude-opus-4-6",
|
||||
name: "Claude Opus 4.6",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
contextWindow: 200000,
|
||||
maxTokens: 64000,
|
||||
},
|
||||
{
|
||||
id: "gpt-4o",
|
||||
name: "GPT-4o",
|
||||
reasoning: false,
|
||||
input: ["text", "image"],
|
||||
contextWindow: 128000,
|
||||
maxTokens: 8192,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "litellm/claude-opus-4-6" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
The default model onboarding writes is `litellm/claude-opus-4-6`.
|
||||
|
||||
## Image generation
|
||||
|
||||
LiteLLM can back the `image_generate` tool through OpenAI-compatible `/images/generations` and
|
||||
`/images/edits` routes. Default image model is `gpt-image-2`; configure a different one under
|
||||
`agents.defaults.imageGenerationModel`:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
litellm: {
|
||||
baseUrl: "http://localhost:4000",
|
||||
apiKey: "${LITELLM_API_KEY}",
|
||||
},
|
||||
},
|
||||
},
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "litellm/gpt-image-2",
|
||||
timeoutMs: 180_000,
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Loopback LiteLLM URLs (`http://localhost:4000`, `127.0.0.1`, `::1`, `host.docker.internal`) work
|
||||
without a global private-network override. For a LAN-hosted proxy, set
|
||||
`models.providers.litellm.request.allowPrivateNetwork: true` because the API key is sent to that host.
|
||||
|
||||
## Advanced
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Virtual keys">
|
||||
Create a dedicated key for OpenClaw with spend limits:
|
||||
|
||||
```bash
|
||||
curl -X POST "http://localhost:4000/key/generate" \
|
||||
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"key_alias": "openclaw",
|
||||
"max_budget": 50.00,
|
||||
"budget_duration": "monthly"
|
||||
}'
|
||||
```
|
||||
|
||||
Use the generated key as `LITELLM_API_KEY`.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Model routing">
|
||||
LiteLLM can route model requests to different backends. Configure in your LiteLLM `config.yaml`:
|
||||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: claude-opus-4-6
|
||||
litellm_params:
|
||||
model: claude-opus-4-6
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
|
||||
- model_name: gpt-4o
|
||||
litellm_params:
|
||||
model: gpt-4o
|
||||
api_key: os.environ/OPENAI_API_KEY
|
||||
```
|
||||
|
||||
OpenClaw keeps requesting `claude-opus-4-6`; LiteLLM handles the routing.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Viewing usage">
|
||||
```bash
|
||||
# Key info
|
||||
curl "http://localhost:4000/key/info" \
|
||||
-H "Authorization: Bearer sk-litellm-key"
|
||||
|
||||
# Spend logs
|
||||
curl "http://localhost:4000/spend/logs" \
|
||||
-H "Authorization: Bearer $LITELLM_MASTER_KEY"
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Proxy behavior notes">
|
||||
- LiteLLM runs on `http://localhost:4000` by default.
|
||||
- OpenClaw connects through LiteLLM's proxy-style OpenAI-compatible `/v1` endpoint.
|
||||
- Native-OpenAI-only request shaping does not apply through a configured LiteLLM base URL:
|
||||
no `service_tier`, no Responses `store`, no prompt-cache hints, no OpenAI reasoning-effort
|
||||
payload shaping.
|
||||
- Hidden OpenClaw attribution headers (`originator`, `version`, `User-Agent`) are only sent to
|
||||
verified native OpenAI endpoints, so they are not injected on a custom LiteLLM base URL.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Note>
|
||||
For general provider configuration and failover behavior, see [Model Providers](/concepts/model-providers).
|
||||
</Note>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="LiteLLM Docs" href="https://docs.litellm.ai" icon="book">
|
||||
Official LiteLLM documentation and API reference.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Overview of all providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration" href="/gateway/configuration" icon="gear">
|
||||
Full config reference.
|
||||
</Card>
|
||||
<Card title="Models" href="/concepts/models" icon="brain">
|
||||
How to choose and configure models.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
214
docs/providers/lmstudio.md
Normal file
214
docs/providers/lmstudio.md
Normal file
@@ -0,0 +1,214 @@
|
||||
---
|
||||
summary: "Run OpenClaw with LM Studio"
|
||||
read_when:
|
||||
- You want to run OpenClaw with open source models via LM Studio
|
||||
- You want to set up and configure LM Studio
|
||||
title: "LM Studio"
|
||||
---
|
||||
|
||||
LM Studio runs llama.cpp (GGUF) or MLX models locally, as a GUI app or the headless `llmster`
|
||||
daemon. For install and product docs, see [lmstudio.ai](https://lmstudio.ai/).
|
||||
|
||||
## Quick start
|
||||
|
||||
<Steps>
|
||||
<Step title="Install and start the server">
|
||||
Install LM Studio (desktop) or `llmster` (headless), then start the server:
|
||||
|
||||
```bash
|
||||
lms server start --port 1234
|
||||
```
|
||||
|
||||
Or run the headless daemon:
|
||||
|
||||
```bash
|
||||
lms daemon up
|
||||
```
|
||||
|
||||
If using the desktop app, enable JIT for smooth model loading; see the
|
||||
[LM Studio JIT and TTL guide](https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict).
|
||||
|
||||
</Step>
|
||||
<Step title="Set an API key if auth is enabled">
|
||||
```bash
|
||||
export LM_API_TOKEN="your-lm-studio-api-token"
|
||||
```
|
||||
|
||||
If LM Studio authentication is disabled, leave the API key blank during setup. See
|
||||
[LM Studio Authentication](https://lmstudio.ai/docs/developer/core/authentication).
|
||||
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard
|
||||
```
|
||||
|
||||
Choose `LM Studio`, then pick a model at the `Default model` prompt.
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
Change the default model later:
|
||||
|
||||
```bash
|
||||
openclaw models set lmstudio/qwen/qwen3.5-9b
|
||||
```
|
||||
|
||||
LM Studio model keys use an `author/model-name` format (e.g. `qwen/qwen3.5-9b`); OpenClaw model refs
|
||||
prepend the provider: `lmstudio/qwen/qwen3.5-9b`. Find the exact key for a model by running the
|
||||
command below and looking at the `key` field:
|
||||
|
||||
```bash
|
||||
curl http://localhost:1234/api/v1/models
|
||||
```
|
||||
|
||||
## Non-interactive onboarding
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive --accept-risk --auth-choice lmstudio
|
||||
```
|
||||
|
||||
Or specify base URL, model, and API key explicitly:
|
||||
|
||||
```bash
|
||||
openclaw onboard \
|
||||
--non-interactive \
|
||||
--accept-risk \
|
||||
--auth-choice lmstudio \
|
||||
--custom-base-url http://localhost:1234/v1 \
|
||||
--lmstudio-api-key "$LM_API_TOKEN" \
|
||||
--custom-model-id qwen/qwen3.5-9b
|
||||
```
|
||||
|
||||
`--custom-model-id` takes the model key as returned by LM Studio (e.g. `qwen/qwen3.5-9b`), without
|
||||
the `lmstudio/` provider prefix. Pass `--lmstudio-api-key` (or set `LM_API_TOKEN`) for authenticated
|
||||
servers; omit it for unauthenticated servers and OpenClaw stores a local non-secret marker instead.
|
||||
`--custom-api-key` is still accepted for compatibility, but `--lmstudio-api-key` is preferred.
|
||||
|
||||
This writes `models.providers.lmstudio` and sets the default model to `lmstudio/<custom-model-id>`.
|
||||
Providing an API key also writes the `lmstudio:default` auth profile.
|
||||
|
||||
Interactive setup can additionally prompt for a preferred load context length and applies it across
|
||||
the discovered models it saves to config.
|
||||
|
||||
## Configuration
|
||||
|
||||
### Streaming usage compatibility
|
||||
|
||||
LM Studio doesn't always emit an OpenAI-shaped `usage` object on streamed responses. OpenClaw
|
||||
recovers token counts from llama.cpp-style `timings.prompt_n` / `timings.predicted_n` metadata
|
||||
instead. Any OpenAI-compatible endpoint resolved as a local endpoint (loopback host) gets this same
|
||||
fallback, which covers other local backends such as vLLM, SGLang, llama.cpp, LocalAI, Jan, TabbyAPI,
|
||||
and text-generation-webui.
|
||||
|
||||
### Thinking compatibility
|
||||
|
||||
When LM Studio's `/api/v1/models` discovery reports model-specific reasoning options, OpenClaw
|
||||
exposes matching `reasoning_effort` values (`none`, `minimal`, `low`, `medium`, `high`, `xhigh`) in
|
||||
model compat metadata. Some LM Studio builds advertise a binary UI option (`allowed_options: ["off",
|
||||
"on"]`) while rejecting those literal values on `/v1/chat/completions`; OpenClaw normalizes that
|
||||
binary shape to the six-level scale before sending requests, including for older saved config that
|
||||
still has `off`/`on` reasoning maps.
|
||||
|
||||
### Explicit configuration
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
lmstudio: {
|
||||
baseUrl: "http://localhost:1234/v1",
|
||||
apiKey: "${LM_API_TOKEN}",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
{
|
||||
id: "qwen/qwen3-coder-next",
|
||||
name: "Qwen 3 Coder Next",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 128000,
|
||||
maxTokens: 8192,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### Disabling preload
|
||||
|
||||
LM Studio supports just-in-time (JIT) model loading, loading models on first request. OpenClaw
|
||||
preloads models through LM Studio's native load endpoint by default, which helps when JIT is
|
||||
disabled. To let LM Studio's JIT, idle TTL, and auto-evict behavior own model lifecycle instead,
|
||||
disable OpenClaw's preload step:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
lmstudio: {
|
||||
baseUrl: "http://localhost:1234/v1",
|
||||
api: "openai-completions",
|
||||
params: { preload: false },
|
||||
models: [{ id: "qwen/qwen3.5-9b" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### LAN or tailnet host
|
||||
|
||||
Use the LM Studio host's reachable address, keep `/v1`, and make sure LM Studio is bound beyond
|
||||
loopback on that machine:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
lmstudio: {
|
||||
baseUrl: "http://gpu-box.local:1234/v1",
|
||||
apiKey: "lmstudio",
|
||||
api: "openai-completions",
|
||||
models: [{ id: "qwen/qwen3.5-9b" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`lmstudio` automatically trusts its configured endpoint for model requests, including loopback,
|
||||
LAN, and tailnet hosts (except metadata/link-local origins). Any custom/local OpenAI-compatible
|
||||
provider entry gets the same exact-origin trust. Requests to a different private host or port still
|
||||
require `models.providers.<id>.request.allowPrivateNetwork: true`; set it to `false` to opt out of
|
||||
the default trust.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### LM Studio not detected
|
||||
|
||||
Make sure LM Studio is running:
|
||||
|
||||
```bash
|
||||
lms server start --port 1234
|
||||
```
|
||||
|
||||
If authentication is enabled, also set `LM_API_TOKEN`. Verify the API is reachable:
|
||||
|
||||
```bash
|
||||
curl http://localhost:1234/api/v1/models
|
||||
```
|
||||
|
||||
### Authentication errors (HTTP 401)
|
||||
|
||||
- Check that `LM_API_TOKEN` matches the key configured in LM Studio.
|
||||
- See [LM Studio Authentication](https://lmstudio.ai/docs/developer/core/authentication).
|
||||
- If the server does not require authentication, leave the key blank during setup.
|
||||
|
||||
## Related
|
||||
|
||||
- [Model selection](/concepts/model-providers)
|
||||
- [Ollama](/providers/ollama)
|
||||
- [Local models](/gateway/local-models)
|
||||
451
docs/providers/minimax.md
Normal file
451
docs/providers/minimax.md
Normal file
@@ -0,0 +1,451 @@
|
||||
---
|
||||
summary: "Use MiniMax models in OpenClaw"
|
||||
read_when:
|
||||
- You want MiniMax models in OpenClaw
|
||||
- You need MiniMax setup guidance
|
||||
title: "MiniMax"
|
||||
---
|
||||
|
||||
The bundled `minimax` plugin registers two providers plus five capabilities: chat, image generation, music generation, video generation, image understanding, speech (T2A v2), and web search.
|
||||
|
||||
| Provider ID | Auth | Capabilities |
|
||||
| ---------------- | ------- | --------------------------------------------------------------------------------------------------- |
|
||||
| `minimax` | API key | Text, image generation, music generation, video generation, image understanding, speech, web search |
|
||||
| `minimax-portal` | OAuth | Text, image generation, music generation, video generation, image understanding, speech |
|
||||
|
||||
<Tip>
|
||||
Referral link for MiniMax Coding Plan (10% off): [MiniMax Coding Plan](https://platform.minimax.io/subscribe/coding-plan?code=DbXJTRClnb&source=link)
|
||||
</Tip>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
| Model | Type | Description |
|
||||
| ------------------------ | ---------------- | ---------------------------------------- |
|
||||
| `MiniMax-M3` | Chat (reasoning) | Default hosted reasoning model |
|
||||
| `MiniMax-M2.7` | Chat (reasoning) | Previous hosted reasoning model |
|
||||
| `MiniMax-M2.7-highspeed` | Chat (reasoning) | Faster M2.7 reasoning tier |
|
||||
| `MiniMax-VL-01` | Vision | Image understanding model |
|
||||
| `image-01` | Image generation | Text-to-image and image-to-image editing |
|
||||
| `music-2.6` | Music generation | Default music model |
|
||||
| `MiniMax-Hailuo-2.3` | Video generation | Text-to-video and image-to-video flows |
|
||||
|
||||
Model refs follow the auth path: `minimax/<model>` for API-key setups, `minimax-portal/<model>` for OAuth setups.
|
||||
|
||||
## Getting started
|
||||
|
||||
<Tabs>
|
||||
<Tab title="OAuth (Coding Plan)">
|
||||
**Best for:** quick setup with MiniMax Coding Plan via OAuth, no API key required.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="International">
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice minimax-global-oauth
|
||||
```
|
||||
|
||||
Resulting provider base URL: `api.minimax.io`.
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider minimax-portal
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
<Tab title="China">
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice minimax-cn-oauth
|
||||
```
|
||||
|
||||
Resulting provider base URL: `api.minimaxi.com`.
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider minimax-portal
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
<Note>
|
||||
OAuth setups use the `minimax-portal` provider id. Model refs follow the form `minimax-portal/MiniMax-M3`.
|
||||
</Note>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="API key">
|
||||
**Best for:** hosted MiniMax with Anthropic-compatible API.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="International">
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice minimax-global-api
|
||||
```
|
||||
|
||||
This configures `api.minimax.io` as the base URL.
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider minimax
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
<Tab title="China">
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice minimax-cn-api
|
||||
```
|
||||
|
||||
This configures `api.minimaxi.com` as the base URL.
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider minimax
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
### Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { MINIMAX_API_KEY: "sk-..." },
|
||||
agents: { defaults: { model: { primary: "minimax/MiniMax-M3" } } },
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
minimax: {
|
||||
baseUrl: "https://api.minimax.io/anthropic",
|
||||
apiKey: "${MINIMAX_API_KEY}",
|
||||
api: "anthropic-messages",
|
||||
models: [
|
||||
{
|
||||
id: "MiniMax-M3",
|
||||
name: "MiniMax M3",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
cost: { input: 0.6, output: 2.4, cacheRead: 0.12, cacheWrite: 0 },
|
||||
contextWindow: 1000000,
|
||||
maxTokens: 131072,
|
||||
},
|
||||
{
|
||||
id: "MiniMax-M2.7",
|
||||
name: "MiniMax M2.7",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0.3, output: 1.2, cacheRead: 0.06, cacheWrite: 0.375 },
|
||||
contextWindow: 204800,
|
||||
maxTokens: 131072,
|
||||
},
|
||||
{
|
||||
id: "MiniMax-M2.7-highspeed",
|
||||
name: "MiniMax M2.7 Highspeed",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0.6, output: 2.4, cacheRead: 0.06, cacheWrite: 0.375 },
|
||||
contextWindow: 204800,
|
||||
maxTokens: 131072,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Warning>
|
||||
MiniMax-M2.x's Anthropic-compatible streaming endpoint emits `reasoning_content` in OpenAI-style delta chunks instead of native Anthropic thinking blocks, which leaks internal reasoning into visible output if thinking is left enabled implicitly. OpenClaw disables M2.x thinking by default unless you explicitly set `thinking` yourself. MiniMax-M3 (and forward-compatible M3.x) is exempt: M3 emits proper Anthropic thinking blocks and requires thinking active to produce visible content, so OpenClaw keeps M3 on the provider's adaptive thinking path. See the Thinking defaults section under Advanced configuration below.
|
||||
</Warning>
|
||||
|
||||
<Note>
|
||||
API-key setups use the `minimax` provider id. Model refs follow the form `minimax/MiniMax-M3`.
|
||||
</Note>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Configure via `openclaw configure`
|
||||
|
||||
<Steps>
|
||||
<Step title="Launch the wizard">
|
||||
```bash
|
||||
openclaw configure
|
||||
```
|
||||
</Step>
|
||||
<Step title="Select Model/auth">
|
||||
Choose **Model/auth** from the menu.
|
||||
</Step>
|
||||
<Step title="Choose a MiniMax auth option">
|
||||
| Auth choice | Description |
|
||||
| ----------------------- | ----------------------------------- |
|
||||
| `minimax-global-oauth` | International OAuth (Coding Plan) |
|
||||
| `minimax-cn-oauth` | China OAuth (Coding Plan) |
|
||||
| `minimax-global-api` | International API key |
|
||||
| `minimax-cn-api` | China API key |
|
||||
</Step>
|
||||
<Step title="Pick your default model">
|
||||
Select your default model when prompted.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Capabilities
|
||||
|
||||
### Image generation
|
||||
|
||||
The MiniMax plugin registers the `image-01` model for the `image_generate` tool on both `minimax` and `minimax-portal`, reusing the same `MINIMAX_API_KEY` or OAuth auth as the text models.
|
||||
|
||||
- Text-to-image generation and image-to-image editing (subject reference), both with aspect ratio control
|
||||
- Up to 9 output images per request, 1 reference image per edit request
|
||||
- Supported aspect ratios: `1:1`, `16:9`, `4:3`, `3:2`, `2:3`, `3:4`, `9:16`, `21:9`
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: { primary: "minimax/image-01" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Image generation always uses MiniMax's dedicated image endpoint (`/v1/image_generation`) and ignores `models.providers.minimax.baseUrl`, since that field configures the chat/Anthropic-compatible base URL instead. Set `MINIMAX_API_HOST=https://api.minimaxi.com` to route image generation through the CN endpoint; the default global endpoint is `https://api.minimax.io`.
|
||||
|
||||
<Note>
|
||||
See [Image Generation](/tools/image-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
### Text-to-speech
|
||||
|
||||
The bundled `minimax` plugin registers MiniMax T2A v2 as a speech provider for `messages.tts`.
|
||||
|
||||
- Default TTS model: `speech-2.8-hd`
|
||||
- Default voice: `English_expressive_narrator`
|
||||
- Bundled model ids: `speech-2.8-hd`, `speech-2.8-turbo`, `speech-2.6-hd`, `speech-2.6-turbo`, `speech-02-hd`, `speech-02-turbo`, `speech-01-hd`, `speech-01-turbo`, `speech-01-240228`
|
||||
- Auth resolution order: `messages.tts.providers.minimax.apiKey`, then `minimax-portal` OAuth/token auth profiles, then Token Plan environment keys (`MINIMAX_OAUTH_TOKEN`, `MINIMAX_CODE_PLAN_KEY`, `MINIMAX_CODING_API_KEY`), then `MINIMAX_API_KEY`
|
||||
- If no TTS host is configured, OpenClaw reuses the configured `minimax-portal` OAuth host and strips Anthropic-compatible path suffixes such as `/anthropic`
|
||||
- Normal audio attachments stay MP3. Voice-note targets (Feishu, Telegram, and other channels that request a voice-note-compatible attachment) are transcoded from MiniMax MP3 to 48kHz Opus with `ffmpeg`, because e.g. the Feishu/Lark file API only accepts `file_type: "opus"` for native audio messages
|
||||
- MiniMax T2A accepts fractional `speed` and `vol`, but `pitch` is sent as an integer; OpenClaw truncates fractional `pitch` values before the API request
|
||||
|
||||
| Setting | Env var | Default | Description |
|
||||
| ---------------------------------------- | ---------------------- | ----------------------------- | -------------------------------- |
|
||||
| `messages.tts.providers.minimax.baseUrl` | `MINIMAX_API_HOST` | `https://api.minimax.io` | MiniMax T2A API host. |
|
||||
| `messages.tts.providers.minimax.model` | `MINIMAX_TTS_MODEL` | `speech-2.8-hd` | TTS model id. |
|
||||
| `messages.tts.providers.minimax.voiceId` | `MINIMAX_TTS_VOICE_ID` | `English_expressive_narrator` | Voice id used for speech output. |
|
||||
| `messages.tts.providers.minimax.speed` | | `1.0` | Playback speed, `0.5..2.0`. |
|
||||
| `messages.tts.providers.minimax.vol` | | `1.0` | Volume, `(0, 10]`. |
|
||||
| `messages.tts.providers.minimax.pitch` | | `0` | Integer pitch shift, `-12..12`. |
|
||||
|
||||
### Music generation
|
||||
|
||||
The bundled MiniMax plugin registers music generation through the shared `music_generate` tool for both `minimax` and `minimax-portal`.
|
||||
|
||||
- Default music model: `minimax/music-2.6` (OAuth: `minimax-portal/music-2.6`)
|
||||
- Also supports `music-2.6-free`, `music-cover`, and `music-cover-free`
|
||||
- Prompt controls: `lyrics`, `instrumental`
|
||||
- Output format: `mp3`
|
||||
- Session-backed runs detach through the shared task/status flow, including `action: "status"`
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
musicGenerationModel: { primary: "minimax/music-2.6" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
See [Music Generation](/tools/music-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
### Video generation
|
||||
|
||||
The bundled MiniMax plugin registers video generation through the shared `video_generate` tool for both `minimax` and `minimax-portal`.
|
||||
|
||||
- Default video model: `minimax/MiniMax-Hailuo-2.3` (OAuth: `minimax-portal/MiniMax-Hailuo-2.3`)
|
||||
- Also supports `MiniMax-Hailuo-2.3-Fast`, `MiniMax-Hailuo-02`, `I2V-01-Director`, `I2V-01-live`, and `I2V-01`
|
||||
- Modes: text-to-video and single-image reference flows
|
||||
- Supports `resolution` (`768P` or `1080P` on Hailuo 2.3/02 models); `aspectRatio` is not supported and is ignored
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: { primary: "minimax/MiniMax-Hailuo-2.3" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
See [Video Generation](/tools/video-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
### Image understanding
|
||||
|
||||
The MiniMax plugin registers image understanding separately from the text catalog:
|
||||
|
||||
| Provider ID | Default image model | PDF text extraction |
|
||||
| ---------------- | ------------------- | ------------------- |
|
||||
| `minimax` | `MiniMax-VL-01` | `MiniMax-M2.7` |
|
||||
| `minimax-portal` | `MiniMax-VL-01` | `MiniMax-M2.7` |
|
||||
|
||||
That is why automatic media routing can use MiniMax image understanding even when the bundled text-provider catalog also includes M3 image-capable chat refs. PDF understanding uses `MiniMax-M2.7` for text extraction only; MiniMax does not register a PDF-to-image conversion path.
|
||||
|
||||
### Web search
|
||||
|
||||
The MiniMax plugin also registers `web_search` through the MiniMax Token Plan search API (`/v1/coding_plan/search`).
|
||||
|
||||
- Provider id: `minimax`
|
||||
- Structured results: titles, URLs, snippets, related queries
|
||||
- Preferred env var: `MINIMAX_CODE_PLAN_KEY`
|
||||
- Accepted env aliases: `MINIMAX_CODING_API_KEY`, `MINIMAX_OAUTH_TOKEN`
|
||||
- Compatibility fallback: `MINIMAX_API_KEY` when it already points at a token-plan credential
|
||||
- Region reuse: `plugins.entries.minimax.config.webSearch.region`, then `MINIMAX_API_HOST`, then MiniMax provider base URLs
|
||||
- Search stays on provider id `minimax`; OAuth CN/global setup can steer region indirectly through `models.providers.minimax-portal.baseUrl` and can provide bearer auth through `MINIMAX_OAUTH_TOKEN`
|
||||
|
||||
Config lives under `plugins.entries.minimax.config.webSearch.*`.
|
||||
|
||||
<Note>
|
||||
See [MiniMax Search](/tools/minimax-search) for full web search configuration and usage.
|
||||
</Note>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Configuration options">
|
||||
| Option | Description |
|
||||
| --- | --- |
|
||||
| `models.providers.minimax.baseUrl` | Prefer `https://api.minimax.io/anthropic` (Anthropic-compatible); `https://api.minimax.io/v1` is optional for OpenAI-compatible payloads |
|
||||
| `models.providers.minimax.api` | Prefer `anthropic-messages`; `openai-completions` is optional for OpenAI-compatible payloads |
|
||||
| `models.providers.minimax.apiKey` | MiniMax API key (`MINIMAX_API_KEY`) |
|
||||
| `models.providers.minimax.models` | Define `id`, `name`, `reasoning`, `contextWindow`, `maxTokens`, `cost` |
|
||||
| `agents.defaults.models` | Alias models you want in the allowlist |
|
||||
| `models.mode` | Keep `merge` if you want to add MiniMax alongside built-ins |
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Thinking defaults">
|
||||
On `api: "anthropic-messages"`, OpenClaw injects `thinking: { type: "disabled" }` for MiniMax M2.x models unless an earlier wrapper already set the `thinking` field in the payload. This prevents M2.x's streaming endpoint from emitting `reasoning_content` in OpenAI-style delta chunks, which would leak internal reasoning into visible output.
|
||||
|
||||
MiniMax-M3 (and M3.x) is exempt: M3 returns an empty `content` array with `stop_reason: "end_turn"` when thinking is disabled, so OpenClaw removes the implicit disabled default for M3 and, when a thinking level is set, forces `thinking: { type: "adaptive" }` instead.
|
||||
|
||||
Available thinking levels per model family:
|
||||
|
||||
| Model family | Levels | Default |
|
||||
| -------------- | ----------------------------------------- | ---------- |
|
||||
| `MiniMax-M3` | `off`, `adaptive` | `adaptive` |
|
||||
| `MiniMax-M2.x` | `off`, `minimal`, `low`, `medium`, `high` | `off` |
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Fast mode">
|
||||
`/fast on` or `params.fastMode: true` rewrites `MiniMax-M2.7` to `MiniMax-M2.7-highspeed` on the Anthropic-compatible stream path (`api: "anthropic-messages"`, provider `minimax` or `minimax-portal`).
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Fallback example">
|
||||
**Best for:** keep your strongest latest-generation model as primary, fail over to MiniMax M2.7. Example below uses Opus as a concrete primary; swap to your preferred latest-gen primary model.
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { MINIMAX_API_KEY: "sk-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"anthropic/claude-opus-4-6": { alias: "primary" },
|
||||
"minimax/MiniMax-M2.7": { alias: "minimax" },
|
||||
},
|
||||
model: {
|
||||
primary: "anthropic/claude-opus-4-6",
|
||||
fallbacks: ["minimax/MiniMax-M2.7"],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Coding Plan usage details">
|
||||
- Coding Plan usage API: `https://api.minimaxi.com/v1/token_plan/remains` or `https://api.minimax.io/v1/token_plan/remains` (requires a coding plan key).
|
||||
- Usage polling derives the host from `models.providers.minimax-portal.baseUrl` or `models.providers.minimax.baseUrl` when configured, so global setups using `https://api.minimax.io/anthropic` poll `api.minimax.io`. Missing or malformed base URLs keep the CN fallback for compatibility.
|
||||
- OpenClaw normalizes MiniMax coding-plan usage to the same `% left` display used by other providers. MiniMax's raw `usage_percent` / `usagePercent` fields are remaining quota, not consumed quota, so OpenClaw inverts them. Count-based fields win when present.
|
||||
- When the API returns `model_remains`, OpenClaw prefers the chat-model entry, derives the window label from `start_time` / `end_time` when needed, and includes the selected model name in the plan label so coding-plan windows are easier to distinguish.
|
||||
- Usage snapshots treat `minimax`, `minimax-cn`, `minimax-portal`, and `minimax-portal-cn` as the same MiniMax quota surface, and prefer stored MiniMax OAuth before falling back to Coding Plan key env vars.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Notes
|
||||
|
||||
- Default chat model: `MiniMax-M3`. Alternate chat models: `MiniMax-M2.7`, `MiniMax-M2.7-highspeed`
|
||||
- Onboarding and direct API-key setup write model definitions for M3 and both M2.7 variants
|
||||
- Image understanding uses the plugin-owned `MiniMax-VL-01` media provider
|
||||
- Update pricing values in `models.json` if you need exact cost tracking
|
||||
- Use `openclaw models list` to confirm the current provider id, then switch with `openclaw models set minimax/MiniMax-M3` or `openclaw models set minimax-portal/MiniMax-M3`
|
||||
|
||||
<Note>
|
||||
See [Model providers](/concepts/model-providers) for provider rules.
|
||||
</Note>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title='"Unknown model: minimax/MiniMax-M3"'>
|
||||
This usually means the **MiniMax provider is not configured** (no matching provider entry and no MiniMax auth profile/env key found). Fix by:
|
||||
|
||||
- Running `openclaw configure` and selecting a **MiniMax** auth option, or
|
||||
- Adding the matching `models.providers.minimax` or `models.providers.minimax-portal` block manually, or
|
||||
- Setting `MINIMAX_API_KEY`, `MINIMAX_OAUTH_TOKEN`, or a MiniMax auth profile so the matching provider can be injected.
|
||||
|
||||
Make sure the model id is **case-sensitive**:
|
||||
|
||||
- API-key path: `minimax/MiniMax-M3`, `minimax/MiniMax-M2.7`, or `minimax/MiniMax-M2.7-highspeed`
|
||||
- OAuth path: `minimax-portal/MiniMax-M3`, `minimax-portal/MiniMax-M2.7`, or `minimax-portal/MiniMax-M2.7-highspeed`
|
||||
|
||||
Then recheck with:
|
||||
|
||||
```bash
|
||||
openclaw models list
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Note>
|
||||
More help: [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq).
|
||||
</Note>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Image generation" href="/tools/image-generation" icon="image">
|
||||
Shared image tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Music generation" href="/tools/music-generation" icon="music">
|
||||
Shared music tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="MiniMax Search" href="/tools/minimax-search" icon="magnifying-glass">
|
||||
Web search configuration via MiniMax Token Plan.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
General troubleshooting and FAQ.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
220
docs/providers/mistral.md
Normal file
220
docs/providers/mistral.md
Normal file
@@ -0,0 +1,220 @@
|
||||
---
|
||||
summary: "Use Mistral models and Voxtral transcription with OpenClaw"
|
||||
read_when:
|
||||
- You want to use Mistral models in OpenClaw
|
||||
- You want Voxtral realtime transcription for Voice Call
|
||||
- You need Mistral API key onboarding and model refs
|
||||
title: "Mistral"
|
||||
---
|
||||
|
||||
The bundled `mistral` plugin registers four contracts: chat completions, media understanding (Voxtral batch transcription), realtime STT for Voice Call (Voxtral Realtime), and memory embeddings (`mistral-embed`).
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | ------------------------------------------- |
|
||||
| Provider id | `mistral` |
|
||||
| Plugin | bundled, enabled by default |
|
||||
| Auth env var | `MISTRAL_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice mistral-api-key` |
|
||||
| Direct CLI flag | `--mistral-api-key <key>` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://api.mistral.ai/v1` |
|
||||
| Default model | `mistral/mistral-large-latest` |
|
||||
| Embedding model | `mistral-embed` |
|
||||
| Voxtral batch | `voxtral-mini-latest` (audio transcription) |
|
||||
| Voxtral realtime | `voxtral-mini-transcribe-realtime-2602` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get your API key">
|
||||
Create an API key in the [Mistral Console](https://console.mistral.ai/).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice mistral-api-key
|
||||
```
|
||||
|
||||
Or pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --mistral-api-key "$MISTRAL_API_KEY"
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
env: { MISTRAL_API_KEY: "sk-..." },
|
||||
agents: { defaults: { model: { primary: "mistral/mistral-large-latest" } } },
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider mistral
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Built-in LLM catalog
|
||||
|
||||
| Model ref | Input | Context | Max output | Notes |
|
||||
| -------------------------------- | ----------- | ------- | ---------- | ---------------------------------------------------------------- |
|
||||
| `mistral/mistral-large-latest` | text, image | 262,144 | 16,384 | Default model |
|
||||
| `mistral/mistral-medium-2508` | text, image | 262,144 | 8,192 | Mistral Medium 3.1 |
|
||||
| `mistral/mistral-medium-3-5` | text, image | 262,144 | 8,192 | Mistral Medium 3.5; adjustable reasoning |
|
||||
| `mistral/mistral-small-latest` | text, image | 128,000 | 16,384 | Mistral Small 4; adjustable reasoning via API `reasoning_effort` |
|
||||
| `mistral/pixtral-large-latest` | text, image | 128,000 | 32,768 | Pixtral |
|
||||
| `mistral/codestral-latest` | text | 256,000 | 4,096 | Coding |
|
||||
| `mistral/devstral-medium-latest` | text | 262,144 | 32,768 | Devstral 2 |
|
||||
| `mistral/magistral-small` | text | 128,000 | 40,000 | Reasoning-enabled |
|
||||
|
||||
Browse the bundled catalog row before changing config:
|
||||
|
||||
```bash
|
||||
openclaw models list --all --provider mistral --plain
|
||||
```
|
||||
|
||||
Smoke-test a model without starting the Gateway:
|
||||
|
||||
```bash
|
||||
openclaw infer model run --local \
|
||||
--model mistral/mistral-medium-3-5 \
|
||||
--prompt "Reply with exactly: mistral-ok" \
|
||||
--json
|
||||
```
|
||||
|
||||
## Audio transcription (Voxtral)
|
||||
|
||||
Use Voxtral for batch audio transcription through the media understanding pipeline:
|
||||
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
enabled: true,
|
||||
models: [{ provider: "mistral", model: "voxtral-mini-latest" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Tip>
|
||||
The media transcription path uses `/v1/audio/transcriptions`. The default audio model for Mistral is `voxtral-mini-latest`.
|
||||
</Tip>
|
||||
|
||||
## Voice Call streaming STT
|
||||
|
||||
The bundled `mistral` plugin registers Voxtral Realtime as a Voice Call streaming STT provider.
|
||||
|
||||
| Setting | Config path | Default |
|
||||
| ------------ | ---------------------------------------------------------------------- | --------------------------------------- |
|
||||
| API key | `plugins.entries.voice-call.config.streaming.providers.mistral.apiKey` | Falls back to `MISTRAL_API_KEY` |
|
||||
| Model | `...mistral.model` | `voxtral-mini-transcribe-realtime-2602` |
|
||||
| Encoding | `...mistral.encoding` | `pcm_mulaw` |
|
||||
| Sample rate | `...mistral.sampleRate` | `8000` |
|
||||
| Target delay | `...mistral.targetStreamingDelayMs` | `800` |
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"voice-call": {
|
||||
config: {
|
||||
streaming: {
|
||||
enabled: true,
|
||||
provider: "mistral",
|
||||
providers: {
|
||||
mistral: {
|
||||
apiKey: "${MISTRAL_API_KEY}",
|
||||
targetStreamingDelayMs: 800,
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
OpenClaw defaults Mistral realtime STT to `pcm_mulaw` at 8 kHz so Voice Call can forward Twilio media frames directly. Use `encoding: "pcm_s16le"` and a matching `sampleRate` only if your upstream stream is already raw PCM.
|
||||
</Note>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Adjustable reasoning">
|
||||
`mistral/mistral-small-latest` and `mistral/mistral-medium-3-5` support [adjustable reasoning](https://docs.mistral.ai/studio-api/conversations/reasoning/adjustable) on the Chat Completions API via `reasoning_effort` (`none` minimizes extra thinking in the output; `high` surfaces full thinking traces before the final answer).
|
||||
|
||||
OpenClaw maps the session **thinking** level to Mistral's API:
|
||||
|
||||
| OpenClaw thinking level | Mistral `reasoning_effort` |
|
||||
| ----------------------------------------------------------------------- | --------------------------- |
|
||||
| **off** / **minimal** | `none` |
|
||||
| **low** / **medium** / **high** / **xhigh** / **adaptive** / **max** | `high` |
|
||||
|
||||
<Warning>
|
||||
Avoid combining Medium 3.5 reasoning mode with `temperature: 0`; the Mistral HTTP API has been reported to reject `reasoning_effort="high"` plus `temperature: 0` with a 400 response. Leave temperature unset, or turn thinking off/minimal so OpenClaw sends `reasoning_effort: "none"` before you set a low temperature.
|
||||
</Warning>
|
||||
|
||||
Example model-scoped config for Medium 3.5 reasoning:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "mistral/mistral-medium-3-5" },
|
||||
models: {
|
||||
"mistral/mistral-medium-3-5": {
|
||||
params: { thinking: "high" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
Other bundled Mistral catalog models do not use this parameter. Keep using `magistral-*` models when you want Mistral's native reasoning-first behavior.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Memory embeddings">
|
||||
Mistral can serve memory embeddings via `/v1/embeddings` (default model: `mistral-embed`):
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
memorySearch: { provider: "mistral" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Auth and base URL">
|
||||
- Mistral auth uses `MISTRAL_API_KEY` (Bearer header).
|
||||
- Provider base URL defaults to `https://api.mistral.ai/v1` and accepts the standard OpenAI-compatible chat-completions request shape.
|
||||
- Onboarding default model is `mistral/mistral-large-latest`.
|
||||
- Override the base URL under `models.providers.mistral.baseUrl` only when Mistral explicitly publishes a regional endpoint you need.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Media understanding" href="/nodes/media-understanding" icon="microphone">
|
||||
Audio transcription setup and provider selection.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
66
docs/providers/models.md
Normal file
66
docs/providers/models.md
Normal file
@@ -0,0 +1,66 @@
|
||||
---
|
||||
summary: "Model providers (LLMs) supported by OpenClaw"
|
||||
read_when:
|
||||
- You want to choose a model provider
|
||||
- You want quick setup examples for LLM auth + model selection
|
||||
title: "Model provider quickstart"
|
||||
---
|
||||
|
||||
Pick a provider, authenticate, then set the default model as `provider/model`.
|
||||
|
||||
## Quick start (two steps)
|
||||
|
||||
1. Authenticate with the provider (usually via `openclaw onboard`).
|
||||
2. Set the default model:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: { defaults: { model: { primary: "anthropic/claude-opus-4-6" } } },
|
||||
}
|
||||
```
|
||||
|
||||
## Supported providers (starter set)
|
||||
|
||||
- [Alibaba Model Studio](/providers/alibaba)
|
||||
- [Amazon Bedrock](/providers/bedrock)
|
||||
- [Anthropic (API + Claude CLI)](/providers/anthropic)
|
||||
- [BytePlus (International)](/concepts/model-providers#byteplus-international)
|
||||
- [Chutes](/providers/chutes)
|
||||
- [Cloudflare AI Gateway](/providers/cloudflare-ai-gateway)
|
||||
- [Cohere](/providers/cohere)
|
||||
- [ComfyUI](/providers/comfy)
|
||||
- [DeepInfra](/providers/deepinfra)
|
||||
- [fal](/providers/fal)
|
||||
- [Fireworks](/providers/fireworks)
|
||||
- [MiniMax](/providers/minimax)
|
||||
- [Mistral](/providers/mistral)
|
||||
- [Moonshot AI (Kimi + Kimi Coding)](/providers/moonshot)
|
||||
- [NovitaAI](/providers/novita)
|
||||
- [OpenAI (API + Codex)](/providers/openai)
|
||||
- [OpenCode (Zen + Go)](/providers/opencode)
|
||||
- [OpenRouter](/providers/openrouter)
|
||||
- [Qianfan](/providers/qianfan)
|
||||
- [Qwen](/providers/qwen)
|
||||
- [Runway](/providers/runway)
|
||||
- [StepFun](/providers/stepfun)
|
||||
- [Synthetic](/providers/synthetic)
|
||||
- [Venice (Venice AI)](/providers/venice)
|
||||
- [Vercel AI Gateway](/providers/vercel-ai-gateway)
|
||||
- [xAI](/providers/xai)
|
||||
- [Z.AI (GLM)](/providers/zai)
|
||||
|
||||
For the full provider catalog and advanced configuration, see
|
||||
[Provider directory](/providers/index) and [Model providers](/concepts/model-providers).
|
||||
|
||||
## Additional provider variants
|
||||
|
||||
- `anthropic-vertex` - install `@openclaw/anthropic-vertex-provider` for implicit Anthropic on Google Vertex support when Vertex credentials are available; no separate onboarding auth choice
|
||||
- `copilot-proxy` - local VS Code Copilot Proxy bridge; use `openclaw onboard --auth-choice copilot-proxy`
|
||||
- `google-gemini-cli` - unofficial Gemini CLI OAuth flow; requires a local `gemini` install (`brew install gemini-cli` or `npm install -g @google/gemini-cli`); default model `google-gemini-cli/gemini-3-flash-preview`; use `openclaw onboard --auth-choice google-gemini-cli` or `openclaw models auth login --provider google-gemini-cli --set-default`
|
||||
|
||||
## Related
|
||||
|
||||
- [Provider directory](/providers/index)
|
||||
- [Model selection](/concepts/model-providers)
|
||||
- [Model failover](/concepts/model-failover)
|
||||
- [Models CLI](/cli/models)
|
||||
435
docs/providers/moonshot.md
Normal file
435
docs/providers/moonshot.md
Normal file
@@ -0,0 +1,435 @@
|
||||
---
|
||||
summary: "Configure Moonshot K2 vs Kimi Coding (separate providers + keys)"
|
||||
read_when:
|
||||
- You want Moonshot K2 (Moonshot Open Platform) vs Kimi Coding setup
|
||||
- You need to understand separate endpoints, keys, and model refs
|
||||
- You want copy/paste config for either provider
|
||||
title: "Moonshot AI"
|
||||
---
|
||||
|
||||
Moonshot provides the Kimi API with OpenAI-compatible endpoints. Set the
|
||||
default model to `moonshot/kimi-k2.6` for the Moonshot Open Platform, or
|
||||
`kimi/kimi-for-coding` for Kimi Coding.
|
||||
|
||||
<Warning>
|
||||
Moonshot and Kimi Coding are **separate providers**, each shipped as a separate external plugin. Keys are not interchangeable, endpoints differ, and model refs differ (`moonshot/...` vs `kimi/...`).
|
||||
</Warning>
|
||||
|
||||
## Built-in model catalog
|
||||
|
||||
[//]: # "moonshot-kimi-k2-ids:start"
|
||||
|
||||
| Model ref | Name | Reasoning | Input | Context | Max output |
|
||||
| --------------------------------- | ---------------------- | --------- | ----------- | ------- | ---------- |
|
||||
| `moonshot/kimi-k2.6` | Kimi K2.6 | No | text, image | 262,144 | 262,144 |
|
||||
| `moonshot/kimi-k2.7-code` | Kimi K2.7 Code | Always on | text, image | 262,144 | 262,144 |
|
||||
| `moonshot/kimi-k2.5` | Kimi K2.5 | No | text, image | 262,144 | 262,144 |
|
||||
| `moonshot/kimi-k2-thinking` | Kimi K2 Thinking | Yes | text | 262,144 | 262,144 |
|
||||
| `moonshot/kimi-k2-thinking-turbo` | Kimi K2 Thinking Turbo | Yes | text | 262,144 | 262,144 |
|
||||
| `moonshot/kimi-k2-turbo` | Kimi K2 Turbo | No | text | 256,000 | 16,384 |
|
||||
|
||||
[//]: # "moonshot-kimi-k2-ids:end"
|
||||
|
||||
Catalog cost estimates use Moonshot's published pay-as-you-go rates: Kimi
|
||||
K2.7 Code is $0.19/MTok cache hit, $0.95/MTok input, $4.00/MTok output; Kimi
|
||||
K2.6 is $0.16/MTok cache hit, $0.95/MTok input, $4.00/MTok output; Kimi K2.5
|
||||
is $0.10/MTok cache hit, $0.60/MTok input, $3.00/MTok output. Other catalog
|
||||
entries keep zero-cost placeholders unless you override them in config.
|
||||
|
||||
Kimi K2.7 Code always uses native thinking. OpenClaw exposes only the `on`
|
||||
thinking state for this model and omits outbound `thinking` and
|
||||
`reasoning_effort` fields, as required by Moonshot. It also omits sampling
|
||||
overrides (`temperature`, `top_p`, `n`, `presence_penalty`,
|
||||
`frequency_penalty`), which K2.7 fixes to provider defaults. Kimi K2.6 remains
|
||||
the onboarding default.
|
||||
|
||||
## Getting started
|
||||
|
||||
Both Moonshot and Kimi Coding are external plugins - install one before
|
||||
onboarding.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Moonshot API">
|
||||
**Best for:** Kimi K2 models via the Moonshot Open Platform.
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the plugin">
|
||||
```bash
|
||||
openclaw plugins install @openclaw/moonshot-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
</Step>
|
||||
<Step title="Choose your endpoint region">
|
||||
| Auth choice | Endpoint | Region |
|
||||
| ---------------------- | ------------------------------ | ------------- |
|
||||
| `moonshot-api-key` | `https://api.moonshot.ai/v1` | International |
|
||||
| `moonshot-api-key-cn` | `https://api.moonshot.cn/v1` | China |
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice moonshot-api-key
|
||||
```
|
||||
|
||||
Or for the China endpoint:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice moonshot-api-key-cn
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "moonshot/kimi-k2.6" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider moonshot
|
||||
```
|
||||
</Step>
|
||||
<Step title="Run a live smoke test">
|
||||
Use an isolated state dir when you want to verify model access and cost
|
||||
tracking without touching your normal sessions:
|
||||
|
||||
```bash
|
||||
OPENCLAW_CONFIG_PATH=/tmp/openclaw-kimi/openclaw.json \
|
||||
OPENCLAW_STATE_DIR=/tmp/openclaw-kimi \
|
||||
openclaw agent --local \
|
||||
--session-id live-kimi-cost \
|
||||
--message 'Reply exactly: KIMI_LIVE_OK' \
|
||||
--thinking off \
|
||||
--json
|
||||
```
|
||||
|
||||
The JSON response should report `provider: "moonshot"` and
|
||||
`model: "kimi-k2.6"`. The assistant transcript entry stores normalized
|
||||
token usage plus estimated cost under `usage.cost` when Moonshot returns
|
||||
usage metadata.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
### Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { MOONSHOT_API_KEY: "sk-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "moonshot/kimi-k2.6" },
|
||||
models: {
|
||||
// moonshot-kimi-k2-aliases:start
|
||||
"moonshot/kimi-k2.6": { alias: "Kimi K2.6" },
|
||||
"moonshot/kimi-k2.7-code": { alias: "Kimi K2.7 Code" },
|
||||
"moonshot/kimi-k2.5": { alias: "Kimi K2.5" },
|
||||
"moonshot/kimi-k2-thinking": { alias: "Kimi K2 Thinking" },
|
||||
"moonshot/kimi-k2-thinking-turbo": { alias: "Kimi K2 Thinking Turbo" },
|
||||
"moonshot/kimi-k2-turbo": { alias: "Kimi K2 Turbo" },
|
||||
// moonshot-kimi-k2-aliases:end
|
||||
},
|
||||
},
|
||||
},
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
moonshot: {
|
||||
baseUrl: "https://api.moonshot.ai/v1",
|
||||
apiKey: "${MOONSHOT_API_KEY}",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
// moonshot-kimi-k2-models:start
|
||||
{
|
||||
id: "kimi-k2.6",
|
||||
name: "Kimi K2.6",
|
||||
reasoning: false,
|
||||
input: ["text", "image"],
|
||||
cost: { input: 0.95, output: 4, cacheRead: 0.16, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 262144,
|
||||
},
|
||||
{
|
||||
id: "kimi-k2.7-code",
|
||||
name: "Kimi K2.7 Code",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
cost: { input: 0.95, output: 4, cacheRead: 0.19, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 262144,
|
||||
},
|
||||
{
|
||||
id: "kimi-k2.5",
|
||||
name: "Kimi K2.5",
|
||||
reasoning: false,
|
||||
input: ["text", "image"],
|
||||
cost: { input: 0.6, output: 3, cacheRead: 0.1, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 262144,
|
||||
},
|
||||
{
|
||||
id: "kimi-k2-thinking",
|
||||
name: "Kimi K2 Thinking",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 262144,
|
||||
},
|
||||
{
|
||||
id: "kimi-k2-thinking-turbo",
|
||||
name: "Kimi K2 Thinking Turbo",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 262144,
|
||||
},
|
||||
{
|
||||
id: "kimi-k2-turbo",
|
||||
name: "Kimi K2 Turbo",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 256000,
|
||||
maxTokens: 16384,
|
||||
},
|
||||
// moonshot-kimi-k2-models:end
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Kimi Coding">
|
||||
**Best for:** code-focused tasks via the Kimi Coding endpoint.
|
||||
|
||||
<Note>
|
||||
Kimi Coding uses a different API key and provider prefix (`kimi/...`) than Moonshot (`moonshot/...`). The stable model ref is `kimi/kimi-for-coding`; legacy refs `kimi/kimi-code` and `kimi/k2p5` remain accepted and normalize to that model id.
|
||||
</Note>
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the plugin">
|
||||
```bash
|
||||
openclaw plugins install @openclaw/kimi-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice kimi-code-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "kimi/kimi-for-coding" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider kimi
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
### Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { KIMI_API_KEY: "sk-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "kimi/kimi-for-coding" },
|
||||
models: {
|
||||
"kimi/kimi-for-coding": { alias: "Kimi" },
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Kimi web search
|
||||
|
||||
The Moonshot plugin also registers **Kimi** as a `web_search` provider, backed by Moonshot web search.
|
||||
|
||||
<Steps>
|
||||
<Step title="Run interactive web search setup">
|
||||
```bash
|
||||
openclaw configure --section web
|
||||
```
|
||||
|
||||
Choose **Kimi** in the web-search section to store
|
||||
`plugins.entries.moonshot.config.webSearch.*`.
|
||||
|
||||
</Step>
|
||||
<Step title="Configure the web search region and model">
|
||||
Interactive setup prompts for:
|
||||
|
||||
| Setting | Options |
|
||||
| ------------------- | -------------------------------------------------------------------- |
|
||||
| API region | `https://api.moonshot.ai/v1` (international) or `https://api.moonshot.cn/v1` (China) |
|
||||
| Web search model | Defaults to `kimi-k2.6` |
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
Config lives under `plugins.entries.moonshot.config.webSearch`:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
moonshot: {
|
||||
config: {
|
||||
webSearch: {
|
||||
apiKey: "sk-...", // or use KIMI_API_KEY / MOONSHOT_API_KEY
|
||||
baseUrl: "https://api.moonshot.ai/v1",
|
||||
model: "kimi-k2.6",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
tools: {
|
||||
web: {
|
||||
search: {
|
||||
provider: "kimi",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Native thinking mode">
|
||||
Kimi K2.7 Code always uses native thinking. Moonshot requires clients to
|
||||
omit the `thinking` field for this model, so OpenClaw exposes only `on` and
|
||||
ignores stale `off` settings. K2.7 also fixes `temperature`, `top_p`, `n`,
|
||||
`presence_penalty`, and `frequency_penalty`; OpenClaw omits configured
|
||||
overrides for those fields.
|
||||
|
||||
Other Moonshot Kimi models support binary native thinking:
|
||||
|
||||
- `thinking: { type: "enabled" }`
|
||||
- `thinking: { type: "disabled" }`
|
||||
|
||||
Configure it per model via `agents.defaults.models.<provider/model>.params`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"moonshot/kimi-k2.6": {
|
||||
params: {
|
||||
thinking: { type: "disabled" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw maps runtime `/think` levels for those models:
|
||||
|
||||
| `/think` level | Moonshot behavior |
|
||||
| -------------------- | -------------------------- |
|
||||
| `/think off` | `thinking.type=disabled` |
|
||||
| Any non-off level | `thinking.type=enabled` |
|
||||
|
||||
<Warning>
|
||||
When Moonshot thinking is enabled, `tool_choice` must be `auto` or `none`. A pinned tool choice (`type: "tool"` or `type: "function"`) forces thinking back to `disabled` instead, so the requested tool still runs; `tool_choice: "required"` is normalized to `auto` instead. This applies to every Moonshot model except Kimi K2.7 Code, whose thinking mode cannot be disabled - its `tool_choice` is normalized to `auto` when incompatible.
|
||||
</Warning>
|
||||
|
||||
Kimi K2.6 also accepts an optional `thinking.keep` field that controls
|
||||
multi-turn retention of `reasoning_content`. Set it to `"all"` to keep full
|
||||
reasoning across turns; omit it (or leave it `null`) to use the server
|
||||
default strategy. OpenClaw only forwards `thinking.keep` for
|
||||
`moonshot/kimi-k2.6` and strips it from other models. Kimi K2.7 Code
|
||||
preserves full reasoning history by default while OpenClaw omits the entire
|
||||
`thinking` field.
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"moonshot/kimi-k2.6": {
|
||||
params: {
|
||||
thinking: { type: "enabled", keep: "all" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tool call id sanitization">
|
||||
Moonshot Kimi serves native tool_call ids shaped like `functions.<name>:<index>`. OpenClaw preserves the first occurrence of each native Kimi id and rewrites later duplicates to deterministic OpenAI-style `call_*` ids. Matching tool results are remapped with the same id so replay remains unique without stripping Kimi's first native id. This behavior is wired into the bundled Moonshot provider and is not a user-configurable setting.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Streaming usage compatibility">
|
||||
Native Moonshot endpoints (`https://api.moonshot.ai/v1` and
|
||||
`https://api.moonshot.cn/v1`) advertise streaming usage compatibility.
|
||||
OpenClaw keys this off the endpoint host, not the provider id, so a custom
|
||||
provider id pointed at the same native Moonshot host inherits the same
|
||||
streaming-usage behavior.
|
||||
|
||||
With the catalog K2.6 pricing, streamed usage that includes input, output,
|
||||
and cache-read tokens is also converted into local estimated USD cost for
|
||||
`/status`, `/usage full`, `/usage cost`, and transcript-backed session
|
||||
accounting.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Endpoint and model ref reference">
|
||||
| Provider | Model ref prefix | Endpoint | Auth env var |
|
||||
| ---------- | ---------------- | ------------------------------ | ------------------- |
|
||||
| Moonshot | `moonshot/` | `https://api.moonshot.ai/v1` | `MOONSHOT_API_KEY` |
|
||||
| Moonshot CN| `moonshot/` | `https://api.moonshot.cn/v1` | `MOONSHOT_API_KEY` |
|
||||
| Kimi Coding| `kimi/` | Kimi Coding endpoint | `KIMI_API_KEY` |
|
||||
| Web search | N/A | Same as Moonshot API region | `KIMI_API_KEY` or `MOONSHOT_API_KEY` |
|
||||
|
||||
- Kimi web search uses `KIMI_API_KEY` or `MOONSHOT_API_KEY`, and defaults to `https://api.moonshot.ai/v1` with model `kimi-k2.6`.
|
||||
- Override pricing and context metadata in `models.providers` if needed.
|
||||
- If Moonshot publishes different context limits for a model, adjust `contextWindow` accordingly.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Web search" href="/tools/web" icon="magnifying-glass">
|
||||
Configuring web search providers including Kimi.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema for providers, models, and plugins.
|
||||
</Card>
|
||||
<Card title="Moonshot Open Platform" href="https://platform.moonshot.ai" icon="globe">
|
||||
Moonshot API key management and documentation.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
83
docs/providers/novita.md
Normal file
83
docs/providers/novita.md
Normal file
@@ -0,0 +1,83 @@
|
||||
---
|
||||
summary: "Use NovitaAI's OpenAI-compatible API with OpenClaw"
|
||||
read_when:
|
||||
- You want to run OpenClaw with NovitaAI models
|
||||
- You need the Novita provider id, key, or endpoint
|
||||
title: "NovitaAI"
|
||||
---
|
||||
|
||||
NovitaAI is a hosted AI infrastructure provider with an OpenAI-compatible API.
|
||||
It ships as a bundled OpenClaw provider (no separate plugin install), so
|
||||
credentials go through the normal model auth flow and model refs look like
|
||||
`novita/deepseek/deepseek-v3-0324`.
|
||||
|
||||
## Setup
|
||||
|
||||
Create an API key at [novita.ai/settings/key-management](https://novita.ai/settings/key-management), then run:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice novita-api-key
|
||||
```
|
||||
|
||||
Or set:
|
||||
|
||||
```bash
|
||||
export NOVITA_API_KEY="<your-novita-api-key>" # pragma: allowlist secret
|
||||
```
|
||||
|
||||
## Defaults
|
||||
|
||||
| Setting | Value |
|
||||
| ------------- | ---------------------------------- |
|
||||
| Provider id | `novita` |
|
||||
| Aliases | `novita-ai`, `novitaai` |
|
||||
| Base URL | `https://api.novita.ai/openai/v1` |
|
||||
| Env var | `NOVITA_API_KEY` |
|
||||
| Default model | `novita/deepseek/deepseek-v3-0324` |
|
||||
|
||||
## Bundled model catalog
|
||||
|
||||
- `novita/moonshotai/kimi-k2.5`
|
||||
- `novita/minimax/minimax-m2.7`
|
||||
- `novita/zai-org/glm-5`
|
||||
- `novita/deepseek/deepseek-v3-0324`
|
||||
- `novita/deepseek/deepseek-r1-0528`
|
||||
- `novita/qwen/qwen3-235b-a22b-fp8`
|
||||
|
||||
This is a starting point, not a live catalog. Your account, region, or
|
||||
Novita's current offering may add, remove, or restrict routes. Check before
|
||||
setting a long-lived default:
|
||||
|
||||
```bash
|
||||
openclaw models list --provider novita
|
||||
```
|
||||
|
||||
## When to choose Novita
|
||||
|
||||
- Hosted open-weight model access with an OpenAI-compatible API.
|
||||
- DeepSeek, Kimi, MiniMax, GLM, or Qwen-family routes through a single provider
|
||||
account.
|
||||
- Another hosted fallback path beside DeepInfra, GMI, OpenRouter, or direct
|
||||
vendor APIs.
|
||||
- Provider-side model hosting instead of maintaining LM Studio, Ollama,
|
||||
SGLang, or vLLM infrastructure.
|
||||
|
||||
Choose a direct vendor provider when you need vendor-native request
|
||||
parameters or support contracts. Choose a local provider when the model must
|
||||
run on your own hardware or network boundary.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- `401`/`403`: verify the key in Novita's key management page and re-run
|
||||
`openclaw onboard --auth-choice novita-api-key` if the stored profile is
|
||||
stale.
|
||||
- Unknown model errors: use the exact `novita/<route-id>` returned by
|
||||
`openclaw models list --provider novita`.
|
||||
- Slow or failed routes: try another Novita model route, or set Novita as a
|
||||
fallback provider for workloads that can tolerate provider-specific
|
||||
variance.
|
||||
|
||||
## Related
|
||||
|
||||
- [Model providers](/concepts/model-providers)
|
||||
- [Provider directory](/providers/index)
|
||||
216
docs/providers/nvidia.md
Normal file
216
docs/providers/nvidia.md
Normal file
@@ -0,0 +1,216 @@
|
||||
---
|
||||
summary: "Use NVIDIA's OpenAI-compatible API in OpenClaw"
|
||||
read_when:
|
||||
- You want to use open models in OpenClaw for free
|
||||
- You need NVIDIA_API_KEY setup
|
||||
- You want to use Nemotron 3 Ultra through NVIDIA
|
||||
title: "NVIDIA"
|
||||
---
|
||||
|
||||
NVIDIA serves open models for free through an OpenAI-compatible API at
|
||||
`https://integrate.api.nvidia.com/v1`, authenticated with an API key from
|
||||
[build.nvidia.com](https://build.nvidia.com/settings/api-keys). OpenClaw
|
||||
defaults the NVIDIA provider to Nemotron 3 Ultra, NVIDIA's 550B total / 55B
|
||||
active reasoning model for long-context agentic work.
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get your API key">
|
||||
Create an API key at [build.nvidia.com](https://build.nvidia.com/settings/api-keys).
|
||||
</Step>
|
||||
<Step title="Export the key and run onboarding">
|
||||
```bash
|
||||
export NVIDIA_API_KEY="nvapi-..."
|
||||
openclaw onboard --auth-choice nvidia-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set an NVIDIA model">
|
||||
```bash
|
||||
openclaw models set nvidia/nvidia/nemotron-3-ultra-550b-a55b
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
For non-interactive setup, pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice nvidia-api-key --nvidia-api-key "nvapi-..."
|
||||
```
|
||||
|
||||
<Warning>
|
||||
`--nvidia-api-key` lands the key in shell history and `ps` output. Prefer the
|
||||
`NVIDIA_API_KEY` environment variable when possible.
|
||||
</Warning>
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { NVIDIA_API_KEY: "nvapi-..." },
|
||||
models: {
|
||||
providers: {
|
||||
nvidia: {
|
||||
baseUrl: "https://integrate.api.nvidia.com/v1",
|
||||
api: "openai-completions",
|
||||
},
|
||||
},
|
||||
},
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "nvidia/nvidia/nemotron-3-ultra-550b-a55b" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Featured catalog
|
||||
|
||||
When an NVIDIA API key is configured, setup and model-selection paths fetch
|
||||
NVIDIA's public featured-model catalog from
|
||||
`https://assets.ngc.nvidia.com/products/api-catalog/featured-models.json` and
|
||||
cache the result for 24 hours (first 32 entries, imported as free text-input
|
||||
rows). New featured models from build.nvidia.com therefore appear in setup and
|
||||
model-selection surfaces without waiting for an OpenClaw release. When the
|
||||
live feed is available, the first returned model is the preselected option
|
||||
during NVIDIA setup.
|
||||
|
||||
The fetch uses a fixed HTTPS host policy for `assets.ngc.nvidia.com`. If no
|
||||
NVIDIA API key is configured, or if the feed is unavailable or malformed,
|
||||
OpenClaw falls back to the bundled catalog and bundled default below.
|
||||
|
||||
## Nemotron 3 Ultra
|
||||
|
||||
Nemotron 3 Ultra is the default NVIDIA model in OpenClaw. NVIDIA's build page for
|
||||
[`nvidia/nemotron-3-ultra-550b-a55b`](https://build.nvidia.com/nvidia/nemotron-3-ultra-550b-a55b)
|
||||
lists it as an available free endpoint with a 1M-token context specification.
|
||||
The bundled catalog records a 16,384-token max output to match NVIDIA's current
|
||||
OpenAI-compatible sample request for the hosted endpoint.
|
||||
|
||||
The bundled Ultra row sends
|
||||
`chat_template_kwargs: { enable_thinking: false, force_nonempty_content: true }`
|
||||
by default so normal chat output stays in the visible answer instead of
|
||||
exposing reasoning text.
|
||||
|
||||
Use Ultra for the highest-capability NVIDIA default. Keep Super selected when
|
||||
you want the smaller Nemotron 3 option, or choose one of the third-party models
|
||||
hosted in NVIDIA's catalog when their context, latency, or behavior fits better.
|
||||
|
||||
## Bundled fallback catalog
|
||||
|
||||
| Model ref | Name | Context | Max output | Notes |
|
||||
| ------------------------------------------ | ---------------------------- | --------- | ---------- | ---------------------------------------- |
|
||||
| `nvidia/nvidia/nemotron-3-ultra-550b-a55b` | NVIDIA Nemotron 3 Ultra 550B | 1,000,000 | 16,384 | Default |
|
||||
| `nvidia/nvidia/nemotron-3-super-120b-a12b` | NVIDIA Nemotron 3 Super 120B | 1,048,576 | 8,192 | |
|
||||
| `nvidia/moonshotai/kimi-k2.5` | Kimi K2.5 | 262,144 | 8,192 | |
|
||||
| `nvidia/minimaxai/minimax-m2.7` | Minimax M2.7 | 196,608 | 8,192 | |
|
||||
| `nvidia/z-ai/glm-5.1` | GLM 5.1 | 202,752 | 8,192 | |
|
||||
| `nvidia/minimaxai/minimax-m2.5` | MiniMax M2.5 | 196,608 | 8,192 | Deprecated; use `minimaxai/minimax-m2.7` |
|
||||
| `nvidia/z-ai/glm5` | GLM-5 | 202,752 | 8,192 | Deprecated; use `z-ai/glm-5.1` |
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Auto-enable behavior">
|
||||
The provider auto-enables when the `NVIDIA_API_KEY` environment variable is
|
||||
set or a key was stored during onboarding. No explicit provider config is
|
||||
required beyond the key.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Catalog and pricing">
|
||||
OpenClaw prefers NVIDIA's public featured-model catalog when NVIDIA auth is
|
||||
configured and caches it for 24 hours. The bundled fallback catalog is static
|
||||
and keeps deprecated shipped refs for upgrade compatibility. Costs default
|
||||
to `0` in source since NVIDIA currently offers free API access for the
|
||||
listed models.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="OpenAI-compatible endpoint">
|
||||
OpenClaw talks to NVIDIA with the `openai-completions` adapter against the
|
||||
standard `/v1` chat completions route. Any OpenAI-compatible tooling should
|
||||
work out of the box with the NVIDIA base URL.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Nemotron 3 Ultra reasoning params">
|
||||
NVIDIA's Ultra sample request uses `chat_template_kwargs.enable_thinking`
|
||||
and `reasoning_budget` for reasoning output. OpenClaw's bundled Ultra row
|
||||
disables template thinking by default for normal chat use. If you need to
|
||||
opt into NVIDIA reasoning output or force other NVIDIA-specific request
|
||||
fields, set per-model params and keep provider-specific overrides scoped to
|
||||
the NVIDIA model:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"nvidia/nvidia/nemotron-3-ultra-550b-a55b": {
|
||||
params: {
|
||||
chat_template_kwargs: { enable_thinking: true },
|
||||
extra_body: { reasoning_budget: 16384 },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`params.chat_template_kwargs` merges into any `chat_template_kwargs`
|
||||
already on the request instead of replacing the whole object.
|
||||
`params.extra_body` is the final OpenAI-compatible request-body override
|
||||
and overwrites colliding payload keys, so use it only for fields NVIDIA
|
||||
documents for the selected endpoint.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Slow custom provider responses">
|
||||
Some NVIDIA-hosted custom models can take longer than the default ~120s
|
||||
model idle watchdog before they emit a first response chunk. For custom
|
||||
NVIDIA provider entries, raise the provider timeout instead of the whole
|
||||
agent runtime timeout; `timeoutSeconds` covers provider HTTP requests and
|
||||
raises the idle/stream watchdog ceiling for that provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
"custom-integrate-api-nvidia-com": {
|
||||
baseUrl: "https://integrate.api.nvidia.com/v1",
|
||||
api: "openai-completions",
|
||||
apiKey: "NVIDIA_API_KEY",
|
||||
timeoutSeconds: 300,
|
||||
},
|
||||
},
|
||||
},
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"custom-integrate-api-nvidia-com/meta/llama-3.1-70b-instruct": {
|
||||
params: { thinking: "off" },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Tip>
|
||||
NVIDIA models are currently free to use. Check
|
||||
[build.nvidia.com](https://build.nvidia.com/) for the latest availability and
|
||||
rate-limit details.
|
||||
</Tip>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config reference for agents, models, and providers.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
117
docs/providers/ollama-cloud.md
Normal file
117
docs/providers/ollama-cloud.md
Normal file
@@ -0,0 +1,117 @@
|
||||
---
|
||||
summary: "Use Ollama Cloud directly with OpenClaw"
|
||||
read_when:
|
||||
- You want to use hosted Ollama models without a local Ollama server
|
||||
- You need the ollama-cloud provider id, key, or endpoint
|
||||
title: "Ollama Cloud"
|
||||
---
|
||||
|
||||
Ollama Cloud is Ollama's hosted model API. The `ollama-cloud` provider calls it
|
||||
directly at `https://ollama.com` over Ollama's native `/api/chat` API, with no
|
||||
local Ollama server and no local Ollama app signed into cloud mode. Use model
|
||||
refs like `ollama-cloud/kimi-k2.6`.
|
||||
|
||||
OpenClaw registers `ollama-cloud` as its own provider id so cloud-only
|
||||
credentials, live catalog discovery, and model selection do not get mixed with
|
||||
a local `ollama` host. For local Ollama, hybrid cloud-plus-local routing,
|
||||
embeddings, and custom host details, see [Ollama](/providers/ollama).
|
||||
|
||||
## Setup
|
||||
|
||||
Create an Ollama Cloud API key at [ollama.com/settings/keys](https://ollama.com/settings/keys), then run:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice ollama-cloud
|
||||
```
|
||||
|
||||
Or set:
|
||||
|
||||
```bash
|
||||
export OLLAMA_API_KEY="<your-ollama-cloud-api-key>" # pragma: allowlist secret
|
||||
```
|
||||
|
||||
Non-interactive onboarding accepts the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice ollama-cloud --ollama-cloud-api-key "<key>"
|
||||
```
|
||||
|
||||
Onboarding sets the default model to `ollama-cloud/kimi-k2.5:cloud`.
|
||||
|
||||
## Defaults
|
||||
|
||||
- Provider: `ollama-cloud`
|
||||
- Base URL: `https://ollama.com`
|
||||
- Env var: `OLLAMA_API_KEY`
|
||||
- API style: Ollama native `/api/chat`
|
||||
- Onboarding default model: `ollama-cloud/kimi-k2.5:cloud`
|
||||
|
||||
## When to choose Ollama Cloud
|
||||
|
||||
- You want hosted Ollama models without running `ollama serve` locally.
|
||||
- You want the same native Ollama chat API shape OpenClaw uses for local
|
||||
Ollama, but pointed at `https://ollama.com`.
|
||||
- You want a simple cloud path for models that are already in Ollama's hosted
|
||||
catalog.
|
||||
- You do not need local model pulls, local GPU control, or LAN-only inference.
|
||||
|
||||
Use [Ollama](/providers/ollama) instead when you want local-only or
|
||||
cloud-plus-local routing through a signed-in Ollama host. Use an
|
||||
OpenAI-compatible provider instead when you need `/v1/chat/completions`
|
||||
semantics or provider-specific OpenAI-style features.
|
||||
|
||||
## Models
|
||||
|
||||
The provider requires an API key; without one it stays inactive. With a key,
|
||||
OpenClaw discovers Ollama Cloud models live from the hosted catalog:
|
||||
|
||||
```bash
|
||||
openclaw models list --provider ollama-cloud
|
||||
openclaw models set ollama-cloud/kimi-k2.6
|
||||
```
|
||||
|
||||
Hosted ids in the live catalog include `deepseek-v4-flash`, `glm-5`,
|
||||
`gpt-oss:20b`, `kimi-k2.6`, and `minimax-m2.7`. When live discovery returns
|
||||
nothing, OpenClaw falls back to the bundled rows `kimi-k2.5:cloud`,
|
||||
`minimax-m2.7:cloud`, `glm-5.1:cloud`, and `glm-5.2:cloud`.
|
||||
|
||||
Model ids are cloud catalog ids, not local pull names. If a model name works in
|
||||
a local Ollama host but is absent from the hosted catalog, use the `ollama`
|
||||
provider with that local host instead.
|
||||
|
||||
## Live test
|
||||
|
||||
For Ollama Cloud API-key smoke tests, point the Ollama live test at the hosted
|
||||
endpoint and choose a model from your current catalog:
|
||||
|
||||
```bash
|
||||
export OLLAMA_API_KEY="<your-ollama-cloud-api-key>" # pragma: allowlist secret
|
||||
|
||||
OPENCLAW_LIVE_TEST=1 \
|
||||
OPENCLAW_LIVE_OLLAMA=1 \
|
||||
OPENCLAW_LIVE_OLLAMA_BASE_URL=https://ollama.com \
|
||||
OPENCLAW_LIVE_OLLAMA_MODEL=kimi-k2.6 \
|
||||
pnpm test:live -- extensions/ollama/ollama.live.test.ts
|
||||
```
|
||||
|
||||
The cloud smoke runs text, native stream, and web search; set
|
||||
`OPENCLAW_LIVE_OLLAMA_WEB_SEARCH=0` to skip web search. It skips embeddings by
|
||||
default for `https://ollama.com` because Ollama Cloud API keys may not
|
||||
authorize `/api/embed`; force them with `OPENCLAW_LIVE_OLLAMA_EMBEDDINGS=1`.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- `Ollama Cloud requires an API key` / `Set OLLAMA_API_KEY` errors: provide a
|
||||
real cloud API key. The local `ollama-local` marker is only for local or
|
||||
private Ollama hosts.
|
||||
- Unknown model errors: run `openclaw models list --provider ollama-cloud` and
|
||||
copy the hosted model id exactly.
|
||||
- Tool-call or raw JSON issues on custom Ollama hosts: check whether you are
|
||||
accidentally using an OpenAI-compatible `/v1` URL. Ollama routes should use
|
||||
the native base URL with no `/v1` suffix.
|
||||
|
||||
## Related
|
||||
|
||||
- [Ollama](/providers/ollama)
|
||||
- [Model providers](/concepts/model-providers)
|
||||
- [All providers](/providers/index)
|
||||
1350
docs/providers/ollama.md
Normal file
1350
docs/providers/ollama.md
Normal file
File diff suppressed because it is too large
Load Diff
1198
docs/providers/openai.md
Normal file
1198
docs/providers/openai.md
Normal file
File diff suppressed because it is too large
Load Diff
129
docs/providers/opencode-go.md
Normal file
129
docs/providers/opencode-go.md
Normal file
@@ -0,0 +1,129 @@
|
||||
---
|
||||
summary: "Use the OpenCode Go catalog with the shared OpenCode setup"
|
||||
read_when:
|
||||
- You want the OpenCode Go catalog
|
||||
- You need the runtime model refs for Go-hosted models
|
||||
title: "OpenCode Go"
|
||||
---
|
||||
|
||||
OpenCode Go is the Go catalog inside [OpenCode](/providers/opencode). It shares
|
||||
the `OPENCODE_API_KEY` credential with the Zen catalog, but keeps its own
|
||||
runtime provider id (`opencode-go`) so upstream per-model routing stays
|
||||
correct.
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | -------------------------------------------------- |
|
||||
| Runtime provider | `opencode-go` |
|
||||
| Auth | `OPENCODE_API_KEY` (alias: `OPENCODE_ZEN_API_KEY`) |
|
||||
| Parent setup | [OpenCode](/providers/opencode) |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Interactive">
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice opencode-go
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a Go model as default">
|
||||
```bash
|
||||
openclaw config set agents.defaults.model.primary "opencode-go/kimi-k2.6"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider opencode-go
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
|
||||
<Tab title="Non-interactive">
|
||||
<Steps>
|
||||
<Step title="Pass the key directly">
|
||||
```bash
|
||||
openclaw onboard --opencode-go-api-key "$OPENCODE_API_KEY"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider opencode-go
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { OPENCODE_API_KEY: "YOUR_API_KEY_HERE" }, // pragma: allowlist secret
|
||||
agents: { defaults: { model: { primary: "opencode-go/kimi-k2.6" } } },
|
||||
}
|
||||
```
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
Run `openclaw models list --provider opencode-go` for the current model list.
|
||||
Bundled rows:
|
||||
|
||||
| Model ref | Name | Context | Max output | Image input |
|
||||
| ------------------------------- | ----------------- | --------- | ---------- | ----------- |
|
||||
| `opencode-go/deepseek-v4-pro` | DeepSeek V4 Pro | 1M | 384K | No |
|
||||
| `opencode-go/deepseek-v4-flash` | DeepSeek V4 Flash | 1M | 384K | No |
|
||||
| `opencode-go/glm-5` | GLM-5 | 202,752 | 32,768 | No |
|
||||
| `opencode-go/glm-5.1` | GLM-5.1 | 202,752 | 32,768 | No |
|
||||
| `opencode-go/glm-5.2` | GLM-5.2 | 1M | 131,072 | No |
|
||||
| `opencode-go/hy3-preview` | HY3 Preview | 262,144 | 32,768 | No |
|
||||
| `opencode-go/kimi-k2.5` | Kimi K2.5 | 262,144 | 65,536 | Yes |
|
||||
| `opencode-go/kimi-k2.6` | Kimi K2.6 | 262,144 | 65,536 | Yes |
|
||||
| `opencode-go/kimi-k2.7-code` | Kimi K2.7 Code | 262,144 | 262,144 | Yes |
|
||||
| `opencode-go/mimo-v2-omni` | MiMo V2 Omni | 262,144 | 32,000 | Yes |
|
||||
| `opencode-go/mimo-v2.5` | MiMo V2.5 | 1M | 128,000 | Yes |
|
||||
| `opencode-go/mimo-v2-pro` | MiMo V2 Pro | 1,048,576 | 32,000 | No |
|
||||
| `opencode-go/mimo-v2.5-pro` | MiMo V2.5 Pro | 1,048,576 | 128,000 | No |
|
||||
| `opencode-go/minimax-m2.5` | MiniMax M2.5 | 204,800 | 65,536 | No |
|
||||
| `opencode-go/minimax-m2.7` | MiniMax M2.7 | 204,800 | 131,072 | No |
|
||||
| `opencode-go/minimax-m3` | MiniMax M3 | 204,800 | 131,072 | No |
|
||||
| `opencode-go/qwen3.5-plus` | Qwen3.5 Plus | 262,144 | 65,536 | Yes |
|
||||
| `opencode-go/qwen3.6-plus` | Qwen3.6 Plus | 262,144 | 65,536 | Yes |
|
||||
| `opencode-go/qwen3.7-max` | Qwen3.7 Max | 1M | 65,536 | No |
|
||||
| `opencode-go/qwen3.7-plus` | Qwen3.7 Plus | 1M | 65,536 | Yes |
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Routing behavior">
|
||||
OpenClaw routes any `opencode-go/...` model ref automatically. No extra
|
||||
provider config is required.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Runtime ref convention">
|
||||
Runtime refs stay explicit: `opencode/...` for Zen, `opencode-go/...` for
|
||||
Go. This keeps upstream per-model routing correct across both catalogs.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Shared credentials">
|
||||
One `OPENCODE_API_KEY` covers both the Zen and Go catalogs. Entering the
|
||||
key during setup stores credentials for both runtime providers.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Tip>
|
||||
See [OpenCode](/providers/opencode) for the shared onboarding overview and the full
|
||||
Zen + Go catalog reference.
|
||||
</Tip>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="OpenCode (parent)" href="/providers/opencode" icon="server">
|
||||
Shared onboarding, catalog overview, and advanced notes.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
156
docs/providers/opencode.md
Normal file
156
docs/providers/opencode.md
Normal file
@@ -0,0 +1,156 @@
|
||||
---
|
||||
summary: "Use OpenCode Zen and Go catalogs with OpenClaw"
|
||||
read_when:
|
||||
- You want OpenCode-hosted model access
|
||||
- You want to pick between the Zen and Go catalogs
|
||||
title: "OpenCode"
|
||||
---
|
||||
|
||||
OpenCode exposes two hosted catalogs in OpenClaw:
|
||||
|
||||
| Catalog | Prefix | Runtime provider |
|
||||
| ------- | ----------------- | ---------------- |
|
||||
| **Zen** | `opencode/...` | `opencode` |
|
||||
| **Go** | `opencode-go/...` | `opencode-go` |
|
||||
|
||||
Both catalogs share one OpenCode API key (`OPENCODE_API_KEY`, alias
|
||||
`OPENCODE_ZEN_API_KEY`). OpenClaw keeps the runtime provider ids split so
|
||||
upstream per-model routing stays correct, but onboarding and docs treat them as
|
||||
one OpenCode setup.
|
||||
|
||||
## Getting started
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Zen catalog">
|
||||
**Best for:** the curated OpenCode multi-model proxy (Claude, GPT, Gemini, GLM,
|
||||
DeepSeek, Kimi, MiniMax, Qwen).
|
||||
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice opencode-zen
|
||||
```
|
||||
|
||||
Or pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --opencode-zen-api-key "$OPENCODE_API_KEY"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a Zen model as the default">
|
||||
```bash
|
||||
openclaw config set agents.defaults.model.primary "opencode/claude-opus-4-6"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider opencode
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Go catalog">
|
||||
**Best for:** the OpenCode-hosted Kimi, GLM, MiniMax, Qwen, and DeepSeek lineup.
|
||||
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice opencode-go
|
||||
```
|
||||
|
||||
Or pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --opencode-go-api-key "$OPENCODE_API_KEY"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a Go model as the default">
|
||||
```bash
|
||||
openclaw config set agents.defaults.model.primary "opencode-go/kimi-k2.6"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider opencode-go
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { OPENCODE_API_KEY: "sk-..." },
|
||||
agents: { defaults: { model: { primary: "opencode/claude-opus-4-6" } } },
|
||||
}
|
||||
```
|
||||
|
||||
## Built-in catalogs
|
||||
|
||||
### Zen
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | --------------------------------------------------------------------------------------------- |
|
||||
| Runtime provider | `opencode` |
|
||||
| Example models | `opencode/claude-opus-4-6`, `opencode/gpt-5.5`, `opencode/gemini-3.1-pro`, `opencode/glm-5.2` |
|
||||
|
||||
Run `openclaw models list --provider opencode` for the full current list, which
|
||||
also includes free-tier rows such as `opencode/big-pickle` and
|
||||
`opencode/deepseek-v4-flash-free`.
|
||||
|
||||
### Go
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | ------------------------------------------------------------------------ |
|
||||
| Runtime provider | `opencode-go` |
|
||||
| Example models | `opencode-go/kimi-k2.6`, `opencode-go/glm-5`, `opencode-go/minimax-m2.5` |
|
||||
|
||||
See [OpenCode Go](/providers/opencode-go) for the full Go model table.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="API key aliases">
|
||||
`OPENCODE_ZEN_API_KEY` is also accepted as an alias for `OPENCODE_API_KEY`.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Shared credentials">
|
||||
Entering one OpenCode key during setup stores credentials for both runtime
|
||||
providers. You do not need to onboard each catalog separately.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Getting an API key">
|
||||
Create an OpenCode account and generate an API key at
|
||||
[opencode.ai/auth](https://opencode.ai/auth). Billing and catalog
|
||||
availability are managed from the OpenCode dashboard.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Gemini replay behavior">
|
||||
Gemini-backed OpenCode refs stay on the proxy-Gemini path, so OpenClaw keeps
|
||||
Gemini thought-signature sanitation there without enabling native Gemini
|
||||
replay validation or bootstrap rewrites.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Non-Gemini replay behavior">
|
||||
Non-Gemini OpenCode refs keep the minimal OpenAI-compatible replay policy.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="OpenCode Go" href="/providers/opencode-go" icon="server">
|
||||
Full Go catalog reference.
|
||||
</Card>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config reference for agents, models, and providers.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
469
docs/providers/openrouter.md
Normal file
469
docs/providers/openrouter.md
Normal file
@@ -0,0 +1,469 @@
|
||||
---
|
||||
summary: "Use OpenRouter's unified API to access many models in OpenClaw"
|
||||
read_when:
|
||||
- You want a single API key for many LLMs
|
||||
- You want to run models via OpenRouter in OpenClaw
|
||||
- You want to use OpenRouter for image generation
|
||||
- You want to use OpenRouter for music generation
|
||||
- You want to use OpenRouter for video generation
|
||||
title: "OpenRouter"
|
||||
---
|
||||
|
||||
OpenRouter routes requests to many models behind one API and one key. It is
|
||||
OpenAI-compatible, so OpenClaw talks to it over the same
|
||||
`openai-completions`-style transport used for other proxy providers.
|
||||
|
||||
## Getting started
|
||||
|
||||
<Tabs>
|
||||
<Tab title="OAuth">
|
||||
<Steps>
|
||||
<Step title="Run OAuth onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice openrouter-oauth
|
||||
```
|
||||
|
||||
OpenClaw opens OpenRouter's browser sign-in flow (PKCE), exchanges the
|
||||
code for an OpenRouter API key, and stores it in the default
|
||||
OpenRouter auth profile. On remote/headless hosts, OpenClaw prints the
|
||||
sign-in URL and asks you to paste the redirect URL after signing in.
|
||||
</Step>
|
||||
<Step title="(Optional) Switch to a specific model">
|
||||
Onboarding defaults to `openrouter/auto`. Pick a concrete model later:
|
||||
|
||||
```bash
|
||||
openclaw models set openrouter/<provider>/<model>
|
||||
```
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
<Tab title="API key">
|
||||
<Steps>
|
||||
<Step title="Get your API key">
|
||||
Create an API key at [openrouter.ai/keys](https://openrouter.ai/keys).
|
||||
</Step>
|
||||
<Step title="Run API-key onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice openrouter-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="(Optional) Switch to a specific model">
|
||||
Onboarding defaults to `openrouter/auto`. Pick a concrete model later:
|
||||
|
||||
```bash
|
||||
openclaw models set openrouter/<provider>/<model>
|
||||
```
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { OPENROUTER_API_KEY: "sk-or-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "openrouter/auto" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Model references
|
||||
|
||||
<Note>
|
||||
Model refs follow the pattern `openrouter/<provider>/<model>`. For the full list of
|
||||
available providers and models, see [/concepts/model-providers](/concepts/model-providers).
|
||||
</Note>
|
||||
|
||||
Bundled fallback models, used when live catalog discovery is unavailable:
|
||||
|
||||
| Model ref | Notes |
|
||||
| --------------------------------- | ---------------------------- |
|
||||
| `openrouter/auto` | OpenRouter automatic routing |
|
||||
| `openrouter/moonshotai/kimi-k2.6` | Kimi K2.6 via MoonshotAI |
|
||||
| `openrouter/moonshotai/kimi-k2.5` | Kimi K2.5 via MoonshotAI |
|
||||
|
||||
Any other `openrouter/<provider>/<model>` ref, including
|
||||
`openrouter/openrouter/fusion` (see [Fusion router](#fusion-router)), resolves
|
||||
dynamically against OpenRouter's live model catalog.
|
||||
|
||||
## Image generation
|
||||
|
||||
OpenRouter can back the `image_generate` tool. Set an OpenRouter image model
|
||||
under `agents.defaults.imageGenerationModel`:
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { OPENROUTER_API_KEY: "sk-or-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "openrouter/google/gemini-3.1-flash-image-preview",
|
||||
timeoutMs: 180_000,
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw sends image requests to OpenRouter's chat-completions image API with
|
||||
`modalities: ["image", "text"]`. Gemini image models additionally receive
|
||||
`aspectRatio` and `resolution` hints through OpenRouter's `image_config`; other
|
||||
image models do not. Use `agents.defaults.imageGenerationModel.timeoutMs` for
|
||||
slower models; the `image_generate` tool's per-call `timeoutMs` still wins.
|
||||
|
||||
## Video generation
|
||||
|
||||
OpenRouter can back the `video_generate` tool through its asynchronous
|
||||
`/videos` API. Set an OpenRouter video model under
|
||||
`agents.defaults.videoGenerationModel`:
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { OPENROUTER_API_KEY: "sk-or-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "openrouter/google/veo-3.1-fast",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw submits text-to-video and image-to-video jobs, polls the returned
|
||||
`polling_url`, and downloads the finished video from OpenRouter's
|
||||
`unsigned_urls` or the job content endpoint. Reference images default to
|
||||
first/last-frame images; images tagged `reference_image` are sent as input
|
||||
references instead. The bundled `google/veo-3.1-fast` default supports 4/6/8
|
||||
second durations, `720P`/`1080P` resolutions, and `16:9`/`9:16` aspect ratios.
|
||||
Video-to-video is not supported: the upstream API only accepts text and image
|
||||
references.
|
||||
|
||||
## Music generation
|
||||
|
||||
OpenRouter can back the `music_generate` tool through chat-completions audio
|
||||
output. Set an OpenRouter audio model under
|
||||
`agents.defaults.musicGenerationModel`:
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { OPENROUTER_API_KEY: "sk-or-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
musicGenerationModel: {
|
||||
primary: "openrouter/google/lyria-3-pro-preview",
|
||||
timeoutMs: 180_000,
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
The bundled OpenRouter music provider defaults to `google/lyria-3-pro-preview`
|
||||
and also exposes `google/lyria-3-clip-preview`. OpenClaw sends `modalities:
|
||||
["text", "audio"]`, streams the response, collects the audio chunks, and saves
|
||||
the result as generated media for channel delivery. Lyria models accept one
|
||||
reference image through the shared `music_generate image=...` parameter.
|
||||
|
||||
## Text-to-speech
|
||||
|
||||
OpenRouter can act as a TTS provider through its OpenAI-compatible
|
||||
`/audio/speech` endpoint.
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "openrouter",
|
||||
providers: {
|
||||
openrouter: {
|
||||
model: "hexgrad/kokoro-82m",
|
||||
speakerVoice: "af_alloy",
|
||||
responseFormat: "mp3",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
If `messages.tts.providers.openrouter.apiKey` is omitted, TTS falls back to
|
||||
`models.providers.openrouter.apiKey`, then `OPENROUTER_API_KEY`.
|
||||
|
||||
## Speech-to-text (inbound audio)
|
||||
|
||||
OpenRouter can transcribe inbound voice/audio attachments through the shared
|
||||
`tools.media.audio` path, using its STT endpoint (`/audio/transcriptions`).
|
||||
This applies to any channel plugin that forwards inbound voice/audio into
|
||||
media understanding preflight.
|
||||
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
enabled: true,
|
||||
models: [{ provider: "openrouter", model: "openai/whisper-large-v3-turbo" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw sends OpenRouter STT requests as JSON with base64 audio under
|
||||
`input_audio` (OpenRouter's STT contract), not as multipart OpenAI form
|
||||
uploads.
|
||||
|
||||
## Fusion router
|
||||
|
||||
OpenRouter Fusion sends one OpenClaw model ref to several OpenRouter models in
|
||||
parallel, has OpenRouter judge their answers, and returns one final response
|
||||
through the normal OpenRouter endpoint. The upstream model slug is
|
||||
`openrouter/fusion`, so the OpenClaw model ref carries both the OpenClaw
|
||||
provider prefix and the upstream OpenRouter namespace:
|
||||
|
||||
```bash
|
||||
openclaw models set openrouter/openrouter/fusion
|
||||
```
|
||||
|
||||
Configure Fusion's panel and judge through the model's `params.extraBody`;
|
||||
those fields forward directly into the OpenRouter chat-completions request
|
||||
body. Fusion works with either OAuth or API-key onboarding; if you use OAuth,
|
||||
omit the `env.OPENROUTER_API_KEY` line below.
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { OPENROUTER_API_KEY: "sk-or-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "openrouter/openrouter/fusion" },
|
||||
models: {
|
||||
"openrouter/openrouter/fusion": {
|
||||
params: {
|
||||
extraBody: {
|
||||
plugins: [
|
||||
{
|
||||
id: "fusion",
|
||||
analysis_models: [
|
||||
"google/gemini-3.5-flash",
|
||||
"moonshotai/kimi-k2.6",
|
||||
"deepseek/deepseek-v4-pro",
|
||||
],
|
||||
model: "google/gemini-3.5-flash",
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`analysis_models` is the parallel panel; `model` inside the Fusion plugin
|
||||
config is the judge model. Do not set top-level `tool_choice` to `"required"`
|
||||
in normal agent/chat turns to try to force Fusion: OpenClaw turns can include
|
||||
its own tool definitions, and a top-level required tool choice may pick one of
|
||||
those instead of the Fusion router. When this Fusion plugin config is present,
|
||||
OpenClaw adds a sanitized system-prompt note listing the configured analysis
|
||||
models and judge model, so the agent can answer questions about its own Fusion
|
||||
panel. Other `extraBody` fields are not copied into the prompt.
|
||||
|
||||
Fusion is slower by design: OpenRouter fans the prompt out to multiple
|
||||
analysis models, then runs a judge/synthesis step, so latency runs higher than
|
||||
a direct single-model request. Use it for deliberate, high-quality answers or
|
||||
escalation paths, not as a latency-sensitive default. Keep the panel small and
|
||||
pick faster analysis/judge models for quicker responses.
|
||||
|
||||
Test a configured ref with a one-shot local call:
|
||||
|
||||
```bash
|
||||
openclaw infer model run --local \
|
||||
--model openrouter/openrouter/fusion \
|
||||
--prompt "Reply with exactly: FUSION_OK" \
|
||||
--json
|
||||
```
|
||||
|
||||
## Authentication and headers
|
||||
|
||||
OpenRouter uses a Bearer token from your API key. OpenRouter OAuth is a PKCE
|
||||
login flow that issues an OpenRouter API key, so OpenClaw stores the result in
|
||||
the same `openrouter:default` API-key auth profile used by manual API-key
|
||||
setup.
|
||||
|
||||
To sign in or rotate the stored key on an existing install without rerunning
|
||||
full onboarding:
|
||||
|
||||
```bash
|
||||
openclaw models auth login --provider openrouter --method oauth
|
||||
openclaw models auth login --provider openrouter --method api-key
|
||||
```
|
||||
|
||||
On verified OpenRouter requests (`https://openrouter.ai/api/v1`), OpenClaw adds
|
||||
OpenRouter's documented app-attribution headers:
|
||||
|
||||
| Header | Value |
|
||||
| ------------------------- | ------------------------------------------------------------------------------------------------------ |
|
||||
| `HTTP-Referer` | `https://openclaw.ai` |
|
||||
| `X-OpenRouter-Title` | `OpenClaw` |
|
||||
| `X-OpenRouter-Categories` | `cli-agent,cloud-agent,programming-app,creative-writing,writing-assistant,general-chat,personal-agent` |
|
||||
|
||||
<Warning>
|
||||
If you repoint the OpenRouter provider at some other proxy or base URL, OpenClaw
|
||||
does **not** inject those OpenRouter-specific headers or Anthropic cache markers.
|
||||
</Warning>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Response caching">
|
||||
OpenRouter response caching is opt-in. Enable it per model:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"openrouter/auto": {
|
||||
params: {
|
||||
responseCache: true,
|
||||
responseCacheTtlSeconds: 300,
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw sends `X-OpenRouter-Cache: true` and, when configured,
|
||||
`X-OpenRouter-Cache-TTL`. `responseCacheClear: true` forces a refresh for
|
||||
the current request and stores the replacement response. Snake_case
|
||||
aliases (`response_cache`, `response_cache_ttl_seconds`,
|
||||
`response_cache_clear`) are accepted, as is `responseCacheTtl` /
|
||||
`response_cache_ttl` without the `Seconds` suffix.
|
||||
|
||||
This is separate from provider prompt caching and from OpenRouter's
|
||||
Anthropic `cache_control` markers. It only applies on verified
|
||||
`openrouter.ai` routes, not custom proxy base URLs.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Anthropic cache markers">
|
||||
On verified OpenRouter routes, Anthropic model refs keep OpenRouter's
|
||||
Anthropic `cache_control` markers for better prompt-cache reuse on
|
||||
system/developer prompt blocks.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Anthropic reasoning prefill">
|
||||
On verified OpenRouter routes, Anthropic model refs with reasoning enabled
|
||||
drop trailing assistant prefill turns before the request reaches
|
||||
OpenRouter, matching Anthropic's requirement that reasoning conversations
|
||||
end with a user turn.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Thinking / reasoning injection">
|
||||
On supported non-`auto` routes, OpenClaw maps the selected thinking level
|
||||
to OpenRouter proxy reasoning payloads. `openrouter/auto` and unsupported
|
||||
model hints skip that injection. Stale `openrouter/hunter-alpha` refs also
|
||||
skip it, because OpenRouter could return final answer text in reasoning
|
||||
fields on that retired route.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="DeepSeek V4 reasoning replay">
|
||||
On verified OpenRouter routes, `openrouter/deepseek/deepseek-v4-flash` and
|
||||
`openrouter/deepseek/deepseek-v4-pro` fill missing `reasoning_content` on
|
||||
replayed assistant turns, keeping thinking/tool conversations in DeepSeek
|
||||
V4's required follow-up shape. OpenClaw sends OpenRouter-supported
|
||||
`reasoning.effort` values for these routes: `xhigh`/`max` map to `xhigh`,
|
||||
every other non-off level maps to `high`.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="OpenAI-only request shaping">
|
||||
OpenRouter runs through the proxy-style OpenAI-compatible path, so native
|
||||
OpenAI-only request shaping such as `serviceTier`, Responses `store`,
|
||||
OpenAI reasoning-compat payloads, and prompt-cache hints is not forwarded.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Gemini-backed routes">
|
||||
Gemini-backed OpenRouter refs stay on the proxy-Gemini path: OpenClaw keeps
|
||||
Gemini thought-signature sanitation there, but does not enable native
|
||||
Gemini replay validation or bootstrap rewrites.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Provider routing metadata">
|
||||
OpenRouter supports a `provider` request object for underlying provider
|
||||
routing. Configure a default policy for all OpenRouter text-model requests
|
||||
with `models.providers.openrouter.params.provider`:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
openrouter: {
|
||||
params: {
|
||||
provider: {
|
||||
sort: "latency",
|
||||
require_parameters: true,
|
||||
data_collection: "deny",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw forwards that object to OpenRouter as the request `provider`
|
||||
payload. Use OpenRouter's documented snake_case fields, including `sort`,
|
||||
`only`, `ignore`, `order`, `allow_fallbacks`, `require_parameters`,
|
||||
`data_collection`, `quantizations`, `max_price`, `preferred_max_latency`,
|
||||
`preferred_min_throughput`, `zdr`, and `enforce_distillable_text`.
|
||||
|
||||
Per-model params override the provider-wide routing object:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"openrouter/anthropic/claude-sonnet-4-6": {
|
||||
params: {
|
||||
provider: {
|
||||
order: ["anthropic"],
|
||||
allow_fallbacks: false,
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
This only applies on OpenRouter chat-completions routes. Direct Anthropic,
|
||||
Google, OpenAI, or custom provider routes ignore OpenRouter routing params.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config reference for agents, models, and providers.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
123
docs/providers/perplexity-provider.md
Normal file
123
docs/providers/perplexity-provider.md
Normal file
@@ -0,0 +1,123 @@
|
||||
---
|
||||
summary: "Perplexity web search provider setup (API key, search modes, filtering)"
|
||||
title: "Perplexity"
|
||||
read_when:
|
||||
- You want to configure Perplexity as a web search provider
|
||||
- You need the Perplexity API key or OpenRouter proxy setup
|
||||
---
|
||||
|
||||
The Perplexity plugin registers a `web_search` provider with two transports: the
|
||||
native Perplexity Search API (structured results with filters) and Perplexity
|
||||
Sonar chat completions, direct or via OpenRouter (AI-synthesized answers with
|
||||
citations).
|
||||
|
||||
<Note>
|
||||
This page covers the Perplexity **provider** setup. For the Perplexity **tool** (how the agent uses it), see [Perplexity search](/tools/perplexity-search).
|
||||
</Note>
|
||||
|
||||
| Property | Value |
|
||||
| ----------- | ---------------------------------------------------------------------- |
|
||||
| Type | Web search provider (not a model provider) |
|
||||
| Auth | `PERPLEXITY_API_KEY` (native) or `OPENROUTER_API_KEY` (via OpenRouter) |
|
||||
| Config path | `plugins.entries.perplexity.config.webSearch.apiKey` |
|
||||
| Overrides | `plugins.entries.perplexity.config.webSearch.baseUrl` / `.model` |
|
||||
| Get a key | [perplexity.ai/settings/api](https://www.perplexity.ai/settings/api) |
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/perplexity-plugin
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set the API key">
|
||||
```bash
|
||||
openclaw configure --section web
|
||||
```
|
||||
|
||||
Or set the key directly:
|
||||
|
||||
```bash
|
||||
openclaw config set plugins.entries.perplexity.config.webSearch.apiKey "pplx-xxxxxxxxxxxx"
|
||||
```
|
||||
|
||||
A key exported as `PERPLEXITY_API_KEY` or `OPENROUTER_API_KEY` in the Gateway
|
||||
environment also works.
|
||||
|
||||
</Step>
|
||||
<Step title="Start searching">
|
||||
`web_search` auto-detects Perplexity once its key is the available search
|
||||
credential; no further setup is required. To pin the provider explicitly:
|
||||
|
||||
```bash
|
||||
openclaw config set tools.web.search.provider perplexity
|
||||
```
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Search modes
|
||||
|
||||
The plugin resolves transport in this order:
|
||||
|
||||
1. `webSearch.baseUrl` or `webSearch.model` set: always routes through Sonar chat completions against that endpoint, regardless of key type.
|
||||
2. Otherwise, key source decides the endpoint: a configured key's prefix picks the transport (config beats environment variables); an environment key uses its matching endpoint directly.
|
||||
|
||||
| Key prefix | Transport | Features |
|
||||
| ---------- | ---------------------------------------------------------- | ------------------------------------------------ |
|
||||
| `pplx-` | Native Perplexity Search API (`https://api.perplexity.ai`) | Structured results, domain/language/date filters |
|
||||
| `sk-or-` | OpenRouter (`https://openrouter.ai/api/v1`), Sonar model | AI-synthesized answers with citations |
|
||||
|
||||
A configured key with any other prefix also uses the native Search API. The
|
||||
chat-completions path defaults to the `perplexity/sonar-pro` model; override it
|
||||
with `plugins.entries.perplexity.config.webSearch.model`.
|
||||
|
||||
## Native API filtering
|
||||
|
||||
| Filter | Description | Transport |
|
||||
| ------------------------------------ | --------------------------------------------------------------- | ----------- |
|
||||
| `count` | Results per search, 1-10 (default 5) | Native only |
|
||||
| `freshness` | Recency window: `day`, `week`, `month`, `year` | Both |
|
||||
| `country` | 2-letter country code (`us`, `de`, `jp`) | Native only |
|
||||
| `language` | ISO 639-1 language code (`en`, `fr`, `zh`) | Native only |
|
||||
| `date_after` / `date_before` | Published-date range in `YYYY-MM-DD` | Native only |
|
||||
| `domain_filter` | Max 20 domains; allowlist or `-`-prefixed denylist, never mixed | Native only |
|
||||
| `max_tokens` / `max_tokens_per_page` | Content budget across all results / per page | Native only |
|
||||
|
||||
Native-only filters return a descriptive error on the chat-completions path.
|
||||
`freshness` cannot be combined with `date_after`/`date_before`.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Environment variable for daemon processes">
|
||||
<Warning>
|
||||
A key exported only in an interactive shell is not visible to a
|
||||
launchd/systemd Gateway daemon unless that environment is explicitly
|
||||
imported. Set the key in `~/.openclaw/.env` or via `env.shellEnv` so the
|
||||
Gateway process can read it. See [Environment variables](/help/environment)
|
||||
for the full precedence order.
|
||||
</Warning>
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="OpenRouter proxy setup">
|
||||
To route Perplexity searches through OpenRouter, set an `OPENROUTER_API_KEY`
|
||||
(prefix `sk-or-`) instead of a native Perplexity key. OpenClaw detects the
|
||||
key and switches to the Sonar transport automatically. Useful if you already
|
||||
have OpenRouter billing set up and want to consolidate providers there.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Perplexity search tool" href="/tools/perplexity-search" icon="magnifying-glass">
|
||||
How the agent invokes Perplexity searches and interprets results.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full configuration reference including plugin entries.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
170
docs/providers/pixverse.md
Normal file
170
docs/providers/pixverse.md
Normal file
@@ -0,0 +1,170 @@
|
||||
---
|
||||
summary: "PixVerse video generation setup in OpenClaw"
|
||||
title: "PixVerse"
|
||||
read_when:
|
||||
- You want to use PixVerse video generation in OpenClaw
|
||||
- You need the PixVerse API key/env setup
|
||||
- You want to make PixVerse the default video provider
|
||||
---
|
||||
|
||||
OpenClaw provides `pixverse` as an official external plugin for hosted PixVerse video generation. The plugin registers the `pixverse` provider against the `videoGenerationProviders` contract.
|
||||
|
||||
| Property | Value |
|
||||
| ------------------ | -------------------------------------------------------------------- |
|
||||
| Provider id | `pixverse` |
|
||||
| Plugin package | `@openclaw/pixverse-provider` |
|
||||
| Auth env var | `PIXVERSE_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice pixverse-api-key` |
|
||||
| Direct CLI flag | `--pixverse-api-key <key>` |
|
||||
| API | PixVerse Platform API v2 (`video_id` submission plus result polling) |
|
||||
| Default model | `pixverse/v6` |
|
||||
| Default API region | International |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the plugin">
|
||||
```bash
|
||||
openclaw plugins install @openclaw/pixverse-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set the API key">
|
||||
```bash
|
||||
openclaw onboard --auth-choice pixverse-api-key
|
||||
```
|
||||
|
||||
The wizard prompts for the International or CN endpoint (see API region
|
||||
below) before writing `region` and `baseUrl` into the provider config.
|
||||
Non-interactive runs (key from `--pixverse-api-key` or `PIXVERSE_API_KEY`)
|
||||
default to International.
|
||||
|
||||
Onboarding also sets `agents.defaults.videoGenerationModel.primary` to
|
||||
`pixverse/v6` when no default video model is configured yet.
|
||||
|
||||
</Step>
|
||||
<Step title="Switch an existing default video provider (optional)">
|
||||
```bash
|
||||
openclaw config set agents.defaults.videoGenerationModel.primary "pixverse/v6"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Generate a video">
|
||||
Ask the agent to generate a video. PixVerse will be used automatically.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Supported modes and models
|
||||
|
||||
The provider exposes PixVerse generation models through OpenClaw's shared video tool.
|
||||
|
||||
| Mode | Models | Reference input |
|
||||
| -------------- | -------------------- | ----------------------- |
|
||||
| Text-to-video | `v6` (default), `c1` | None |
|
||||
| Image-to-video | `v6` (default), `c1` | 1 local or remote image |
|
||||
|
||||
Local image references are uploaded to PixVerse before the image-to-video request. Remote image URLs are passed through the PixVerse image upload endpoint as `image_url`.
|
||||
|
||||
| Option | Supported values |
|
||||
| --------------- | -------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Duration | 1-15 seconds (default 5) |
|
||||
| Resolution | `360P`, `540P`, `720P`, `1080P` (default `540P`; `480P` requests map to `540P`) |
|
||||
| Aspect ratio | `16:9` (default), `4:3`, `1:1`, `3:4`, `9:16`, `2:3`, `3:2`, `21:9`; text-to-video only, image-to-video follows the source image |
|
||||
| Generated audio | `audio: true` |
|
||||
|
||||
<Note>
|
||||
PixVerse image template generation is not exposed through `image_generate` yet. That API is template-id driven, while OpenClaw's shared image-generation contract does not currently have a PixVerse-specific typed option bag.
|
||||
</Note>
|
||||
|
||||
## Provider options
|
||||
|
||||
The video provider accepts these optional provider-specific keys:
|
||||
|
||||
| Option | Type | Effect |
|
||||
| ------------------------------------ | ------ | --------------------------------------------- |
|
||||
| `seed` | number | Deterministic seed, 0 to 2147483647 |
|
||||
| `negativePrompt` / `negative_prompt` | string | Negative prompt |
|
||||
| `quality` | string | PixVerse quality such as `720p` |
|
||||
| `motionMode` / `motion_mode` | string | Image-to-video motion mode (default `normal`) |
|
||||
| `cameraMovement` / `camera_movement` | string | PixVerse camera movement preset |
|
||||
| `templateId` / `template_id` | number | Activated PixVerse template id |
|
||||
|
||||
## Configuration
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "pixverse/v6",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="API region">
|
||||
| Region value | PixVerse API base URL |
|
||||
| --------------- | --------------------------------------------- |
|
||||
| `international` | `https://app-api.pixverse.ai/openapi/v2` |
|
||||
| `cn` | `https://app-api.pixverseai.cn/openapi/v2` |
|
||||
|
||||
Set `models.providers.pixverse.region` manually when your key belongs to a
|
||||
specific PixVerse platform region, or run
|
||||
`openclaw onboard --auth-choice pixverse-api-key` to choose one in the
|
||||
setup wizard:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
pixverse: {
|
||||
region: "cn", // "international" or "cn"
|
||||
baseUrl: "https://app-api.pixverseai.cn/openapi/v2",
|
||||
models: [],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Custom base URL">
|
||||
Set `models.providers.pixverse.baseUrl` only when routing through a trusted compatible proxy.
|
||||
`baseUrl` takes precedence over `region`.
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
pixverse: {
|
||||
baseUrl: "https://app-api.pixverse.ai/openapi/v2",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Task polling">
|
||||
PixVerse returns a `video_id` from the generation request. OpenClaw polls
|
||||
`/openapi/v2/video/result/{video_id}` every 5 seconds until the task
|
||||
succeeds, fails, or hits the timeout (default 5 minutes; override with
|
||||
`agents.defaults.videoGenerationModel.timeoutMs`).
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared tool parameters, provider selection, and async behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/config-agents#agent-defaults" icon="gear">
|
||||
Agent default settings including video generation model.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
144
docs/providers/qianfan.md
Normal file
144
docs/providers/qianfan.md
Normal file
@@ -0,0 +1,144 @@
|
||||
---
|
||||
summary: "Use Qianfan's unified API to access many models in OpenClaw"
|
||||
read_when:
|
||||
- You want a single API key for many LLMs
|
||||
- You need Baidu Qianfan setup guidance
|
||||
title: "Qianfan"
|
||||
---
|
||||
|
||||
Qianfan is Baidu's MaaS platform: a unified, OpenAI-compatible API that routes requests to many models behind a single endpoint and API key. OpenClaw ships it as the official external plugin `@openclaw/qianfan-provider`.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | ---------------------------------------- |
|
||||
| Provider | `qianfan` |
|
||||
| Auth | `QIANFAN_API_KEY` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://qianfan.baidubce.com/v2` |
|
||||
| Default model | `qianfan/deepseek-v3.2` |
|
||||
|
||||
## Install plugin
|
||||
|
||||
Install the official plugin, then restart Gateway:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/qianfan-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Create a Baidu Cloud account">
|
||||
Sign up or log in at the [Qianfan Console](https://console.bce.baidu.com/qianfan/ais/console/apiKey) and ensure you have Qianfan API access enabled.
|
||||
</Step>
|
||||
<Step title="Generate an API key">
|
||||
Create a new application or select an existing one, then generate an API key. Baidu Cloud keys use the `bce-v3/ALTAK-...` format.
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice qianfan-api-key
|
||||
```
|
||||
|
||||
Non-interactive runs read the key from `--qianfan-api-key <key>` or
|
||||
`QIANFAN_API_KEY`. Onboarding writes the provider config, adds the
|
||||
`QIANFAN` alias for the default model, and sets `qianfan/deepseek-v3.2`
|
||||
as the default model when none is configured.
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider qianfan
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
| Model ref | Input | Context | Max output | Reasoning | Notes |
|
||||
| ------------------------------------ | ----------- | ------- | ---------- | --------- | ------------- |
|
||||
| `qianfan/deepseek-v3.2` | text | 98,304 | 32,768 | Yes | Default model |
|
||||
| `qianfan/ernie-5.0-thinking-preview` | text, image | 119,000 | 64,000 | Yes | Multimodal |
|
||||
|
||||
The catalog is static; there is no live model discovery.
|
||||
|
||||
<Tip>
|
||||
You only need to override `models.providers.qianfan` when you need a custom base URL or model metadata.
|
||||
</Tip>
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { QIANFAN_API_KEY: "bce-v3/ALTAK-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "qianfan/deepseek-v3.2" },
|
||||
models: {
|
||||
"qianfan/deepseek-v3.2": { alias: "QIANFAN" },
|
||||
},
|
||||
},
|
||||
},
|
||||
models: {
|
||||
providers: {
|
||||
qianfan: {
|
||||
baseUrl: "https://qianfan.baidubce.com/v2",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
{
|
||||
id: "deepseek-v3.2",
|
||||
name: "DEEPSEEK V3.2",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 98304,
|
||||
maxTokens: 32768,
|
||||
},
|
||||
{
|
||||
id: "ernie-5.0-thinking-preview",
|
||||
name: "ERNIE-5.0-Thinking-Preview",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 119000,
|
||||
maxTokens: 64000,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
Model refs use the `qianfan/` prefix (for example `qianfan/deepseek-v3.2`).
|
||||
</Note>
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Transport and compatibility">
|
||||
Qianfan runs through the OpenAI-compatible transport path, not native OpenAI request shaping. Standard OpenAI SDK features work, but provider-specific parameters may not be forwarded.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Troubleshooting">
|
||||
- Ensure your API key starts with `bce-v3/ALTAK-` and has Qianfan API access enabled in the Baidu Cloud console.
|
||||
- If models are not listed, confirm your account has the Qianfan service activated.
|
||||
- Only change the base URL if you use a custom endpoint or proxy.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full OpenClaw configuration reference.
|
||||
</Card>
|
||||
<Card title="Agent setup" href="/concepts/agent" icon="robot">
|
||||
Configuring agent defaults and model assignments.
|
||||
</Card>
|
||||
<Card title="Qianfan API docs" href="https://cloud.baidu.com/doc/qianfan-api/s/3m7of64lb" icon="arrow-up-right-from-square">
|
||||
Official Qianfan API documentation.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
128
docs/providers/qwen-oauth.md
Normal file
128
docs/providers/qwen-oauth.md
Normal file
@@ -0,0 +1,128 @@
|
||||
---
|
||||
summary: "Use the Qwen Portal provider id with OpenClaw"
|
||||
read_when:
|
||||
- You want to configure the qwen-oauth provider id
|
||||
- You previously used Qwen Portal OAuth credentials
|
||||
- You need the Qwen Portal endpoint or migration guidance
|
||||
title: "Qwen OAuth / Portal"
|
||||
---
|
||||
|
||||
`qwen-oauth` is the Qwen Portal provider id, registered by the Qwen plugin
|
||||
(`@openclaw/qwen-provider`). It targets the Qwen Portal endpoint at
|
||||
`https://portal.qwen.ai/v1` and keeps older Qwen OAuth / portal setups
|
||||
addressable through a distinct provider id, separate from the canonical `qwen`
|
||||
provider.
|
||||
|
||||
Choose `qwen-oauth` if you already have a working Qwen Portal token, are
|
||||
migrating a legacy Qwen OAuth or Qwen CLI workflow, or need to test the Qwen
|
||||
Portal endpoint specifically. For new setups, prefer
|
||||
[Qwen](/providers/qwen) with the Standard ModelStudio endpoint: it covers new
|
||||
API-key setups, broader endpoint choices, Standard pay-as-you-go, Coding Plan,
|
||||
and the full Qwen plugin catalog.
|
||||
|
||||
## Setup
|
||||
|
||||
Install the Qwen plugin if you have not already:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/qwen-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
Provide your portal token through onboarding:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-oauth
|
||||
```
|
||||
|
||||
Non-interactive runs read the token from `--qwen-oauth-token <token>`, or set:
|
||||
|
||||
```bash
|
||||
export QWEN_API_KEY="<your-qwen-portal-token>" # pragma: allowlist secret
|
||||
```
|
||||
|
||||
Onboarding stores the token under a `qwen-oauth` auth profile, seeds the portal
|
||||
model catalog, and sets `qwen-oauth/qwen3.5-plus` as the default model when
|
||||
none is configured.
|
||||
|
||||
## Defaults
|
||||
|
||||
- Provider: `qwen-oauth`
|
||||
- Aliases: `qwen-portal`, `qwen-cli`
|
||||
- Base URL: `https://portal.qwen.ai/v1`
|
||||
- Env var: `QWEN_API_KEY`
|
||||
- API style: OpenAI-compatible
|
||||
- Default model: `qwen-oauth/qwen3.5-plus`
|
||||
|
||||
## How this differs from Qwen
|
||||
|
||||
OpenClaw has two Qwen-facing provider ids:
|
||||
|
||||
| Provider | Endpoint family | Best for |
|
||||
| ------------ | -------------------------------------------------------- | -------------------------------------------------------------------------------------- |
|
||||
| `qwen` | Qwen Cloud / Alibaba DashScope and Coding Plan endpoints | New API-key setups, Standard pay-as-you-go, Coding Plan, multimodal DashScope features |
|
||||
| `qwen-oauth` | Qwen Portal endpoint at `portal.qwen.ai/v1` | Existing Qwen Portal tokens and legacy Qwen OAuth / CLI setups |
|
||||
|
||||
Both providers use OpenAI-compatible request shapes, but they are separate auth
|
||||
surfaces. A token stored for `qwen-oauth` should not be treated as a DashScope
|
||||
or ModelStudio key, and a new DashScope key should use the canonical `qwen`
|
||||
provider instead.
|
||||
|
||||
## Models
|
||||
|
||||
The Qwen plugin seeds this static catalog for the Qwen Portal endpoint. All
|
||||
entries use a 65,536-token max output; availability depends on the current Qwen
|
||||
Portal account and token.
|
||||
|
||||
| Model ref | Input | Context | Notes |
|
||||
| --------------------------------- | ----------- | --------- | ------------- |
|
||||
| `qwen-oauth/qwen3.5-plus` | text, image | 1,000,000 | Default model |
|
||||
| `qwen-oauth/qwen3.6-plus` | text, image | 1,000,000 | |
|
||||
| `qwen-oauth/qwen3-max-2026-01-23` | text | 262,144 | |
|
||||
| `qwen-oauth/qwen3-coder-next` | text | 262,144 | |
|
||||
| `qwen-oauth/qwen3-coder-plus` | text | 1,000,000 | |
|
||||
| `qwen-oauth/MiniMax-M2.5` | text | 1,000,000 | Reasoning |
|
||||
| `qwen-oauth/glm-5` | text | 202,752 | |
|
||||
| `qwen-oauth/glm-4.7` | text | 202,752 | |
|
||||
| `qwen-oauth/kimi-k2.5` | text, image | 262,144 | |
|
||||
|
||||
If your account uses ModelStudio / DashScope API keys instead, configure the
|
||||
canonical `qwen` provider:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-standard-api-key
|
||||
openclaw models set qwen/qwen3-coder-plus
|
||||
```
|
||||
|
||||
## Migration
|
||||
|
||||
Legacy Qwen Portal OAuth profiles are not refreshable; `openclaw doctor` flags
|
||||
them. If a portal profile stops working, re-run onboarding with a current token
|
||||
or switch to the Standard Qwen provider:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-standard-api-key
|
||||
```
|
||||
|
||||
Standard global ModelStudio uses:
|
||||
|
||||
```text
|
||||
https://dashscope-intl.aliyuncs.com/compatible-mode/v1
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- Portal OAuth refresh failures: legacy Qwen Portal OAuth profiles are not
|
||||
refreshable. Re-run onboarding with a current token.
|
||||
- Wrong endpoint errors: confirm the model ref starts with `qwen-oauth/` when
|
||||
using a portal token. Use `qwen/` refs only for the canonical Qwen provider.
|
||||
- `QWEN_API_KEY` confusion: both Qwen pages mention this env var, but onboarding
|
||||
stores credentials under the selected provider id. Prefer onboarding when you
|
||||
keep both `qwen` and `qwen-oauth` available on the same machine.
|
||||
|
||||
## Related
|
||||
|
||||
- [Qwen](/providers/qwen)
|
||||
- [Alibaba Model Studio](/providers/alibaba)
|
||||
- [Model providers](/concepts/model-providers)
|
||||
- [All providers](/providers/index)
|
||||
345
docs/providers/qwen.md
Normal file
345
docs/providers/qwen.md
Normal file
@@ -0,0 +1,345 @@
|
||||
---
|
||||
summary: "Use Qwen Cloud through its OpenClaw plugin"
|
||||
read_when:
|
||||
- You want to use Qwen with OpenClaw
|
||||
- You previously used Qwen OAuth
|
||||
title: "Qwen"
|
||||
---
|
||||
|
||||
Qwen Cloud is an official external OpenClaw provider plugin with canonical id `qwen`. It targets Qwen Cloud / Alibaba DashScope Standard and Coding Plan endpoints, keeps legacy `modelstudio` ids working as a compatibility alias, and exposes the Qwen Portal token flow as a separate provider, [`qwen-oauth`](/providers/qwen-oauth).
|
||||
|
||||
| Property | Value |
|
||||
| ---------------------- | ------------------------------------------ |
|
||||
| Provider | `qwen` |
|
||||
| Portal provider | [`qwen-oauth`](/providers/qwen-oauth) |
|
||||
| Preferred env var | `QWEN_API_KEY` |
|
||||
| Also accepted (compat) | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` |
|
||||
| API style | OpenAI-compatible |
|
||||
|
||||
<Tip>
|
||||
For `qwen3.6-plus`, use a **Standard (pay-as-you-go)** endpoint. It is not available on the Coding Plan endpoints.
|
||||
</Tip>
|
||||
|
||||
## Install plugin
|
||||
|
||||
`qwen` ships as an official external plugin, not bundled with core. Install it and restart Gateway:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/qwen-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
Choose your plan type and follow the setup steps.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Coding Plan (subscription)">
|
||||
**Best for:** subscription-based access through the Qwen Coding Plan.
|
||||
|
||||
<Steps>
|
||||
<Step title="Get your API key">
|
||||
Create or copy an API key from [home.qwencloud.com/api-keys](https://home.qwencloud.com/api-keys).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
For the **Global** endpoint:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-api-key
|
||||
```
|
||||
|
||||
For the **China** endpoint:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-api-key-cn
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "qwen/qwen3.5-plus" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider qwen
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Note>
|
||||
Legacy `modelstudio-*` auth-choice ids and `modelstudio/...` model refs still
|
||||
work as compatibility aliases, but new setup flows should prefer the canonical
|
||||
`qwen-*` auth-choice ids and `qwen/...` model refs. If you define an exact
|
||||
custom `models.providers.modelstudio` entry with another `api` value, that
|
||||
custom provider owns `modelstudio/...` refs instead of the Qwen compatibility
|
||||
alias.
|
||||
</Note>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Standard (pay-as-you-go)">
|
||||
**Best for:** pay-as-you-go access through the Standard Model Studio endpoint, including models like `qwen3.6-plus` that are not available on the Coding Plan.
|
||||
|
||||
<Steps>
|
||||
<Step title="Get your API key">
|
||||
Create or copy an API key from [home.qwencloud.com/api-keys](https://home.qwencloud.com/api-keys).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
For the **Global** endpoint:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-standard-api-key
|
||||
```
|
||||
|
||||
For the **China** endpoint:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-standard-api-key-cn
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "qwen/qwen3.5-plus" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider qwen
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Note>
|
||||
Legacy `modelstudio-*` auth-choice ids and `modelstudio/...` model refs still
|
||||
work as compatibility aliases, but new setup flows should prefer the canonical
|
||||
`qwen-*` auth-choice ids and `qwen/...` model refs. If you define an exact
|
||||
custom `models.providers.modelstudio` entry with another `api` value, that
|
||||
custom provider owns `modelstudio/...` refs instead of the Qwen compatibility
|
||||
alias.
|
||||
</Note>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Qwen OAuth / Portal">
|
||||
**Best for:** a Qwen Portal token against `https://portal.qwen.ai/v1`.
|
||||
|
||||
See [Qwen OAuth / Portal](/providers/qwen-oauth) for the dedicated provider
|
||||
page and migration notes.
|
||||
|
||||
<Steps>
|
||||
<Step title="Provide your portal token">
|
||||
```bash
|
||||
openclaw onboard --auth-choice qwen-oauth
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "qwen-oauth/qwen3.5-plus" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider qwen-oauth
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Note>
|
||||
`qwen-oauth` uses the same `QWEN_API_KEY` env var name as the Qwen Cloud
|
||||
provider, but stores auth under the `qwen-oauth` provider id when configured
|
||||
through OpenClaw onboarding.
|
||||
</Note>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
## Plan types and endpoints
|
||||
|
||||
| Plan | Region | Auth choice | Endpoint |
|
||||
| -------------------------- | ------ | -------------------------- | ------------------------------------------------ |
|
||||
| Coding Plan (subscription) | China | `qwen-api-key-cn` | `coding.dashscope.aliyuncs.com/v1` |
|
||||
| Coding Plan (subscription) | Global | `qwen-api-key` | `coding-intl.dashscope.aliyuncs.com/v1` |
|
||||
| Qwen Portal | Global | `qwen-oauth` | `portal.qwen.ai/v1` |
|
||||
| Standard (pay-as-you-go) | China | `qwen-standard-api-key-cn` | `dashscope.aliyuncs.com/compatible-mode/v1` |
|
||||
| Standard (pay-as-you-go) | Global | `qwen-standard-api-key` | `dashscope-intl.aliyuncs.com/compatible-mode/v1` |
|
||||
|
||||
The provider auto-selects the endpoint based on your auth choice. Canonical
|
||||
choices use the `qwen-*` family; `modelstudio-*` remains compatibility-only.
|
||||
Override with a custom `baseUrl` in config.
|
||||
|
||||
<Tip>
|
||||
**Manage keys:** [home.qwencloud.com/api-keys](https://home.qwencloud.com/api-keys) |
|
||||
**Docs:** [docs.qwencloud.com](https://docs.qwencloud.com/developer-guides/getting-started/introduction)
|
||||
</Tip>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
OpenClaw ships this Qwen static catalog. The catalog is endpoint-aware: Coding
|
||||
Plan configs omit models that only work on the Standard endpoint.
|
||||
|
||||
| Model ref | Input | Context | Notes |
|
||||
| --------------------------- | ----------- | --------- | ----------------------- |
|
||||
| `qwen/qwen3.5-plus` | text, image | 1,000,000 | Default model |
|
||||
| `qwen/qwen3.6-plus` | text, image | 1,000,000 | Standard endpoints only |
|
||||
| `qwen/qwen3-max-2026-01-23` | text | 262,144 | Qwen Max line |
|
||||
| `qwen/qwen3-coder-next` | text | 262,144 | Coding |
|
||||
| `qwen/qwen3-coder-plus` | text | 1,000,000 | Coding |
|
||||
| `qwen/MiniMax-M2.5` | text | 1,000,000 | Reasoning enabled |
|
||||
| `qwen/glm-5` | text | 202,752 | GLM |
|
||||
| `qwen/glm-4.7` | text | 202,752 | GLM |
|
||||
| `qwen/kimi-k2.5` | text, image | 262,144 | Moonshot AI via Alibaba |
|
||||
| `qwen-oauth/qwen3.5-plus` | text, image | 1,000,000 | Qwen Portal default |
|
||||
|
||||
<Note>
|
||||
Availability can still vary by endpoint and billing plan even when a model is
|
||||
present in the static catalog.
|
||||
</Note>
|
||||
|
||||
## Thinking controls
|
||||
|
||||
`qwen/MiniMax-M2.5` is the only reasoning-enabled model in the built-in
|
||||
catalog. For reasoning models on the `qwen` family, the provider maps
|
||||
OpenClaw thinking levels to DashScope's top-level `enable_thinking` request
|
||||
flag: disabled thinking sends `enable_thinking: false`, any other level sends
|
||||
`enable_thinking: true`. Custom models can opt into an alternate chat-template
|
||||
thinking payload by setting `compat.thinkingFormat: "qwen-chat-template"` on
|
||||
the model entry.
|
||||
|
||||
## Multimodal add-ons
|
||||
|
||||
The `qwen` plugin exposes multimodal capabilities on the **Standard** DashScope
|
||||
endpoints only, not the Coding Plan endpoints:
|
||||
|
||||
- **Image and video understanding** via `qwen-vl-max-latest`
|
||||
- **Wan video generation** via `wan2.6-t2v` (default), `wan2.6-i2v`, `wan2.6-r2v`, `wan2.6-r2v-flash`, `wan2.7-r2v`
|
||||
|
||||
Media understanding is auto-resolved from the configured Qwen auth; no extra
|
||||
config is needed. Make sure you are on a Standard (pay-as-you-go) endpoint for
|
||||
media understanding to work.
|
||||
|
||||
To make Qwen the default video provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: { primary: "qwen/wan2.6-t2v" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Video-generation limits: 1 output video per request, up to 1 input image
|
||||
(image-to-video), up to 4 input videos (video-to-video), max 10 seconds
|
||||
duration. Supports `size`, `aspectRatio`, `resolution`, `audio`, and
|
||||
`watermark`. Reference image/video inputs require remote http(s) URLs; local
|
||||
file paths are rejected up front because the DashScope video endpoint does not
|
||||
accept uploaded local buffers for those references.
|
||||
|
||||
<Note>
|
||||
See [Video generation](/tools/video-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Qwen 3.6 Plus availability">
|
||||
`qwen3.6-plus` is available on the Standard (pay-as-you-go) endpoints:
|
||||
|
||||
- China: `dashscope.aliyuncs.com/compatible-mode/v1`
|
||||
- Global: `dashscope-intl.aliyuncs.com/compatible-mode/v1`
|
||||
|
||||
If the Coding Plan endpoints return an "unsupported model" error for
|
||||
`qwen3.6-plus`, switch to Standard (pay-as-you-go) instead of the Coding Plan
|
||||
endpoint/key pair.
|
||||
|
||||
OpenClaw's Qwen static catalog does not advertise `qwen3.6-plus` on Coding
|
||||
Plan endpoints, but an explicitly configured `qwen/qwen3.6-plus` entry under
|
||||
`models.providers.qwen.models` is honored on Coding Plan base URLs, so you
|
||||
can opt that model in if Aliyun enables it on your subscription. The
|
||||
upstream API still decides whether the call succeeds.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Video generation region routing">
|
||||
OpenClaw maps the configured Qwen region to the matching DashScope AIGC host
|
||||
before submitting a video job:
|
||||
|
||||
- Global/Intl: `https://dashscope-intl.aliyuncs.com`
|
||||
- China: `https://dashscope.aliyuncs.com`
|
||||
|
||||
A normal `models.providers.qwen.baseUrl` pointing at either the Coding Plan
|
||||
or Standard Qwen hosts still routes video generation to the matching
|
||||
regional DashScope video endpoint.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Streaming usage compatibility">
|
||||
Native Qwen endpoints advertise streaming usage compatibility on the shared
|
||||
`openai-completions` transport, so DashScope-compatible custom provider ids
|
||||
targeting the same native hosts inherit the same behavior without requiring
|
||||
the built-in `qwen` provider id specifically. This applies to both Coding
|
||||
Plan and Standard endpoints:
|
||||
|
||||
- `https://coding.dashscope.aliyuncs.com/v1`
|
||||
- `https://coding-intl.dashscope.aliyuncs.com/v1`
|
||||
- `https://dashscope.aliyuncs.com/compatible-mode/v1`
|
||||
- `https://dashscope-intl.aliyuncs.com/compatible-mode/v1`
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Capability plan">
|
||||
The `qwen` plugin is being positioned as the vendor home for the full Qwen
|
||||
Cloud surface, not just coding/text models.
|
||||
|
||||
- **Text/chat models:** available through the plugin
|
||||
- **Tool calling, structured output, thinking:** inherited from the OpenAI-compatible transport
|
||||
- **Image generation:** planned at the provider-plugin layer
|
||||
- **Image/video understanding:** available through the plugin on the Standard endpoint
|
||||
- **Speech/audio:** planned at the provider-plugin layer
|
||||
- **Memory embeddings/reranking:** planned through the embedding adapter surface
|
||||
- **Video generation:** available through the plugin through the shared video-generation capability
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment and daemon setup">
|
||||
If the Gateway runs as a daemon (launchd/systemd), make sure `QWEN_API_KEY` is
|
||||
available to that process (for example, in `~/.openclaw/.env` or via
|
||||
`env.shellEnv`).
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Alibaba Model Studio" href="/providers/alibaba" icon="cloud">
|
||||
Bundled Wan video generation provider on the same DashScope platform.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
General troubleshooting and FAQ.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
103
docs/providers/runway.md
Normal file
103
docs/providers/runway.md
Normal file
@@ -0,0 +1,103 @@
|
||||
---
|
||||
summary: "Runway video generation setup in OpenClaw"
|
||||
title: "Runway"
|
||||
read_when:
|
||||
- You want to use Runway video generation in OpenClaw
|
||||
- You need the Runway API key/env setup
|
||||
- You want to make Runway the default video provider
|
||||
---
|
||||
|
||||
OpenClaw ships a bundled `runway` provider for hosted video generation, enabled by default, registered against the `videoGenerationProviders` contract.
|
||||
|
||||
| Property | Value |
|
||||
| --------------- | ----------------------------------------------------------------- |
|
||||
| Provider id | `runway` |
|
||||
| Plugin | bundled, `enabledByDefault: true` |
|
||||
| Auth env vars | `RUNWAYML_API_SECRET` (canonical) or `RUNWAY_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice runway-api-key` |
|
||||
| Direct CLI flag | `--runway-api-key <key>` |
|
||||
| API | Runway task-based video generation (`GET /v1/tasks/{id}` polling) |
|
||||
| Default model | `runway/gen4.5` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set the API key">
|
||||
```bash
|
||||
openclaw onboard --auth-choice runway-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set Runway as the default video provider">
|
||||
```bash
|
||||
openclaw config set agents.defaults.videoGenerationModel.primary "runway/gen4.5"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Generate a video">
|
||||
Ask the agent to generate a video. Runway will be used automatically.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Supported modes and models
|
||||
|
||||
The provider exposes seven Runway models split across three modes. The same model id can serve more than one mode (for example `gen4.5` works for both text-to-video and image-to-video).
|
||||
|
||||
| Mode | Models | Reference input |
|
||||
| -------------- | ---------------------------------------------------------------------- | ----------------------- |
|
||||
| Text-to-video | `gen4.5` (default), `veo3.1`, `veo3.1_fast`, `veo3` | None |
|
||||
| Image-to-video | `gen4.5`, `gen4_turbo`, `gen3a_turbo`, `veo3.1`, `veo3.1_fast`, `veo3` | 1 local or remote image |
|
||||
| Video-to-video | `gen4_aleph` | 1 local or remote video |
|
||||
|
||||
Local image and video references are supported via data URIs.
|
||||
|
||||
| Aspect ratios | Allowed values |
|
||||
| --------------------- | ------------------------------------------- |
|
||||
| Text-to-video | `16:9`, `9:16` |
|
||||
| Image and video edits | `1:1`, `16:9`, `9:16`, `3:4`, `4:3`, `21:9` |
|
||||
|
||||
<Warning>
|
||||
Video-to-video currently requires `runway/gen4_aleph`. Other Runway model ids reject video reference inputs.
|
||||
</Warning>
|
||||
|
||||
<Note>
|
||||
Picking a Runway model id from the wrong column produces an explicit error before the API request leaves OpenClaw. The provider validates `model` against the mode's allowlist (`TEXT_ONLY_MODELS`, `IMAGE_MODELS`, `VIDEO_MODELS`) in `extensions/runway/video-generation-provider.ts`.
|
||||
</Note>
|
||||
|
||||
## Configuration
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "runway/gen4.5",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Environment variable aliases">
|
||||
OpenClaw recognizes both `RUNWAYML_API_SECRET` (canonical) and `RUNWAY_API_KEY`.
|
||||
Either variable authenticates the Runway provider.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Task polling">
|
||||
Runway uses a task-based API. After submitting a generation request, OpenClaw
|
||||
polls `GET /v1/tasks/{id}` until the video is ready. No additional
|
||||
configuration is needed for the polling behavior.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared tool parameters, provider selection, and async behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/config-agents#agent-defaults" icon="gear">
|
||||
Agent default settings including video generation model.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
68
docs/providers/senseaudio.md
Normal file
68
docs/providers/senseaudio.md
Normal file
@@ -0,0 +1,68 @@
|
||||
---
|
||||
summary: "SenseAudio batch speech-to-text for inbound voice notes"
|
||||
read_when:
|
||||
- You want SenseAudio speech-to-text for audio attachments
|
||||
- You need the SenseAudio API key env var or audio config path
|
||||
title: "SenseAudio"
|
||||
---
|
||||
|
||||
SenseAudio transcribes inbound audio and voice-note attachments through OpenClaw's shared `tools.media.audio` pipeline. OpenClaw posts multipart audio to the OpenAI-compatible transcription endpoint and injects the returned text as `{{Transcript}}` plus an `[Audio]` block.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | ------------------------------------------------ |
|
||||
| Provider id | `senseaudio` |
|
||||
| Plugin | bundled, `enabledByDefault: true` |
|
||||
| Contract | `mediaUnderstandingProviders` (audio) |
|
||||
| Auth env var | `SENSEAUDIO_API_KEY` |
|
||||
| Default model | `senseaudio-asr-pro-1.5-260319` |
|
||||
| Default URL | `https://api.senseaudio.cn/v1` |
|
||||
| Website | [senseaudio.cn](https://senseaudio.cn) |
|
||||
| Docs | [senseaudio.cn/docs](https://senseaudio.cn/docs) |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set your API key">
|
||||
```bash
|
||||
export SENSEAUDIO_API_KEY="..."
|
||||
```
|
||||
</Step>
|
||||
<Step title="Enable the audio provider">
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
enabled: true,
|
||||
models: [{ provider: "senseaudio", model: "senseaudio-asr-pro-1.5-260319" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Send a voice note">
|
||||
Send an audio message through any connected channel. OpenClaw uploads the
|
||||
audio to SenseAudio and uses the transcript in the reply pipeline.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Options
|
||||
|
||||
| Option | Path | Description |
|
||||
| ---------- | ------------------------------------- | ----------------------------------- |
|
||||
| `model` | `tools.media.audio.models[].model` | SenseAudio ASR model id |
|
||||
| `language` | `tools.media.audio.models[].language` | Optional language hint |
|
||||
| `prompt` | `tools.media.audio.prompt` | Optional transcription prompt |
|
||||
| `baseUrl` | `tools.media.audio.baseUrl` or model | Override the OpenAI-compatible base |
|
||||
| `headers` | `tools.media.audio.request.headers` | Extra request headers |
|
||||
|
||||
<Note>
|
||||
SenseAudio is batch STT only in OpenClaw. Voice Call realtime transcription
|
||||
continues to use providers with streaming STT support.
|
||||
</Note>
|
||||
|
||||
## Related
|
||||
|
||||
- [Media understanding (audio)](/nodes/audio)
|
||||
- [Model providers](/concepts/model-providers)
|
||||
161
docs/providers/sglang.md
Normal file
161
docs/providers/sglang.md
Normal file
@@ -0,0 +1,161 @@
|
||||
---
|
||||
summary: "Run OpenClaw with SGLang (OpenAI-compatible self-hosted server)"
|
||||
read_when:
|
||||
- You want to run OpenClaw against a local SGLang server
|
||||
- You want OpenAI-compatible /v1 endpoints with your own models
|
||||
title: "SGLang"
|
||||
---
|
||||
|
||||
SGLang serves open-weight models via an OpenAI-compatible HTTP API. OpenClaw connects to SGLang using the `openai-completions` provider family with auto-discovery of available models.
|
||||
|
||||
| Property | Value |
|
||||
| ------------------------- | ------------------------------------------------------------ |
|
||||
| Provider id | `sglang` |
|
||||
| Plugin | bundled, `enabledByDefault: true` |
|
||||
| Auth env var | `SGLANG_API_KEY` (any non-empty value if server has no auth) |
|
||||
| Onboarding flag | `--auth-choice sglang` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Default base URL | `http://127.0.0.1:30000/v1` |
|
||||
| Default model placeholder | `sglang/Qwen/Qwen3-8B` |
|
||||
| Streaming usage | Yes (`supportsStreamingUsage: true`) |
|
||||
| Pricing | Marked external-free (`modelPricing.external: false`) |
|
||||
|
||||
OpenClaw also **auto-discovers** available models from SGLang when you opt in with `SGLANG_API_KEY`. Use `sglang/*` in `agents.defaults.models` to keep discovery dynamic when you also configure a custom SGLang base URL. See [Model discovery (implicit provider)](#model-discovery-implicit-provider) below.
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Start SGLang">
|
||||
Launch SGLang with an OpenAI-compatible server. Your base URL should expose
|
||||
`/v1` endpoints (for example `/v1/models`, `/v1/chat/completions`). SGLang
|
||||
commonly runs on:
|
||||
|
||||
- `http://127.0.0.1:30000/v1`
|
||||
|
||||
</Step>
|
||||
<Step title="Set an API key">
|
||||
Any value works if no auth is configured on your server:
|
||||
|
||||
```bash
|
||||
export SGLANG_API_KEY="sglang-local"
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Run onboarding or set a model directly">
|
||||
```bash
|
||||
openclaw onboard
|
||||
```
|
||||
|
||||
Or configure the model manually:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "sglang/your-model-id" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Model discovery (implicit provider)
|
||||
|
||||
When `SGLANG_API_KEY` is set (or an auth profile exists) and you **do not**
|
||||
define `models.providers.sglang`, OpenClaw queries:
|
||||
|
||||
- `GET http://127.0.0.1:30000/v1/models`
|
||||
|
||||
and converts the returned IDs into model entries.
|
||||
|
||||
<Note>
|
||||
If you set `models.providers.sglang` explicitly, OpenClaw uses your declared
|
||||
models by default. Add `"sglang/*": {}` to `agents.defaults.models` when you
|
||||
want OpenClaw to query that configured provider's `/models` endpoint and include
|
||||
all advertised SGLang models.
|
||||
</Note>
|
||||
|
||||
## Explicit configuration (manual models)
|
||||
|
||||
Use explicit config when:
|
||||
|
||||
- SGLang runs on a different host/port.
|
||||
- You want to pin `contextWindow`/`maxTokens` values.
|
||||
- Your server requires a real API key (or you want to control headers).
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
sglang: {
|
||||
baseUrl: "http://127.0.0.1:30000/v1",
|
||||
apiKey: "${SGLANG_API_KEY}",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
{
|
||||
id: "your-model-id",
|
||||
name: "Local SGLang Model",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 128000,
|
||||
maxTokens: 8192,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Proxy-style behavior">
|
||||
SGLang is treated as a proxy-style OpenAI-compatible `/v1` backend, not a
|
||||
native OpenAI endpoint.
|
||||
|
||||
| Behavior | SGLang |
|
||||
|----------|--------|
|
||||
| OpenAI-only request shaping | Not applied |
|
||||
| `service_tier`, Responses `store`, prompt-cache hints | Not sent |
|
||||
| Reasoning-compat payload shaping | Not applied |
|
||||
| Hidden attribution headers (`originator`, `version`, `User-Agent`) | Not injected on custom SGLang base URLs |
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Troubleshooting">
|
||||
**Server not reachable**
|
||||
|
||||
Verify the server is running and responding:
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:30000/v1/models
|
||||
```
|
||||
|
||||
**Auth errors**
|
||||
|
||||
If requests fail with auth errors, set a real `SGLANG_API_KEY` that matches
|
||||
your server configuration, or configure the provider explicitly under
|
||||
`models.providers.sglang`.
|
||||
|
||||
<Tip>
|
||||
If you run SGLang without authentication, any non-empty value for
|
||||
`SGLANG_API_KEY` is sufficient to opt in to model discovery.
|
||||
</Tip>
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema including provider entries.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
226
docs/providers/stepfun.md
Normal file
226
docs/providers/stepfun.md
Normal file
@@ -0,0 +1,226 @@
|
||||
---
|
||||
summary: "Use StepFun models with OpenClaw"
|
||||
read_when:
|
||||
- You want StepFun models in OpenClaw
|
||||
- You need StepFun setup guidance
|
||||
title: "StepFun"
|
||||
---
|
||||
|
||||
StepFun ships as an external official plugin (`@openclaw/stepfun-provider`) with two provider ids:
|
||||
|
||||
- `stepfun` for the standard endpoint
|
||||
- `stepfun-plan` for the Step Plan endpoint
|
||||
|
||||
<Warning>
|
||||
Standard and Step Plan are **separate providers** with different endpoints and model ref prefixes (`stepfun/...` vs `stepfun-plan/...`). Use a China key with the `.com` endpoints and a global key with the `.ai` endpoints.
|
||||
</Warning>
|
||||
|
||||
## Install plugin
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/stepfun-provider
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Region and endpoint overview
|
||||
|
||||
| Endpoint | China (`.com`) | Global (`.ai`) |
|
||||
| --------- | -------------------------------------- | ------------------------------------- |
|
||||
| Standard | `https://api.stepfun.com/v1` | `https://api.stepfun.ai/v1` |
|
||||
| Step Plan | `https://api.stepfun.com/step_plan/v1` | `https://api.stepfun.ai/step_plan/v1` |
|
||||
|
||||
Auth env var: `STEPFUN_API_KEY`
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
Standard (`stepfun`):
|
||||
|
||||
| Model ref | Context | Max output | Notes |
|
||||
| ------------------------ | ------- | ---------- | ---------------------- |
|
||||
| `stepfun/step-3.5-flash` | 262,144 | 65,536 | Default standard model |
|
||||
|
||||
Step Plan (`stepfun-plan`):
|
||||
|
||||
| Model ref | Context | Max output | Notes |
|
||||
| ---------------------------------- | ------- | ---------- | -------------------------- |
|
||||
| `stepfun-plan/step-3.5-flash` | 262,144 | 65,536 | Default Step Plan model |
|
||||
| `stepfun-plan/step-3.5-flash-2603` | 262,144 | 65,536 | Additional Step Plan model |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Standard">
|
||||
Best for general-purpose use via the standard StepFun endpoint.
|
||||
|
||||
<Steps>
|
||||
<Step title="Choose your endpoint region">
|
||||
| Auth choice | Endpoint | Region |
|
||||
| -------------------------------- | ----------------------------- | -------------- |
|
||||
| `stepfun-standard-api-key-intl` | `https://api.stepfun.ai/v1` | International |
|
||||
| `stepfun-standard-api-key-cn` | `https://api.stepfun.com/v1` | China |
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice stepfun-standard-api-key-intl
|
||||
```
|
||||
|
||||
China endpoint:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice stepfun-standard-api-key-cn
|
||||
```
|
||||
</Step>
|
||||
<Step title="Non-interactive alternative">
|
||||
```bash
|
||||
openclaw onboard --auth-choice stepfun-standard-api-key-intl \
|
||||
--stepfun-api-key "$STEPFUN_API_KEY"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider stepfun
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
Default model: `stepfun/step-3.5-flash`
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Step Plan">
|
||||
Best for the Step Plan reasoning endpoint.
|
||||
|
||||
<Steps>
|
||||
<Step title="Choose your endpoint region">
|
||||
| Auth choice | Endpoint | Region |
|
||||
| ------------------------------ | ------------------------------------------ | -------------- |
|
||||
| `stepfun-plan-api-key-intl` | `https://api.stepfun.ai/step_plan/v1` | International |
|
||||
| `stepfun-plan-api-key-cn` | `https://api.stepfun.com/step_plan/v1` | China |
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice stepfun-plan-api-key-intl
|
||||
```
|
||||
|
||||
China endpoint:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice stepfun-plan-api-key-cn
|
||||
```
|
||||
</Step>
|
||||
<Step title="Non-interactive alternative">
|
||||
```bash
|
||||
openclaw onboard --auth-choice stepfun-plan-api-key-intl \
|
||||
--stepfun-api-key "$STEPFUN_API_KEY"
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify models are available">
|
||||
```bash
|
||||
openclaw models list --provider stepfun-plan
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
Default model: `stepfun-plan/step-3.5-flash`
|
||||
Alternate model: `stepfun-plan/step-3.5-flash-2603`
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
A single auth flow writes region-matched profiles for both `stepfun` and `stepfun-plan`, so both surfaces are discovered together after one onboarding run.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Full config: Standard provider">
|
||||
```json5
|
||||
{
|
||||
env: { STEPFUN_API_KEY: "your-key" },
|
||||
agents: { defaults: { model: { primary: "stepfun/step-3.5-flash" } } },
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
stepfun: {
|
||||
baseUrl: "https://api.stepfun.ai/v1",
|
||||
api: "openai-completions",
|
||||
apiKey: "${STEPFUN_API_KEY}",
|
||||
models: [
|
||||
{
|
||||
id: "step-3.5-flash",
|
||||
name: "Step 3.5 Flash",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 65536,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Full config: Step Plan provider">
|
||||
```json5
|
||||
{
|
||||
env: { STEPFUN_API_KEY: "your-key" },
|
||||
agents: { defaults: { model: { primary: "stepfun-plan/step-3.5-flash" } } },
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
"stepfun-plan": {
|
||||
baseUrl: "https://api.stepfun.ai/step_plan/v1",
|
||||
api: "openai-completions",
|
||||
apiKey: "${STEPFUN_API_KEY}",
|
||||
models: [
|
||||
{
|
||||
id: "step-3.5-flash",
|
||||
name: "Step 3.5 Flash",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 65536,
|
||||
},
|
||||
{
|
||||
id: "step-3.5-flash-2603",
|
||||
name: "Step 3.5 Flash 2603",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 262144,
|
||||
maxTokens: 65536,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Notes">
|
||||
- `step-3.5-flash-2603` is currently exposed only on `stepfun-plan`.
|
||||
- Use `openclaw models list` and `openclaw models set <provider/model>` to inspect or switch models.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Overview of all providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema for providers, models, and plugins.
|
||||
</Card>
|
||||
<Card title="Models CLI" href="/concepts/models" icon="brain">
|
||||
How to choose and configure models.
|
||||
</Card>
|
||||
<Card title="StepFun Platform" href="https://platform.stepfun.com" icon="globe">
|
||||
StepFun API key management and documentation.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
154
docs/providers/synthetic.md
Normal file
154
docs/providers/synthetic.md
Normal file
@@ -0,0 +1,154 @@
|
||||
---
|
||||
summary: "Use Synthetic's Anthropic-compatible API in OpenClaw"
|
||||
read_when:
|
||||
- You want to use Synthetic as a model provider
|
||||
- You need a Synthetic API key or base URL setup
|
||||
title: "Synthetic"
|
||||
---
|
||||
|
||||
[Synthetic](https://synthetic.new) exposes Anthropic-compatible endpoints.
|
||||
OpenClaw bundles it as the `synthetic` provider and uses the Anthropic
|
||||
Messages API.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ------------------------------------- |
|
||||
| Provider | `synthetic` |
|
||||
| Auth | `SYNTHETIC_API_KEY` |
|
||||
| API | Anthropic Messages |
|
||||
| Base URL | `https://api.synthetic.new/anthropic` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Get a `SYNTHETIC_API_KEY` from your Synthetic account, or let onboarding
|
||||
prompt you for one.
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice synthetic-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the default model">
|
||||
Onboarding sets the default model to:
|
||||
```text
|
||||
synthetic/hf:MiniMaxAI/MiniMax-M2.5
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Warning>
|
||||
OpenClaw's Anthropic client appends `/v1` to the base URL automatically, so use
|
||||
`https://api.synthetic.new/anthropic` (not `/anthropic/v1`). If Synthetic
|
||||
changes its base URL, override `models.providers.synthetic.baseUrl`.
|
||||
</Warning>
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { SYNTHETIC_API_KEY: "sk-..." },
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "synthetic/hf:MiniMaxAI/MiniMax-M2.5" },
|
||||
models: { "synthetic/hf:MiniMaxAI/MiniMax-M2.5": { alias: "MiniMax M2.5" } },
|
||||
},
|
||||
},
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
synthetic: {
|
||||
baseUrl: "https://api.synthetic.new/anthropic",
|
||||
apiKey: "${SYNTHETIC_API_KEY}",
|
||||
api: "anthropic-messages",
|
||||
models: [
|
||||
{
|
||||
id: "hf:MiniMaxAI/MiniMax-M2.5",
|
||||
name: "MiniMax M2.5",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 192000,
|
||||
maxTokens: 65536,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
All Synthetic models use cost `0` (input/output/cache).
|
||||
|
||||
| Model ID | Context window | Max tokens | Reasoning | Input |
|
||||
| ------------------------------------------------------ | -------------- | ---------- | --------- | ------------ |
|
||||
| `hf:MiniMaxAI/MiniMax-M2.5` | 192,000 | 65,536 | no | text |
|
||||
| `hf:moonshotai/Kimi-K2-Thinking` | 256,000 | 8,192 | yes | text |
|
||||
| `hf:zai-org/GLM-4.7` | 198,000 | 128,000 | no | text |
|
||||
| `hf:deepseek-ai/DeepSeek-R1-0528` | 128,000 | 8,192 | no | text |
|
||||
| `hf:deepseek-ai/DeepSeek-V3-0324` | 128,000 | 8,192 | no | text |
|
||||
| `hf:deepseek-ai/DeepSeek-V3.1` | 128,000 | 8,192 | no | text |
|
||||
| `hf:deepseek-ai/DeepSeek-V3.1-Terminus` | 128,000 | 8,192 | no | text |
|
||||
| `hf:deepseek-ai/DeepSeek-V3.2` | 159,000 | 8,192 | no | text |
|
||||
| `hf:meta-llama/Llama-3.3-70B-Instruct` | 128,000 | 8,192 | no | text |
|
||||
| `hf:meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8` | 524,000 | 8,192 | no | text |
|
||||
| `hf:moonshotai/Kimi-K2-Instruct-0905` | 256,000 | 8,192 | no | text |
|
||||
| `hf:moonshotai/Kimi-K2.5` | 256,000 | 8,192 | yes | text + image |
|
||||
| `hf:openai/gpt-oss-120b` | 128,000 | 8,192 | no | text |
|
||||
| `hf:Qwen/Qwen3-235B-A22B-Instruct-2507` | 256,000 | 8,192 | no | text |
|
||||
| `hf:Qwen/Qwen3-Coder-480B-A35B-Instruct` | 256,000 | 8,192 | no | text |
|
||||
| `hf:Qwen/Qwen3-VL-235B-A22B-Instruct` | 250,000 | 8,192 | no | text + image |
|
||||
| `hf:zai-org/GLM-4.5` | 128,000 | 128,000 | no | text |
|
||||
| `hf:zai-org/GLM-4.6` | 198,000 | 128,000 | no | text |
|
||||
| `hf:zai-org/GLM-5` | 256,000 | 128,000 | yes | text + image |
|
||||
| `hf:deepseek-ai/DeepSeek-V3` | 128,000 | 8,192 | no | text |
|
||||
| `hf:Qwen/Qwen3-235B-A22B-Thinking-2507` | 256,000 | 8,192 | yes | text |
|
||||
|
||||
<Tip>
|
||||
Model refs use the form `synthetic/<modelId>`. Use
|
||||
`openclaw models list --provider synthetic` to see all models available on your
|
||||
account.
|
||||
</Tip>
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Model allowlist">
|
||||
If you enable a model allowlist (`agents.defaults.models`), add every
|
||||
Synthetic model you plan to use. Models not in the allowlist are hidden
|
||||
from the agent.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Base URL override">
|
||||
If Synthetic changes its API endpoint, override the base URL:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
synthetic: {
|
||||
baseUrl: "https://new-api.synthetic.new/anthropic",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw still appends `/v1` automatically.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Provider rules, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema including provider settings.
|
||||
</Card>
|
||||
<Card title="Synthetic" href="https://synthetic.new" icon="arrow-up-right-from-square">
|
||||
Synthetic dashboard and API docs.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
136
docs/providers/tencent.md
Normal file
136
docs/providers/tencent.md
Normal file
@@ -0,0 +1,136 @@
|
||||
---
|
||||
summary: "Tencent Cloud TokenHub setup for Hy3 preview"
|
||||
title: "Tencent Cloud (TokenHub)"
|
||||
read_when:
|
||||
- You want to use Tencent Hy3 preview with OpenClaw
|
||||
- You need the TokenHub API key setup
|
||||
---
|
||||
|
||||
Install the official Tencent Cloud provider plugin to access Tencent Hy3 preview through the TokenHub endpoint (`tencent-tokenhub`) using an OpenAI-compatible API.
|
||||
|
||||
| Property | Value |
|
||||
| --------------- | ---------------------------------------- |
|
||||
| Provider id | `tencent-tokenhub` |
|
||||
| Package | `@openclaw/tencent-provider` |
|
||||
| Auth env var | `TOKENHUB_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice tokenhub-api-key` |
|
||||
| Direct CLI flag | `--tokenhub-api-key <key>` |
|
||||
| API | OpenAI-compatible (`openai-completions`) |
|
||||
| Base URL | `https://tokenhub.tencentmaas.com/v1` |
|
||||
| Default model | `tencent-tokenhub/hy3-preview` |
|
||||
|
||||
## Quick start
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the plugin">
|
||||
```bash
|
||||
openclaw plugins install @openclaw/tencent-provider
|
||||
```
|
||||
</Step>
|
||||
<Step title="Create a TokenHub API key">
|
||||
Create an API key in Tencent Cloud TokenHub. If you choose a limited access scope for the key, include **Hy3 preview** in the allowed models.
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
<CodeGroup>
|
||||
|
||||
```bash Onboarding
|
||||
openclaw onboard --auth-choice tokenhub-api-key
|
||||
```
|
||||
|
||||
```bash Direct flag
|
||||
openclaw onboard --non-interactive \
|
||||
--auth-choice tokenhub-api-key \
|
||||
--tokenhub-api-key "$TOKENHUB_API_KEY"
|
||||
```
|
||||
|
||||
```bash Env only
|
||||
export TOKENHUB_API_KEY=...
|
||||
```
|
||||
|
||||
</CodeGroup>
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model">
|
||||
```bash
|
||||
openclaw models list --provider tencent-tokenhub
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Non-interactive setup
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice tokenhub-api-key \
|
||||
--tokenhub-api-key "$TOKENHUB_API_KEY" \
|
||||
--skip-health \
|
||||
--accept-risk
|
||||
```
|
||||
|
||||
<Note>
|
||||
`--accept-risk` is required alongside `--non-interactive`.
|
||||
</Note>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
| Model ref | Name | Input | Context | Max output | Notes |
|
||||
| ------------------------------ | ---------------------- | ----- | ------- | ---------- | -------------------------- |
|
||||
| `tencent-tokenhub/hy3-preview` | Hy3 preview (TokenHub) | text | 256,000 | 64,000 | Default; reasoning-enabled |
|
||||
|
||||
Hy3 preview is Tencent Hunyuan's large MoE language model for reasoning, long-context instruction following, code, and agent workflows. It supports standard chat-completions tool calling plus `reasoning_effort`.
|
||||
|
||||
<Tip>
|
||||
The model id is `hy3-preview`. Do not confuse it with Tencent's `HY-3D-*` models, which are 3D generation APIs and are not the OpenClaw chat model configured by this provider.
|
||||
</Tip>
|
||||
|
||||
## Tiered pricing
|
||||
|
||||
The provider catalog ships tiered cost metadata that scales with input window length, so cost estimates are populated without manual overrides.
|
||||
|
||||
| Input tokens range | Input rate | Output rate | Cache read |
|
||||
| ------------------ | ---------- | ----------- | ---------- |
|
||||
| 0 - 16,000 | 0.176 | 0.587 | 0.059 |
|
||||
| 16,000 - 32,000 | 0.235 | 0.939 | 0.088 |
|
||||
| 32,000+ | 0.293 | 1.173 | 0.117 |
|
||||
|
||||
Rates are per million tokens in USD as advertised by Tencent. Override pricing under `models.providers.tencent-tokenhub` only when you need a different surface.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Endpoint override">
|
||||
OpenClaw's built-in catalog uses Tencent Cloud's `https://tokenhub.tencentmaas.com/v1` endpoint. Override it only if your TokenHub account or region requires a different one:
|
||||
|
||||
```bash
|
||||
openclaw config set models.providers.tencent-tokenhub.baseUrl "https://your-endpoint/v1"
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment availability for the daemon">
|
||||
If the Gateway runs as a managed service (launchd, systemd, Docker), `TOKENHUB_API_KEY` must be visible to that process. Set it in `~/.openclaw/.env` or via `env.shellEnv` so launchd, systemd, or Docker exec environments can read it.
|
||||
|
||||
<Warning>
|
||||
Keys exported only in an interactive shell are not visible to managed gateway processes. Use the env file or config seam for persistent availability.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema including provider settings.
|
||||
</Card>
|
||||
<Card title="Tencent TokenHub" href="https://cloud.tencent.com/product/tokenhub" icon="arrow-up-right-from-square">
|
||||
Tencent Cloud's TokenHub product page.
|
||||
</Card>
|
||||
<Card title="Hy3 preview model card" href="https://huggingface.co/tencent/Hy3-preview" icon="square-poll-horizontal">
|
||||
Tencent Hunyuan Hy3 preview details and benchmarks.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
143
docs/providers/together.md
Normal file
143
docs/providers/together.md
Normal file
@@ -0,0 +1,143 @@
|
||||
---
|
||||
summary: "Together AI setup (auth + model selection)"
|
||||
title: "Together AI"
|
||||
read_when:
|
||||
- You want to use Together AI with OpenClaw
|
||||
- You need the API key env var or CLI auth choice
|
||||
---
|
||||
|
||||
[Together AI](https://together.ai) provides access to leading open-source
|
||||
models including Llama, DeepSeek, Kimi, and more through a unified API.
|
||||
OpenClaw bundles it as the `together` provider.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ----------------------------- |
|
||||
| Provider | `together` |
|
||||
| Auth | `TOGETHER_API_KEY` |
|
||||
| API | OpenAI-compatible |
|
||||
| Base URL | `https://api.together.xyz/v1` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get an API key">
|
||||
Create an API key at
|
||||
[api.together.ai/settings/api-keys](https://api.together.ai/settings/api-keys).
|
||||
</Step>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice together-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: {
|
||||
primary: "together/meta-llama/Llama-3.3-70B-Instruct-Turbo",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
### Non-interactive example
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice together-api-key \
|
||||
--together-api-key "$TOGETHER_API_KEY"
|
||||
```
|
||||
|
||||
<Note>
|
||||
Onboarding sets `together/meta-llama/Llama-3.3-70B-Instruct-Turbo` as the
|
||||
default model.
|
||||
</Note>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
Cost is USD per million tokens.
|
||||
|
||||
| Model ref | Name | Input | Context | Max output | Cost (in/out) | Notes |
|
||||
| -------------------------------------------------- | ---------------------------- | ----------- | ------- | ---------- | ------------- | ------------------- |
|
||||
| `together/meta-llama/Llama-3.3-70B-Instruct-Turbo` | Llama 3.3 70B Instruct Turbo | text | 131,072 | 8,192 | 0.88 / 0.88 | Default model |
|
||||
| `together/moonshotai/Kimi-K2.6` | Kimi K2.6 FP4 | text, image | 262,144 | 32,768 | 1.20 / 4.50 | Reasoning model |
|
||||
| `together/deepseek-ai/DeepSeek-V4-Pro` | DeepSeek V4 Pro | text | 512,000 | 8,192 | 2.10 / 4.40 | Reasoning model |
|
||||
| `together/Qwen/Qwen2.5-7B-Instruct-Turbo` | Qwen2.5 7B Instruct Turbo | text | 32,768 | 8,192 | 0.30 / 0.30 | Fast, non-reasoning |
|
||||
| `together/zai-org/GLM-5.1` | GLM 5.1 FP4 | text | 202,752 | 8,192 | 1.40 / 4.40 | Reasoning model |
|
||||
|
||||
## Video generation
|
||||
|
||||
The bundled `together` plugin also registers video generation through the
|
||||
shared `video_generate` tool.
|
||||
|
||||
| Property | Value |
|
||||
| -------------------- | ----------------------------------------------------------------------------------------- |
|
||||
| Default video model | `Wan-AI/Wan2.2-T2V-A14B` |
|
||||
| Other models | `Wan-AI/Wan2.2-I2V-A14B`, `minimax/Hailuo-02`, `Kwai/Kling-2.1-Master` |
|
||||
| Modes | text-to-video; image-to-video only with `Wan-AI/Wan2.2-I2V-A14B` (single reference image) |
|
||||
| Duration | 1-10 seconds |
|
||||
| Supported parameters | `size` (parsed as `<width>x<height>`); `aspectRatio`/`resolution` are not read |
|
||||
|
||||
To use Together as the default video provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "together/Wan-AI/Wan2.2-T2V-A14B",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Tip>
|
||||
See [Video generation](/tools/video-generation) for the shared tool parameters,
|
||||
provider selection, and failover behavior.
|
||||
</Tip>
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Environment note">
|
||||
If the Gateway runs as a daemon (launchd/systemd), make sure
|
||||
`TOGETHER_API_KEY` is available to that process (for example, in
|
||||
`~/.openclaw/.env` or via `env.shellEnv`).
|
||||
|
||||
<Warning>
|
||||
Keys set only in your interactive shell are not visible to daemon-managed
|
||||
gateway processes. Use `~/.openclaw/.env` or `env.shellEnv` config for
|
||||
persistent availability.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Troubleshooting">
|
||||
- Verify your key works: `openclaw models list --provider together`
|
||||
- If models are not appearing, confirm the API key is set in the correct
|
||||
environment for your Gateway process.
|
||||
- Model refs use the form `together/<model-id>`.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model providers" href="/concepts/model-providers" icon="layers">
|
||||
Provider rules, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video generation tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full config schema including provider settings.
|
||||
</Card>
|
||||
<Card title="Together AI" href="https://together.ai" icon="arrow-up-right-from-square">
|
||||
Together AI dashboard, API docs, and pricing.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
273
docs/providers/venice.md
Normal file
273
docs/providers/venice.md
Normal file
@@ -0,0 +1,273 @@
|
||||
---
|
||||
summary: "Use Venice AI privacy-focused models in OpenClaw"
|
||||
read_when:
|
||||
- You want privacy-focused inference in OpenClaw
|
||||
- You want Venice AI setup guidance
|
||||
title: "Venice AI"
|
||||
---
|
||||
|
||||
[Venice AI](https://venice.ai) provides privacy-focused inference: open models run
|
||||
with no logging, plus anonymized proxy access to Claude, GPT, Gemini, and Grok.
|
||||
All endpoints are OpenAI-compatible (`/v1`).
|
||||
|
||||
## Privacy modes
|
||||
|
||||
| Mode | Behavior | Models |
|
||||
| -------------- | ---------------------------------------------------------------- | ------------------------------------------------------------- |
|
||||
| **Private** | Prompts/responses are never stored or logged. Ephemeral. | Llama, Qwen, DeepSeek, Kimi, MiniMax, Venice Uncensored, etc. |
|
||||
| **Anonymized** | Proxied through Venice with metadata stripped before forwarding. | Claude, GPT, Gemini, Grok |
|
||||
|
||||
<Warning>
|
||||
Anonymized models are not fully private. Venice strips metadata before forwarding, but the underlying provider (OpenAI, Anthropic, Google, xAI) still processes the request. Use Private models when full privacy is required.
|
||||
</Warning>
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the plugin">
|
||||
```bash
|
||||
openclaw plugins install @openclaw/venice-provider
|
||||
```
|
||||
</Step>
|
||||
<Step title="Get your API key">
|
||||
1. Sign up at [venice.ai](https://venice.ai)
|
||||
2. Go to **Settings > API Keys > Create new key**
|
||||
3. Copy your API key (format: `vapi_xxxxxxxxxxxx`)
|
||||
</Step>
|
||||
<Step title="Configure OpenClaw">
|
||||
<Tabs>
|
||||
<Tab title="Interactive (recommended)">
|
||||
```bash
|
||||
openclaw onboard --auth-choice venice-api-key
|
||||
```
|
||||
|
||||
Prompts for the API key (or reuses an existing `VENICE_API_KEY`), lists available Venice models, and sets your default model.
|
||||
</Tab>
|
||||
<Tab title="Environment variable">
|
||||
```bash
|
||||
export VENICE_API_KEY="vapi_xxxxxxxxxxxx"
|
||||
```
|
||||
</Tab>
|
||||
<Tab title="Non-interactive">
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--auth-choice venice-api-key \
|
||||
--venice-api-key "vapi_xxxxxxxxxxxx"
|
||||
```
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
</Step>
|
||||
<Step title="Verify setup">
|
||||
```bash
|
||||
openclaw agent --model venice/kimi-k2-5 --message "Hello, are you working?"
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Model selection
|
||||
|
||||
- **Default**: `venice/kimi-k2-5` (private, reasoning, vision).
|
||||
- **Strongest anonymized option**: `venice/claude-opus-4-6`.
|
||||
|
||||
```bash
|
||||
openclaw models set venice/kimi-k2-5
|
||||
openclaw models list --all --provider venice
|
||||
```
|
||||
|
||||
You can also run `openclaw configure` and pick **Model/auth provider > Venice AI**.
|
||||
|
||||
<Tip>
|
||||
| Use case | Model | Why |
|
||||
| ------------------------- | ---------------------------------- | ------------------------------------------ |
|
||||
| General chat (default) | `kimi-k2-5` | Strong private reasoning plus vision |
|
||||
| Best overall quality | `claude-opus-4-6` | Strongest anonymized Venice option |
|
||||
| Privacy + coding | `qwen3-coder-480b-a35b-instruct` | Private coding model with large context |
|
||||
| Fast + cheap | `qwen3-4b` | Lightweight reasoning model |
|
||||
| Complex private tasks | `deepseek-v3.2` | Strong reasoning; tool calling disabled |
|
||||
| Uncensored | `venice-uncensored` | No content restrictions |
|
||||
</Tip>
|
||||
|
||||
## Built-in catalog (38 models)
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Private models (26) — fully private, no logging">
|
||||
| Model ID | Name | Context | Notes |
|
||||
| -------------------------------------- | ------------------------------------- | ------- | --------------------------- |
|
||||
| `kimi-k2-5` | Kimi K2.5 | 256k | Default, reasoning, vision |
|
||||
| `kimi-k2-thinking` | Kimi K2 Thinking | 256k | Reasoning |
|
||||
| `llama-3.3-70b` | Llama 3.3 70B | 128k | General |
|
||||
| `llama-3.2-3b` | Llama 3.2 3B | 128k | General |
|
||||
| `hermes-3-llama-3.1-405b` | Hermes 3 Llama 3.1 405B | 128k | General, tools disabled |
|
||||
| `qwen3-235b-a22b-thinking-2507` | Qwen3 235B Thinking | 128k | Reasoning |
|
||||
| `qwen3-235b-a22b-instruct-2507` | Qwen3 235B Instruct | 128k | General |
|
||||
| `qwen3-coder-480b-a35b-instruct` | Qwen3 Coder 480B | 256k | Coding |
|
||||
| `qwen3-coder-480b-a35b-instruct-turbo` | Qwen3 Coder 480B Turbo | 256k | Coding |
|
||||
| `qwen3-5-35b-a3b` | Qwen3.5 35B A3B | 256k | Reasoning, vision |
|
||||
| `qwen3-next-80b` | Qwen3 Next 80B | 256k | General |
|
||||
| `qwen3-vl-235b-a22b` | Qwen3 VL 235B (Vision) | 256k | Vision |
|
||||
| `qwen3-4b` | Venice Small (Qwen3 4B) | 32k | Fast, reasoning |
|
||||
| `deepseek-v3.2` | DeepSeek V3.2 | 160k | Reasoning, tools disabled |
|
||||
| `venice-uncensored` | Venice Uncensored (Dolphin-Mistral) | 32k | Uncensored, tools disabled |
|
||||
| `mistral-31-24b` | Venice Medium (Mistral) | 128k | Vision |
|
||||
| `google-gemma-3-27b-it` | Google Gemma 3 27B Instruct | 198k | Vision |
|
||||
| `openai-gpt-oss-120b` | OpenAI GPT OSS 120B | 128k | General |
|
||||
| `nvidia-nemotron-3-nano-30b-a3b` | NVIDIA Nemotron 3 Nano 30B | 128k | General |
|
||||
| `olafangensan-glm-4.7-flash-heretic` | GLM 4.7 Flash Heretic | 128k | Reasoning |
|
||||
| `zai-org-glm-4.6` | GLM 4.6 | 198k | General |
|
||||
| `zai-org-glm-4.7` | GLM 4.7 | 198k | Reasoning |
|
||||
| `zai-org-glm-4.7-flash` | GLM 4.7 Flash | 128k | Reasoning |
|
||||
| `zai-org-glm-5` | GLM 5 | 198k | Reasoning |
|
||||
| `minimax-m21` | MiniMax M2.1 | 198k | Reasoning |
|
||||
| `minimax-m25` | MiniMax M2.5 | 198k | Reasoning |
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Anonymized models (12) — via Venice proxy">
|
||||
| Model ID | Name | Context | Notes |
|
||||
| -------------------------------- | -------------------------------- | ------- | ---------------------------- |
|
||||
| `claude-opus-4-6` | Claude Opus 4.6 (via Venice) | 1M | Reasoning, vision |
|
||||
| `claude-sonnet-4-6` | Claude Sonnet 4.6 (via Venice) | 1M | Reasoning, vision |
|
||||
| `openai-gpt-54` | GPT-5.4 (via Venice) | 1M | Reasoning, vision |
|
||||
| `openai-gpt-53-codex` | GPT-5.3 Codex (via Venice) | 400k | Reasoning, vision, coding |
|
||||
| `openai-gpt-52` | GPT-5.2 (via Venice) | 256k | Reasoning |
|
||||
| `openai-gpt-52-codex` | GPT-5.2 Codex (via Venice) | 256k | Reasoning, vision, coding |
|
||||
| `openai-gpt-4o-2024-11-20` | GPT-4o (via Venice) | 128k | Vision |
|
||||
| `openai-gpt-4o-mini-2024-07-18` | GPT-4o Mini (via Venice) | 128k | Vision |
|
||||
| `gemini-3-1-pro-preview` | Gemini 3.1 Pro (via Venice) | 1M | Reasoning, vision |
|
||||
| `gemini-3-pro-preview` | Gemini 3 Pro (via Venice) | 198k | Reasoning, vision |
|
||||
| `gemini-3-flash-preview` | Gemini 3 Flash (via Venice) | 256k | Reasoning, vision |
|
||||
| `grok-41-fast` | Grok 4.1 Fast (via Venice) | 1M | Reasoning, vision |
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
Grok-backed Venice models (`grok-41-fast` and similar) get the same tool-schema
|
||||
compat patch as the native xAI provider, since they share the same upstream
|
||||
tool-call format.
|
||||
|
||||
## Model discovery
|
||||
|
||||
The bundled catalog above is a manifest-backed seed list. At runtime OpenClaw
|
||||
refreshes it from the Venice `/models` API and falls back to the seed list if
|
||||
the API is unreachable. The `/models` endpoint is public (no auth needed for
|
||||
listing), but inference requires a valid API key.
|
||||
|
||||
## DeepSeek V4 replay behavior
|
||||
|
||||
If Venice exposes DeepSeek V4 models such as `deepseek-v4-pro` or
|
||||
`deepseek-v4-flash`, OpenClaw fills the required `reasoning_content` replay
|
||||
field on assistant messages when Venice omits it, and strips `thinking`/
|
||||
`reasoning`/`reasoning_effort` from the request payload (Venice rejects
|
||||
DeepSeek's native `thinking` control on these models). This replay fix is
|
||||
separate from the native DeepSeek provider's own thinking controls.
|
||||
|
||||
## Streaming and tool support
|
||||
|
||||
| Feature | Support |
|
||||
| ---------------- | ------------------------------------------------- |
|
||||
| Streaming | All models |
|
||||
| Function calling | Most models; disabled per-model where noted above |
|
||||
| Vision/Images | Models marked "Vision" above |
|
||||
| JSON mode | Via `response_format` |
|
||||
|
||||
## Pricing
|
||||
|
||||
Venice uses a credit-based system. Anonymized models cost roughly the same as
|
||||
direct API pricing plus a small Venice fee. See
|
||||
[venice.ai/pricing](https://venice.ai/pricing) for current rates.
|
||||
|
||||
## Usage examples
|
||||
|
||||
```bash
|
||||
# Default private model
|
||||
openclaw agent --model venice/kimi-k2-5 --message "Quick health check"
|
||||
|
||||
# Claude Opus via Venice (anonymized)
|
||||
openclaw agent --model venice/claude-opus-4-6 --message "Summarize this task"
|
||||
|
||||
# Uncensored model
|
||||
openclaw agent --model venice/venice-uncensored --message "Draft options"
|
||||
|
||||
# Vision model with image
|
||||
openclaw agent --model venice/qwen3-vl-235b-a22b --message "Review attached image"
|
||||
|
||||
# Coding model
|
||||
openclaw agent --model venice/qwen3-coder-480b-a35b-instruct --message "Refactor this function"
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="API key not recognized">
|
||||
```bash
|
||||
echo $VENICE_API_KEY
|
||||
openclaw models list | grep venice
|
||||
```
|
||||
|
||||
Confirm the key starts with `vapi_`.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Model not available">
|
||||
Run `openclaw models list --all --provider venice` to see currently
|
||||
available models; the catalog changes as Venice adds or retires models.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Connection issues">
|
||||
Venice API is at `https://api.venice.ai/api/v1`. Confirm your network allows HTTPS to that host.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Note>
|
||||
More help: [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq).
|
||||
</Note>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Config file example">
|
||||
```json5
|
||||
{
|
||||
env: { VENICE_API_KEY: "vapi_..." },
|
||||
agents: { defaults: { model: { primary: "venice/kimi-k2-5" } } },
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
venice: {
|
||||
baseUrl: "https://api.venice.ai/api/v1",
|
||||
apiKey: "${VENICE_API_KEY}",
|
||||
api: "openai-completions",
|
||||
models: [
|
||||
{
|
||||
id: "kimi-k2-5",
|
||||
name: "Kimi K2.5",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 256000,
|
||||
maxTokens: 65536,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Venice AI" href="https://venice.ai" icon="globe">
|
||||
Venice AI homepage and account signup.
|
||||
</Card>
|
||||
<Card title="API documentation" href="https://docs.venice.ai" icon="book">
|
||||
Venice API reference and developer docs.
|
||||
</Card>
|
||||
<Card title="Pricing" href="https://venice.ai/pricing" icon="credit-card">
|
||||
Current Venice credit rates and plans.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
126
docs/providers/vercel-ai-gateway.md
Normal file
126
docs/providers/vercel-ai-gateway.md
Normal file
@@ -0,0 +1,126 @@
|
||||
---
|
||||
summary: "Vercel AI Gateway setup (auth + model selection)"
|
||||
title: "Vercel AI gateway"
|
||||
read_when:
|
||||
- You want to use Vercel AI Gateway with OpenClaw
|
||||
- You need the API key env var or CLI auth choice
|
||||
---
|
||||
|
||||
The [Vercel AI Gateway](https://vercel.com/ai-gateway) provides a unified API to
|
||||
access hundreds of models through a single endpoint.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | -------------------------------------- |
|
||||
| Provider | `vercel-ai-gateway` |
|
||||
| Package | `@openclaw/vercel-ai-gateway-provider` |
|
||||
| Auth | `AI_GATEWAY_API_KEY` |
|
||||
| API | Anthropic Messages compatible |
|
||||
| Base URL | `https://ai-gateway.vercel.sh` |
|
||||
| Model catalog | Auto-discovered via `/v1/models` |
|
||||
|
||||
<Tip>
|
||||
OpenClaw auto-discovers the Gateway `/v1/models` catalog, so both the
|
||||
`/models vercel-ai-gateway` chat command and
|
||||
`openclaw models list --provider vercel-ai-gateway` include current model
|
||||
refs such as `vercel-ai-gateway/openai/gpt-5.5` and
|
||||
`vercel-ai-gateway/moonshotai/kimi-k2.6`.
|
||||
</Tip>
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Install the plugin">
|
||||
```bash
|
||||
openclaw plugins install @openclaw/vercel-ai-gateway-provider
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set the API key">
|
||||
```bash
|
||||
openclaw onboard --auth-choice ai-gateway-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "vercel-ai-gateway/anthropic/claude-opus-4.6" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider vercel-ai-gateway
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Non-interactive example
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice ai-gateway-api-key \
|
||||
--ai-gateway-api-key "$AI_GATEWAY_API_KEY"
|
||||
```
|
||||
|
||||
## Model ID shorthand
|
||||
|
||||
OpenClaw normalizes Claude shorthand model refs at runtime:
|
||||
|
||||
| Shorthand input | Normalized model ref |
|
||||
| ----------------------------------- | --------------------------------------------- |
|
||||
| `vercel-ai-gateway/claude-opus-4.6` | `vercel-ai-gateway/anthropic/claude-opus-4.6` |
|
||||
| `vercel-ai-gateway/opus-4.6` | `vercel-ai-gateway/anthropic/claude-opus-4-6` |
|
||||
|
||||
<Tip>
|
||||
Use either form in your configuration; OpenClaw resolves the canonical
|
||||
`anthropic/...` ref automatically.
|
||||
</Tip>
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Environment variable for daemon processes">
|
||||
If the OpenClaw Gateway runs as a daemon (launchd/systemd), make sure
|
||||
`AI_GATEWAY_API_KEY` is available to that process.
|
||||
|
||||
<Warning>
|
||||
A key exported only in an interactive shell will not be visible to a
|
||||
launchd/systemd daemon unless that environment is explicitly imported. Set
|
||||
the key in `~/.openclaw/.env` or via `env.shellEnv` to ensure the gateway
|
||||
process can read it.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Provider routing">
|
||||
Vercel AI Gateway routes each request to the upstream provider named in the
|
||||
model ref prefix. For example, `vercel-ai-gateway/anthropic/claude-opus-4.6`
|
||||
routes through Anthropic, `vercel-ai-gateway/openai/gpt-5.5` routes through
|
||||
OpenAI, and `vercel-ai-gateway/moonshotai/kimi-k2.6` routes through
|
||||
MoonshotAI. One `AI_GATEWAY_API_KEY` authenticates all upstream providers.
|
||||
</Accordion>
|
||||
<Accordion title="Thinking levels">
|
||||
`/think` options follow the upstream model prefix when OpenClaw recognizes
|
||||
it. `vercel-ai-gateway/anthropic/...` uses the Claude thinking profile,
|
||||
including the adaptive default for Claude 4.6 models. Trusted
|
||||
`vercel-ai-gateway/openai/...` refs (`gpt-5.2` and newer, plus Codex
|
||||
variants down to `gpt-5.1-codex`) expose `/think xhigh`. Other namespaced
|
||||
refs keep the standard reasoning levels unless their catalog metadata
|
||||
declares more.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
General troubleshooting and FAQ.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
359
docs/providers/vllm.md
Normal file
359
docs/providers/vllm.md
Normal file
@@ -0,0 +1,359 @@
|
||||
---
|
||||
summary: "Run OpenClaw with vLLM (OpenAI-compatible local server)"
|
||||
read_when:
|
||||
- You want to run OpenClaw against a local vLLM server
|
||||
- You want OpenAI-compatible /v1 endpoints with your own models
|
||||
title: "vLLM"
|
||||
---
|
||||
|
||||
vLLM serves open-source (and some custom) models through an **OpenAI-compatible** HTTP API. OpenClaw connects using the `openai-completions` API and can **auto-discover** models when you opt in with `VLLM_API_KEY`.
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | ------------------------------------------ |
|
||||
| Provider ID | `vllm` |
|
||||
| API | `openai-completions` (OpenAI-compatible) |
|
||||
| Auth | `VLLM_API_KEY` environment variable |
|
||||
| Default base URL | `http://127.0.0.1:8000/v1` |
|
||||
| Streaming usage | Supported (`stream_options.include_usage`) |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Start vLLM with an OpenAI-compatible server">
|
||||
Your base URL must expose `/v1` endpoints (`/v1/models`, `/v1/chat/completions`). vLLM commonly runs on:
|
||||
|
||||
```text
|
||||
http://127.0.0.1:8000/v1
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Set the API key environment variable">
|
||||
Any non-empty value works if your server does not enforce auth:
|
||||
|
||||
```bash
|
||||
export VLLM_API_KEY="vllm-local"
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Select a model">
|
||||
Replace with one of your vLLM model IDs:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "vllm/your-model-id" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider vllm
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
For non-interactive setup (CI, scripting), pass the base URL, key, and model directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice vllm \
|
||||
--custom-base-url "http://127.0.0.1:8000/v1" \
|
||||
--custom-api-key "vllm-local" \
|
||||
--custom-model-id "your-model-id"
|
||||
```
|
||||
|
||||
</Tip>
|
||||
|
||||
## Model discovery (implicit provider)
|
||||
|
||||
When `VLLM_API_KEY` is set (or an auth profile exists) and `models.providers.vllm` is **not** defined, OpenClaw queries `GET http://127.0.0.1:8000/v1/models` and converts the returned IDs into model entries.
|
||||
|
||||
<Note>
|
||||
If you set `models.providers.vllm` explicitly, OpenClaw uses only your declared models. Add `"vllm/*": {}` to `agents.defaults.models` to make OpenClaw also query that configured provider's `/models` endpoint and include all advertised vLLM models.
|
||||
</Note>
|
||||
|
||||
## Explicit configuration
|
||||
|
||||
Configure explicitly when vLLM runs on a different host or port, you want to pin `contextWindow`/`maxTokens`, your server requires a real API key, or you connect to a trusted loopback, LAN, or Tailscale endpoint:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
vllm: {
|
||||
baseUrl: "http://127.0.0.1:8000/v1",
|
||||
apiKey: "${VLLM_API_KEY}",
|
||||
api: "openai-completions",
|
||||
timeoutSeconds: 300, // Optional: extend request timeout for slow local models
|
||||
models: [
|
||||
{
|
||||
id: "your-model-id",
|
||||
name: "Local vLLM Model",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 128000,
|
||||
maxTokens: 8192,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
To keep the provider dynamic without listing every model, add a wildcard to the visible model catalog:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"vllm/*": {},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Proxy-style behavior">
|
||||
vLLM is treated as a proxy-style OpenAI-compatible `/v1` backend, not a native OpenAI endpoint:
|
||||
|
||||
| Behavior | Applied? |
|
||||
| --------------------------------------- | -------------------------------- |
|
||||
| Native OpenAI request shaping | No |
|
||||
| `service_tier` | Not sent |
|
||||
| Responses `store` | Not sent |
|
||||
| Prompt-cache hints | Not sent |
|
||||
| OpenAI reasoning-compat payload shaping | Not applied |
|
||||
| Hidden OpenClaw attribution headers | Not injected on custom base URLs |
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Qwen thinking controls">
|
||||
For Qwen models, set `compat.thinkingFormat: "qwen-chat-template"` on the model row when the server expects Qwen chat-template kwargs. These models expose a binary `/think` profile (`off`, `on`) because Qwen chat-template thinking is an on/off flag, not an OpenAI-style effort ladder.
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
vllm: {
|
||||
models: [
|
||||
{
|
||||
id: "Qwen/Qwen3-8B",
|
||||
name: "Qwen3 8B",
|
||||
reasoning: true,
|
||||
compat: { thinkingFormat: "qwen-chat-template" },
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
OpenClaw maps `/think off` to:
|
||||
|
||||
```json
|
||||
{
|
||||
"chat_template_kwargs": {
|
||||
"enable_thinking": false,
|
||||
"preserve_thinking": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Non-`off` thinking levels send `enable_thinking: true`. If your endpoint expects DashScope-style top-level flags instead, use `compat.thinkingFormat: "qwen"` to send `enable_thinking` at the request root.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Nemotron 3 thinking controls">
|
||||
For `vllm/nemotron-3-*` models with thinking off, the bundled plugin sends:
|
||||
|
||||
```json
|
||||
{
|
||||
"chat_template_kwargs": {
|
||||
"enable_thinking": false,
|
||||
"force_nonempty_content": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
To customize these values, set `chat_template_kwargs` under the model params. If you also set `params.extra_body.chat_template_kwargs`, that value wins because `extra_body` is the last request-body override.
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"vllm/nemotron-3-super": {
|
||||
params: {
|
||||
chat_template_kwargs: {
|
||||
enable_thinking: false,
|
||||
force_nonempty_content: true,
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Qwen tool calls appear as text">
|
||||
First confirm vLLM was started with the right tool-call parser and chat template for the model. vLLM documents `hermes` for Qwen2.5 models and `qwen3_xml` for Qwen3-Coder models.
|
||||
|
||||
Symptoms: skills/tools never run, the assistant prints raw JSON/XML such as `{"name":"read","arguments":...}`, or vLLM returns an empty `tool_calls` array when OpenClaw sends `tool_choice: "auto"`.
|
||||
|
||||
Some Qwen/vLLM combinations return structured tool calls only when the request uses `tool_choice: "required"`. Force it per model with `params.extra_body`:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"vllm/Qwen-Qwen2.5-Coder-32B-Instruct": {
|
||||
params: {
|
||||
extra_body: {
|
||||
tool_choice: "required",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Replace the model id with the exact id from `openclaw models list --provider vllm`, or apply the same override from the CLI:
|
||||
|
||||
```bash
|
||||
openclaw config set agents.defaults.models '{"vllm/Qwen-Qwen2.5-Coder-32B-Instruct":{"params":{"extra_body":{"tool_choice":"required"}}}}' --strict-json --merge
|
||||
```
|
||||
|
||||
This is an opt-in workaround: it forces every turn with tools to make a tool call, so use it only for a dedicated model entry where that is acceptable. Do not set it as a global default for all vLLM models, and do not pair it with a proxy that converts arbitrary assistant text into executable tool calls.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Custom base URL">
|
||||
If your vLLM server runs on a non-default host or port, set `baseUrl` in the explicit provider config:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
vllm: {
|
||||
baseUrl: "http://192.168.1.50:9000/v1",
|
||||
apiKey: "${VLLM_API_KEY}",
|
||||
api: "openai-completions",
|
||||
timeoutSeconds: 300,
|
||||
models: [
|
||||
{
|
||||
id: "my-custom-model",
|
||||
name: "Remote vLLM Model",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
contextWindow: 64000,
|
||||
maxTokens: 4096,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Slow first response or remote server timeout">
|
||||
For large local models, remote LAN hosts, or tailnet links, set a provider-scoped request timeout:
|
||||
|
||||
```json5
|
||||
{
|
||||
models: {
|
||||
providers: {
|
||||
vllm: {
|
||||
baseUrl: "http://192.168.1.50:8000/v1",
|
||||
apiKey: "${VLLM_API_KEY}",
|
||||
api: "openai-completions",
|
||||
timeoutSeconds: 300,
|
||||
models: [{ id: "your-model-id", name: "Local vLLM Model" }],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
`timeoutSeconds` applies to vLLM model HTTP requests only: connection setup, response headers, body streaming, and the total guarded-fetch abort. It also raises the LLM idle/stream watchdog ceiling above the implicit ~120s default for this provider. Prefer this over increasing `agents.defaults.timeoutSeconds`, which controls the whole agent run.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Server not reachable">
|
||||
Check that the vLLM server is running and accessible:
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:8000/v1/models
|
||||
```
|
||||
|
||||
If you see a connection error, verify the host, port, and that vLLM started in OpenAI-compatible server mode. OpenClaw trusts the exact configured `models.providers.vllm.baseUrl` origin for guarded model requests on loopback, LAN, and Tailscale endpoints. Metadata/link-local origins remain blocked without explicit opt-in. Set `models.providers.vllm.request.allowPrivateNetwork: true` only when vLLM requests must reach another private origin, or `false` to opt out of exact-origin trust.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Auth errors on requests">
|
||||
If requests fail with auth errors, set a real `VLLM_API_KEY` that matches your server configuration, or configure the provider explicitly under `models.providers.vllm`.
|
||||
|
||||
<Tip>
|
||||
If your vLLM server does not enforce auth, any non-empty value for `VLLM_API_KEY` works as an opt-in signal for OpenClaw.
|
||||
</Tip>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="No models discovered">
|
||||
Auto-discovery requires `VLLM_API_KEY` to be set. If you have defined `models.providers.vllm`, OpenClaw uses only your declared models unless `agents.defaults.models` includes `"vllm/*": {}`.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tools render as raw text">
|
||||
If a Qwen model prints JSON/XML tool syntax instead of executing a skill:
|
||||
|
||||
- Start vLLM with the correct parser/template for that model.
|
||||
- Confirm the exact model id with `openclaw models list --provider vllm`.
|
||||
- Add a dedicated per-model `params.extra_body.tool_choice: "required"` override only if `tool_choice: "auto"` still returns empty or text-only tool calls.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Warning>
|
||||
More help: [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq).
|
||||
</Warning>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="OpenAI" href="/providers/openai" icon="bolt">
|
||||
Native OpenAI provider and OpenAI-compatible route behavior.
|
||||
</Card>
|
||||
<Card title="OAuth and auth" href="/gateway/authentication" icon="key">
|
||||
Auth details and credential reuse rules.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
Common issues and how to resolve them.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
184
docs/providers/volcengine.md
Normal file
184
docs/providers/volcengine.md
Normal file
@@ -0,0 +1,184 @@
|
||||
---
|
||||
summary: "Volcano Engine setup (Doubao models, coding endpoints, and Seed Speech TTS)"
|
||||
title: "Volcengine (Doubao)"
|
||||
read_when:
|
||||
- You want to use Volcano Engine or Doubao models with OpenClaw
|
||||
- You need the Volcengine API key setup
|
||||
- You want to use Volcengine Speech text-to-speech
|
||||
---
|
||||
|
||||
The Volcengine provider gives access to Doubao models and third-party models hosted on Volcano Engine, with separate endpoints for general and coding workloads. The same bundled plugin also registers Volcengine Speech as a TTS provider.
|
||||
|
||||
| Detail | Value |
|
||||
| ---------- | ---------------------------------------------------------- |
|
||||
| Providers | `volcengine` (general + TTS), `volcengine-plan` (coding) |
|
||||
| Model auth | `VOLCANO_ENGINE_API_KEY` |
|
||||
| TTS auth | `VOLCENGINE_TTS_API_KEY` or `BYTEPLUS_SEED_SPEECH_API_KEY` |
|
||||
| API | OpenAI-compatible models, BytePlus Seed Speech TTS |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Set the API key">
|
||||
Run interactive onboarding:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice volcengine-api-key
|
||||
```
|
||||
|
||||
This registers both the general (`volcengine`) and coding (`volcengine-plan`) providers from a single API key.
|
||||
|
||||
</Step>
|
||||
<Step title="Set a default model">
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
model: { primary: "volcengine-plan/ark-code-latest" },
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider volcengine
|
||||
openclaw models list --provider volcengine-plan
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
For non-interactive setup (CI, scripting), pass the key directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --non-interactive \
|
||||
--mode local \
|
||||
--auth-choice volcengine-api-key \
|
||||
--volcengine-api-key "$VOLCANO_ENGINE_API_KEY"
|
||||
```
|
||||
|
||||
</Tip>
|
||||
|
||||
## Providers and endpoints
|
||||
|
||||
| Provider | Endpoint | Use case |
|
||||
| ----------------- | ----------------------------------------- | -------------- |
|
||||
| `volcengine` | `ark.cn-beijing.volces.com/api/v3` | General models |
|
||||
| `volcengine-plan` | `ark.cn-beijing.volces.com/api/coding/v3` | Coding models |
|
||||
|
||||
<Note>
|
||||
Both providers are configured from a single API key. Setup registers both automatically, and the coding provider's model picker also reuses the general provider's auth (`volcengine-plan` is an auth alias of `volcengine`).
|
||||
</Note>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
<Tabs>
|
||||
<Tab title="General (volcengine)">
|
||||
| Model ref | Name | Input | Context |
|
||||
| -------------------------------------------- | ------------------------------- | ----------- | ------- |
|
||||
| `volcengine/deepseek-v3-2-251201` | DeepSeek V3.2 | text, image | 128,000 |
|
||||
| `volcengine/doubao-seed-1-8-251228` | Doubao Seed 1.8 | text, image | 256,000 |
|
||||
| `volcengine/doubao-seed-code-preview-251028` | doubao-seed-code-preview-251028 | text, image | 256,000 |
|
||||
| `volcengine/glm-4-7-251222` | GLM 4.7 | text, image | 200,000 |
|
||||
| `volcengine/kimi-k2-5-260127` | Kimi K2.5 | text, image | 256,000 |
|
||||
</Tab>
|
||||
<Tab title="Coding (volcengine-plan)">
|
||||
| Model ref | Name | Input | Context |
|
||||
| ------------------------------------------------- | ------------------------ | ----- | ------- |
|
||||
| `volcengine-plan/ark-code-latest` | Ark Coding Plan | text | 256,000 |
|
||||
| `volcengine-plan/doubao-seed-code` | Doubao Seed Code | text | 256,000 |
|
||||
| `volcengine-plan/doubao-seed-code-preview-251028` | Doubao Seed Code Preview | text | 256,000 |
|
||||
| `volcengine-plan/glm-4.7` | GLM 4.7 Coding | text | 200,000 |
|
||||
| `volcengine-plan/kimi-k2-thinking` | Kimi K2 Thinking | text | 256,000 |
|
||||
| `volcengine-plan/kimi-k2.5` | Kimi K2.5 Coding | text | 256,000 |
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
Both catalogs are static (no `/models` discovery call) and support OpenAI-compatible streamed usage accounting. Tool schemas for both providers automatically drop `minLength`, `maxLength`, `minItems`, `maxItems`, `minContains`, and `maxContains` keywords, since the Volcengine tool-call API rejects them.
|
||||
|
||||
## Text-to-speech
|
||||
|
||||
Volcengine TTS uses the BytePlus Seed Speech HTTP API (`voice.ap-southeast-1.bytepluses.com`) and is configured separately from the OpenAI-compatible Doubao model API key. In the BytePlus console, open Seed Speech > Settings > API Keys, copy the API key, then set:
|
||||
|
||||
```bash
|
||||
export VOLCENGINE_TTS_API_KEY="byteplus_seed_speech_api_key"
|
||||
export VOLCENGINE_TTS_RESOURCE_ID="seed-tts-1.0"
|
||||
```
|
||||
|
||||
Then enable it in `openclaw.json`:
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "volcengine",
|
||||
providers: {
|
||||
volcengine: {
|
||||
apiKey: "byteplus_seed_speech_api_key",
|
||||
voice: "en_female_anna_mars_bigtts",
|
||||
speedRatio: 1.0,
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Available fields under `messages.tts.providers.volcengine`: `apiKey`, `voice`, `speedRatio` (0.2-3.0), `emotion`, `cluster`, `resourceId`, `appKey`, and `baseUrl`. `!emotion=<value>` also works as an inline voice directive when voice-setting overrides are allowed.
|
||||
|
||||
For voice-note targets, OpenClaw requests provider-native `ogg_opus`. For normal audio attachments, it requests `mp3`. Provider aliases `bytedance` and `doubao` also resolve to this speech provider.
|
||||
|
||||
The default resource id is `seed-tts-1.0`, the entitlement BytePlus grants to newly created Seed Speech API keys by default. If your project has TTS 2.0 entitlement, set `VOLCENGINE_TTS_RESOURCE_ID=seed-tts-2.0`.
|
||||
|
||||
<Warning>
|
||||
`VOLCANO_ENGINE_API_KEY` is for the ModelArk/Doubao model endpoints and is not a Seed Speech API key. TTS needs a Seed Speech API key from the BytePlus Speech Console, or a legacy Speech Console AppID/token pair.
|
||||
</Warning>
|
||||
|
||||
Legacy AppID/token auth remains supported for older Speech Console applications:
|
||||
|
||||
```bash
|
||||
export VOLCENGINE_TTS_APPID="speech_app_id"
|
||||
export VOLCENGINE_TTS_TOKEN="speech_access_token"
|
||||
export VOLCENGINE_TTS_CLUSTER="volcano_tts"
|
||||
```
|
||||
|
||||
Other optional TTS env vars: `VOLCENGINE_TTS_VOICE`, `VOLCENGINE_TTS_APP_KEY`, and `VOLCENGINE_TTS_BASE_URL` override the corresponding `messages.tts.providers.volcengine` config fields when set.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Default model after onboarding">
|
||||
`openclaw onboard --auth-choice volcengine-api-key` sets `volcengine-plan/ark-code-latest` as the default model while also registering the general `volcengine` catalog.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Model picker fallback behavior">
|
||||
During onboarding/configure model selection, the Volcengine auth choice prefers both `volcengine/*` and `volcengine-plan/*` rows. If those models are not loaded yet, OpenClaw falls back to the unfiltered catalog instead of showing an empty provider-scoped picker.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Environment variables for daemon processes">
|
||||
If the Gateway runs as a daemon (launchd/systemd), make sure model and TTS env vars such as `VOLCANO_ENGINE_API_KEY`, `VOLCENGINE_TTS_API_KEY`, `BYTEPLUS_SEED_SPEECH_API_KEY`, `VOLCENGINE_TTS_APPID`, and `VOLCENGINE_TTS_TOKEN` are available to that process (for example, in `~/.openclaw/.env` or via `env.shellEnv`).
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Warning>
|
||||
When running OpenClaw as a background service, environment variables set in your interactive shell are not automatically inherited. See the daemon note above.
|
||||
</Warning>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration" href="/gateway/configuration" icon="gear">
|
||||
Full config reference for agents, models, and providers.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
Common issues and debugging steps.
|
||||
</Card>
|
||||
<Card title="FAQ" href="/help/faq" icon="circle-question">
|
||||
Frequently asked questions about OpenClaw setup.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
179
docs/providers/vydra.md
Normal file
179
docs/providers/vydra.md
Normal file
@@ -0,0 +1,179 @@
|
||||
---
|
||||
summary: "Use Vydra image, video, and speech in OpenClaw"
|
||||
read_when:
|
||||
- You want Vydra media generation in OpenClaw
|
||||
- You need Vydra API key setup guidance
|
||||
title: "Vydra"
|
||||
---
|
||||
|
||||
The bundled Vydra plugin adds:
|
||||
|
||||
- Image generation via `vydra/grok-imagine`
|
||||
- Video generation via `vydra/veo3` (text-to-video) and `vydra/kling` (image-to-video)
|
||||
- Speech synthesis via Vydra's ElevenLabs-backed TTS route
|
||||
|
||||
OpenClaw uses the same `VYDRA_API_KEY` for all three capabilities.
|
||||
|
||||
| Property | Value |
|
||||
| --------------- | ------------------------------------------------------------------------- |
|
||||
| Provider id | `vydra` |
|
||||
| Plugin | bundled, `enabledByDefault: true` |
|
||||
| Auth env var | `VYDRA_API_KEY` |
|
||||
| Onboarding flag | `--auth-choice vydra-api-key` |
|
||||
| Direct CLI flag | `--vydra-api-key <key>` |
|
||||
| Contracts | `imageGenerationProviders`, `videoGenerationProviders`, `speechProviders` |
|
||||
| Base URL | `https://www.vydra.ai/api/v1` (use the `www` host) |
|
||||
|
||||
<Warning>
|
||||
Use `https://www.vydra.ai/api/v1` as the base URL. Vydra's apex host (`https://vydra.ai/api/v1`) currently redirects to `www`. Some HTTP clients drop `Authorization` on that cross-host redirect, which turns a valid API key into a misleading auth failure. The bundled plugin normalizes any configured `vydra.ai` base URL to `www.vydra.ai` to avoid that.
|
||||
</Warning>
|
||||
|
||||
## Setup
|
||||
|
||||
<Steps>
|
||||
<Step title="Run interactive onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice vydra-api-key
|
||||
```
|
||||
|
||||
Or set the env var directly:
|
||||
|
||||
```bash
|
||||
export VYDRA_API_KEY="vydra_live_..."
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Choose a default capability">
|
||||
Pick one or more of the capabilities below (image, video, or speech) and apply the matching configuration.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Capabilities
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Image generation">
|
||||
Default and only bundled image model:
|
||||
|
||||
- `vydra/grok-imagine`
|
||||
|
||||
Set it as the default image provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "vydra/grok-imagine",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Bundled support is text-to-image only, at most one image per request. Vydra's hosted edit routes expect remote image URLs, and the bundled plugin does not add a Vydra-specific upload bridge.
|
||||
|
||||
<Note>
|
||||
See [Image Generation](/tools/image-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Video generation">
|
||||
Registered video models:
|
||||
|
||||
- `vydra/veo3` for text-to-video (rejects image reference inputs)
|
||||
- `vydra/kling` for image-to-video (requires exactly one remote image URL)
|
||||
|
||||
Set Vydra as the default video provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "vydra/veo3",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- `vydra/kling` rejects local file uploads up front; only a remote image URL reference works.
|
||||
- Vydra's `kling` HTTP route has been inconsistent about whether it requires `image_url` or `video_url`; the bundled provider sends the same remote image URL in both fields.
|
||||
- The bundled plugin stays conservative and does not forward undocumented style knobs such as aspect ratio, resolution, watermark, or generated audio.
|
||||
|
||||
<Note>
|
||||
See [Video Generation](/tools/video-generation) for shared tool parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Video live tests">
|
||||
Provider-specific live coverage:
|
||||
|
||||
```bash
|
||||
OPENCLAW_LIVE_TEST=1 \
|
||||
OPENCLAW_LIVE_VYDRA_VIDEO=1 \
|
||||
pnpm test:live -- extensions/vydra/vydra.live.test.ts
|
||||
```
|
||||
|
||||
The bundled Vydra live file covers:
|
||||
|
||||
- `vydra/veo3` text-to-video
|
||||
- `vydra/kling` image-to-video using a remote image URL
|
||||
|
||||
Override the remote image fixture when needed:
|
||||
|
||||
```bash
|
||||
export OPENCLAW_LIVE_VYDRA_KLING_IMAGE_URL="https://example.com/reference.png"
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Speech synthesis">
|
||||
Set Vydra as the speech provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
provider: "vydra",
|
||||
providers: {
|
||||
vydra: {
|
||||
apiKey: "${VYDRA_API_KEY}",
|
||||
voiceId: "21m00Tcm4TlvDq8ikWAM",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Defaults:
|
||||
|
||||
- Model: `elevenlabs/tts`
|
||||
- Voice id: `21m00Tcm4TlvDq8ikWAM` ("Rachel")
|
||||
|
||||
The bundled plugin exposes this one known-good default voice and returns MP3 audio files.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Provider directory" href="/providers/index" icon="list">
|
||||
Browse all available providers.
|
||||
</Card>
|
||||
<Card title="Image generation" href="/tools/image-generation" icon="image">
|
||||
Shared image tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/config-agents#agent-defaults" icon="gear">
|
||||
Agent defaults and model configuration.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
547
docs/providers/xai.md
Normal file
547
docs/providers/xai.md
Normal file
@@ -0,0 +1,547 @@
|
||||
---
|
||||
summary: "Use xAI Grok models in OpenClaw"
|
||||
read_when:
|
||||
- You want to use Grok models in OpenClaw
|
||||
- You are configuring xAI auth or model ids
|
||||
title: "xAI"
|
||||
---
|
||||
|
||||
OpenClaw ships a bundled `xai` provider plugin for Grok models. The
|
||||
recommended path is Grok OAuth with an eligible SuperGrok or X Premium
|
||||
subscription. Gateway, config, routing, and tools stay local; only Grok
|
||||
requests go to xAI's API.
|
||||
|
||||
OAuth does not require an xAI API key or the Grok Build app. xAI may still
|
||||
show Grok Build on the consent screen because OpenClaw uses xAI's shared
|
||||
OAuth client.
|
||||
|
||||
## Setup
|
||||
|
||||
<Steps>
|
||||
<Step title="New install">
|
||||
Run onboarding with daemon install, then pick xAI/Grok OAuth at the
|
||||
model/auth step:
|
||||
|
||||
```bash
|
||||
openclaw onboard --install-daemon
|
||||
```
|
||||
|
||||
On a VPS or over SSH, select xAI OAuth directly; it uses device-code
|
||||
verification and does not need a localhost callback:
|
||||
|
||||
```bash
|
||||
openclaw onboard --install-daemon --auth-choice xai-oauth
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Existing install">
|
||||
Sign in to xAI only; do not rerun full onboarding just to connect Grok:
|
||||
|
||||
```bash
|
||||
openclaw models auth login --provider xai --method oauth
|
||||
```
|
||||
|
||||
Apply Grok as the default model separately:
|
||||
|
||||
```bash
|
||||
openclaw models set xai/grok-4.3
|
||||
```
|
||||
|
||||
Rerun full onboarding only if you intentionally want to change Gateway,
|
||||
daemon, channel, workspace, or other setup choices.
|
||||
|
||||
</Step>
|
||||
<Step title="API-key path">
|
||||
API-key setup still works for xAI Console keys and for media surfaces
|
||||
that need key-backed provider config:
|
||||
|
||||
```bash
|
||||
openclaw models auth login --provider xai --method api-key
|
||||
export XAI_API_KEY=xai-...
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Pick a model">
|
||||
```json5
|
||||
{
|
||||
agents: { defaults: { model: { primary: "xai/grok-4.3" } } },
|
||||
}
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Note>
|
||||
OpenClaw uses the xAI Responses API as the bundled xAI transport. The same
|
||||
credential from `openclaw models auth login --provider xai --method oauth` or
|
||||
`--method api-key` also powers `web_search` (provider id `grok`), `x_search`,
|
||||
`code_execution`, speech/transcription, and xAI image/video generation. If you
|
||||
store an xAI key under `plugins.entries.xai.config.webSearch.apiKey`, the
|
||||
bundled xAI model provider reuses it as a fallback too.
|
||||
</Note>
|
||||
|
||||
## OAuth troubleshooting
|
||||
|
||||
- For SSH, Docker, VPS, or other remote setups, use
|
||||
`openclaw models auth login --provider xai --method oauth`; it uses
|
||||
device-code verification, not a localhost callback.
|
||||
- If sign-in succeeds but Grok is not the default model, run
|
||||
`openclaw models set xai/grok-4.3`.
|
||||
- Inspect saved xAI auth profiles:
|
||||
|
||||
```bash
|
||||
openclaw models auth list --provider xai
|
||||
openclaw models status
|
||||
```
|
||||
|
||||
- xAI decides which accounts can receive OAuth API tokens. If an account is
|
||||
not eligible, use the API-key path or check the subscription on xAI's side.
|
||||
|
||||
<Tip>
|
||||
Use `xai-oauth` when signing in from SSH, Docker, or a VPS. OpenClaw prints a
|
||||
URL and short code; finish sign-in in any local browser while the remote
|
||||
process polls xAI for the completed token exchange.
|
||||
</Tip>
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
Selectable ids in model pickers. The plugin still resolves older Grok 3,
|
||||
Grok 4, Grok 4 Fast, Grok 4.1 Fast, and Grok Code ids for existing configs;
|
||||
see [legacy compatibility aliases](#legacy-compatibility-aliases).
|
||||
|
||||
| Family | Model ids |
|
||||
| -------------- | ------------------------------------------------------------------------ |
|
||||
| Grok Build 0.1 | `grok-build-0.1` |
|
||||
| Grok 4.3 | `grok-4.3` |
|
||||
| Grok 4.20 Beta | `grok-4.20-beta-latest-reasoning`, `grok-4.20-beta-latest-non-reasoning` |
|
||||
|
||||
<Tip>
|
||||
Use `grok-4.3` for general chat and `grok-build-0.1` for build/coding-focused
|
||||
workloads unless you need a Grok 4.20 beta alias.
|
||||
</Tip>
|
||||
|
||||
## Feature coverage
|
||||
|
||||
The bundled plugin maps xAI's current public API surface onto OpenClaw's
|
||||
shared provider and tool contracts. Capabilities that do not fit the shared
|
||||
contract, such as streaming TTS and realtime voice, are not exposed.
|
||||
|
||||
| xAI capability | OpenClaw surface | Status |
|
||||
| -------------------------- | --------------------------------------- | ------------------------------------------------------------------- |
|
||||
| Chat / Responses | `xai/<model>` model provider | Yes |
|
||||
| Server-side web search | `web_search` provider `grok` | Yes |
|
||||
| Server-side X search | `x_search` tool | Yes |
|
||||
| Server-side code execution | `code_execution` tool | Yes |
|
||||
| Images | `image_generate` | Yes |
|
||||
| Videos | `video_generate` | Yes |
|
||||
| Batch text-to-speech | `messages.tts.provider: "xai"` / `tts` | Yes |
|
||||
| Streaming TTS | - | Not exposed; OpenClaw's TTS contract returns complete audio buffers |
|
||||
| Batch speech-to-text | `tools.media.audio` media understanding | Yes |
|
||||
| Streaming speech-to-text | Voice Call `streaming.provider: "xai"` | Yes |
|
||||
| Realtime voice | - | Not exposed yet; needs a different session/WebSocket contract |
|
||||
| Files / batches | Generic model API compatibility only | Not a first-class OpenClaw tool |
|
||||
|
||||
<Note>
|
||||
OpenClaw uses xAI's REST image/video/TTS/STT APIs for media generation and
|
||||
batch transcription, xAI's streaming STT WebSocket for live voice-call
|
||||
transcription, and the Responses API for chat, search, and code-execution
|
||||
tools.
|
||||
</Note>
|
||||
|
||||
### Fast-mode mappings
|
||||
|
||||
`/fast on` or `agents.defaults.models["xai/<model>"].params.fastMode: true`
|
||||
rewrites native xAI requests as follows:
|
||||
|
||||
| Source model | Fast-mode target |
|
||||
| ------------- | ------------------ |
|
||||
| `grok-3` | `grok-3-fast` |
|
||||
| `grok-3-mini` | `grok-3-mini-fast` |
|
||||
| `grok-4` | `grok-4-fast` |
|
||||
| `grok-4-0709` | `grok-4-fast` |
|
||||
|
||||
### Legacy compatibility aliases
|
||||
|
||||
Legacy aliases normalize to the canonical bundled ids:
|
||||
|
||||
| Legacy alias | Canonical id |
|
||||
| --------------------------------------------------------------------------- | ------------------------------------- |
|
||||
| `grok-code-fast-1`, `grok-code-fast`, `grok-code-fast-1-0825` | `grok-build-0.1` |
|
||||
| `grok-4-fast-reasoning` | `grok-4-fast` |
|
||||
| `grok-4-1-fast-reasoning` | `grok-4-1-fast` |
|
||||
| `grok-4.20-reasoning`, `grok-4.20-experimental-beta-0304-reasoning` | `grok-4.20-beta-latest-reasoning` |
|
||||
| `grok-4.20-non-reasoning`, `grok-4.20-experimental-beta-0304-non-reasoning` | `grok-4.20-beta-latest-non-reasoning` |
|
||||
|
||||
## Features
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Web search">
|
||||
The bundled `grok` web-search provider prefers xAI OAuth, then falls back
|
||||
to `XAI_API_KEY` or a plugin web-search key:
|
||||
|
||||
```bash
|
||||
openclaw models auth login --provider xai --method oauth
|
||||
openclaw config set tools.web.search.provider grok
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Video generation">
|
||||
The bundled `xai` plugin registers video generation through the shared
|
||||
`video_generate` tool.
|
||||
|
||||
- Default video model: `xai/grok-imagine-video`
|
||||
- Modes: text-to-video, image-to-video, reference-image generation, remote
|
||||
video edit, and remote video extension
|
||||
- Aspect ratios: `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `3:2`, `2:3`
|
||||
- Resolutions: `480P`, `720P`
|
||||
- Duration: 1-15 seconds for generation/image-to-video, 1-10 seconds when
|
||||
using `reference_image` roles, 2-10 seconds for extension
|
||||
- Reference-image generation: set `imageRoles` to `reference_image` for
|
||||
every supplied image; xAI accepts up to 7 such images
|
||||
- Default operation timeout: 600 seconds unless `video_generate.timeoutMs`
|
||||
or `agents.defaults.videoGenerationModel.timeoutMs` is set
|
||||
|
||||
<Warning>
|
||||
Local video buffers are not accepted. Use remote `http(s)` URLs for video
|
||||
edit/extend inputs. Image-to-video accepts local image buffers because
|
||||
OpenClaw encodes those as data URLs for xAI.
|
||||
</Warning>
|
||||
|
||||
To use xAI as the default video provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
videoGenerationModel: {
|
||||
primary: "xai/grok-imagine-video",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
See [Video Generation](/tools/video-generation) for shared tool
|
||||
parameters, provider selection, and failover behavior.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Image generation">
|
||||
The bundled `xai` plugin registers image generation through the shared
|
||||
`image_generate` tool.
|
||||
|
||||
- Default image model: `xai/grok-imagine-image`
|
||||
- Additional model: `xai/grok-imagine-image-quality`
|
||||
- Modes: text-to-image and reference-image edit
|
||||
- Reference inputs: one `image` or up to five `images`
|
||||
- Aspect ratios: `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `2:3`, `3:2`
|
||||
- Resolutions: `1K`, `2K`
|
||||
- Count: up to 4 images
|
||||
- Default operation timeout: 600 seconds unless `image_generate.timeoutMs`
|
||||
or `agents.defaults.imageGenerationModel.timeoutMs` is set
|
||||
|
||||
OpenClaw asks xAI for `b64_json` image responses so generated media can be
|
||||
stored and delivered through the normal channel attachment path. Local
|
||||
reference images are converted to data URLs; remote `http(s)` references
|
||||
pass through unchanged.
|
||||
|
||||
To use xAI as the default image provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
imageGenerationModel: {
|
||||
primary: "xai/grok-imagine-image",
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
xAI also documents `quality`, `mask`, `user`, and additional native ratios
|
||||
such as `1:2`, `2:1`, `9:20`, and `20:9`. OpenClaw forwards only the shared
|
||||
cross-provider image controls today; these native-only knobs are not
|
||||
exposed through `image_generate`.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Text-to-speech">
|
||||
The bundled `xai` plugin registers text-to-speech through the shared `tts`
|
||||
provider surface.
|
||||
|
||||
- Voices: `eve`, `ara`, `rex`, `sal`, `leo`, `una`
|
||||
- Default voice: `eve`
|
||||
- Formats: `mp3`, `wav`, `pcm`, `mulaw`, `alaw`
|
||||
- Language: BCP-47 code or `auto`
|
||||
- Speed: provider-native speed override
|
||||
- Native Opus voice-note format is not supported
|
||||
|
||||
To use xAI as the default TTS provider:
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
provider: "xai",
|
||||
providers: {
|
||||
xai: {
|
||||
voiceId: "eve",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
<Note>
|
||||
OpenClaw uses xAI's batch `/v1/tts` endpoint. xAI also offers streaming
|
||||
TTS over WebSocket, but the OpenClaw speech provider contract currently
|
||||
expects a complete audio buffer before reply delivery.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Speech-to-text">
|
||||
The bundled `xai` plugin registers batch speech-to-text through OpenClaw's
|
||||
media-understanding transcription surface.
|
||||
|
||||
- Default model: `grok-stt`
|
||||
- Endpoint: xAI REST `/v1/stt`
|
||||
- Input path: multipart audio file upload
|
||||
- Used wherever inbound audio transcription reads `tools.media.audio`,
|
||||
including Discord voice-channel segments and channel audio attachments
|
||||
|
||||
To force xAI for inbound audio transcription:
|
||||
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
audio: {
|
||||
models: [
|
||||
{
|
||||
type: "provider",
|
||||
provider: "xai",
|
||||
model: "grok-stt",
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Language can be supplied through the shared audio media config or per-call
|
||||
transcription request. Prompt hints are accepted by the shared OpenClaw
|
||||
surface, but the xAI REST STT integration forwards only file, model, and
|
||||
language because those map cleanly to the current public xAI
|
||||
endpoint.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Streaming speech-to-text">
|
||||
The bundled `xai` plugin also registers a realtime transcription provider
|
||||
for live voice-call audio.
|
||||
|
||||
- Endpoint: xAI WebSocket `wss://api.x.ai/v1/stt`
|
||||
- Default encoding: `mulaw`
|
||||
- Default sample rate: `8000`
|
||||
- Default endpointing: `800ms`
|
||||
- Interim transcripts: enabled by default
|
||||
|
||||
Voice Call's Twilio media stream sends G.711 mu-law audio frames, so the
|
||||
xAI provider forwards those frames directly without transcoding:
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
"voice-call": {
|
||||
config: {
|
||||
streaming: {
|
||||
enabled: true,
|
||||
provider: "xai",
|
||||
providers: {
|
||||
xai: {
|
||||
apiKey: "${XAI_API_KEY}",
|
||||
endpointingMs: 800,
|
||||
language: "en",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Provider-owned config lives under
|
||||
`plugins.entries.voice-call.config.streaming.providers.xai`. Supported
|
||||
keys are `apiKey`, `baseUrl`, `sampleRate`, `encoding` (`pcm`, `mulaw`, or
|
||||
`alaw`), `interimResults`, `endpointingMs`, and `language`.
|
||||
|
||||
<Note>
|
||||
This streaming provider is for Voice Call's realtime transcription path.
|
||||
Discord voice records short segments and uses the batch
|
||||
`tools.media.audio` transcription path instead.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="x_search configuration">
|
||||
The bundled xAI plugin exposes `x_search` as an OpenClaw tool for
|
||||
searching X (formerly Twitter) content via Grok.
|
||||
|
||||
Config path: `plugins.entries.xai.config.xSearch`
|
||||
|
||||
| Key | Type | Default | Description |
|
||||
| ----------------- | ------- | ------------------------------ | ------------------------------------- |
|
||||
| `enabled` | boolean | `true` (if key available) | Enable or disable x_search |
|
||||
| `model` | string | `grok-4-1-fast-non-reasoning` | Model used for x_search requests |
|
||||
| `baseUrl` | string | - | xAI Responses base URL override |
|
||||
| `inlineCitations` | boolean | - | Include inline citations in results |
|
||||
| `maxTurns` | number | - | Maximum conversation turns |
|
||||
| `timeoutSeconds` | number | `30` | Request timeout in seconds |
|
||||
| `cacheTtlMinutes` | number | `15` | Cache time-to-live in minutes |
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
xai: {
|
||||
config: {
|
||||
xSearch: {
|
||||
enabled: true,
|
||||
model: "grok-4-1-fast-non-reasoning",
|
||||
baseUrl: "https://api.x.ai/v1",
|
||||
inlineCitations: true,
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Code execution configuration">
|
||||
The bundled xAI plugin exposes `code_execution` as an OpenClaw tool for
|
||||
remote code execution in xAI's sandbox environment.
|
||||
|
||||
Config path: `plugins.entries.xai.config.codeExecution`
|
||||
|
||||
| Key | Type | Default | Description |
|
||||
| ---------------- | ------- | -------------------------- | ---------------------------------------- |
|
||||
| `enabled` | boolean | `true` (if key available) | Enable or disable code execution |
|
||||
| `model` | string | `grok-4-1-fast` | Model used for code execution requests |
|
||||
| `maxTurns` | number | - | Maximum conversation turns |
|
||||
| `timeoutSeconds` | number | `30` | Request timeout in seconds |
|
||||
|
||||
<Note>
|
||||
This is remote xAI sandbox execution, not local [`exec`](/tools/exec).
|
||||
</Note>
|
||||
|
||||
```json5
|
||||
{
|
||||
plugins: {
|
||||
entries: {
|
||||
xai: {
|
||||
config: {
|
||||
codeExecution: {
|
||||
enabled: true,
|
||||
model: "grok-4-1-fast",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Known limits">
|
||||
- xAI auth can use an API key, environment variable, plugin config
|
||||
fallback, or OAuth with an eligible xAI account. OAuth uses device-code
|
||||
verification without a localhost callback. xAI decides which accounts
|
||||
can receive OAuth API tokens, and the consent page may show Grok Build
|
||||
even though OpenClaw does not require the Grok Build app.
|
||||
- OpenClaw does not currently expose the xAI multi-agent model family. xAI
|
||||
serves these models through the Responses API, but they do not accept
|
||||
the client-side or custom tools used by OpenClaw's shared agent loop.
|
||||
See the
|
||||
[xAI multi-agent limitations](https://docs.x.ai/developers/model-capabilities/text/multi-agent#limitations).
|
||||
- xAI Realtime voice is not registered as an OpenClaw provider yet. It
|
||||
needs a different bidirectional voice session contract than batch STT
|
||||
or streaming transcription.
|
||||
- xAI image `quality`, image `mask`, and extra native-only aspect ratios
|
||||
are not exposed until the shared `image_generate` tool has
|
||||
corresponding cross-provider controls.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Advanced notes">
|
||||
- OpenClaw applies xAI-specific tool-schema and tool-call compatibility
|
||||
fixes automatically on the shared runner path.
|
||||
- Native xAI requests default `tool_stream: true`. Set
|
||||
`agents.defaults.models["xai/<model>"].params.tool_stream` to `false`
|
||||
to disable it.
|
||||
- The bundled xAI wrapper strips unsupported strict tool-schema flags and
|
||||
reasoning *effort* payload keys before sending native xAI requests. Only
|
||||
`grok-4.3` / `grok-4.3-*` advertise configurable reasoning effort; all
|
||||
other reasoning-capable xAI models still request
|
||||
`include: ["reasoning.encrypted_content"]` so prior encrypted reasoning
|
||||
can be replayed on follow-up turns.
|
||||
- `web_search`, `x_search`, and `code_execution` are exposed as OpenClaw
|
||||
tools. OpenClaw attaches only the specific xAI built-in each tool needs
|
||||
to that tool's request instead of attaching every native tool to every
|
||||
chat turn.
|
||||
- Grok `web_search` reads `plugins.entries.xai.config.webSearch.baseUrl`.
|
||||
`x_search` reads `plugins.entries.xai.config.xSearch.baseUrl`, then
|
||||
falls back to the Grok web-search base URL.
|
||||
- `x_search` and `code_execution` are owned by the bundled xAI plugin
|
||||
rather than hardcoded into the core model runtime.
|
||||
- `code_execution` is remote xAI sandbox execution, not local
|
||||
[`exec`](/tools/exec).
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Live testing
|
||||
|
||||
The xAI media paths are covered by unit tests and opt-in live suites. Export
|
||||
`XAI_API_KEY` in the process environment before running live probes.
|
||||
|
||||
```bash
|
||||
pnpm test extensions/xai
|
||||
OPENCLAW_LIVE_TEST=1 OPENCLAW_LIVE_TEST_QUIET=1 pnpm test:live -- extensions/xai/xai.live.test.ts
|
||||
OPENCLAW_LIVE_TEST=1 OPENCLAW_LIVE_TEST_QUIET=1 OPENCLAW_LIVE_IMAGE_GENERATION_PROVIDERS=xai pnpm test:live -- test/image-generation.runtime.live.test.ts
|
||||
```
|
||||
|
||||
The provider-specific live file synthesizes normal TTS, telephony-friendly PCM
|
||||
TTS, transcribes audio through xAI batch STT, streams the same PCM through xAI
|
||||
realtime STT, generates text-to-image output, and edits a reference image.
|
||||
The shared image live file verifies the same xAI provider through OpenClaw's
|
||||
runtime selection, fallback, normalization, and media attachment path.
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Video generation" href="/tools/video-generation" icon="video">
|
||||
Shared video tool parameters and provider selection.
|
||||
</Card>
|
||||
<Card title="All providers" href="/providers/index" icon="grid-2">
|
||||
The broader provider overview.
|
||||
</Card>
|
||||
<Card title="Troubleshooting" href="/help/troubleshooting" icon="wrench">
|
||||
Common issues and fixes.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
300
docs/providers/xiaomi.md
Normal file
300
docs/providers/xiaomi.md
Normal file
@@ -0,0 +1,300 @@
|
||||
---
|
||||
summary: "Use Xiaomi MiMo pay-as-you-go and Token Plan models with OpenClaw"
|
||||
read_when:
|
||||
- You want Xiaomi MiMo models in OpenClaw
|
||||
- You need Xiaomi MiMo auth or Token Plan setup
|
||||
title: "Xiaomi MiMo"
|
||||
---
|
||||
|
||||
Xiaomi MiMo is the API platform for **MiMo** models. The bundled `xiaomi`
|
||||
plugin (`enabledByDefault: true`, no install step) registers two text
|
||||
providers plus a speech (TTS) provider:
|
||||
|
||||
- `xiaomi` - pay-as-you-go keys (`sk-...`)
|
||||
- `xiaomi-token-plan` - Token Plan keys (`tp-...`) with regional endpoint presets
|
||||
|
||||
| Property | Value |
|
||||
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Provider ids | `xiaomi` (pay-as-you-go), `xiaomi-token-plan` (Token Plan) |
|
||||
| Auth env vars | `XIAOMI_API_KEY`, `XIAOMI_TOKEN_PLAN_API_KEY` |
|
||||
| Onboarding flags | `--auth-choice xiaomi-api-key`, `--auth-choice xiaomi-token-plan-cn`, `--auth-choice xiaomi-token-plan-sgp`, `--auth-choice xiaomi-token-plan-ams` |
|
||||
| Direct CLI flags | `--xiaomi-api-key <key>`, `--xiaomi-token-plan-api-key <key>` |
|
||||
| API | OpenAI-compatible chat completions (`openai-completions`) |
|
||||
| Speech contract | `speechProviders: ["xiaomi"]` |
|
||||
| Base URLs | Pay-as-you-go: `https://api.xiaomimimo.com/v1`; Token Plan: `token-plan-{cn,sgp,ams}.xiaomimimo.com/v1` |
|
||||
| Default models | `xiaomi/mimo-v2-flash`, `xiaomi-token-plan/mimo-v2.5-pro` |
|
||||
| TTS default | `mimo-v2.5-tts`, voice `mimo_default`; voicedesign model `mimo-v2.5-tts-voicedesign` |
|
||||
|
||||
## Getting started
|
||||
|
||||
<Steps>
|
||||
<Step title="Get the right key">
|
||||
Create a pay-as-you-go key in the [Xiaomi MiMo console](https://platform.xiaomimimo.com/#/console/api-keys), or open your Token Plan subscription page and copy the regional OpenAI-compatible base URL plus the matching `tp-...` key.
|
||||
</Step>
|
||||
|
||||
<Step title="Run onboarding">
|
||||
Pay-as-you-go:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice xiaomi-api-key
|
||||
```
|
||||
|
||||
Token Plan:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice xiaomi-token-plan-sgp
|
||||
```
|
||||
|
||||
Or pass the keys directly:
|
||||
|
||||
```bash
|
||||
openclaw onboard --auth-choice xiaomi-api-key --xiaomi-api-key "$XIAOMI_API_KEY"
|
||||
openclaw onboard --auth-choice xiaomi-token-plan-sgp --xiaomi-token-plan-api-key "$XIAOMI_TOKEN_PLAN_API_KEY"
|
||||
```
|
||||
|
||||
</Step>
|
||||
<Step title="Verify the model is available">
|
||||
```bash
|
||||
openclaw models list --provider xiaomi
|
||||
openclaw models list --provider xiaomi-token-plan
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
Onboarding validates the key shape and warns when a `tp-...` key is entered into the pay-as-you-go path, or an `sk-...` key is entered into the Token Plan path.
|
||||
</Tip>
|
||||
|
||||
## Pay-as-you-go catalog
|
||||
|
||||
| Model ref | Input | Context | Max output | Reasoning | Notes |
|
||||
| ---------------------- | ----------- | --------- | ---------- | --------- | ------------- |
|
||||
| `xiaomi/mimo-v2-flash` | text | 262,144 | 8,192 | No | Default model |
|
||||
| `xiaomi/mimo-v2-pro` | text | 1,048,576 | 32,000 | Yes | Large context |
|
||||
| `xiaomi/mimo-v2-omni` | text, image | 262,144 | 32,000 | Yes | Multimodal |
|
||||
|
||||
## Token Plan catalog
|
||||
|
||||
Choose the Token Plan auth choice that matches the regional base URL shown in Xiaomi's subscription UI:
|
||||
|
||||
| Auth choice | Base URL |
|
||||
| ----------------------- | ------------------------------------------ |
|
||||
| `xiaomi-token-plan-cn` | `https://token-plan-cn.xiaomimimo.com/v1` |
|
||||
| `xiaomi-token-plan-sgp` | `https://token-plan-sgp.xiaomimimo.com/v1` |
|
||||
| `xiaomi-token-plan-ams` | `https://token-plan-ams.xiaomimimo.com/v1` |
|
||||
|
||||
| Model ref | Input | Context | Max output | Reasoning | Notes |
|
||||
| --------------------------------- | ----------- | --------- | ---------- | --------- | ------------- |
|
||||
| `xiaomi-token-plan/mimo-v2.5-pro` | text | 1,048,576 | 131,072 | Yes | Default model |
|
||||
| `xiaomi-token-plan/mimo-v2.5` | text, image | 1,048,576 | 131,072 | Yes | Multimodal |
|
||||
|
||||
`xiaomi-token-plan` needs a regional base URL to resolve. The supported path
|
||||
is a bundled Token Plan onboarding choice or an explicit
|
||||
`models.providers.xiaomi-token-plan` config block with `baseUrl` set; the
|
||||
provider is not offered without one of those.
|
||||
|
||||
## Reasoning models
|
||||
|
||||
`mimo-v2-pro`, `mimo-v2-omni`, `mimo-v2.5`, and `mimo-v2.5-pro` support
|
||||
OpenClaw's [`/think` directive](/tools/thinking) with levels `off`,
|
||||
`minimal`, `low`, `medium`, `high`, `xhigh`, and `max` (default `high`).
|
||||
`mimo-v2-flash` has no reasoning support.
|
||||
|
||||
## Text-to-speech
|
||||
|
||||
The bundled `xiaomi` plugin also registers Xiaomi MiMo as a speech provider
|
||||
for `messages.tts`. It calls Xiaomi's chat-completions TTS contract with the
|
||||
text as an `assistant` message and optional style guidance as a `user`
|
||||
message.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ---------------------------------------- |
|
||||
| TTS id | `xiaomi` (`mimo` alias) |
|
||||
| Auth | `XIAOMI_API_KEY` |
|
||||
| API | `POST /v1/chat/completions` with `audio` |
|
||||
| Default | `mimo-v2.5-tts`, voice `mimo_default` |
|
||||
| Output | MP3 by default; WAV when configured |
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
auto: "always",
|
||||
provider: "xiaomi",
|
||||
providers: {
|
||||
xiaomi: {
|
||||
apiKey: "xiaomi_api_key",
|
||||
model: "mimo-v2.5-tts",
|
||||
speakerVoice: "mimo_default",
|
||||
format: "mp3",
|
||||
style: "Bright, natural, conversational tone.",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Built-in voices: `mimo_default`, `default_zh`, `default_en`, `Mia`, `Chloe`,
|
||||
`Milo`, `Dean`. Preset-voice models (`mimo-v2.5-tts`, `mimo-v2-tts`) use
|
||||
`audio.voice`, so OpenClaw sends `speakerVoice` for those models.
|
||||
|
||||
The voicedesign model `mimo-v2.5-tts-voicedesign` generates the voice from a
|
||||
natural-language style prompt instead of a preset voice id. Set `style` to
|
||||
the desired voice description; OpenClaw sends it as the `user` message, sends
|
||||
the spoken text as the `assistant` message, and omits `audio.voice` for this
|
||||
model.
|
||||
|
||||
```json5
|
||||
{
|
||||
messages: {
|
||||
tts: {
|
||||
provider: "xiaomi",
|
||||
providers: {
|
||||
xiaomi: {
|
||||
model: "mimo-v2.5-tts-voicedesign",
|
||||
format: "wav",
|
||||
style: "Warm, natural female voice with clear pronunciation.",
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
For channels that request a voice-note synthesis target (Discord, Feishu,
|
||||
Matrix, Telegram, and WhatsApp), OpenClaw transcodes Xiaomi output to 48kHz
|
||||
mono Opus with `ffmpeg` before delivery.
|
||||
|
||||
## Config example
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { XIAOMI_API_KEY: "your-key" },
|
||||
agents: { defaults: { model: { primary: "xiaomi/mimo-v2-flash" } } },
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
xiaomi: {
|
||||
baseUrl: "https://api.xiaomimimo.com/v1",
|
||||
api: "openai-completions",
|
||||
apiKey: "XIAOMI_API_KEY",
|
||||
models: [
|
||||
{
|
||||
id: "mimo-v2-flash",
|
||||
name: "Xiaomi MiMo V2 Flash",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
contextWindow: 262144,
|
||||
maxTokens: 8192,
|
||||
},
|
||||
{
|
||||
id: "mimo-v2-pro",
|
||||
name: "Xiaomi MiMo V2 Pro",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
contextWindow: 1048576,
|
||||
maxTokens: 32000,
|
||||
},
|
||||
{
|
||||
id: "mimo-v2-omni",
|
||||
name: "Xiaomi MiMo V2 Omni",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
contextWindow: 262144,
|
||||
maxTokens: 32000,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Pricing and compat flags come from the bundled plugin manifest, so the config example omits `cost` and `compat` to avoid diverging from runtime behavior.
|
||||
|
||||
Token Plan:
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { XIAOMI_TOKEN_PLAN_API_KEY: "tp-your-key" },
|
||||
agents: { defaults: { model: { primary: "xiaomi-token-plan/mimo-v2.5-pro" } } },
|
||||
models: {
|
||||
mode: "merge",
|
||||
providers: {
|
||||
"xiaomi-token-plan": {
|
||||
baseUrl: "https://token-plan-sgp.xiaomimimo.com/v1",
|
||||
api: "openai-completions",
|
||||
apiKey: "XIAOMI_TOKEN_PLAN_API_KEY",
|
||||
models: [
|
||||
{
|
||||
id: "mimo-v2.5-pro",
|
||||
name: "Xiaomi MiMo V2.5 Pro",
|
||||
reasoning: true,
|
||||
input: ["text"],
|
||||
contextWindow: 1048576,
|
||||
maxTokens: 131072,
|
||||
},
|
||||
{
|
||||
id: "mimo-v2.5",
|
||||
name: "Xiaomi MiMo V2.5",
|
||||
reasoning: true,
|
||||
input: ["text", "image"],
|
||||
contextWindow: 1048576,
|
||||
maxTokens: 131072,
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Pricing comes from the bundled manifest (Token Plan models include tiered cache-read pricing), so the config example omits `cost`.
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Auto-injection behavior">
|
||||
The `xiaomi` provider is auto-enabled when `XIAOMI_API_KEY` is set in your environment or an auth profile exists. `xiaomi-token-plan` needs a regional base URL, so the supported path is the bundled Token Plan onboarding choice or an explicit `models.providers.xiaomi-token-plan` config block.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Model details">
|
||||
- **mimo-v2-flash** - lightweight and fast, ideal for general-purpose text tasks. No reasoning support.
|
||||
- **mimo-v2-pro** - supports reasoning with a 1M token context window for long-document workloads.
|
||||
- **mimo-v2-omni** - reasoning-enabled multimodal model that accepts both text and image inputs.
|
||||
- **mimo-v2.5-pro** - Token Plan default with Xiaomi's current V2.5 reasoning stack.
|
||||
- **mimo-v2.5** - Token Plan multimodal V2.5 route.
|
||||
|
||||
<Note>
|
||||
Pay-as-you-go models use the `xiaomi/` prefix. Token Plan models use the `xiaomi-token-plan/` prefix.
|
||||
</Note>
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Troubleshooting">
|
||||
- If models do not appear, confirm the relevant key env var or auth profile is present and valid.
|
||||
- For Token Plan, confirm the chosen onboarding region matches the subscription page base URL and that the key starts with `tp-`.
|
||||
- When the Gateway runs as a daemon, ensure the key is available to that process (for example in `~/.openclaw/.env` or via `env.shellEnv`).
|
||||
|
||||
<Warning>
|
||||
Keys set only in your interactive shell are not visible to daemon-managed gateway processes. Use `~/.openclaw/.env` or `env.shellEnv` config for persistent availability.
|
||||
</Warning>
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Thinking levels" href="/tools/thinking" icon="brain">
|
||||
`/think` directive syntax and level mapping.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full OpenClaw configuration reference.
|
||||
</Card>
|
||||
<Card title="Xiaomi MiMo console" href="https://platform.xiaomimimo.com" icon="arrow-up-right-from-square">
|
||||
Xiaomi MiMo dashboard and API key management.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
266
docs/providers/zai.md
Normal file
266
docs/providers/zai.md
Normal file
@@ -0,0 +1,266 @@
|
||||
---
|
||||
summary: "Use Z.AI (GLM models) with OpenClaw"
|
||||
read_when:
|
||||
- You want Z.AI / GLM models in OpenClaw
|
||||
- You need a simple ZAI_API_KEY setup
|
||||
title: "Z.AI"
|
||||
---
|
||||
|
||||
Z.AI is the API platform for **GLM** models. It provides REST APIs for GLM and
|
||||
uses API keys for authentication. Create your API key in the Z.AI console.
|
||||
OpenClaw uses the `zai` provider with a Z.AI API key.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | -------------------------------------------- |
|
||||
| Provider | `zai` |
|
||||
| Package | `@openclaw/zai-provider` |
|
||||
| Auth | `ZAI_API_KEY` (legacy alias: `Z_AI_API_KEY`) |
|
||||
| API | Z.AI Chat Completions (Bearer auth) |
|
||||
|
||||
## GLM models
|
||||
|
||||
GLM is a model family, not a separate provider. In OpenClaw, GLM models use
|
||||
refs such as `zai/glm-5.2`: provider `zai`, model id `glm-5.2`.
|
||||
|
||||
## Getting started
|
||||
|
||||
Install the provider plugin first:
|
||||
|
||||
```bash
|
||||
openclaw plugins install @openclaw/zai-provider
|
||||
```
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Auto-detect endpoint">
|
||||
**Best for:** most users. OpenClaw probes supported Z.AI endpoints with your API key and applies the correct base URL automatically.
|
||||
|
||||
<Steps>
|
||||
<Step title="Run onboarding">
|
||||
```bash
|
||||
openclaw onboard --auth-choice zai-api-key
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is listed">
|
||||
```bash
|
||||
openclaw models list --all --provider zai
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Explicit regional endpoint">
|
||||
**Best for:** users who want to force a specific Coding Plan or general API surface.
|
||||
|
||||
<Steps>
|
||||
<Step title="Pick the right onboarding choice">
|
||||
```bash
|
||||
# Coding Plan Global (recommended for Coding Plan users)
|
||||
openclaw onboard --auth-choice zai-coding-global
|
||||
|
||||
# Coding Plan CN (China region)
|
||||
openclaw onboard --auth-choice zai-coding-cn
|
||||
|
||||
# General API
|
||||
openclaw onboard --auth-choice zai-global
|
||||
|
||||
# General API CN (China region)
|
||||
openclaw onboard --auth-choice zai-cn
|
||||
```
|
||||
</Step>
|
||||
<Step title="Verify the model is listed">
|
||||
```bash
|
||||
openclaw models list --all --provider zai
|
||||
```
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
### Endpoints
|
||||
|
||||
| Onboarding choice | Base URL | Default model |
|
||||
| ------------------- | --------------------------------------------- | ------------- |
|
||||
| `zai-global` | `https://api.z.ai/api/paas/v4` | `glm-5.1` |
|
||||
| `zai-cn` | `https://open.bigmodel.cn/api/paas/v4` | `glm-5.1` |
|
||||
| `zai-coding-global` | `https://api.z.ai/api/coding/paas/v4` | `glm-5.2` |
|
||||
| `zai-coding-cn` | `https://open.bigmodel.cn/api/coding/paas/v4` | `glm-5.2` |
|
||||
|
||||
`zai-api-key` auto-detects one of these four by probing your key against each
|
||||
endpoint's chat-completions API, checking general endpoints (`zai-global`,
|
||||
then `zai-cn`) before Coding Plan endpoints (`zai-coding-global`, then
|
||||
`zai-coding-cn`), and stopping at the first endpoint that accepts a request.
|
||||
Use an explicit `--auth-choice` to force a Coding Plan endpoint if your key
|
||||
works on both.
|
||||
|
||||
## Config example
|
||||
|
||||
<Tip>
|
||||
`zai-api-key` lets OpenClaw detect the matching Z.AI endpoint from the key and
|
||||
apply the correct base URL automatically. Use the explicit regional choices when
|
||||
you want to force a specific Coding Plan or general API surface.
|
||||
</Tip>
|
||||
|
||||
```json5
|
||||
{
|
||||
env: { ZAI_API_KEY: "sk-..." },
|
||||
models: {
|
||||
providers: {
|
||||
zai: {
|
||||
// GLM-5.2 uses the Coding Plan endpoint.
|
||||
baseUrl: "https://api.z.ai/api/coding/paas/v4",
|
||||
},
|
||||
},
|
||||
},
|
||||
agents: { defaults: { model: { primary: "zai/glm-5.2" } } },
|
||||
}
|
||||
```
|
||||
|
||||
## Built-in catalog
|
||||
|
||||
The `zai` provider plugin ships its catalog in the plugin manifest, so read-only
|
||||
listing can show known GLM rows without loading provider runtime:
|
||||
|
||||
```bash
|
||||
openclaw models list --all --provider zai
|
||||
```
|
||||
|
||||
The manifest-backed catalog currently includes:
|
||||
|
||||
| Model ref | Notes |
|
||||
| -------------------- | ------------------------------- |
|
||||
| `zai/glm-5.2` | Coding Plan default; 1M context |
|
||||
| `zai/glm-5.1` | General API default |
|
||||
| `zai/glm-5` | |
|
||||
| `zai/glm-5-turbo` | |
|
||||
| `zai/glm-5v-turbo` | |
|
||||
| `zai/glm-4.7` | |
|
||||
| `zai/glm-4.7-flash` | |
|
||||
| `zai/glm-4.7-flashx` | |
|
||||
| `zai/glm-4.6` | |
|
||||
| `zai/glm-4.6v` | |
|
||||
| `zai/glm-4.5` | |
|
||||
| `zai/glm-4.5-air` | |
|
||||
| `zai/glm-4.5-flash` | |
|
||||
| `zai/glm-4.5v` | |
|
||||
|
||||
<Tip>
|
||||
GLM models are available as `zai/<model>` (example: `zai/glm-5`).
|
||||
</Tip>
|
||||
|
||||
<Note>
|
||||
Coding Plan setup defaults to `zai/glm-5.2`; general API setup keeps
|
||||
`zai/glm-5.1`. On the Coding Plan endpoints, auto-detection falls back to
|
||||
`glm-5.1` and then `glm-4.7` when the key/plan does not expose GLM-5.2. GLM
|
||||
versions and availability can change; run `openclaw models list --all --provider zai`
|
||||
to see the catalog known to your installed version.
|
||||
</Note>
|
||||
|
||||
## Thinking levels
|
||||
|
||||
<Tabs>
|
||||
<Tab title="GLM-5.2">
|
||||
Full range: `off`, `low`, `high`, `max` (default `off`). OpenClaw maps
|
||||
`low` and `high` to Z.AI's `high` reasoning effort, and `max` to Z.AI's
|
||||
`max` effort, via `reasoning_effort` on the request payload.
|
||||
</Tab>
|
||||
<Tab title="Other GLM models">
|
||||
Binary toggle only: `off` and `low` (shown as `on` in pickers), default
|
||||
`off`. Setting thinking to `off` sends `thinking: { type: "disabled" }`;
|
||||
any other level leaves the request payload untouched (Z.AI's own default
|
||||
reasoning behavior applies).
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
Setting thinking to `off` avoids responses that spend the output budget on
|
||||
`reasoning_content` before visible text.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Forward-resolving unknown GLM-5 models">
|
||||
Unknown `glm-5*` ids still forward-resolve on the provider path by
|
||||
synthesizing provider-owned metadata from the `glm-4.7` template when the id
|
||||
matches the current GLM-5 family shape.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tool-call streaming">
|
||||
`tool_stream` is enabled by default for Z.AI tool-call streaming. To disable it:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"zai/<model>": {
|
||||
params: { tool_stream: false },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Preserved thinking">
|
||||
Preserved thinking is opt-in because Z.AI requires the full historical
|
||||
`reasoning_content` to be replayed, which increases prompt tokens. Enable it
|
||||
per model:
|
||||
|
||||
```json5
|
||||
{
|
||||
agents: {
|
||||
defaults: {
|
||||
models: {
|
||||
"zai/glm-5.2": {
|
||||
params: { preserveThinking: true },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
When enabled and thinking is on, OpenClaw sends
|
||||
`thinking: { type: "enabled", clear_thinking: false }` and replays prior
|
||||
`reasoning_content` for the same OpenAI-compatible transcript. The snake_case
|
||||
`preserve_thinking` param key works as an alias.
|
||||
|
||||
Advanced users can still override the exact provider payload with
|
||||
`params.extra_body.thinking`.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Image understanding">
|
||||
The Z.AI plugin registers image understanding.
|
||||
|
||||
| Property | Value |
|
||||
| ------------- | ----------- |
|
||||
| Model | `glm-4.6v` |
|
||||
|
||||
Image understanding is auto-resolved from the configured Z.AI auth — no
|
||||
additional config is needed.
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Auth details">
|
||||
- Z.AI uses Bearer auth with your API key.
|
||||
- The `zai-api-key` onboarding choice auto-detects the matching Z.AI endpoint by probing supported endpoints with your key.
|
||||
- Use the explicit regional choices (`zai-coding-global`, `zai-coding-cn`, `zai-global`, `zai-cn`) when you want to force a specific API surface.
|
||||
- The legacy env var `Z_AI_API_KEY` is still accepted; OpenClaw copies it to `ZAI_API_KEY` at startup if `ZAI_API_KEY` is unset.
|
||||
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Model selection" href="/concepts/model-providers" icon="layers">
|
||||
Choosing providers, model refs, and failover behavior.
|
||||
</Card>
|
||||
<Card title="Configuration reference" href="/gateway/configuration-reference" icon="gear">
|
||||
Full OpenClaw config schema, including provider and model settings.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
Reference in New Issue
Block a user