Kimi K3 for OpenClaw: API vs Self-Hosted GPU (2026)
Kimi K3 is a 2.8T-parameter Mixture-of-Experts model with a 1M-token context window, released 16 July 2026 with open weights following on 27 July 2026. Kimi API pricing: $0.30 per 1M cached input, $3 per 1M input, $15 per 1M output.
Cost breakeven
An 8×H100 self-host node runs roughly $17,500/month at on-demand pricing. Breakeven against the API sits near 1.17 billion output tokens per month. Below that volume the API wins by a wide margin.
Hardware reality
K3 in INT4 quantization needs ~1.4 TB of VRAM — an 8×H100 80GB or 8×MI300X 192GB node. Consumer GPUs cannot serve the full model; distilled 32B–70B derivatives will be the realistic self-host path for smaller budgets.
Hybrid setup
Route reasoning and long-context requests to the Kimi K3 API; run a small quantized model locally for high-volume or private queries. OpenClaw supports per-model routing rules out of the box.
English | Deutsch | Español | Français | 中文 | Português | Bahasa | Italiano