Model Stacks for AI Assistants
A model stack is a curated set of AI models that your OpenClaw assistant routes across intelligently. Instead of locking you to one model, the stack hands every message to OpenRouter’s auto-router, which picks the best model from the set for that specific request, a hard coding question might go to the strongest model, a quick reply to a faster, cheaper one. Each stack also has a built-in cost-vs-quality lean, so the routing matches the stack’s intent.
Stacks at a glance
Section titled “Stacks at a glance”Free models only. Explore without spending a credit.
Efficient paid models. Quality stays high, spend stays low.
Balanced quality and speed for everyday work. Includes everything in Hustler.
The best models available. For when quality is everything. Includes everything in Professional.
How we choose models
Section titled “How we choose models”Every stack has different criteria. Before making any changes we cross-reference our own fleet routing telemetry (which models the auto-router actually picks, at what cost), OpenRouter’s category rankings (real-world usage for agent and coding work), each model’s live OpenRouter endpoint specs (pricing, throughput, availability, output-token limits), and independent evaluations from Artificial Analysis.
| Stack | Capability bar | Cost bar |
|---|---|---|
| Wanderer | Solid task completion | Must be :free on OpenRouter |
| Hustler | High success rate, strong value | Lowest cost-per-run with acceptable quality |
| Professional | High success rate, fast throughput | Mid cost, latency matters here |
| Operator | Best available, leads on reasoning | Cost is secondary |
Every candidate is also checked per provider endpoint for output-token limits that would overflow the CX23 VPS — the same model can be safe on one host and unusable on another. See VPS specs and model limits.
Model details and benchmarks
Section titled “Model details and benchmarks”CX23-safe free coding model with tool support. Single Cohere endpoint, so availability is monitored.
Added Aug 2026 (Model Scout) as a faster free rung. Single Nvidia endpoint with variable free-tier uptime.
The strongest free model in the pool. Single endpoint with variable uptime, so it backs up rather than leads.
The resilient free floor and the image path for this stack — the only vision-capable free model in the pool.
The reliable cheap floor. Consistent availability and broad task coverage.
Ultra-low latency, built-in reasoning. Replaced the retired Gemini 2.5 Flash Lite (Model Scout, Aug 2026).
Cheap, capable, vision-capable. Single Alibaba endpoint with a safe 1M/131K context-output profile.
Near-frontier agent quality at reseller prices — one of the two most-used coding models on OpenRouter. Routing avoids its oversized-output endpoints (see VPS limits).
Leads AutomationBench-AA and sits on the intelligence-vs-latency Pareto frontier. Promotional $0.75 / $3.75 per M through end of 2026.
OpenAI's cost-sensitive GPT-5.6 tier. Prices dropped sharply in Aug 2026. Safe 1M/128K context-output profile.
The open-weight flagship behind OpenRouter's "Ox Alpha" stealth launch. Ties Kimi K3 on Artificial Analysis's Intelligence Index.
OpenAI's balanced GPT-5.6 tier for harder coding, reasoning, and agentic work.
The non-OpenAI ceiling, replacing Qwen 3.7 Max. Single Alibaba endpoint; independent benchmark data is still limited.
Replaces Sonnet 4.6 — cheaper and stronger. Frontier instruction-following and complex reasoning across coding and agents.
OpenAI's current flagship for complex professional work, coding, and agentic reasoning. Price dropped in Aug 2026.
The top capability rung and the image model for this stack. Every provider reports a safe 128K max output on its 1M context.
How auto-routing works
Section titled “How auto-routing works”For each message, OpenRouter’s auto-router looks at the models in your stack and picks the best fit for that request, weighing capability against cost using the stack’s lean (Operator leans hard toward quality; Wanderer and Hustler lean toward savings; Professional sits in the middle). Harder tasks pull a stronger model; simple ones get a faster, cheaper one. You don’t configure anything, it happens per message, automatically, within the set you’ve chosen. If the router itself ever has a hiccup, a reliable safety model in the stack catches it, so there’s no downtime.
Credit usage
Section titled “Credit usage”Costs vary by model and message length. Rough per-message estimates:
| Stack | Typical cost per message |
|---|---|
| Wanderer | $0 |
| Hustler | $0.001–$0.005 |
| Professional | $0.003–$0.015 |
| Operator | $0.015–$0.08 |
Check your Dashboard usage tab to monitor actual spend.
Switching stacks
Section titled “Switching stacks”Open your Dashboard, go to Settings, and select a new stack. The change applies immediately, no redeploy needed. You can also ask your bot to use any individual model for a specific conversation without changing your stack.
Model compatibility
Section titled “Model compatibility”Stack models are tested for compatibility with the CX23 VPS before any changes are made. If you use the custom model picker, be aware that some newer 1M context models have a token specification issue that causes immediate overflow errors. See VPS specs and model limits for the full details.