Weave Router 2.0: Achieve Astra quality at half the cost with intelligent model routing
Open-source model router matches GPT-6 Astra performance on coding tasks while cutting costs 40-70% and improving speed 2-2.5x by routing each request to the optimal LLM.
- Weave Router 2.0 achieves 62.1% pass rate on Terminal-Bench vs. Astra's 60.6% at 52% lower cost.
- Intelligent routing using Hidden Markov Models makes routing decisions in <50ms per request.
- Drop-in proxy: single endpoint change enables cost optimization without client code changes.
Weave Router 2.0, an open-source model routing platform, achieves GPT-6 Astra-level performance on coding tasks while cutting costs by 40-70% and improving speed by 2-2.5x. By intelligently routing each request to the optimal model (Claude Fable, Astra, Gemini, or smaller alternatives), the router eliminates the cost-performance tradeoff that has constrained AI agent adoption.
The problem: one model can't be optimal for everything
As of September 2026, frontier models occupy wildly different points on capability, cost, speed, and context window:
- Claude Opus 5 ($15/MTok input) - highest quality, slowest, most expensive
- GPT-6 Astra ($2/MTok input) - frontier quality, moderate cost
- Claude Fable 5.1 ($0.10/MTok input) - strong coding, 10x cheaper
- GPT-6 Luna ($0.10/MTok input) - budget tier, acceptable quality
No single model maximizes quality, cost, and speed simultaneously. Organizations pick one and live with the tradeoff—or waste money routing everything through the most expensive option.
Weave Router solves this by routing each request type to the appropriate model. A simple question goes to Fable. A complex architecture review goes to Astra. A terminal command validation goes to Luna. The router learns which model serves each task best.
How it works
Weave Router uses a Hidden Markov Model to trace session state, understanding not just what the user is asking now but how the conversation arrived at that point. A classifier then maps sessions to model buckets, making routing decisions in under 50ms per request.
The router presents a universal API supporting Anthropic, OpenAI, Google Gemini, and DeepSeek. Clients point at the router endpoint instead of providers directly—a single-line change. The router handles authentication, caching tradeoffs, cost optimization, and failover.
Performance vs. GPT-6 Astra
Weave Router 2.0 benchmarks against OpenAI's frontier model:
| Benchmark | Metric | Weave Router | GPT-6 Astra | Weave Advantage |
|---|---|---|---|---|
| Terminal-Bench 4.0 | Pass rate | 62.1% | 60.6% | ✓ 1.5% higher |
| Cost per task | $5.22 | $10.03 | ✓ 52% cheaper | |
| Time per task | 20.5 min | 44.1 min | ✓ 2.2x faster | |
| SWE-Atlas Codebase QnA | Pass rate | 62.1% | 66.1% | Astra 4% ahead |
| Cost per task | $2.31 | $5.04 | ✓ 54% cheaper | |
| Time per task | 6.8 min | 16.7 min | ✓ 2.5x faster |
The key insight: Weave Router matches Astra on coding tasks while costing roughly half as much and running 2-2.5x faster. This isn't because Weave is inherently better—it's because routing to the right model for each task is more efficient than using one model for everything.
Enterprise features
- OAuth 2.1 authentication with per-user identity
- Role-based access control and API keys
- Audit logging of every request
- IP allowlists and Cloud Armor integration
- Drop-in proxy: works with any LLM client without code changes
Who benefits
Startups and enterprises building AI coding agents can now deploy production systems at Fable-tier cost while achieving Astra-level quality—a shift that makes widespread agent adoption economically feasible.
Platform teams can offer model routing as a managed service to internal teams, enforcing cost budgets and compliance policies transparently.
Research teams can experiment with model combinations and routing strategies without committing to a single provider's premium tier.
Open source and adoption
Weave Router 2.0 is Apache 2.0 licensed on GitHub (github.com/weave-os/router) with both self-hosted and managed versions. The project ranked #1 Product of the Day on Product Hunt and supports deployment on AWS (EKS), GCP (GKE), and local Docker.
Why it matters: Model routing represents a fundamental shift in AI economics. For two years, organizations choosing between quality and cost made do with one frontier model for all tasks. Weave Router and similar systems now let them achieve both—routing intelligently rather than paying a premium for unused capability. This pricing shift, combined with 2x speed improvements, removes a major barrier to embedding AI agents throughout enterprise software. Expect model routing to become as standard as load balancing.
The Appboxs newsroom covers launches, funding, acquisitions, pricing changes and AI across the SaaS and no-code world. Every story links to its primary sources. Have a tip, a correction or a story we should cover? Send it through our contact page.