Toolbox/Running agents in production

Gateways & routing

LLM gateways, proxies and model routers

10 repositories, most starred first
State of the category

diegosouzapw/OmniRoute OmniRoute is the de facto leader by adoption and breadth — it gives a single local endpoint that auto-fallbacks across hundreds of providers with quota‑aware routing and token compression. Pick it if your primary goal is maximum reliability and cost-smoothing across many providers; other projects exist because teams often prioritize smaller footprint, extreme performance, local-only privacy, provider automation, or special workflow integrations instead.

Lightweight production
Teams that need a small, easy-to-run, production-ready gateway with virtual keys, spend tracking and guardrails prioritize minimal operational complexity over maximal provider breadth.
High-scale routing
Some deployments need ultra-low-latency, high-throughput routing or massive provider catalogs (and different routing strategies) that trade extra features for raw scale and performance.
Local & privacy-first
Organizations with strict data residency or offline needs require gateways that run entirely locally or integrate local-model runtimes rather than proxying cloud APIs.
Provider automation
Teams looking to access niche or free provider endpoints rely on browser-account rotation, session pooling and automation to unlock models that aren’t available via standard APIs.
Workflow-specific routing
Specialized use cases like IDE/coding integrations, cost-aware answer selection, or human‑in‑the‑loop governance need custom routing and approval flows that a general gateway won’t provide out of the box.
Find your fit

Which one matches your setup?

Answer any of the questions — the shortlist updates as you go. Recommendations come from the capability passports below, nothing else.

Where should reviews happen?
Can code leave your infrastructure?
What matters most?
Model access?
Pick at least one answer to get a shortlist.
Side by side

Comparison matrix

Axes are extracted from each project's docs by our review pipeline; the maturity score is computed from stars, growth and commit activity — not an opinion. Click a column to sort.

Runs inModelsContextCost to run
61.0k +3816/7d
CLIIDECoding-agent pluginWeb appBYOKOpenAIAnthropicGeminiFixed providerEndpoint onlyGuardrails & evalsSelf-hostableFree to start; includes keyless free providers
58.0k +550/7d
CLIWeb appBYOKNo repo contextBasic guardrailsSelf-hostableSelf-hosted — your API keys (pay providers)
12.9k +55/7d
CLIWeb appBYOKOpenAIAnthropicGeminiFixed providerDiff onlyGuardrails & filtersSelf-hostableUses your provider API keys; self-hosted or hosted enterprise plans.
772
CLIIDECoding-agent pluginFixed providerDiff/file onlyNo noise controlsThird-party cloudRequires a Duel API key (dashboard subscription)
724 +6/7d
CLIFixed providerDiff onlyRouting + tagsSelf-hostableFree, self-hosted (no paid cloud required)
626
CLIWeb appCoding-agent pluginIDEBYOKOpenAIAnthropicOllama / localFixed providerDiff onlyBasic filtersRuns fully localYour API keys for upstream providers; local models run locally (no provider fees).
178 +8/7d
CLIWeb appFixed providerDiff onlyContent filter + guardsSelf-hostableFree via Qwen accounts; self-hosted
178 +8/7d
CLIWeb appFixed providerDiff onlyContent filter + spam guardSelf-hostable (Qwen)Free via your Qwen accounts (self-hosted)
29 +4/7d
CLIWeb appBYOKOpenAIAnthropicGeminiNo repo accessConfidence gating & auditSelf-hosted gatewayYour providers' API keys; provider billing applies
28 +4/7d
Web appCLIBYOKRequest context onlyGovernance gatingFully self-hostableFree (open-source, self-hosted)
At a glance

Capability profiles

Six axes, 0–5 each. The shape tells you the strategy: a wide hexagon is a generalist, a spike is a specialist. Showing the 8 most established — the rest are in the full catalog.

ContextNoiseCustomPrivacyModelsMaturity

Provides a single local endpoint that auto-fallbacks across 290+ providers with quota-aware routing and stacked token compression (RTK+Caveman) so you rarely hit limits while saving tokens.

ContextNoiseCustomPrivacyModelsMaturity

A lightweight, production-ready self-hosted AI gateway that unifies 100+ LLM providers into a single OpenAI-compatible API with virtual keys, spend tracking, guardrails, and load balancing.

ContextNoiseCustomPrivacyModelsMaturity

A tiny, high-performance LLM gateway that routes to 1,600+ models while providing built-in guardrails, retries, and load‑balancing.

ContextNoiseCustomPrivacyModelsMaturity

An IDE-native routing layer that runs prompts across multiple models and picks the cheapest answer that still wins.

ContextNoiseCustomPrivacyModelsMaturity

Lets you run Hermes Agent, OpenClaw, and OpenCode simultaneously on a single WeChat account by acting as the sole iLink poller and proxying requests to local gateway endpoints.

ContextNoiseCustomPrivacyModelsMaturity

Multi-provider gateway that routes Claude Code and other coding agents across a broad provider catalog with ordered fallbacks and local-model runtimes.

ContextNoiseCustomPrivacyModelsMaturity

A drop-in OpenAI-compatible gateway that uses browser-automated Qwen accounts to provide free Qwen models with multi-account rotation, session pooling, and streaming support.

ContextNoiseCustomPrivacyModelsMaturity

Provides a drop-in OpenAI-compatible gateway that uses browser automation and multi-account rotation to let you use Qwen models in any OpenAI-compatible client.

All repositories (10)

Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models Kimi, Claude, GP…

61.0k
+3,8167d
TypeScript
6 mo

The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format wit…

58.0k
+5507d
Python
3 yrs

A blazing fast AI Gateway with integrated guardrails. Route to 1,600+ LLMs, 50+ AI Guardrails with 1 fast & fr…

12.9k
+557d
TypeScript
3 yrs

CLI, SDK, and IDE plugins for Duel Agents

772
-87d
TypeScript
3 mo

Run Hermes Agent and OpenClaw on the same WeChat account

724
+67d
Python
4 mo

Open-source multi-provider AI gateway for Claude Code and other coding agents, with model routing, streaming,…

626
Python
1 wk

Drop-in OpenAI-compatible API gateway for Qwen AI models. Use your Qwen account (chat.qwen.ai) as a free AI AP…

178
+87d
TypeScript
3 mo

Drop-in OpenAI-compatible API gateway for Qwen AI models. Use your Qwen account (chat.qwen.ai) as a free AI AP…

178
+87d
TypeScript
3 mo

Enchant Version of 9Router. 304+ Providers (API-key, OAuth, free-tier, and 39 web-cookie providers), 6 combo s…

29
+47d
JavaScript
2 mo

Human as Agent(人即智能体)· Human-as-LLM 人工代理网关:把工程师变成模型,OpenAI 兼容 /v1 接入 Agent 调度池,涉密/需人工任务路由给真实工程师。为 AI Agent 提…

28
+47d
JavaScript
3 wk
Don't want to self-host?

Hosted alternatives

If running your own reviewer is more ops than you want, these managed services cover the same job.

OpenRouter

A hosted gateway to hundreds of models behind one API, with routing and fallbacks handled, so you swap providers without running your own proxy.

Try OpenRouter
Portkey

Managed AI gateway with routing, caching, retries and observability, the hosted counterpart to running an open-source gateway in your stack.

Try Portkey
Vercel AI Gateway

A hosted unified endpoint across providers with fallbacks and usage tracking, so multi-model access is managed rather than proxied by you.

Try Vercel AI Gateway

More in Running agents in production