A curated directory and decision tree for comparing AI gateways and LLM proxies by cost, security, and self-hosting options. The project includes a reproducible cost benchmark, feature matrices, and guidance for selecting the right gateway based on specific needs like cost optimization, compliance, or enterprise requirements.
⚡ Awesome AI Gateway — curated comparison of 100+ AI gateways & LLM proxies (LiteLLM, OpenRouter, Portkey, Kong, Higress, new-api, Bifrost) by cost, security, compliance & self-hosting. Decision tree + reproducible benchmarks. Open source, bilingual, updated daily.
README
Awesome AI Gateway 
Pick the right AI gateway for your need in ~10 seconds — then trust the answer. A decision tree, a reproducible cost benchmark, and independent evidence for what we exclude. Organized by what you actually need, not by vendor.
Built the hard way: I burned $788 on AI coding in a single day — one flagship model ate 78% of it, just because I'd defaulted everything to the priciest option. So I mapped the whole gateway landscape. → the story
Languages: English · 简体中文
Contents
- 🔥 Top gateways (by stars)
- Browse by need
- Decide & compare
- Signal & reference
🔥 Top gateways (by stars)
The most-starred, most-authoritative projects in the list — a fast orientation. Stars auto-refresh daily; full context (features, license, caveats) is in each linked section. ⚠️ flags a project whose coding-subscription / account-pool routing can carry provider-ToS or account-ban risk.
| Gateway | Stars | What it is | Jump to |
|---|---|---|---|
| LiteLLM | ⭐ 54.2k | The default OSS proxy + SDK — OpenAI format to 100+ providers | Self-hosted |
| Kong | ⭐ 43.8k | Mature API gateway with AI plugins (semantic cache, guard) | Enterprise |
| new-api | ⭐ 42.9k | The most active relay/billing hub for teams | China |
| CLIProxyAPI ⚠️ | ⭐ 43.9k | Wraps coding-CLI subscriptions (Claude Code, Codex…) into APIs | Self-hosted |
| Claude Code Router | ⭐ 36k | Route Claude Code and agent CLIs to any model/provider | Smart routing |
| one-api | ⭐ 35.8k | The original LLM API management / distribution system | China |
| sub2api ⚠️ | ⭐ 33.2k | Pools subscription accounts behind one endpoint | China |
| MLflow AI Gateway | ⭐ 27.1k | Unified endpoints + governance in the MLflow platform | Observability |
| 9router ⚠️ | ⭐ 22.9k | BYOK local proxy, subscription→cheap→free fallback | Self-hosted |
| Apache APISIX | ⭐ 16.9k | Cloud-native API + AI gateway (ai-proxy plugins) |
Enterprise |
| aisuite | ⭐ 15k | Andrew Ng's unified multi-provider client (a library) | Self-hosted |
| OmniRoute ⚠️ | ⭐ 22.2k | Coding-agent token-saver across 231+ providers | Self-hosted |
| Portkey Gateway | ⭐ 12.5k | Fast TypeScript gateway, 1,600+ models, 50+ guardrails | Self-hosted |
| Higress | ⭐ 8.9k | Alibaba's AI-native gateway on Envoy/Istio | China |
| NVIDIA Dynamo | ⭐ 7.5k | Datacenter-scale, KV-cache-aware inference routing | K8s |
| Bifrost | ⭐ 6.6k | Go gateway, lowest independently-measured overhead | Self-hosted |
Hosted leaders (SaaS, no GitHub stars): OpenRouter (400+ models, ~5.5% fee) · Vercel AI Gateway & Cloudflare AI Gateway (0% markup) → Cost-first.
Stars measure popularity, not fitness for your need — that's what the category sections and the evidence-based scorecard are for.
💰 Cost-first: cheapest multi-model access
Pain point: "I want many models for the least money and zero ops."
- OpenRouter — The dominant model marketplace: 400+ models behind one OpenAI-compatible API, pay-as-you-go with automatic failover; ~5.5% fee when buying credits. $113M Series B (May 2026), ~8M users.
- Vercel AI Gateway — Hundreds of models at provider list price (0% markup), $5/month free credits, zero-data-retention option; pairs naturally with the AI SDK.
- Cloudflare AI Gateway — Free control plane in front of your own provider keys: caching, dynamic routing, unified billing, and dollar-denominated spend limits (2026 beta).
- Requesty — EU-friendly OpenRouter alternative: 400+ models, sub-20ms failover, ~5% markup.
- Eden AI — Unified API for 500+ models plus vision/OCR/speech; EU-based, ~5.5% platform fee.
- Helicone AI Gateway (cloud) — Passthrough billing at 0% markup with observability bundled.
- GPT-Load ⭐ 6.3k — High-performance Go proxy that rotates pools of API keys across channels to maximize quota usage.
- Loop Gateway — OpenAI-compatible proxy that meters every request in Bitcoin sats instead of dollars. 311 models via OpenRouter at a 15% markup. No accounts, no email, no card; top up over Lightning, get a bearer token. Three auth rails (prepaid bearer, L402, Cashu). Hosted at api.loopxxi.com. New & unverified (anonymous; its public GitHub repo has since been removed, so treat it as a closed hosted relay) — it resells frontier models through the operator's own OpenRouter account at a 15% markup, and account-less + crypto-prepaid means no recourse if it swaps models or vanishes; confirm fidelity with canary_check.py and only top up what you can afford to lose.
- nullsink (repo) — Account-less, metered proxy for frontier-model APIs, paid in Monero or Bitcoin. No accounts, no email, no card; mint a bearer token, prepay on-chain, and point the official SDKs at one base URL. ~10% markup taken once at top-up; no IP logging, no request logs; payment and token kept unlinkable. Self-hostable single binary (TypeScript/Bun, AGPL-3.0), live at nullsink.is. New & unverified (repo created 2026-06, 4★) — account-less + crypto-prepaid + no logs means no recourse if it swaps models or vanishes; confirm fidelity with canary_check.py and only top up what you can afford to lose.
- AIMLAPI — One OpenAI/Anthropic-compatible endpoint fronting 400+ models (chat, image, video, audio, embeddings); prepaid, OpenRouter-style aggregator.
- RunAPI — Hosted OpenAI-compatible relay (
https://runapi.ai/v1) spanning LLM + media-generation (image, video, music/audio) behind one developer account; ships an active CLI and SDKs (runapi-ai) though the gateway itself is closed/hosted. New & unverified (self-submitted) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production. - TeamoRouter — Hosted SaaS gateway speaking OpenAI, Anthropic Messages and Gemini APIs over 500+ providers, with quality/latency/cost-aware "agentic routing" and a channel-dilution check that reroutes when a provider's output quality degrades; native Claude Code / Codex support, MCP, pay-as-you-go plus plans (Alipay/WeChat). New & unverified (community-recommended; the routing / quality-detection behavior is the operator's own description) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
- KeepRouter — OpenAI- and Anthropic-compatible gateway: one key fronts 50+ models (Claude, GPT, Gemini, Mistral, Qwen, Kimi, GLM, DeepSeek, MiMo, MiniMax). Native /v1/messages means it works with Claude Code and the Anthropic SDK, not just the OpenAI SDK. Prepaid pay-as-you-go, billed at cost with 0% token markup (top-up fee 8%+$0.35); one genuinely free $0 model plus trial credit for new accounts. Bilingual EN/简体中文, Merchant of Record Paddle; not available in mainland China. New & unverified — confirm fidelity with canary_check.py before relying on it in production.
- AI快站 (aifast.club) — Operator-submitted OpenAI- and Anthropic-compatible relay aimed at the mainland-China market (Cursor / Claude Code / Codex / Dify integration docs). Its current model list and rates are published live at
/api/ratio_config, with service status at/api/status— read those rather than a fixed snapshot. New & unverified (closed-source, self-submitted) — confirm model fidelity with canary_check.py before relying on it in production. - NovAI — Hosted OpenAI-compatible relay fronting Chinese frontier models (DeepSeek, Qwen, GLM, Kimi, MiniMax, Doubao, Hunyuan) under one key; unlike the other Chinese-model relays here it also exposes image (Doubao Seedream) and video (Doubao Seedance) generation on the same endpoint, per-token pay-as-you-go with trial credit for new accounts. New & unverified (self-submitted) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
- Novita AI — Unified API to 200+ open-source models (DeepSeek/Qwen/Llama…) with load balancing, autoscaling and failover; also a GPU cloud.
- FlintAPI (repo) — Hosted OpenAI-compatible relay over Chinese LLMs (DeepSeek, Qwen, Kimi, GLM, MiniMax); the operator now describes it as a smart-routing engine that dispatches each prompt to a best-suited model rather than a plain aggregator, with free trial credit for new accounts. New and unverified (routing behavior is the operator's own description) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
- FlowBar — Hosted OpenAI-compatible relay reselling dozens of models (GPT, Claude, Gemini, DeepSeek, Qwen, GLM, Kimi) below OpenRouter, with USD/CNY/crypto payment. New and unverified — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
- Meshs One — Hosted OpenAI-compatible relay fronting Chinese frontier models (DeepSeek-V4, Qwen3.7-Max, MiniMax-M3) under one key, per-token pay-as-you-go (its
/v1endpoint returns anew_api_error, so it appears to run on new-api). New & unverified — closed-source and brand-new; confirm model fidelity (e.g. with canary_check.py) before relying on it in production. - CoderPlan — Hosted OpenAI-compatible relay fronting Claude/GPT/Gemini/DeepSeek/Grok for the China market, per-token pay-as-you-go, ¥10 minimum top-up with Alipay/WeChat; Hong Kong/Singapore nodes (API base
api.coderplan.ai/v1, which returns anew_api_error, so it appears to run on new-api). New and unverified — confirm model fidelity (e.g. with canary_check.py) before relying on it in production. - lxg2it ModelRouter (repo) — Solo-built, OpenAI-compatible router over 7+ providers (Anthropic, OpenAI, Google, Cerebras, Groq, Grok, GLM) with tiered automatic fallback that selects the cheapest available model. Free tier plus a paid tier advertised at 0% markup on Anthropic (a deposit fee may apply — verify current pricing). New and unverified — the public repo is a thin, unlicensed core stub still seeing fresh commits (routing logic lives in the closed hosted service; no license file) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
- OpenPaths (repo) — Hosted OpenAI-compatible router auto-routing across 15+ providers (OpenAI, Anthropic, Gemini, Groq, xAI, DeepSeek, Mistral) under one API, spanning chat, image, video, music, speech, embeddings and transcription. New & unverified — despite the "open source" framing the GitHub repo is a no-code, unlicensed marketing mirror (canonical code lives on the third-party Codex Infinity platform and is agent-maintained), so treat it as a closed hosted relay and confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
- Glama Gateway — OpenAI-compatible gateway to 100+ models with consolidated billing, caching and logging (OSS core glama-ai/lightport).
- RouterPlex — Hosted OpenAI-compatible gateway to 25+ models (GPT, Claude, Gemini, DeepSeek, Qwen, Kimi and others) across 11 providers; prepaid, billed per token at published vendor list rates, no subscription. $5 free credit for new accounts. New and unverified, closed-source — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
- TierUp — Hosted OpenAI-compatible gateway exposing four fixed performance tiers (tier-1…tier-4) instead of model names, each mapped server-side to a current best-value model; routes through OpenRouter and prices ~50% under the underlying models' retail, transparently subsidized during an early product-market-fit phase (solo-built, ~zero production users, tier 1 currently free). New & unverified — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
💡 Squeeze more from any gateway: enable semantic caching (Kong, Bifrost, Zuplo), set spend limits (Cloudflare, Zuplo, Pydantic/Logfire), and route easy prompts to cheap models (see Smart routing).
🆓 Free tiers that still work — the verified-limits table
One of the most-asked questions in the ecosystem, and the internet's answers are mostly stale. Every row below was re-verified against the provider's own docs (machine-readable, verified 2026-07-09, CI-enforced ≤30-day re-review). "Unpublished" means the provider now hides the numbers behind a login — we say so instead of quoting third-hand figures.
| Provider | What's free (verified limits) | Notable free models | Card? | The catch |
|---|---|---|---|---|
OpenRouter :free |
50 req/day (<$10 lifetime top-up) → 1,000 req/day ($10+ once); 20 req/min shared across all :free models |
rotating :free pool |
❌ | free routes hit third-party providers that may train on your data (per-provider policy — check the free/paid routing settings) |
| Google Gemini API | Flash-class models free; per-model RPM/RPD unpublished since 2026 (login-only in AI Studio); daily quota resets midnight PT | Gemini 3.5 Flash · 3.1 Flash-Lite · 2.5 Flash · Gemma 4 | ❌ | free-tier content is "used to improve our products" (their pricing page's words) |
| Groq | per-model: Llama-3.3-70B 30 RPM / 1K req/day / 100K tok/day · GPT-OSS-120B 30 RPM / 1K RPD / 200K TPD · Llama-3.1-8B 14.4K RPD / 500K TPD | GPT-OSS-120B/20B · Llama-4-Scout · Llama-3.3-70B · Qwen3-32B | ❌ (card only to upgrade) | daily token caps burn fast on 70B+; limits are org-level |
| Cerebras | 5 RPM / 30K TPM / 1M tokens/day, flat across models | GPT-OSS-120B · GLM-4.7 · Gemma-4-31B | ❔ not stated | 5 RPM = one interactive user; "all models" marketing vs 3 models on the actual table |
| GitHub Models | any GitHub account: low-tier 15 RPM / 150 req/day, high-tier 10 RPM / 50 req/day; 8K in / 4K out per request | GPT-5 · o4-mini · Llama 4 · Phi-4 · DeepSeek-R1 | ❌ | explicitly experimentation-only; the 8K/4K per-request caps rule out long context |
| Cloudflare Workers AI | 10,000 neurons/day (≈4M input tok/day on Llama-3.2-1B, ≈314K on GPT-OSS-120B — our arithmetic from official rates) | Llama-3.3-70B · GPT-OSS-120B/20B · DeepSeek-R1-distill | ❔ not stated | neurons meter compute — output burns 5–10× faster than input |
| Mistral | free Experiment mode; exact limits unpublished (per-account, Admin Console) | Large · Small · Codestral (reported) | ❌ | trains on your data by default — free users must manually toggle it off |
| Cohere | trial key: 1,000 calls/month; Chat 20 RPM | Command A · Command R+ | ❌ | evaluation scale only |
| SambaNova | 20 RPM / 20 req/day / 200K tok/day | DeepSeek-V3.1 · Llama-3.3-70B · GPT-OSS-120B | ❌ | 20 req/day = demo only (very fast inference though) |
| Hugging Face | $0.10/month provider-passthrough credits (PRO: $2/mo) | 200+ provider-routed (DeepSeek-V3 …) | ❌ | $0.10 ≈ a handful of big-model requests |
| Z.ai (GLM) | GLM Flash models $0 in AND out; rate limits unpublished (per-key, login) | GLM-4.7-Flash · GLM-4.5-Flash · GLM-4.6V-Flash | ❔ not stated | opaque shifting concurrency limits; China-HQ provider — weigh your data sensitivity |
Trial credits ≠ free tiers (they expire): NVIDIA build.nvidia.com (1,000 requests at signup, +4,000 with a business email — staff-forum figures; prototyping-only license) and Alibaba Cloud Model Studio international (per-model quotas, hard 90-day expiry, Singapore region). Recently discontinued — ignore stale listicles: Together AI killed its -free models (now $5 minimum prepaid, "does not currently offer free trials"), Moonshot/Kimi requires a $1 top-up to start, and xAI's much-cited data-sharing credits are no longer documented on any public page. Full evidence per row: data/free_tiers.json. A wrong row is a bug — report it.
🔓 Self-hosted open source
Pain point: "My keys, my infra, no per-token middleman fee."
- LiteLLM ⭐ 54.2k — The default choice: Python SDK + proxy server speaking OpenAI format to 100+ providers, with virtual keys, budgets, load balancing and guardrails.
- Portkey Gateway ⭐ 12.5k — Fast TypeScript gateway (1,600+ models, 50+ guardrails) that also powers Portkey's commercial LLMOps platform.
- CLIProxyAPI ⭐ 43.9k — Go gateway that wraps coding-agent CLI subscriptions (Claude Code, Codex, Gemini, Grok, Antigravity) into OpenAI/Gemini/Claude/Codex-compatible APIs with multi-account pools, round-robin load balancing and a management API; one of the highest-starred OSS gateways in the space. BYO accounts — but routing OAuth coding-tier subscriptions through an API can violate provider ToS, so weigh account-ban risk.
- 9router ⭐ 22.9k — MIT self-hosted BYOK local proxy that auto-routes across 40+ providers with subscription→cheap→free fallback, multi-account load balancing and token compression; cost-first and very popular, but its free/OAuth coding-tier routing (Claude Code, Codex, Kiro) carries provider-ToS/account-ban risk.
- OmniRoute ⭐ 22.2k — MIT self-hosted TypeScript gateway: one endpoint to 231+ providers (50+ free), plugging Claude Code / Codex / Cursor / Cline / Copilot into free Claude/GPT/Gemini with stacked token compression (15–95% savings), 17 routing strategies, smart auto-fallback and MCP/A2A. A 2026 breakout of the coding-agent "token-saver" wave — genuine code (not a relay farm), but its free/OAuth coding-tier routing carries provider-ToS/account-ban risk.
- Chat Nio (CoAI) ⭐ 9.2k — Multi-tenant "one-stop" gateway with a built-in admin + credit/subscription billing panel over 200+ models / 35+ providers, priority-based load balancing and model caching — the same commercial-panel genre as the new-api / one-api / VoAPI entries here.
- TensorZero ⭐ 11.7k — ⚠️ Archived June 2026 (company wound down; repo read-only, Apache-2.0 code + community forks remain). Rust gateway unified with observability, evals, experimentation and optimization.
- Bifrost ⭐ 6.6k — Go gateway from Maxim AI claiming ~50x LiteLLM throughput; adaptive load balancing, cluster mode, MCP support.
- Traceloop Hub ⭐ 221 — High-scale gateway written in Rust from the Traceloop team (OpenLLMetry / OTel-for-LLMs); OpenTelemetry-native observability built in.
- Helicone ⭐ 6k — Observability-first platform (YC W23) with a Rust ai-gateway ⭐ 614.
- Plano ⭐ 6.9k — AI-native proxy and data plane for agents (formerly Arch Gateway / archgw).
- AxonHub ⭐ 4.7k — Go gateway: call 100+ LLMs from any SDK behind one OpenAI/Anthropic-compatible endpoint, with built-in failover, load balancing, cost control and end-to-end tracing. BYOK self-hosted.
- Manifest ⭐ 7.3k — Self-hosted TypeScript router (MIT): one OpenAI-compatible
/autoendpoint (plus/v1/messagesfor Anthropic clients) over 300+ models / 31+ providers, mixing API keys, local models (Ollama/LM Studio) and coding-tier subscriptions with complexity/header-based routing, cost tracking, budgets and failover. BYO accounts — routing OAuth coding-tier subscriptions through an API can carry provider-ToS risk. - LLM Gateway ⭐ 1.4k — Open-source OpenRouter alternative: route, manage and analyze requests across providers.
- APIPark ⭐ 1.8k — Cloud-native LLM API management and distribution platform.
- Pydantic AI Gateway ⭐ 192 — BYOK gateway with cost caps and OTel; ⚠️ repo archived, now folded into Pydantic Logfire.
- OptiLLM ⭐ 4.2k — Optimizing inference proxy that boosts accuracy via test-time compute techniques.
- aisuite ⭐ 15k — Andrew Ng's unified multi-provider client. A library rather than a deployable proxy — fits when you don't want network hops.
- Shepherd Model Gateway (SMG) ⭐ 408 — Engine-agnostic gateway in Rust: one OpenAI/Anthropic-compatible endpoint over vLLM/SGLang/TRT-LLM + cloud providers, with KV-cache-aware routing and WASM plugins.
- RelayPlane ⭐ 192 — MIT, local-first proxy (npm): 11 providers behind one endpoint with per-request cost attribution and hard daily/hourly budget caps.
- SentryNode Gateway ⭐ 0 — Open-core (Apache-2.0) AI proxy for cost governance / FinOps routing: adaptive model routing, budget caps and audit logging. Early-stage; the public repo currently ships a demo scaffold.
- GoModel ⭐ 1k — Lightweight single-binary Go gateway (open-source LiteLLM alternative) exposing one OpenAI/Anthropic-compatible API across 18+ providers with caching, guardrails and usage/cost tracking; fast-growing, though its throughput-vs-LiteLLM figures are vendor-run.
- OpenGateLLM ⭐ 172 — Production-grade open-source GenAI gateway from France's Etalab (powers the government's "Albert" assistant): one OpenAI-compatible API over self-hosted + provider models, with auth, rate limits and usage tracking. Distinct public-sector / EU-sovereignty angle.
- ⚠️ Stale but historically notable: BricksLLM ⭐ 1.2k (PII masking, per-key limits; inactive since early 2025), Glide ⭐ 159 (inactive since 2024).
🏢 Enterprise & compliance
Pain point: "Audit logs, PII redaction, RBAC, on-prem, and the EU AI Act (enforceable Aug 2026)."
- Kong AI Gateway ⭐ 43.8k — Mature API gateway with AI plugins: semantic caching/routing, prompt guard, token rate-limiting; Konnect for managed control plane.
- Apache APISIX ⭐ 16.9k — Cloud-native API + AI gateway with
ai-proxy/ai-proxy-multiplugins. - Envoy AI Gateway ⭐ 1.9k — CNCF-aligned GenAI access on Envoy Gateway, backed by Tetrate and Bloomberg.
- kgateway ⭐ 5.6k — CNCF API/AI gateway, the base of Solo.io's commercial Gloo AI Gateway.
- TrueFoundry AI Gateway — Enterprise gateway with routing, guardrails and RBAC, deployable into your K8s/VPC.
- nexos.ai — Enterprise AI gateway/orchestration from the Nord Security founders (€30M Series A, Oct 2025).
- Tyk AI Studio — AI governance suite: budgets, model catalogs, guardrails on Tyk's gateway.
- Gravitee Agent Mesh — LLM Proxy, MCP Proxy and A2A support inside Gravitee APIM.
- WSO2 AI Gateway — Egress management for LLM traffic: model routing, semantic caching, guardrails.
- F5 AI Gateway — Containerized AI traffic gateway; data-leakage detection via the LeakSignal acquisition (announced Jul 2025).
- IBM API Connect AI Gateway — Policy enforcement, masking and audit for LLM traffic.
- MuleSoft AI / Omni Gateway — Governs LLM, MCP and agent traffic alongside classic APIs.
- Lunar.dev ⭐ 469 — Egress consumption gateway repositioned around MCP/agent governance.
- KrakenD AI Gateway — High-performance, stateless Go API gateway (krakend/krakend-ce ⭐ 2.7k) with an AI proxy + prompt-security layer.
- Broadcom Layer7 AI Gateway — LLM traffic governance, threat protection and quotas on the mature Layer7 API platform.
- Cequence AI Gateway — API-security-first AI gateway: discovery, guardrails and threat protection for LLM/agent traffic.
- Axway Amplify AI Gateway — Centralized control plane on Axway's Amplify platform governing LLM/MCP/agent traffic with business-logic model routing, RBAC, spend caps, prompt-injection controls and RAG integration, from a 10× Gartner MQ API-management Leader.
- Red Hat Connectivity Link — Kubernetes-native gateway (built on the Kuadrant project, successor to 3scale) unifying AI gateway, API management and multicluster connectivity; powers OpenShift AI Models-as-a-Service as the front door governing external and self-hosted LLM endpoints.
- Sensedia AI Gateway — Gartner-recognized APIM vendor's agnostic AI gateway governing LLMs, MCP servers and AI agents with multi-model routing, guardrails, cost controls and observability across a multi-cloud control plane.
- Ambassador Edge Stack — Envoy-based, Kubernetes-native API gateway (OSS core emissary-ingress ⭐ 4.5k) whose AI Gateway layer adds LLM-provider routing, token rate-limiting and fallback — a peer to Kong/Tyk/APISIX in the API-vendor cohort.
☁️ First-party gateways (cloud & model vendors)
Pain point: "We're already committed to one cloud — give us the native path."
- AWS Bedrock — Multi-model access via the unified Converse API, cross-region inference, and AgentCore Gateway for tools/MCP.
- Azure API Management — GenAI gateway — Token limits, semantic caching and load balancing in front of Azure OpenAI / AI Foundry.
- Google Apigee + Vertex AI — LLM gateway patterns on Apigee with Vertex Model Garden as the managed hub.
- Cloudflare AI Gateway — See Cost-first; the strongest free first-party option.
- Vercel AI Gateway — GA, 0% markup, ZDR option; the default for Next.js/AI SDK shops.
- Databricks Unity AI Gateway — Mosaic AI Gateway folded into Unity Catalog, adding agent + MCP governance.
- Tencent Cloud AI Gateway — Tencent's first-party cloud-native intelligent gateway bundling LLM + MCP + Agent gateways with protocol conversion, cost/performance-based routing, and unified access to Hunyuan + third-party models.
🇨🇳 China ecosystem
Pain point: "Domestic models (Qwen/DeepSeek/GLM/Kimi), CNY payment, key distribution & billing for teams."
- new-api ⭐ 42.9k — The most active one-api fork, now a "unified AI model hub": protocol conversion, billing, Rerank/Realtime endpoints. AGPL-3.0.
- one-api ⭐ 35.8k — The original LLM API management & distribution system (OpenAI/Azure/Claude/Gemini/DeepSeek/Doubao…); development has slowed.
- Higress ⭐ 8.9k — Alibaba's AI-native gateway on Envoy/Istio, first-class Tongyi/DeepSeek support; hosted version at higress.ai.
- GPT-Load ⭐ 6.3k — Smart API-key rotation multi-channel proxy in Go.
- one-hub ⭐ 2.9k — one-api fork with better non-OpenAI function calling and stats.
- simple-one-api ⭐ 2.3k — Single binary adapting Qianfan/Spark/Hunyuan/MiniMax/DeepSeek to the OpenAI interface.
- Octopus ⭐ 2.3k — Personal LLM API aggregation gateway unifying multiple providers behind one endpoint, with load balancing and OpenAI/Anthropic protocol conversion (Go + Next.js).
- Veloera ⭐ 1.6k — Newer relay platform in the one-api/new-api lineage.
- uni-api ⭐ 1.2k — Lightweight single-config unified API manager, no frontend.
- APIPark ⭐ 1.8k — China-origin, cloud-native AI & API gateway with an open developer portal.
- VoAPI ⭐ 1.1k — Polished new-api-lineage relay/billing panel (Go), focused on UI and operations.
- done-hub ⭐ 793 — one-api/new-api fork with richer billing and channel management.
- sub2api ⭐ 33.2k — Go relay platform that pools Claude/OpenAI/Gemini/Antigravity subscription accounts (OAuth, session keys, API keys) behind one OpenAI/Anthropic-compatible endpoint, adding cost-sharing "carpool" billing (Stripe/Alipay/WeChat), key distribution and per-token rate limits. One of 2026's fastest-rising China-ecosystem relays — but account-pooling sits adjacent to the resold-relay category this list excludes; BYO accounts and vet before use.
- AI Proxy ⭐ 510 — Self-hosted Go gateway from the Sealos team that accepts OpenAI/Claude/Gemini protocols, converts between them, and adds multi-channel routing, load balancing, rate limiting, multi-tenant isolation, and a caching/web-search/reasoning plugin layer.
- metapi ⭐ 3.1k — Self-hosted "router of routers": aggregates your accounts across new-api/one
Comments (0)
Sign in to join the discussion.
No comments yet
Be the first to share your take.