0
0
via GitHub · Posted Jul 13, 2026 · 1 min read

Awesome AI Gateway — 100+ LLM Proxy Comparison

cuihuan/awesome-ai-gateway
Tool

⚡ Awesome AI Gateway — curated comparison of 100+ AI gateways & LLM proxies (LiteLLM, OpenRouter, Portkey, Kong, Higress, new-api, Bifrost) by cost, security, compliance & self-hosting. Decision tree + reproducible benchmarks. Open source, bilingual, updated daily.

45Stars
22Forks
5Open issues
HTML CC0-1.0 v1.1.0 Updated 1 week ago

A curated directory and decision tree for comparing AI gateways and LLM proxies by cost, security, and self-hosting options. The project includes a reproducible cost benchmark, feature matrices, and guidance for selecting the right gateway based on specific needs like cost optimization, compliance, or enterprise requirements.

0 comments

README

Awesome AI Gateway Awesome

GitHub stars Evaluation set Data verified CI PRs Welcome License: CC0

Pick the right AI gateway for your need in ~10 seconds — then trust the answer. A decision tree, a reproducible cost benchmark, and independent evidence for what we exclude. Organized by what you actually need, not by vendor.

Built the hard way: I burned $788 on AI coding in a single day — one flagship model ate 78% of it, just because I'd defaulted everything to the priciest option. So I mapped the whole gateway landscape. → the story

Languages: English · 简体中文

Contents

🔥 Top gateways (by stars)

The most-starred, most-authoritative projects in the list — a fast orientation. Stars auto-refresh daily; full context (features, license, caveats) is in each linked section. ⚠️ flags a project whose coding-subscription / account-pool routing can carry provider-ToS or account-ban risk.

Gateway Stars What it is Jump to
LiteLLM ⭐ 54.2k The default OSS proxy + SDK — OpenAI format to 100+ providers Self-hosted
Kong ⭐ 43.8k Mature API gateway with AI plugins (semantic cache, guard) Enterprise
new-api ⭐ 42.9k The most active relay/billing hub for teams China
CLIProxyAPI ⚠️ ⭐ 43.9k Wraps coding-CLI subscriptions (Claude Code, Codex…) into APIs Self-hosted
Claude Code Router ⭐ 36k Route Claude Code and agent CLIs to any model/provider Smart routing
one-api ⭐ 35.8k The original LLM API management / distribution system China
sub2api ⚠️ ⭐ 33.2k Pools subscription accounts behind one endpoint China
MLflow AI Gateway ⭐ 27.1k Unified endpoints + governance in the MLflow platform Observability
9router ⚠️ ⭐ 22.9k BYOK local proxy, subscription→cheap→free fallback Self-hosted
Apache APISIX ⭐ 16.9k Cloud-native API + AI gateway (ai-proxy plugins) Enterprise
aisuite ⭐ 15k Andrew Ng's unified multi-provider client (a library) Self-hosted
OmniRoute ⚠️ ⭐ 22.2k Coding-agent token-saver across 231+ providers Self-hosted
Portkey Gateway ⭐ 12.5k Fast TypeScript gateway, 1,600+ models, 50+ guardrails Self-hosted
Higress ⭐ 8.9k Alibaba's AI-native gateway on Envoy/Istio China
NVIDIA Dynamo ⭐ 7.5k Datacenter-scale, KV-cache-aware inference routing K8s
Bifrost ⭐ 6.6k Go gateway, lowest independently-measured overhead Self-hosted

Hosted leaders (SaaS, no GitHub stars): OpenRouter (400+ models, ~5.5% fee) · Vercel AI Gateway & Cloudflare AI Gateway (0% markup) → Cost-first.

Stars measure popularity, not fitness for your need — that's what the category sections and the evidence-based scorecard are for.

💰 Cost-first: cheapest multi-model access

Pain point: "I want many models for the least money and zero ops."

  • OpenRouter — The dominant model marketplace: 400+ models behind one OpenAI-compatible API, pay-as-you-go with automatic failover; ~5.5% fee when buying credits. $113M Series B (May 2026), ~8M users.
  • Vercel AI Gateway — Hundreds of models at provider list price (0% markup), $5/month free credits, zero-data-retention option; pairs naturally with the AI SDK.
  • Cloudflare AI Gateway — Free control plane in front of your own provider keys: caching, dynamic routing, unified billing, and dollar-denominated spend limits (2026 beta).
  • Requesty — EU-friendly OpenRouter alternative: 400+ models, sub-20ms failover, ~5% markup.
  • Eden AI — Unified API for 500+ models plus vision/OCR/speech; EU-based, ~5.5% platform fee.
  • Helicone AI Gateway (cloud) — Passthrough billing at 0% markup with observability bundled.
  • GPT-Load ⭐ 6.3k — High-performance Go proxy that rotates pools of API keys across channels to maximize quota usage.
  • Loop Gateway — OpenAI-compatible proxy that meters every request in Bitcoin sats instead of dollars. 311 models via OpenRouter at a 15% markup. No accounts, no email, no card; top up over Lightning, get a bearer token. Three auth rails (prepaid bearer, L402, Cashu). Hosted at api.loopxxi.com. New & unverified (anonymous; its public GitHub repo has since been removed, so treat it as a closed hosted relay) — it resells frontier models through the operator's own OpenRouter account at a 15% markup, and account-less + crypto-prepaid means no recourse if it swaps models or vanishes; confirm fidelity with canary_check.py and only top up what you can afford to lose.
  • nullsink (repo) — Account-less, metered proxy for frontier-model APIs, paid in Monero or Bitcoin. No accounts, no email, no card; mint a bearer token, prepay on-chain, and point the official SDKs at one base URL. ~10% markup taken once at top-up; no IP logging, no request logs; payment and token kept unlinkable. Self-hostable single binary (TypeScript/Bun, AGPL-3.0), live at nullsink.is. New & unverified (repo created 2026-06, 4★) — account-less + crypto-prepaid + no logs means no recourse if it swaps models or vanishes; confirm fidelity with canary_check.py and only top up what you can afford to lose.
  • AIMLAPI — One OpenAI/Anthropic-compatible endpoint fronting 400+ models (chat, image, video, audio, embeddings); prepaid, OpenRouter-style aggregator.
  • RunAPI — Hosted OpenAI-compatible relay (https://runapi.ai/v1) spanning LLM + media-generation (image, video, music/audio) behind one developer account; ships an active CLI and SDKs (runapi-ai) though the gateway itself is closed/hosted. New & unverified (self-submitted) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • TeamoRouter — Hosted SaaS gateway speaking OpenAI, Anthropic Messages and Gemini APIs over 500+ providers, with quality/latency/cost-aware "agentic routing" and a channel-dilution check that reroutes when a provider's output quality degrades; native Claude Code / Codex support, MCP, pay-as-you-go plus plans (Alipay/WeChat). New & unverified (community-recommended; the routing / quality-detection behavior is the operator's own description) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • KeepRouter — OpenAI- and Anthropic-compatible gateway: one key fronts 50+ models (Claude, GPT, Gemini, Mistral, Qwen, Kimi, GLM, DeepSeek, MiMo, MiniMax). Native /v1/messages means it works with Claude Code and the Anthropic SDK, not just the OpenAI SDK. Prepaid pay-as-you-go, billed at cost with 0% token markup (top-up fee 8%+$0.35); one genuinely free $0 model plus trial credit for new accounts. Bilingual EN/简体中文, Merchant of Record Paddle; not available in mainland China. New & unverified — confirm fidelity with canary_check.py before relying on it in production.
  • AI快站 (aifast.club) — Operator-submitted OpenAI- and Anthropic-compatible relay aimed at the mainland-China market (Cursor / Claude Code / Codex / Dify integration docs). Its current model list and rates are published live at /api/ratio_config, with service status at /api/status — read those rather than a fixed snapshot. New & unverified (closed-source, self-submitted) — confirm model fidelity with canary_check.py before relying on it in production.
  • NovAI — Hosted OpenAI-compatible relay fronting Chinese frontier models (DeepSeek, Qwen, GLM, Kimi, MiniMax, Doubao, Hunyuan) under one key; unlike the other Chinese-model relays here it also exposes image (Doubao Seedream) and video (Doubao Seedance) generation on the same endpoint, per-token pay-as-you-go with trial credit for new accounts. New & unverified (self-submitted) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • Novita AI — Unified API to 200+ open-source models (DeepSeek/Qwen/Llama…) with load balancing, autoscaling and failover; also a GPU cloud.
  • FlintAPI (repo) — Hosted OpenAI-compatible relay over Chinese LLMs (DeepSeek, Qwen, Kimi, GLM, MiniMax); the operator now describes it as a smart-routing engine that dispatches each prompt to a best-suited model rather than a plain aggregator, with free trial credit for new accounts. New and unverified (routing behavior is the operator's own description) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • FlowBar — Hosted OpenAI-compatible relay reselling dozens of models (GPT, Claude, Gemini, DeepSeek, Qwen, GLM, Kimi) below OpenRouter, with USD/CNY/crypto payment. New and unverified — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • Meshs One — Hosted OpenAI-compatible relay fronting Chinese frontier models (DeepSeek-V4, Qwen3.7-Max, MiniMax-M3) under one key, per-token pay-as-you-go (its /v1 endpoint returns a new_api_error, so it appears to run on new-api). New & unverified — closed-source and brand-new; confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • CoderPlan — Hosted OpenAI-compatible relay fronting Claude/GPT/Gemini/DeepSeek/Grok for the China market, per-token pay-as-you-go, ¥10 minimum top-up with Alipay/WeChat; Hong Kong/Singapore nodes (API base api.coderplan.ai/v1, which returns a new_api_error, so it appears to run on new-api). New and unverified — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • lxg2it ModelRouter (repo) — Solo-built, OpenAI-compatible router over 7+ providers (Anthropic, OpenAI, Google, Cerebras, Groq, Grok, GLM) with tiered automatic fallback that selects the cheapest available model. Free tier plus a paid tier advertised at 0% markup on Anthropic (a deposit fee may apply — verify current pricing). New and unverified — the public repo is a thin, unlicensed core stub still seeing fresh commits (routing logic lives in the closed hosted service; no license file) — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • OpenPaths (repo) — Hosted OpenAI-compatible router auto-routing across 15+ providers (OpenAI, Anthropic, Gemini, Groq, xAI, DeepSeek, Mistral) under one API, spanning chat, image, video, music, speech, embeddings and transcription. New & unverified — despite the "open source" framing the GitHub repo is a no-code, unlicensed marketing mirror (canonical code lives on the third-party Codex Infinity platform and is agent-maintained), so treat it as a closed hosted relay and confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • Glama Gateway — OpenAI-compatible gateway to 100+ models with consolidated billing, caching and logging (OSS core glama-ai/lightport).
  • RouterPlex — Hosted OpenAI-compatible gateway to 25+ models (GPT, Claude, Gemini, DeepSeek, Qwen, Kimi and others) across 11 providers; prepaid, billed per token at published vendor list rates, no subscription. $5 free credit for new accounts. New and unverified, closed-source — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.
  • TierUp — Hosted OpenAI-compatible gateway exposing four fixed performance tiers (tier-1…tier-4) instead of model names, each mapped server-side to a current best-value model; routes through OpenRouter and prices ~50% under the underlying models' retail, transparently subsidized during an early product-market-fit phase (solo-built, ~zero production users, tier 1 currently free). New & unverified — confirm model fidelity (e.g. with canary_check.py) before relying on it in production.

💡 Squeeze more from any gateway: enable semantic caching (Kong, Bifrost, Zuplo), set spend limits (Cloudflare, Zuplo, Pydantic/Logfire), and route easy prompts to cheap models (see Smart routing).

🆓 Free tiers that still work — the verified-limits table

One of the most-asked questions in the ecosystem, and the internet's answers are mostly stale. Every row below was re-verified against the provider's own docs (machine-readable, verified 2026-07-09, CI-enforced ≤30-day re-review). "Unpublished" means the provider now hides the numbers behind a login — we say so instead of quoting third-hand figures.

Provider What's free (verified limits) Notable free models Card? The catch
OpenRouter :free 50 req/day (<$10 lifetime top-up) → 1,000 req/day ($10+ once); 20 req/min shared across all :free models rotating :free pool free routes hit third-party providers that may train on your data (per-provider policy — check the free/paid routing settings)
Google Gemini API Flash-class models free; per-model RPM/RPD unpublished since 2026 (login-only in AI Studio); daily quota resets midnight PT Gemini 3.5 Flash · 3.1 Flash-Lite · 2.5 Flash · Gemma 4 free-tier content is "used to improve our products" (their pricing page's words)
Groq per-model: Llama-3.3-70B 30 RPM / 1K req/day / 100K tok/day · GPT-OSS-120B 30 RPM / 1K RPD / 200K TPD · Llama-3.1-8B 14.4K RPD / 500K TPD GPT-OSS-120B/20B · Llama-4-Scout · Llama-3.3-70B · Qwen3-32B ❌ (card only to upgrade) daily token caps burn fast on 70B+; limits are org-level
Cerebras 5 RPM / 30K TPM / 1M tokens/day, flat across models GPT-OSS-120B · GLM-4.7 · Gemma-4-31B ❔ not stated 5 RPM = one interactive user; "all models" marketing vs 3 models on the actual table
GitHub Models any GitHub account: low-tier 15 RPM / 150 req/day, high-tier 10 RPM / 50 req/day; 8K in / 4K out per request GPT-5 · o4-mini · Llama 4 · Phi-4 · DeepSeek-R1 explicitly experimentation-only; the 8K/4K per-request caps rule out long context
Cloudflare Workers AI 10,000 neurons/day (≈4M input tok/day on Llama-3.2-1B, ≈314K on GPT-OSS-120B — our arithmetic from official rates) Llama-3.3-70B · GPT-OSS-120B/20B · DeepSeek-R1-distill ❔ not stated neurons meter compute — output burns 5–10× faster than input
Mistral free Experiment mode; exact limits unpublished (per-account, Admin Console) Large · Small · Codestral (reported) trains on your data by default — free users must manually toggle it off
Cohere trial key: 1,000 calls/month; Chat 20 RPM Command A · Command R+ evaluation scale only
SambaNova 20 RPM / 20 req/day / 200K tok/day DeepSeek-V3.1 · Llama-3.3-70B · GPT-OSS-120B 20 req/day = demo only (very fast inference though)
Hugging Face $0.10/month provider-passthrough credits (PRO: $2/mo) 200+ provider-routed (DeepSeek-V3 …) $0.10 ≈ a handful of big-model requests
Z.ai (GLM) GLM Flash models $0 in AND out; rate limits unpublished (per-key, login) GLM-4.7-Flash · GLM-4.5-Flash · GLM-4.6V-Flash ❔ not stated opaque shifting concurrency limits; China-HQ provider — weigh your data sensitivity

Trial credits ≠ free tiers (they expire): NVIDIA build.nvidia.com (1,000 requests at signup, +4,000 with a business email — staff-forum figures; prototyping-only license) and Alibaba Cloud Model Studio international (per-model quotas, hard 90-day expiry, Singapore region). Recently discontinued — ignore stale listicles: Together AI killed its -free models (now $5 minimum prepaid, "does not currently offer free trials"), Moonshot/Kimi requires a $1 top-up to start, and xAI's much-cited data-sharing credits are no longer documented on any public page. Full evidence per row: data/free_tiers.json. A wrong row is a bug — report it.

🔓 Self-hosted open source

Pain point: "My keys, my infra, no per-token middleman fee."

  • LiteLLM ⭐ 54.2k — The default choice: Python SDK + proxy server speaking OpenAI format to 100+ providers, with virtual keys, budgets, load balancing and guardrails.
  • Portkey Gateway ⭐ 12.5k — Fast TypeScript gateway (1,600+ models, 50+ guardrails) that also powers Portkey's commercial LLMOps platform.
  • CLIProxyAPI ⭐ 43.9k — Go gateway that wraps coding-agent CLI subscriptions (Claude Code, Codex, Gemini, Grok, Antigravity) into OpenAI/Gemini/Claude/Codex-compatible APIs with multi-account pools, round-robin load balancing and a management API; one of the highest-starred OSS gateways in the space. BYO accounts — but routing OAuth coding-tier subscriptions through an API can violate provider ToS, so weigh account-ban risk.
  • 9router ⭐ 22.9k — MIT self-hosted BYOK local proxy that auto-routes across 40+ providers with subscription→cheap→free fallback, multi-account load balancing and token compression; cost-first and very popular, but its free/OAuth coding-tier routing (Claude Code, Codex, Kiro) carries provider-ToS/account-ban risk.
  • OmniRoute ⭐ 22.2k — MIT self-hosted TypeScript gateway: one endpoint to 231+ providers (50+ free), plugging Claude Code / Codex / Cursor / Cline / Copilot into free Claude/GPT/Gemini with stacked token compression (15–95% savings), 17 routing strategies, smart auto-fallback and MCP/A2A. A 2026 breakout of the coding-agent "token-saver" wave — genuine code (not a relay farm), but its free/OAuth coding-tier routing carries provider-ToS/account-ban risk.
  • Chat Nio (CoAI) ⭐ 9.2k — Multi-tenant "one-stop" gateway with a built-in admin + credit/subscription billing panel over 200+ models / 35+ providers, priority-based load balancing and model caching — the same commercial-panel genre as the new-api / one-api / VoAPI entries here.
  • TensorZero ⭐ 11.7k — ⚠️ Archived June 2026 (company wound down; repo read-only, Apache-2.0 code + community forks remain). Rust gateway unified with observability, evals, experimentation and optimization.
  • Bifrost ⭐ 6.6k — Go gateway from Maxim AI claiming ~50x LiteLLM throughput; adaptive load balancing, cluster mode, MCP support.
  • Traceloop Hub ⭐ 221 — High-scale gateway written in Rust from the Traceloop team (OpenLLMetry / OTel-for-LLMs); OpenTelemetry-native observability built in.
  • Helicone ⭐ 6k — Observability-first platform (YC W23) with a Rust ai-gateway ⭐ 614.
  • Plano ⭐ 6.9k — AI-native proxy and data plane for agents (formerly Arch Gateway / archgw).
  • AxonHub ⭐ 4.7k — Go gateway: call 100+ LLMs from any SDK behind one OpenAI/Anthropic-compatible endpoint, with built-in failover, load balancing, cost control and end-to-end tracing. BYOK self-hosted.
  • Manifest ⭐ 7.3k — Self-hosted TypeScript router (MIT): one OpenAI-compatible /auto endpoint (plus /v1/messages for Anthropic clients) over 300+ models / 31+ providers, mixing API keys, local models (Ollama/LM Studio) and coding-tier subscriptions with complexity/header-based routing, cost tracking, budgets and failover. BYO accounts — routing OAuth coding-tier subscriptions through an API can carry provider-ToS risk.
  • LLM Gateway ⭐ 1.4k — Open-source OpenRouter alternative: route, manage and analyze requests across providers.
  • APIPark ⭐ 1.8k — Cloud-native LLM API management and distribution platform.
  • Pydantic AI Gateway ⭐ 192 — BYOK gateway with cost caps and OTel; ⚠️ repo archived, now folded into Pydantic Logfire.
  • OptiLLM ⭐ 4.2k — Optimizing inference proxy that boosts accuracy via test-time compute techniques.
  • aisuite ⭐ 15k — Andrew Ng's unified multi-provider client. A library rather than a deployable proxy — fits when you don't want network hops.
  • Shepherd Model Gateway (SMG) ⭐ 408 — Engine-agnostic gateway in Rust: one OpenAI/Anthropic-compatible endpoint over vLLM/SGLang/TRT-LLM + cloud providers, with KV-cache-aware routing and WASM plugins.
  • RelayPlane ⭐ 192 — MIT, local-first proxy (npm): 11 providers behind one endpoint with per-request cost attribution and hard daily/hourly budget caps.
  • SentryNode Gateway ⭐ 0 — Open-core (Apache-2.0) AI proxy for cost governance / FinOps routing: adaptive model routing, budget caps and audit logging. Early-stage; the public repo currently ships a demo scaffold.
  • GoModel ⭐ 1k — Lightweight single-binary Go gateway (open-source LiteLLM alternative) exposing one OpenAI/Anthropic-compatible API across 18+ providers with caching, guardrails and usage/cost tracking; fast-growing, though its throughput-vs-LiteLLM figures are vendor-run.
  • OpenGateLLM ⭐ 172 — Production-grade open-source GenAI gateway from France's Etalab (powers the government's "Albert" assistant): one OpenAI-compatible API over self-hosted + provider models, with auth, rate limits and usage tracking. Distinct public-sector / EU-sovereignty angle.
  • ⚠️ Stale but historically notable: BricksLLM ⭐ 1.2k (PII masking, per-key limits; inactive since early 2025), Glide ⭐ 159 (inactive since 2024).

🏢 Enterprise & compliance

Pain point: "Audit logs, PII redaction, RBAC, on-prem, and the EU AI Act (enforceable Aug 2026)."

  • Kong AI Gateway ⭐ 43.8k — Mature API gateway with AI plugins: semantic caching/routing, prompt guard, token rate-limiting; Konnect for managed control plane.
  • Apache APISIX ⭐ 16.9k — Cloud-native API + AI gateway with ai-proxy / ai-proxy-multi plugins.
  • Envoy AI Gateway ⭐ 1.9k — CNCF-aligned GenAI access on Envoy Gateway, backed by Tetrate and Bloomberg.
  • kgateway ⭐ 5.6k — CNCF API/AI gateway, the base of Solo.io's commercial Gloo AI Gateway.
  • TrueFoundry AI Gateway — Enterprise gateway with routing, guardrails and RBAC, deployable into your K8s/VPC.
  • nexos.ai — Enterprise AI gateway/orchestration from the Nord Security founders (€30M Series A, Oct 2025).
  • Tyk AI Studio — AI governance suite: budgets, model catalogs, guardrails on Tyk's gateway.
  • Gravitee Agent Mesh — LLM Proxy, MCP Proxy and A2A support inside Gravitee APIM.
  • WSO2 AI Gateway — Egress management for LLM traffic: model routing, semantic caching, guardrails.
  • F5 AI Gateway — Containerized AI traffic gateway; data-leakage detection via the LeakSignal acquisition (announced Jul 2025).
  • IBM API Connect AI Gateway — Policy enforcement, masking and audit for LLM traffic.
  • MuleSoft AI / Omni Gateway — Governs LLM, MCP and agent traffic alongside classic APIs.
  • Lunar.dev ⭐ 469 — Egress consumption gateway repositioned around MCP/agent governance.
  • KrakenD AI Gateway — High-performance, stateless Go API gateway (krakend/krakend-ce ⭐ 2.7k) with an AI proxy + prompt-security layer.
  • Broadcom Layer7 AI Gateway — LLM traffic governance, threat protection and quotas on the mature Layer7 API platform.
  • Cequence AI Gateway — API-security-first AI gateway: discovery, guardrails and threat protection for LLM/agent traffic.
  • Axway Amplify AI Gateway — Centralized control plane on Axway's Amplify platform governing LLM/MCP/agent traffic with business-logic model routing, RBAC, spend caps, prompt-injection controls and RAG integration, from a 10× Gartner MQ API-management Leader.
  • Red Hat Connectivity Link — Kubernetes-native gateway (built on the Kuadrant project, successor to 3scale) unifying AI gateway, API management and multicluster connectivity; powers OpenShift AI Models-as-a-Service as the front door governing external and self-hosted LLM endpoints.
  • Sensedia AI Gateway — Gartner-recognized APIM vendor's agnostic AI gateway governing LLMs, MCP servers and AI agents with multi-model routing, guardrails, cost controls and observability across a multi-cloud control plane.
  • Ambassador Edge Stack — Envoy-based, Kubernetes-native API gateway (OSS core emissary-ingress ⭐ 4.5k) whose AI Gateway layer adds LLM-provider routing, token rate-limiting and fallback — a peer to Kong/Tyk/APISIX in the API-vendor cohort.

☁️ First-party gateways (cloud & model vendors)

Pain point: "We're already committed to one cloud — give us the native path."

  • AWS Bedrock — Multi-model access via the unified Converse API, cross-region inference, and AgentCore Gateway for tools/MCP.
  • Azure API Management — GenAI gateway — Token limits, semantic caching and load balancing in front of Azure OpenAI / AI Foundry.
  • Google Apigee + Vertex AI — LLM gateway patterns on Apigee with Vertex Model Garden as the managed hub.
  • Cloudflare AI Gateway — See Cost-first; the strongest free first-party option.
  • Vercel AI Gateway — GA, 0% markup, ZDR option; the default for Next.js/AI SDK shops.
  • Databricks Unity AI Gateway — Mosaic AI Gateway folded into Unity Catalog, adding agent + MCP governance.
  • Tencent Cloud AI Gateway — Tencent's first-party cloud-native intelligent gateway bundling LLM + MCP + Agent gateways with protocol conversion, cost/performance-based routing, and unified access to Hunyuan + third-party models.

🇨🇳 China ecosystem

Pain point: "Domestic models (Qwen/DeepSeek/GLM/Kimi), CNY payment, key distribution & billing for teams."

  • new-api ⭐ 42.9k — The most active one-api fork, now a "unified AI model hub": protocol conversion, billing, Rerank/Realtime endpoints. AGPL-3.0.
  • one-api ⭐ 35.8k — The original LLM API management & distribution system (OpenAI/Azure/Claude/Gemini/DeepSeek/Doubao…); development has slowed.
  • Higress ⭐ 8.9k — Alibaba's AI-native gateway on Envoy/Istio, first-class Tongyi/DeepSeek support; hosted version at higress.ai.
  • GPT-Load ⭐ 6.3k — Smart API-key rotation multi-channel proxy in Go.
  • one-hub ⭐ 2.9k — one-api fork with better non-OpenAI function calling and stats.
  • simple-one-api ⭐ 2.3k — Single binary adapting Qianfan/Spark/Hunyuan/MiniMax/DeepSeek to the OpenAI interface.
  • Octopus ⭐ 2.3k — Personal LLM API aggregation gateway unifying multiple providers behind one endpoint, with load balancing and OpenAI/Anthropic protocol conversion (Go + Next.js).
  • Veloera ⭐ 1.6k — Newer relay platform in the one-api/new-api lineage.
  • uni-api ⭐ 1.2k — Lightweight single-config unified API manager, no frontend.
  • APIPark ⭐ 1.8k — China-origin, cloud-native AI & API gateway with an open developer portal.
  • VoAPI ⭐ 1.1k — Polished new-api-lineage relay/billing panel (Go), focused on UI and operations.
  • done-hub ⭐ 793 — one-api/new-api fork with richer billing and channel management.
  • sub2api ⭐ 33.2k — Go relay platform that pools Claude/OpenAI/Gemini/Antigravity subscription accounts (OAuth, session keys, API keys) behind one OpenAI/Anthropic-compatible endpoint, adding cost-sharing "carpool" billing (Stripe/Alipay/WeChat), key distribution and per-token rate limits. One of 2026's fastest-rising China-ecosystem relays — but account-pooling sits adjacent to the resold-relay category this list excludes; BYO accounts and vet before use.
  • AI Proxy ⭐ 510 — Self-hosted Go gateway from the Sealos team that accepts OpenAI/Claude/Gemini protocols, converts between them, and adds multi-channel routing, load balancing, rate limiting, multi-tenant isolation, and a caching/web-search/reasoning plugin layer.
  • metapi ⭐ 3.1k — Self-hosted "router of routers": aggregates your accounts across new-api/one

Comments (0)

Sign in to join the discussion.

No comments yet

Be the first to share your take.