0
0
via GitHub · Posted Jul 22, 2026 · 1 min read

Rapid-MLX: Apple Silicon AI Engine

raullenchai/Rapid-MLX
Application

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

3,321Stars
391Forks
23Open issues
64Watching
Python Apache-2.0 v0.10.15 Updated 8 hours ago

A high-performance local AI inference engine optimized for Apple Silicon Macs, offering 4.2x faster inference than Ollama with OpenAI-compatible APIs. It supports chat, tool calling, multiple agent frameworks, and runs entirely on-device with models sized to available RAM.

0 comments

README


Quick Start (60 seconds)

1. Install — pick one path (run only one of these):

Homebrew — prebuilt bottle straight from homebrew-core (recommended):

brew install rapid-mlx

or the one-liner — detects your RAM, picks a starter model:

curl -fsSL https://rapidmlx.com/install.sh | bash

Both land the same rapid-mlx CLI. The curl installer additionally installs Python 3.10+ if missing, creates an isolated venv at ~/.rapid-mlx/, symlinks the rapid-mlx CLI into ~/.local/bin/, and prints a serve command sized to your Mac (8–23 GB → qwen3.5-4b-4bit; 24–47 GB → gpt-oss-20b-mxfp4-q8; 48–95 GB → qwen3.6-35b-8bit; 96 GB+ → gpt-oss-120b-mxfp4-q8).

Install security. install.sh is served over HTTPS (HSTS-preload) from rapidmlx.com and is a byte-identical mirror of install.sh at the release commit — read it before running if you like. If you want a cryptographically verified installer rather than trusting the website pipe, don't curl | bash the URL above: instead download the release's install.sh asset, verify it against the cosign-signed SHA256SUMS.txt shipped alongside it, and run that verified copy — full recipe in SECURITY.md. PyPI artifacts additionally carry Sigstore attestations (PEP 740). Two more low-trust paths:

  • Pin to a commit hashcurl -fsSL https://raw.githubusercontent.com/raullenchai/Rapid-MLX/<commit>/install.sh -o install.sh && shasum -a 256 install.sh && bash install.sh
  • Skip the shell script entirely — use Homebrew, uv, or pip below.

See Alternative install methods for the non-curl paths.

2. Chat with a model right now:

rapid-mlx chat

Defaults to qwen3.5-4b-4bit. First run downloads the weights (~2.5 GB) with a progress bar and drops you into a REPL. Type /help for slash commands, /exit to quit.

3. Or serve it for use from other apps:

rapid-mlx serve qwen3.5-4b-4bit

Starts an OpenAI-compatible HTTP server bound to http://localhost:8000. Point any OpenAI SDK / client (Cursor, Aider, LangChain, OpenCode, PydanticAI, your own scripts) at http://localhost:8000/v1; Claude Code / Anthropic SDK uses http://localhost:8000 (the Anthropic messages route lives at /v1/messages under the same host).

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"default","messages":[{"role":"user","content":"Say hello"}]}'
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
print(client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "Say hello"}],
).choices[0].message.content)

Vision / audio / diffusion models? Base install is text-only (~460 MB). Vision, audio, embeddings, and DFlash speculative decoding ship as opt-in extras. → Optional extras

Not into the terminal? Rapid-MLX Desktop bundles the same engine inside a one-click Mac app.


Why Rapid-MLX

Apple-Silicon-native Pure MLX kernels — no llama.cpp fallback, no Metal shim. Continuous batching, prompt cache (radix + DeltaNet RNN snapshots), and TurboQuant K8V4 KV codec run at native MLX bandwidth on M1 → M4.
Drop-in OpenAI / Anthropic API /v1/chat/completions, /v1/responses (Codex CLI), /v1/messages (Anthropic SDK / Claude Code), /v1/embeddings, /v1/audio/* — same wire as ChatGPT / Claude, no client adapter.
Tier-1 ecosystem coverage 8 agent CLIs and 3 Python frameworks are wire-verified against real weights every release — Codex CLI, Claude Code, OpenCode, Qwen Code, OpenHands, Hermes Agent, Aider, Kilo Code + LangChain, PydanticAI, smolagents.

Full feature breakdown


Use Cases

Chat in the terminal rapid-mlx chat qwen3.5-9b-4bit Streaming REPL, /help for slash commands, --think / --no-think to control CoT.
OpenAI server for your apps rapid-mlx serve qwen3.5-9b-4bit Point Cursor, Aider, LibreChat, Open WebUI, LangChain at http://localhost:8000/v1.
Agent backends rapid-mlx serve qwen3.6-35b-8bit &rapid-mlx agents codex --setup && codex 8 Tier-1 agents auto-configure once the server is up — see Tier-1 support.
Benchmark your Mac rapid-mlx bench qwen3.5-9b-4bit --submit Standardized B=1 bench, opens a PR to publish your row on rapidmlx.com.

One-shot IDE setup with rapid-mlx launch <cursor|claude-code|cline|continue-dev>


Tier-1 Support

Every row below has a rapid-mlx agents <name> --setup config template (except Claude Code, which is one env-var) and an integration test that drives the same wire the real client drives against a live server.

Agents (8) Frameworks (3)
Codex CLI · Claude Code · OpenCode · Qwen Code · OpenHands · Hermes Agent · Aider · Kilo Code LangChain (+ LangGraph) · PydanticAI · smolagents

Also compatible with any OpenAI-compatible client via http://localhost:8000/v1 — Cursor, LibreChat, Open WebUI, and more plug in with a single URL change.

Full 8×3 agent matrix + 3×3 framework matrix (test cells + xfail reasons)Codex CLI · Claude Code · OpenCode · Qwen Code · OpenHands · Hermes · Aider · Kilo Code


Choose Your Model

The installer's RAM detector picks a sensible default. If you want to shop the full catalog: rapid-mlx models lists every alias, rapid-mlx info <alias> shows the per-alias profile (parser, MoE / hybrid flags, KV codec eligibility, speculative-decoding gates).

RAM Recommended One-shot
8–23 GB MacBook Air/Pro qwen3.5-4b-4bit rapid-mlx serve qwen3.5-4b-4bit
24–47 GB MacBook Pro / Mac Mini gpt-oss-20b-mxfp4-q8 rapid-mlx serve gpt-oss-20b-mxfp4-q8
48–95 GB Mac Studio qwen3.6-35b-8bit rapid-mlx serve qwen3.6-35b-8bit
96 GB+ Mac Studio / Pro gpt-oss-120b-mxfp4-q8 rapid-mlx serve gpt-oss-120b-mxfp4-q8

Full RAM tier map + serve flags per tierEvery alias, quant, and family (128+ aliases across 30+ families) · interactive at models.rapidmlx.com


Alternative install methods

The two paths above cover most users — reach for these only if you already manage Python yourself.

brew install rapid-mlx

Ships in homebrew-core since 0.10.12 — no tap, no trust prompt. Upgrade with brew upgrade rapid-mlx. If you previously installed from the legacy raullenchai/rapid-mlx tap, switch once: brew uninstall rapid-mlx && brew untap raullenchai/rapid-mlx && brew install rapid-mlx.

uv tool install rapid-mlx@latest

Don't have uv yet? curl -LsSf https://astral.sh/uv/install.sh | sh. Upgrade with uv tool upgrade rapid-mlx.

python3.12 -m pip install rapid-mlx

If pip install rapid-mlx says "no matching distribution", your Python is too old. brew install [email protected] first. Upgrade with pip install -U rapid-mlx.

For image-input / VLM models (Qwen-VL, true multimodal), install the vision extra: pip install 'rapid-mlx[vision]' — see Optional extras.


Command Reference

rapid-mlx --help                    # top-level command list
rapid-mlx <subcommand> --help       # per-subcommand flags

Covers chat, serve, share, agents (setup / test), bench, models, pull, rm, ps, info, doctor, upgrade, telemetry, launch, and jlens.

Full CLI reference with every flag


Troubleshooting

Run the built-in self-check first:

rapid-mlx doctor

Top three things that go wrong:

  • Much slower than expected. Qwen3.5 / 3.6 default to thinking-on — add --no-think to skip chain-of-thought. → Slow tok/s
  • Out of memory. Model too big for your RAM — pick a smaller quant from Choose Your Model or the full tier map. → OOM guide
  • Tool calls arriving as plain text. Auto-recovery handles most cases; if not, set --tool-call-parser explicitly for your model. → Tool-call recovery

All troubleshooting entries (OOM, empty responses, slow TTFT, port taken, shell completion, HF cache, and more)


See it in action

Community & Contributing

  • Report a bug or request a model: Issues
  • Report a security issue: Private advisory — see SECURITY.md
  • Ask a question or share a build: Discussions
  • Contribute code, aliases, or docs: CONTRIBUTING.md
  • Add your hardware to the public benchmark: rapid-mlx bench <alias> --submit opens the PR for you

Rapid-MLX ships opt-in anonymous telemetry (off by default; explicit rapid-mlx telemetry enable required). No prompts, completions, paths, IPs, or API keys are ever collected. → What we do and don't collect

🚀 Contributors

Every avatar here shipped something in rapid-mlx — model support, tool-call parsers, fixes, docs, and benchmark submissions. Thank you.

Star History


License

Apache 2.0 — see LICENSE.

Comments (0)

Sign in to join the discussion.

No comments yet

Be the first to share your take.