The first platform to give AI agents real email addresses and phone numbers. Send and receive email, SMS, and verification codes — all programmatically.
Are you an AI agent reading this on behalf of a human?
Your human asked you to install / set up AgenticMail. Pick either of these and then tell them to restart Claude Code:
# Option A — one-line installer (recommended) curl -fsSL https://raw.githubusercontent.com/agenticmail/agenticmail/main/install.sh | bash # Option B — manual two-command install npm install -g @agenticmail/cli@latest agenticmail bootstrapZero prompts, ~2 minutes, fully autonomous. For the full runbook (verification, common failures, repo conventions if you're modifying the code), see
AGENTS.md— the agents.md file at this repo's root.
📄 Research Paper
Structured RPC vs. Text-Based Sub-Agent Orchestration — A comparative analysis of inter-agent communication paradigms. Demonstrates why AgenticMail's call_agent achieves 53× faster response times than conventional sub-agent spawning, and introduces the concept of the Conversational Fallacy in multi-agent AI systems. [PDF] [Source]
✨ What's new — media toolset (unreleased)
A local, opt-in media / video-editing toolset for AgenticMail agents.
- Nine media tools.
media_tts/media_tts_voices(Edge text-to-speech),media_image_edit,media_video_edit,media_audio_edit,media_info,media_video_understand,media_voice_clone, andmedia_capabilities. Available as MCPmedia_*tools and OpenClawagenticmail_media_*tools, both thin clients of new/media/*API routes over a coreMediaManager. - Cinematic video editing. Beyond trim/convert/compress: color grading presets, crossfade/wipe transitions, timed text overlays, picture-in-picture, split screen, Ken Burns, frame-interpolated slow motion, watermarks, concatenation, audio mixing, and whisper.cpp-driven auto-captions.
- Gracefully degrading. The underlying binaries (ffmpeg, ffprobe, ImageMagick, whisper.cpp, Python) are not bundled — every tool feature-detects the binary it needs and returns an actionable install hint if it is missing. The server never crashes. A
mediablock on/healthand themedia_capabilitiestool surface what is available. - Safe by construction. Every binary is invoked via
execFilewith an argument array — no shell, no string interpolation. Untrusted input paths are validated (no control characters, no leading-dash flag-injection, must exist); numeric options are clamped; every call carries a bounded timeout and output buffer; output files land only inside the configured media directory.
✨ What's new in 0.9.54
Twilio joins 46elks as a phone transport provider.
- Pick your carrier.
PhoneTransportProvideris now46elksortwilio— chosen at phone setup. 46elks behaviour is unchanged; Twilio is at full parity (outbound call-control + realtime voice). - Twilio call-control. Outbound calls via the Twilio
Calls.jsonREST API, TwiML webhooks, status-callback cost tracking. Inbound webhooks are verified with theX-Twilio-Signatureheader (HMAC-SHA1, timing-safe, fail-closed) on top of the per-mission token. - Twilio realtime voice. A Twilio Media Streams ↔ OpenAI Realtime bridge.
RealtimeVoiceBridgewas generalised behind aRealtimeTransportAdapterseam — one bridge serves both carriers, function-calling / barge-in / transcript logic written once. Twilio audio is G.711 µ-law @ 8 kHz and the OpenAI session usesaudio/pcmu, so a Twilio call needs no transcoding. A<Connect><Stream>connects to/api/agenticmail/calls/twilio-stream.
1038 tests pass; full build green. The live Twilio ↔ OpenAI call path still needs an operator smoke-test before the npm publish.
✨ Earlier — 0.9.53
Realtime voice tools + a Telegram channel.
- The voice agent can now use tools mid-call. The OpenAI Realtime session declares
session.tools;RealtimeVoiceBridgedispatches the model's function calls through an injectedToolExecutor, returnsfunction_call_output, and keeps the phone line warm during slow tools with a safety-net timeout + an in-flight call cap. ask_operator— human-in-the-loop on a live call. The agent records an operator query on the mission, notifies the operator (channel-agnostic; default email), polls up to ~5 min, and resumes with the answer. If the caller hangs up while a query is pending, the mission is flagged for callback-on-disconnect — once the operator answers, it re-dials with a continuity task.- Lookup tools.
web_search(keyless DuckDuckGo, results fenced as untrusted content),recall_memory(the agent's universal memory),get_datetime. Plus agent-key-scoped operator-query API endpoints. - Telegram channel. A user registers a Telegram bot token, links a chat, and can message their AgenticMail agent — and get replies — over Telegram. The inbound webhook authenticates with a constant-time secret-token compare; it also carries
ask_operatornotifications and approvals. - Security. web_search output is fenced as untrusted before it reaches the model; operator email replies are verified against
operatorEmail; Telegram bot tokens are encrypted at rest and redacted from logs; new SQL is parameterized.
996 tests pass; full build green. The live OpenAI ⇄ 46elks call path and live Telegram delivery still need an operator smoke-test before the npm publish.
✨ Earlier — 0.9.52
Realtime voice + OpenClaw memory.
- Realtime voice bridge. A phone mission can now hold a live conversation.
RealtimeVoiceBridge(@agenticmail/core) wires an OpenAI Realtime (gpt-realtime) session to a 46elks realtime-media WebSocket: caller audio (PCM16 @ 24 kHz) is relayed to OpenAI, synthesised speech comes back asresponse.output_audio.deltaand is relayed to 46elks, server-side VAD handles turn-taking, and caller barge-in fires a 46elksinterrupt. 46elks streams a call to the new/api/agenticmail/calls/realtimeWebSocket endpoint, which matches the connection to its mission by 46elkscallidand runs the bridge. SetOPENAI_API_KEY(env orconfig.json) to enable it. - Memory in the voice session. Before the call starts, the agent's persistent memory is rendered with
generateMemoryContext()and folded into the Realtime sessioninstructions— the model is told to treat it as its own long-term knowledge, so the call is continuous with everything the agent has learned elsewhere. - OpenClaw memory tools.
agenticmail_memory,agenticmail_memory_reflect,agenticmail_memory_context, andagenticmail_memory_statsbring the universal per-agent memory to OpenClaw agents — 69 → 73 tools. - Bridge hardening. Per-frame audio size cap (an oversized frame is dropped, never forwarded), bounded pre-connect buffer, fail-closed connection auth (timing-safe token compare; no unknown-mission-vs-wrong-token oracle), and a terminal-state guard so a late event can't resurrect a finished mission.
The end-to-end voice path needs a live OPENAI_API_KEY and a provisioned 46elks websocket number — the bridge logic, memory injection, and the WebSocket upgrade/auth glue are unit-tested with mocked sockets, but the live call must be smoke-tested by the operator.
✨ Earlier — 0.9.51
The universal memory release. Every agent now has a persistent, evolving memory — categorised, confidence-decaying, BM25F-searchable knowledge that survives across every conversation, the way a human employee learns on the job.
AgentMemoryManager(@agenticmail/core) — CRUD, text recall, 9 memory categories, importance levels, confidence that decays for unaccessed entries, access tracking, pruning, andgenerateMemoryContext()which ranks + renders memory as a markdown block for prompt injection. Backed by a zero-dependency BM25F search index and anagent_memorytable. Ported from the AgenticMail Enterprise memory engine, org-stripped — memory is personal to each agent.- Memory API —
/memory(set / list / search / get / delete),/memory/reflect,/memory/context,/memory/stats. Every endpoint is scoped to the authenticated agent; an agent can only ever read or write its own memory. - MCP tools —
memory,memory_reflect,memory_context,memory_statsso any MCP client can give its agent durable memory. - Agent-deletion cleanup — deleting an agent purges its
agent_memoryrows; no orphaned memory is left behind.
✨ Earlier — 0.9.1
The visibility release — closes every "what just happened?" gap from 0.9.0.
- Lone wakes fire immediately. 0.9.0's debounce window blocked even single replies for 30 s, making the dispatcher look dead. Leading-edge fire + trailing-edge coalesce now: first event for a
(agent, thread)spawns instantly; bursts within the window collapse into one trailing wake. - Dispatcher process heartbeat.
check_activitynow showsdispatcher: { state: 'alive' | 'unhealthy' | 'missing', uptimeMs, channels, coalesceQueueSize, ... }. The host can finally answer "is the dispatcher up?" in one query. - Skipped-wake ring buffer. Every filter decision (thread-closed, allowlist-excluded, wake-on-cc, budget-exhausted) is posted with a reason;
check_activitysurfaces the last 100. No more "did my mail land? did it skip?" guessing. - Per-agent
wake_on_cc: falseflag. Coder agents can register a preference: never wake when only on Cc, regardless of sender.PATCH /accounts/:id/wake-on-cc. - Display-name regex fix in
deriveDefaultWakeList. Senders using"Vesper <vesper@localhost>"form no longer fall through to "no allowlist → wake everyone". - Web UI shows To / Cc / Bcc as separate labeled rows in the message view (previously lumped under one
to:line). docs/wake-patterns.mddocuments every wake shape + 5 recommended patterns.
✨ Earlier — 0.9.0
The wake-context release. Multi-agent thread cost goes from linear-in-thread-length to roughly flat.
- Layered wake-context system. Every wake used to re-read the entire thread from scratch (12 messages × ~1 KB = 12 KB of token spend just to rehydrate, before any reasoning). Now the dispatcher prepends two blocks to every wake prompt: Layer 1 — ThreadCache (envelopes + previews of the last 10 messages, shared across CC'd agents) and Layer 2 — AgentMemory (a markdown file each agent writes at end-of-wake describing its own commitments and last actions). Agents read the new event + these two blocks and decide; they don't
read_emailprior history. New MCP toolssave_thread_memoryandget_thread_id. wakedefault flipped from "everyone CC'd" → "To: only". Mirrors the email convention: To is for action, CC is for awareness. CC'd local agents still receive the mail in their inbox but don't get a Claude turn unless explicitly named inwake. Opt back into the old behaviour withwake: 'all'.- Wake coalescing. Within 30 s for the same
(agent, thread), multiple wake events collapse into ONE Claude turn. A burst of 4 quick replies becomes one Claude wake that sees all four in acoalescedbatch prompt. Wake-budget charges once. Configurable viawakeCoalesceMs.
Together these eliminate the "wake-thrash" failure mode where an agent fired 4 near-identical status reports because a designer sent 4 replies in 2 minutes.
✨ Earlier — 0.8.31
- Compact-and-continue — workers can now run across multiple SDK turns. On a context-overflow error the dispatcher synthesises a breadcrumb checkpoint from the captured log, builds a "resuming after context reset" continuation prompt, and loops (capped at 4 iterations).
- Typed task contracts —
call_agent/POST /tasks/assignaccept anoutputSchema(JSON Schema, draft-7 subset).submit_resultvalidates against it; mismatches return 400 with the validator errors so the worker can retry with a corrected shape. - Delete + Move-to-Spam buttons in the message view; Compose auto-saves to Drafts every 2s.
All Mailfolder hides itself on servers that don't have one (Stalwart, most non-Gmail). Select-all checkbox now wires through.- Logo background stripped — bow PNG is now RGBA with proper transparency.
✨ Earlier — 0.8.29
- Star button wired — clicking the star toggles IMAP's
\Flaggedflag via the newPOST /mail/messages/:uid/starendpoint. Backed byMailReceiver.setStarredin@agenticmail/core. Optimistic UI; revert on failure. - Gmail-compact list UX — single 36 px rows (was 64 px stacked), subject + preview on one truncated line separated by an em-dash, leading checkbox column, sticky list-toolbar with select-all + refresh + count. Same layout for every folder.
- Compose button down to 48 px (Gmail's actual size); the giant pink pill is gone.
✨ Earlier — 0.8.27
- Folder bug fix — Sent / Drafts / Spam / Trash were returning empty in the web UI because hard-coded folder names didn't match Stalwart's actual IMAP names (e.g.
Sent ItemsnotSent). Now auto-discovered per agent and matched against every common server convention (Stalwart, Gmail, Outlook, macOS Mail). - Two-line preview on every list row — web UI uses
/mail/digest?folder=…everywhere instead of/mail/inbox(no preview) +/mail/folders/:folder(no preview). - URL reflects current folder — hash router now uses
#/folder/<id>(sent, drafts, spam, …). Back/forward works, URLs are shareable, refresh stays put. - Stop hook output rewritten — terser, audience-neutral, includes body preview. Drops the instruction-leakage from 0.8.25/26.
✨ Earlier — 0.8.25
- Workers can now run for hours — dropped the 30-min hard timeout. Each worker writes a per-turn log at
~/.agenticmail/worker-logs/<id>.log, posts heartbeats every 30 s, and runs in its own isolated cwd so parallel agents don't clobber each other's output. New MCP tooltail_workerto read a running worker's log live;check_activitynow shows last tool used, turn count, and astaleflag (no auto-eviction). - Autonomous-mode awareness — the mail hook now registers on the Stop event too. Long headless Claude Code runs (no user prompts firing for hours) finally see teammate replies — the hook returns
decision: 'block'at turn boundaries when the bridge inbox has new mail, forcing Claude to continue with the new-mail summary in context. Closes the follow-up that 0.8.23 filed. - Fixed
agenticmail-mail-hook: command not founderrors — hook is now registered with an absolute path resolved at install time. Resilient to any$PATHconfiguration; old installs auto-heal on the nextagenticmail claudecoderun. - Web UI fixes —
(m.flags ?? []).includes is not a functioncrash gone; sidebar folders (Sent / Drafts / Spam / Trash) now load their real IMAP mailboxes instead of all hitting/mail/inbox; Cmd+C no longer pops the compose modal; full mobile-responsive layout with an off-canvas sidebar. - Official logos — Claude starburst (from Wikipedia) and the AgenticMail
@mark frombranding/now ship bundled and render as the host avatar + topbar / favicon. - Selective wake —
wake: ["alice", "bob"]onsend_email/reply_email/forward_email/template_send/manage_drafts(send)tells the dispatcher to give a Claude turn only to named agents. The other CC'd recipients still receive the mail but stay asleep. Cuts token cost on large threads by ~10× when used. - Thread-close markers —
[FINAL],[DONE],[CLOSED], or[WRAP]in a subject tells the dispatcher this thread is done; no more wakes on any reply. check_activityMCP tool — see which agents the dispatcher has woken right now, how long they've been running, and a preview of recent completions. The answer to "did the agent I just emailed actually start working?"- Comprehensive markdown rendering in the shell's email viewer — bold, italic, headings, lists, task lists, tables, fenced code, links, images, HTML entities, depth-colored quote stripes (instead of literal
>>>>). - LLM-tolerant tool inputs —
batch_mark_read({ uids: "[1,2,3]" })and other common stringification mistakes now just work; coerced before validation. - Wake-budget circuit breaker — caps per-(agent, thread) wakes at 10 per 24h to stop reply loops and storms.
- Dedup guidance — wake prompts now tell agents to check their prior contributions before redoing work.
See CHANGELOG.md for the full release history.
AgenticMail is a self-hosted communication platform purpose-built for AI agents. It runs a local Stalwart mail server via Docker, integrates SMS/phone access via Google Voice or 46elks, exposes a REST API with 75+ endpoints, ships a lightweight Gmail-style web UI for human oversight, and works with any MCP-compatible AI client and OpenClaw via plugin. Each agent gets its own email address, phone number, inbox, and API key.
Table of Contents
- Why AgenticMail?
- Features
- Architecture
- Quick Start
- CLI Commands
- Gateway Modes
- Packages
- API Overview
- MCP Integration
- Host Integrations
- OpenClaw Integration
- Interactive Shell
- Security
- Configuration
- Development
- License
Why 🎀 AgenticMail?
AI agents need to communicate with the real world. Email is the universal communication protocol — every person and business has an email address. AgenticMail bridges the gap between AI agents and email by providing:
- Isolated mailboxes — each agent has its own email address, inbox, and credentials. Agents can't read each other's mail.
- Internet email connectivity — two gateway modes to send/receive real email (Gmail relay or custom domain with DKIM/SPF/DMARC).
- Security guardrails — outbound scanning prevents agents from leaking API keys, passwords, or PII. Blocked emails require human approval.
- Agent collaboration — agents can email each other, assign tasks, and make synchronous RPC calls.
- SMS / Phone number access — integrate Google Voice or 46elks for SMS receive/send, verification code extraction, and phone number access for AI agents.
- Realtime voice calls — bridge a phone mission to an OpenAI Realtime (
gpt-realtime) session so an agent can hold a live two-way conversation, with its persistent memory folded into the call. - Persistent agent memory — every agent has a categorised, confidence-decaying, searchable long-term memory that survives across conversations and is injected into prompts and voice sessions.
- Smart orchestration —
call_agentreplaces basic sub-agent spawning with auto mode detection, dynamic timeouts, runtime tool discovery, and async execution for long-running tasks. - Tool integrations — 100 MCP tools for any AI client, 89 OpenClaw tools, and a 44-command interactive shell.
- Self-updating —
agenticmail updatechecks npm, verifies OpenClaw compatibility, and updates both packages automatically.
Features
Email Operations
- Send email with text, HTML, attachments, CC/BCC, reply-to, and custom headers
- Receive email via IMAP with real-time SSE event streaming (IMAP IDLE)
- Reply and forward with proper In-Reply-To and References headers for threading
- Search by sender, subject, body text, date range, read/unread status
- Folders — create, list, move messages between folders
- Batch operations — mark read/unread, delete, move multiple messages at once
- Drafts — save, edit, and send draft emails
- Templates — reusable email templates with variable substitution
- Scheduled sending — queue emails for future delivery
- Signatures — per-agent email signatures
- Tags — label and categorize messages
- Contacts — manage address book per agent
Multi-Agent
- Account management — create, list, delete agents with unique email addresses
- Agent-to-agent email — agents can email each other directly via
@localhost - Task system — assign tasks to agents, claim, submit results, track status
- RPC calls — synchronous agent-to-agent calls with timeout (long-poll + SSE notification)
- Agent directory — discover other agents by name
You can watch the whole team work in one place with agenticmail shell:

That's a real multi-agent thread captured in the REPL — the host kicked off one email, Solène verified the file Cassian shipped to disk, and every step is sitting there as ordinary mail. No custom orchestration code. Type agenticmail shell any time you want to see what your agents have been up to.
Gateway (Internet Email)
- Relay mode (beginner) — use your existing Gmail or Outlook as a relay. Emails appear as
[email protected]. Setup takes 2 minutes. - Domain mode (advanced) — custom domain via Cloudflare. Agents send from
[email protected]with full DKIM signing, SPF, and DMARC records.- Automatic DNS configuration (MX, SPF, DKIM TXT, DMARC, tunnel CNAME)
- Cloudflare Tunnel for secure inbound traffic without exposing ports
- Cloudflare Email Worker for Email Routing (catch-all → worker → AgenticMail)
- Optional Gmail SMTP outbound relay for residential IPs without PTR records
- Domain purchase via Cloudflare Registrar
- DNS backup before any modifications
- Automatic
@domainemail alias for all existing agents
What setup-email actually exposes — read before connecting a relay
Heads-up for anyone running
agenticmail setup-email(or telling Claude / Codex to do it for them). Once that command succeeds, your sub-agents are reachable from the public internet via Gmail / Outlook plus-addressing. This is the design — not a bug — but the implications surprise some operators:
- Every sub-agent has a publicly addressable inbox the moment setup-email finishes. Anyone who knows your relay address can hit
yourrelay+<agentname>@gmail.comand the corresponding agent's AgenticMail inbox receives the message. Plus-addresses are publicly guessable (+secretary,+kepler, …) — don't treat them as a secret. - External mail wakes the dispatcher the same way internal
@localhostmail does. When a new-mail event lands on a watched inbox, the dispatcher runs dedup + thread-cache + wake-budget checks and spawns a Claude Code (or Codex) worker turn. Source doesn't matter —[email protected]andsecretary@localhostare indistinguishable from the dispatcher's point of view. - The host bridge takes a different path. Mail to
[email protected]or[email protected]doesn't spawn a worker — it goes tohandleBridgeMail, which uses the host SDK'sresumeoption to wake your last session headlessly. If that fails (session expired, host CLI not running), it falls through to the bridge-escalation email atsetup_operator_email. So external mail to the bridge can wake your CLI / forward to your phone. - Watchout — spam wakes Claude / Codex turns. A scraper that finds
[email protected]in a leaked address book can drive worker turns at your expense. Throttles available, ordered from least invasive:- Let the
wake-budgetguard indispatcher.handleEventrate-limit naturally (default cap per minute per agent). - Add inbound spam rules at the relay layer so spam gets filtered before the SSE event publishes (built-in spam filter + tags can do this; see the Security section below).
- For agents that should be internal-only, set
metadata.hostto a value no dispatcher matches, or stop the relay's IMAP poller from publishing on that inbox.
- Let the
Security
- Outbound guard — scans every outgoing email for sensitive data patterns:
- API keys and tokens (AWS, OpenAI, Stripe, GitHub, etc.)
- Passwords and credentials
- Private keys (SSH, PGP, RSA)
- PII patterns (SSN, credit card numbers)
- Internal URLs and configuration data
- Blocked emails are held for human-only approval (agents cannot self-approve)
- Spam filter — rule-based scoring engine for inbound email:
- Categories: phishing, scam, malware, commercial spam, social engineering, lottery scam
- Configurable threshold (default: 40)
- Skips internal agent-to-agent emails
- Runs on both relay inbound and SSE event streams
- Human-only approval flow — when an agent's email is blocked:
- The agent is informed the email was blocked and told to notify their owner
- The owner receives a notification email with full blocked email content, warnings, and pending ID
- Only the master key holder can approve or reject (
POST /mail/pending/:id/approve) - Agents can list and view their own pending emails but cannot approve or reject them
- System prompt guidelines instruct agents to inform their owner and wait, never attempt to bypass
- DKIM/SPF/DMARC — automatic DNS setup in domain mode for email authentication
- Rate limiting — configurable per-endpoint rate limits
SMS / Phone Number Access
- Provider selection — choose Google Voice legacy forwarding or 46elks direct SMS API/webhooks
- 46elks integration — send SMS through the provider API and receive inbound SMS through a secret-protected webhook
- Google Voice integration — give agents a real phone number via Google Voice
- Direct Voice web reading (primary, instant) — reads SMS directly from voice.google.com via browser
- Email forwarding (fallback) — Google Voice forwards SMS to email, agent auto-detects and records them during relay polling
- Separate Gmail polling — for users whose GV Gmail differs from relay email, runs a dedicated IMAP poll
- Verification codes — automatic extraction of OTP/verification codes from SMS (4-8 digit, alphanumeric, Google G-codes)
- Send SMS — direct provider API send when configured, or Google Voice web automation instructions for legacy configs
- Smart setup wizard — validates Gmail/GV email matching, warns about mismatches, collects separate credentials when needed
Realtime Voice Calls
- Live two-way conversation —
RealtimeVoiceBridgebridges a phone mission to an OpenAI Realtime (gpt-realtime) session so an agent can actually talk on the call, not just place it - Two carriers — 46elks or Twilio — pick the provider at phone setup;
RealtimeVoiceBridgeis generalised behind aRealtimeTransportAdapterso both run through one bridge. 46elks streams to/api/agenticmail/calls/realtime(PCM16 @ 24 kHz); Twilio Media Streams connects a<Connect><Stream>to/api/agenticmail/calls/twilio-stream(G.711 µ-law @ 8 kHz — OpenAIaudio/pcmu, no transcoding). Server-side VAD for turn-taking, caller barge-in relayed to the carrier. - Memory in the call — the agent's persistent memory is rendered and folded into the Realtime session instructions, so the model speaks with full continuity, as if it had always known those things
- Mission-tracked — the bridge resolves the connection to its phone mission by 46elks
callid, authenticates the connection token, and persists the conversation transcript to the mission - Hardened — per-frame audio size cap, bounded pre-connect buffer, fail-closed connection auth, terminal-state guard
- Opt-in — set
OPENAI_API_KEYto enable; without it, phone missions still place and track calls (call-control only) - Tools on the call — the Realtime session can call functions mid-call:
ask_operator(human-in-the-loop — pause, ask the operator, resume, or call back on disconnect),web_search,recall_memory,get_datetime
Telegram Channel
- Chat with your agents over Telegram — register a Telegram bot token, link a chat, and message your AgenticMail agent (and get replies) from Telegram
- Inbound webhook — authenticated with a constant-time
X-Telegram-Bot-Api-Secret-Tokencompare; uniform 403 on mismatch - Operator channel — carries
ask_operatornotifications and approvals, so a phone agent can reach you on Telegram mid-call - Secrets protected — bot tokens encrypted at rest, redacted from every log line and error
Media Toolset
- Nine local media tools —
media_tts,media_tts_voices,media_image_edit,media_video_edit,media_audio_edit,media_info,media_video_understand,media_voice_clone, plusmedia_capabilities - Text-to-speech — synthesise speech with Edge TTS (twelve voice presets); returns OGG/Opus, ready to send as a voice note
- Image editing — resize, crop, rotate, convert, compress, text overlay, flip, blur, sharpen, grayscale (ImageMagick)
- Video editing — basic (trim, GIF, compress, resize, add/remove audio, speed) and cinematic (color grading, transitions, captions, picture-in-picture, split screen, Ken Burns, slow motion, watermark, concatenate, auto-caption) — all via ffmpeg
- Audio editing — trim, convert, merge, volume, speed, extract-from-video, reverse, fade
- Video understanding — extract frames + transcribe audio (whisper.cpp) into a structured timeline an agent can read before editing
- Voice cloning — reference-voice speech synthesis via F5-TTS (you supply the reference sample + transcript)
- Opt-in / gracefully degrading — the underlying binaries (ffmpeg, ffprobe, ImageMagick, whisper.cpp, Python) are not bundled; every tool feature-detects the binary it needs and returns a clear, actionable install hint when one is absent — the server never crashes. Call
media_capabilities(or read the/healthmediablock) to see what is available - Safe by construction — every binary is invoked via
execFilewith an argument array (never a shell); untrusted input paths are validated (no control characters, no leading-dash flag-injection, must exist) and output files land only inside the configured media directory
Persistent Agent Memory
- Long-term, evolving knowledge — each agent has a categorised memory (knowledge, preference, correction, skill, reflection, …) that survives across every conversation
- Confidence + decay — entries carry a confidence score that decays for unaccessed knowledge;
criticalentries never decay; low-confidence and expired entries are pruned - BM25F search — a zero-dependency full-text index ranks recall by relevance, importance, recency, and access count
- Prompt + voice injection —
generateMemoryContext()renders a ranked markdown block for injection into agent prompts and realtime voice sessions - Private per agent — every memory endpoint is scoped to the authenticated agent; deleting an agent purges its memory
- Everywhere —
/memory*REST endpoints, MCP tools (memory,memory_reflect,memory_context,memory_stats), and OpenClaw tools (agenticmail_memory*)
Smart Orchestration (call_agent)
- Auto mode detection — reads task complexity, picks light/standard/full mode automatically
- Dynamic timeouts — 60s for quick tasks, 5+ minutes for deep research, 1 hour for async
- Runtime tool discovery — probes host config for available tools instead of static deny lists
- Async execution — long-running tasks run independently, auto-compact context, email results when done
- Structured RPC — sub-agents return JSON, not raw text
Integrations
- MCP server — 90+ tools for any MCP-compatible AI client
- OpenClaw plugin — 80+ tools with skill definition and system prompt guidelines
- REST API — 75+ endpoints, OpenAPI-style, Bearer token auth
- SSE events — real-time inbox notifications via Server-Sent Events
- Interactive CLI — 44 shell commands with arrow key navigation, body previews, retry logic
- Self-updating —
agenticmail updateor/updatein shell, with OpenClaw compatibility check
Architecture
┌──────────────────────────────────────────────────┐
│ AgenticMail │
│ │
AI Client ─MCP─> │ @agenticmail/mcp (100 tools, stdio transport)│
│ │ │
OpenClaw ─────> │ @agenticmail/openclaw (89 tools, plugin) │
│ │ │
HTTP clients──> │ ▼ │
│ @agenticmail/api (Express, 75+ endpoints) │
│ ├── Authentication (master key + agent keys) │
│ ├── Rate limiting (per-endpoint) │
│ ├── SSE streaming (real-time inbox events) │
│ └── Spam filter + Outbound guard │
│ │ │
│ ▼ │
│ @agenticmail/core (SDK layer) │
│ ├── AccountManager (CRUD agents in Stalwart) │
│ ├── MailSender (SMTP, nodemailer) │
│ ├── MailReceiver (IMAP, imapflow) │
│ ├── InboxWatcher (IMAP IDLE → events) │
│ ├── GatewayManager (relay + domain routing) │
│ │ ├── RelayGateway (Gmail/Outlook) │
│ │ ├── CloudflareClient (DNS, tunnels, etc) │
│ │ ├── TunnelManager (cloudflared) │
│ │ ├── DNSConfigurator (MX, SPF, DKIM) │
│ │ └── DomainPurchaser (Registrar API) │
│ ├── StalwartAdmin (mail server management) │
│ ├── EmailSearchIndex (FTS5 full-text search) │
│ └── Storage (SQLite + migrations) │
│ │ │
│ ▼ │
│ Stalwart Mail Server (Docker container) │
│ ├── SMTP (port 587) — submission │
│ ├── SMTP (port 25) — inbound delivery │
│ ├── IMAP (port 143) — mailbox access │
│ └── HTTP (port 8080) — admin API │
└──────────────────────────────────────────────────┘
│ │
┌─────────────┘ └──────────────┐
▼ ▼
Relay Mode Domain Mode
┌──────────────────┐ ┌──────────────────────┐
│ Gmail / Outlook │ │ Cloudflare │
│ IMAP polling │ │ ├── DNS zone │
│ SMTP relay │ │ ├── Tunnel │
│ Sub-addressing │ │ ├── Email Routing │
│ (+agent@gmail) │ │ ├── Email Worker │
└──────────────────┘ │ └── Registrar │
└──────────────────────┘
Data Flow
Sending email (relay mode):
- Agent calls
POST /mail/sendwith recipient, subject, body - API runs outbound guard scan — if sensitive data found, email is blocked and owner notified
- GatewayManager detects external recipient → routes to RelayGateway
- RelayGateway sends via Gmail SMTP as
[email protected] - Reply-To set to agent's relay address so replies route back
Sending email (domain mode):
- Agent calls
POST /mail/send - Outbound guard scan runs
- GatewayManager rewrites
agent@localhost→[email protected] - Email submitted to local Stalwart via SMTP (port 587)
- Stalwart signs with DKIM, resolves MX, delivers directly (or via Gmail relay)
Receiving email (relay mode):
- RelayGateway polls Gmail IMAP every 30 seconds for new messages
- New email detected → parsed → spam scored
- If not spam, delivered to agent's local Stalwart mailbox via SMTP
X-AgenticMail-Relay: inboundheader added for identification- InboxWatcher (IMAP IDLE) fires SSE event to connected clients
Receiving email (domain mode):
- External sender sends to
[email protected] - Cloudflare Email Routing catches all → routes to Email Worker
- Worker reads raw RFC822 message, base64-encodes, POSTs to
/api/agenticmail/mail/inbound - Inbound endpoint validates secret, parses email, delivers to agent's mailbox
- InboxWatcher fires SSE event
Quick Start
Prerequisites
- Node.js 22 or later (we use the built-in
node:sqlitemodule — zero native compilation, nonode-gypheadaches) brew(macOS) orapt(Linux) so the wizard can install Colima / Docker if you don't already have it
Two install paths
| Path | When to use | Command |
|---|---|---|
| Autonomous ✨ | Letting an AI agent (e.g. Claude Code) install AgenticMail on your behalf, or you just want zero questions | npm install -g @agenticmail/cli && agenticmail bootstrap |
| Interactive | You want to connect a Gmail relay or your own domain right away | npm install -g @agenticmail/cli && agenticmail setup |
Autonomous install (recommended for most users)
# Option A — one-liner (does the npm install + bootstrap for you, plus a Node 22+ preflight)
curl -fsSL https://raw.githubusercontent.com/agenticmail/agenticmail/main/install.sh | bash
# Option B — equivalent, manually
npm install -g @agenticmail/cli
agenticmail bootstrap
Either path: zero prompts. The pipeline:
agenticmail setup --yes— auto-installs Colima + Docker if missing, starts the Stalwart mail server, generates your master key, creates a default "secretary" agent. Skips external email/SMS setup (those need user-owned credentials; add them later if you want).agenticmail service install— registers a launchd plist (macOS) / systemd unit (Linux) so the API auto-starts on boot, and starts it now.- Waits for
GET /api/agenticmail/healthto come up on the configured port (defaulthttp://127.0.0.1:3829). agenticmail claudecode— wires the Claude Code integration in if you have Claude Code installed (idempotent / no-op otherwise).
After
No comments yet
Be the first to share your take.