PullMD

Release Docker Pulls CI License MCP

Self-hosted URL-to-Markdown service for humans and AI agents.

PullMD takes any web URL and returns clean, readable Markdown — no navigation, no ads, no boilerplate. It auto-detects Reddit and Hacker News threads (with full comment trees), uses Cloudflare's native Markdown when available, runs Mozilla Readability + Trafilatura on static HTML, and as a last resort renders JavaScript-heavy pages via headless Chromium (Playwright sidecar) before extracting.

As of v3, PullMD goes beyond web pages: it also converts documents (PDF, Office, EPUB), images, audio, and YouTube videos to Markdown, and emits a leaner, token-efficient body by default. See What's new in v3 below.

It ships as:

  • a PWA frontend with raw/rendered and live-frontmatter view toggles, one-tap sharing of the output to other apps (Web Share API), a download button that saves the result as a .md file under the server-suggested name, dark/paper themes, history, archive, share links, and conversion of local HTML files (drag-and-drop on desktop, file picker on desktop and mobile)
  • a REST API at GET /api?url=…
  • an MCP server at POST /mcp (Streamable-HTTP transport, stateless)
  • a Claude Code skill as a downloadable zip

Every conversion gets an 8-hex share id that works as a stable live-endpoint: GET /s/:id returns the cached markdown and re-fetches from the source if older than one hour. Use the share id as a fixed URL that always returns fresh content — useful for subreddit feeds and similar.


What's new in v3

PullMD v3 grows from a web-page reader into a general anything-to-Markdown service for agents, with a leaner default output. Everything beyond plain web extraction is opt-in and degrades gracefully - left unconfigured, v3 handles web pages exactly like v2, just with a cleaner body by default.

  • Clean body by default - the Markdown body is now just # Title + content. The source URL, fetch date, and all metadata moved into the YAML frontmatter, so nothing is duplicated and you spend fewer tokens. Reddit posts follow the same rule: subreddit, author, upvotes, and publish date live in the frontmatter (subreddit, author, upvotes, published), not the body. This is the one breaking change: set PULLMD_SOURCE_HEADER=true to restore the old inline header, and use PULLMD_FRONTMATTER_FIELDS to trim which fields are emitted. See MIGRATION.md.
  • Documents → Markdown - PDF, Word, PowerPoint, Excel, EPUB and more, by URL or upload (POST /api/file, drag-and-drop in the PWA).
  • High-quality PDF tables (OCR) - an opt-in, vendor-neutral OCR tier (?pdf=ocr) for table-grade PDF conversion, with automatic fallback to the free path.
  • Images & audio → Markdown - opt-in captioning and transcription via any OpenAI-compatible or local model; runs inside pullmd, no extra container required.
  • YouTube transcripts - title, description and transcript with clickable timecodes, no API key required.
  • Richer frontmatter - extraction source, quality, and (for media/OCR) model + token/page usage for cost tracking, plus a configurable field allowlist.

Self-hosters upgrading from v2.x: the clean-body change is the only breaking one - MIGRATION.md has the one-line opt-out. Everything else is additive.

Added in the 3.x line since then:

  • Hacker News pipeline (3.1) - items, comment permalinks and listings through a purpose-built converter, plus Web Share and an instant frontmatter toggle in the PWA.
  • X-Transcript-Status (3.2) - tells a transient YouTube rate-limit apart from a genuinely missing transcript.
  • SSRF protection (3.3) - private, loopback, link-local, CGNAT and cloud-metadata targets are rejected by default, on every fetch path and every redirect hop.
  • Query-scoped extraction (3.4) - ?query= returns only the sections relevant to a question, with a max_tokens budget.
  • Site recipes opened up (3.5/3.6) - JSON-LD-to-frontmatter, a contributor guide, and select.content so a recipe can name the article body outright.
  • Coverage guard (3.7) - recovers pages where extraction kept only a sliver of the body; see PULLMD_COVERAGE_GUARD.
  • Account controls (3.8) - a non-admin can clear entries from their own history, self-registration can be closed with PULLMD_ALLOW_SIGNUP, and scripts/admin.js create-user creates accounts from the shell.
  • Download button (3.9) - the PWA saves a result as a .md file, named by the server via X-Suggested-Filename and optionally date-prefixed.

Quick start

Pre-built multi-arch images (linux/amd64, linux/arm64) live on Docker Hub. Drop the compose file somewhere and run:

mkdir pullmd && cd pullmd
curl -O https://raw.githubusercontent.com/AeternaLabsHQ/pullmd/main/docker-compose.yml
docker compose up -d
# → http://localhost:3000

That's it. No .env needed: every variable has a sensible default and PullMD listens on port 3000. Add a .env next to the compose file to override anything (see Configuration).

docker-compose.yml (zero-config, abridged)

services:
  pullmd:
    image: aeternalabshq/pullmd:latest
    container_name: pullmd
    restart: unless-stopped
    ports:
      - "${PORT:-3000}:3000"
    environment:
      - PUBLIC_URL=${PUBLIC_URL:-http://localhost:${PORT:-3000}}
      - TRAFILATURA_URL=http://trafilatura:8001/extract
      - PLAYWRIGHT_URL=http://playwright:8002/render
      - MARKITDOWN_URL=http://markitdown:8003/convert
      - CACHE_DB=/data/cache.db
    volumes:
      - ./data:/data
    networks:
      - pullmd-internal
    depends_on:
      - trafilatura
      - playwright
      - markitdown

  trafilatura:
    image: aeternalabshq/pullmd-trafilatura:latest
    container_name: pullmd-trafilatura
    restart: unless-stopped
    networks:
      - pullmd-internal

  playwright:
    image: aeternalabshq/pullmd-playwright:latest
    container_name: pullmd-playwright
    restart: unless-stopped
    networks:
      - pullmd-internal

  markitdown:
    image: aeternalabshq/pullmd-markitdown:latest
    container_name: pullmd-markitdown
    restart: unless-stopped
    mem_limit: ${MARKITDOWN_MEM_LIMIT:-1g}
    networks:
      - pullmd-internal

networks:
  pullmd-internal:
    driver: bridge

Abridged for readability — the docker-compose.yml in the repo additionally passes every optional .env variable through to the containers (Reddit credentials, auth, media/OCR keys, YouTube options, output shaping). Use the curl -O command above rather than copying this block, or .env overrides beyond the basics won't reach the containers.

Note: the Playwright sidecar adds ~3.7 GB to your image cache (Chromium + Firefox + WebKit binaries from the official Playwright base image). It's optional — leave PLAYWRIGHT_URL unset and the playwright service block off, and PullMD silently degrades to static extraction with a fallback note in the metadata.

Note: the MarkItDown sidecar is optional. Leave MARKITDOWN_URL unset and remove the markitdown service block to disable document conversion. Web-page URLs always work without it.

Mirror on GHCR: ghcr.io/aeternalabshq/{pullmd,pullmd-trafilatura,pullmd-playwright,pullmd-markitdown}. Replace the image: lines if you prefer GitHub's registry.

Behind Traefik

For deployments behind Traefik with TLS, use docker-compose.traefik.yml instead. Same images, but with Traefik labels and the proxy external network. Set HOST_DOMAIN in .env:

curl -O https://raw.githubusercontent.com/AeternaLabsHQ/pullmd/main/docker-compose.traefik.yml
echo "HOST_DOMAIN=pullmd.example.com" > .env
docker compose -f docker-compose.traefik.yml up -d

Local development (no Docker)

git clone https://github.com/AeternaLabsHQ/pullmd.git
cd pullmd
npm install
npm start             # http://localhost:3000
npm test              # node --test

Configuration

All variables go in .env (copy from .env.example):

v3.0.0 output format change: the markdown body is clean by default - just # Title followed by content. The source URL, fetch date, and all extraction metadata remain in the YAML frontmatter unchanged - the body no longer duplicates them. Set PULLMD_SOURCE_HEADER=true to restore the old inline header. Use PULLMD_FRONTMATTER_FIELDS to pick which frontmatter fields are emitted (handy for trimming tokens in agent pipelines).

Variable Required Purpose
HOST_DOMAIN Traefik variant only Public hostname without scheme. Used by Traefik routing and as fallback for PUBLIC_URL. Unused by the default compose.
PUBLIC_URL no Full public origin embedded in /help and the skill zip. Defaults to https://${HOST_DOMAIN}.
TRAFILATURA_URL no URL of the Trafilatura sidecar's /extract endpoint. Unset → skip Trafilatura, Readability only.
PLAYWRIGHT_URL no URL of the Playwright sidecar's /render endpoint. Unset → skip Playwright fallback for JS pages.
MARKITDOWN_URL no URL of the MarkItDown sidecar's /convert endpoint. Unset → document-conversion path disabled; POST /api/file returns 502.
PULLMD_VISION_API_KEY / …_BASE_URL / …_MODEL no Image captioning via an OpenAI-compatible vision endpoint. Enabled when the key is set. _MODEL defaults to gpt-4o-mini.
PULLMD_STT_API_KEY / …_BASE_URL / …_MODEL no Audio transcription via an OpenAI-compatible /audio/transcriptions endpoint. Enabled when the key is set. _MODEL defaults to whisper-1.
PULLMD_LLM_API_KEY / …_BASE_URL no Shared fallback credentials for vision + STT when the per-modality vars are unset. Key and base URL only - there is no PULLMD_LLM_MODEL, and setting one is ignored (the server warns at startup).
PULLMD_PDF_OCR_API_KEY / …_BASE_URL / …_MODEL no Opt-in high-quality PDF→Markdown via an OCR provider that preserves tables (reference: Mistral OCR mistral-ocr-latest). Triggered per request with ?pdf=ocr or a recipe fetch.pdf: ocr. Default PDF handling stays the free markitdown path. _MODEL defaults to mistral-ocr-latest.
MARKITDOWN_YOUTUBE no Set to true to route YouTube URLs through the markitdown sidecar (returns title + description + transcript). No API key required. Default: off.
MARKITDOWN_YT_TIMECODES no (sidecar) Default timecode format in transcripts: links (YouTube timestamp links, default), plain (bare [MM:SS] labels), none (transcript text only). Overridable per-request via ?yt_timecodes=.
MARKITDOWN_YT_CHUNK no (sidecar) Transcript block size in seconds (default 30). 0 keeps the original per-snippet granularity. Overridable per-request via ?yt_chunk=.
MARKITDOWN_YT_LANGS no (sidecar) Comma-separated preferred transcript languages (e.g. de,en). Falls back to the first available language if none of the preferred ones exist.
MARKITDOWN_YT_PROXY no (sidecar) HTTP(S) proxy URL for YouTube requests. Datacenter IP addresses are often rate-limited by YouTube's transcript API; a residential or ISP proxy can help.
REDDIT_CLIENT_ID no OAuth credentials for Reddit. Without them, PullMD uses the public JSON API (lower rate limit).
REDDIT_CLIENT_SECRET no
REDDIT_USER_AGENT no Reddit requires a unique UA. Default: PullMD/1.0 (URL-to-Markdown service).
DISABLE_PUBLIC_HISTORY no When true, hides the global recent-conversions list and archive (/api/history + /api/archive return 403, frontend hides the section). /s/:id share links keep working. Default: false.
PULLMD_USER_AGENT no Pin a single outbound User-Agent for every web fetch. Disables rotation. Useful for CI or when one specific UA is known to work.
PULLMD_UA_FEED_URL no URL of a JSON feed of current real-world UAs. Default: WinFuture23/real-world-user-agents. Set to an empty string to disable live refresh and rely on the built-in seed pool.
PULLMD_AUTH_MODE no disabled (default) / single-admin / multi-user. See "Authentication" below.
PULLMD_ALLOW_SIGNUP no Self-registration in multi-user mode. Default: on. false / 0 / no / off closes /signup (404) and removes the "create an account" link from the login page. Accounts can still be created with node scripts/admin.js create-user <email>.
PULLMD_ADMIN_EMAIL required when AUTH_MODE != disabled, on first startup Bootstrap email for the first admin user.
PULLMD_ADMIN_PASSWORD required when AUTH_MODE != disabled, on first startup Bootstrap password (min 8 chars).
PULLMD_AUTH_TOKEN no Legacy bearer token compat (single-admin mode only, deprecated).
PULLMD_SOURCE_HEADER no Set to true to restore the legacy inline source header in the body (# Title + **domain** · date + url; for Reddit the **r/sub** · u/user · N ↑ line). Default (unset): clean body - just the H1 title; source/date/post meta live in the frontmatter.
PULLMD_FRONTMATTER_FIELDS no Comma-separated allowlist of frontmatter fields to emit (e.g. title,url,source,llm_tokens). Unset = all fields. Trims tokens. Unknown names are ignored with a startup warning.
PULLMD_ALLOWED_HOSTS no Comma-separated CIDRs and/or exact hostnames that may be fetched even though they resolve into a blocked range. Empty by default = every internal target is blocked. See SSRF protection.
PULLMD_SITE_RECIPES no Path to a JSON file of extra site recipes, merged on top of the built-ins. Alternative to data/site-recipes.json.
PULLMD_FILENAME_DATE_PREFIX no Prefix template for the suggested download filename (X-Suggested-Filename). Unset = no prefix. Tokens YYYY MM DD HH mm ss are substituted in local time, all other characters pass through; anything outside A-Za-z0-9._- ends up as a hyphen. Example: YYYY-MM-DD-HH-mm-ss- gives 2026-08-01-13-33-42-YT-some-talk-dQw4w9WgXcQ.md.
PULLMD_COVERAGE_GUARD no Set to off to disable the coverage guard. Default (unset): on. The guard notices when an extraction kept only a sliver of the page - the failure mode of page-builder one-pagers, whose chapters sit in flat sibling containers that Readability's single-candidate scoring discards - and re-converts the container holding the body instead. It only ever grows the result, records source: coverage-guard, and explains itself in metadata.extractorReason.

PUBLIC_URL matters for self-hosting: the help page and downloadable skill embed it as the canonical endpoint. Set it correctly and your users get a copy-paste setup that points at your instance.

PullMD rotates its outbound User-Agent for the web fetch path from a pool of current desktop browsers, refreshed every 48 hours from a live feed of real-world UAs maintained by @WinFuture23. A built-in seed pool ensures rotation works even when the feed is unreachable. Set PULLMD_USER_AGENT to pin a single UA, or PULLMD_UA_FEED_URL to point at your own feed. The Reddit path keeps its dedicated REDDIT_USER_AGENT because Reddit's API expects a stable, identifying UA.

DISABLE_PUBLIC_HISTORY=true is the privacy switch for shared instances (multi-tenant VPS, office deployments). Conversions still get cached and assigned share IDs; users just can't see what other users have fetched. Anyone with a known /s/:id link still gets their markdown back. Use this as a stopgap until per-user scoping lands.


Authentication (v2.0+)

Version pinning: :latest tracks the newest release (v3). v3's only breaking change is the clean-body output format — to stay on the v2.x output format instead, pin the explicit major tag:

services:
  pullmd:
    image: aeternalabshq/pullmd:2

PullMD ships with three auth modes. Pick one with PULLMD_AUTH_MODE:

Mode Behavior
disabled Default. No auth, everything open. Existing v1.x behavior.
single-admin One user, credentials from env vars. No self-signup. For homelab.
multi-user Self-signup at /signup (unless PULLMD_ALLOW_SIGNUP is off), login at /login, per-user data isolation.

In single-admin and multi-user modes, PULLMD_ADMIN_EMAIL + PULLMD_ADMIN_PASSWORD bootstrap the first admin user on first startup. After that, changing these env vars does not change the password — use the admin CLI:

docker compose exec pullmd node scripts/admin.js reset-password [email protected]

Create an account without opening self-registration (useful when PULLMD_ALLOW_SIGNUP is off):

docker compose exec pullmd node scripts/admin.js create-user [email protected]

Both commands read the password from stdin, so they need it attached: docker compose exec, docker exec -it, or a pipe (echo "…" | docker exec -i <container> node scripts/admin.js …). A plain docker exec without -i aborts with an error and exit code 2 instead of doing nothing.

Auth boundary

Endpoint Auth required (when mode != disabled)
/, /help, static assets, /pullmd.zip no
/login, /signup, /api/me (auth surface) no
/s/:id (share links) no
/api, /api/stream yes
POST /api/html, POST /api/file yes
/mcp yes
/api/history, /api/archive yes
DELETE /api/cache/:id, DELETE /api/cache yes
/api/stats, /api/storage, /api/config (aggregate) no

Cache deletes are scoped to the caller. An admin (and every caller in disabled mode) removes the shared, URL-deduped cache row, which affects every user's history. A regular user only unlinks the entry from their own history - the shared row and its /s/:id share link keep working. The response says which happened via "scope": "user" | "global".

Authentication paths

  1. Session cookiesPOST /login sets pullmd_session (HttpOnly, SameSite=Lax, Secure over HTTPS, 90-day TTL with sliding expiry). The PWA uses this automatically.
  2. API keys — generate at /settings, send via Authorization: Bearer pmd_<32-char-base62>. Stored as SHA-256 hashes; only shown once at creation.
  3. Legacy PULLMD_AUTH_TOKEN — deprecated. single-admin mode only. Maps to admin user. Kept for migration compatibility; slated for removal in a future major release.

Migration from v1.x

See MIGRATION.md for the full upgrade checklist. The TL;DR: leave PULLMD_AUTH_MODE unset and v2.0 behaves exactly like v1.x.

OAuth 2.1 (claude.ai Web Connector)

PullMD ships with a full OAuth 2.1 Authorization Code flow so the claude.ai web app's Custom Connector feature can authenticate users against your PullMD instance. All endpoints needed by the spec are implemented: Dynamic Client Registration (RFC 7591), PKCE-S256 (RFC 7636), Authorization Server Metadata (RFC 8414), Protected Resource Metadata (RFC 9728), and Token Revocation (RFC 7009).

Setup:

  1. Set PULLMD_AUTH_MODE to single-admin or multi-user (OAuth requires Phase-1 auth).
  2. Set OAUTH_JWT_SECRET to a 32+ character random string (openssl rand -hex 32).
  3. Set PUBLIC_URL to your instance's public origin (e.g. https://pullmd.example.com).
  4. In claude.ai → Settings → Connectors → Add custom connector, point it at https://pullmd.example.com/mcp — claude.ai discovers everything else automatically via the well-known endpoints.
  5. The first time the user clicks the connector, they'll be redirected to PullMD's /login, then to a consent screen, then back to claude.ai.

Tokens:

  • Access tokens are JWTs (HS256), TTL 1 hour, audience-bound to your /mcp URL.
  • Refresh tokens are opaque (pmd_rt_…), TTL 30 days, rotated on every refresh, with reuse-detection that invalidates the entire refresh chain on replay.
  • Revoke a token via POST /oauth/revoke (RFC 7009).

Scope: Currently a single mcp:full scope (URL conversion + history read). Granular scopes are tracked for a future minor release.

Shipped in v2.3.0 (issues #6 and #10).


AI-agent integration

Three install paths. Once your instance is running, ${PULLMD_URL}/help shows the same boxes with your URL pre-filled. Replace ${PULLMD_URL} below with your hostname (e.g. https://pullmd.example.com).

1. Universal prompt

Drop into any chat agent (ChatGPT, Claude, Gemini, …):

When you need to read a web page, fetch it via PullMD instead of your
built-in fetch/browse tool - not just when that one fails:

  GET ${PULLMD_URL}/api?url=<URL>

Returns clean Markdown (text/markdown). Optional query params:

  comments=false        skip Reddit comments
  comment_depth=N       comment nesting depth (default 3)
  frontmatter=true      prepend YAML metadata block
  format=text           strip Markdown, return plain text
  nocache=true          bypass the 1h cache and refetch
  render=force|skip     override the auto Playwright fallback
  pdf=ocr               high-quality PDF conversion (tables)
  query=<text>          got a specific question? pass it and get back only
                        the matching sections (70-95% fewer tokens)
  max_tokens=N          budget for query= (default 600, 64-20000)
  lang=de|en            language for the comments section header

Response headers worth checking:
  X-Source       reddit | hackernews | cloudflare | readability | trafilatura |
                 playwright | markitdown | youtube | pdf-ocr | ...
  X-Quality      0.0-1.0 extraction confidence
  X-Share-Id     8-hex permalink, openable as /s/<id>
  X-Transcript-Status  youtube only: ok | none | blocked | error
                 (blocked/error = transient, not cached — retry later)

Reddit URLs are auto-detected (incl. redd.it short links and /s/ shares).
Hacker News URLs are auto-detected too — items, comment permalinks, and the
front/newest/ask/show/jobs listings.
Use this whenever you would otherwise fetch raw HTML — the markdown is
much cleaner and saves significant context window space. When you only
need specific information rather than the whole document, pass your
question as query= and get back just the relevant sections.

2. Claude Code skill

pullmd.zip is auto-built with your URL embedded:

curl -O ${PULLMD_URL}/pullmd.zip
mkdir -p ~/.claude/skills
unzip pullmd.zip -d ~/.claude/skills/
# Restart Claude Code; the skill activates on web-reading requests.

Upgrading from pre-v3? The skill was renamed from web-reader to pullmd in v3.0.0. Installing the new zip does not replace an existing install — remove the old one first, or both skills will be active side by side:

rm -rf ~/.claude/skills/web-reader

(The old download URL /web-reader.zip keeps working as a redirect.)

3. MCP server

Remote MCP server at ${PULLMD_URL}/mcp (Streamable-HTTP transport, stateless). Three tools: read_url, get_share, list_recent. Server-side updates reach every client automatically — no local install needed.

Claude Code (CLI):

claude mcp add --transport http pullmd ${PULLMD_URL}/mcp

Claude Desktop / Cursor / other MCP hosts — JSON config:

{
  "mcpServers": {
    "pullmd": {
      "type": "http",
      "url": "${PULLMD_URL}/mcp"
    }
  }
}

Once registered, the three tools surface natively in the agent — no prompt instructions needed, the LLM picks them up via their schema descriptions.

MCP client compatibility

Client Bearer (Authorization: Bearer pmd_...) OAuth (v2.3+) Notes
Claude Code CLI Recommended. Generate a key at /settings.
Cursor Same as CLI.
Claude Desktop Connector UI lacks a header field — use OAuth.
claude.ai (web) Requires OAuth.

OAuth shipped in v2.3.0 — enable it via OAUTH_JWT_SECRET (see above) and Claude Desktop / claude.ai connect natively. The reverse-proxy workaround below is only needed on instances that keep OAuth disabled.

Claude Desktop limitation (without OAuth)

The Claude Desktop "Add custom connector" UI accepts URL + OAuth Client ID/Secret but no custom-header field. Additionally, claude_desktop_config.json entries with "type": "http" are silently rewritten to {} after Desktop launches (current Desktop only honors stdio servers in that file).

If OAuth is not enabled on your instance, the practical workaround for Claude Desktop users is a reverse proxy that accepts the auth token as either a bearer header (for CLI) or as a URL path prefix (for Desktop, which has no header field).

Caddy workaround for Claude Desktop

Contributed by @WinFuture23:

@bearer header Authorization "Bearer {$AUTH_TOKEN}"
handle @bearer { reverse_proxy pullmd:3000 }

@token_path path /{$AUTH_TOKEN}/* /{$AUTH_TOKEN}
handle @token_path {
    uri strip_prefix /{$AUTH_TOKEN}
    reverse_proxy pullmd:3000
}

Then in Claude Desktop's connector dialog, use the URL with the token path prefix: https://your-instance.com/<TOKEN>/mcp. CLI clients keep using the Authorization header as normal.

This is a stopgap pattern; enabling the built-in OAuth removes the need for it.


API

Endpoint Returns
GET /api?url=… Markdown (or JSON / plain text via format=). Also handles direct links to documents (PDF, Office, EPUB, …) when the markitdown sidecar is configured.
GET /api/stream?url=… Server-Sent Events stream of extraction-stage status, ending in a result event. Used by the PWA.
POST /api/html Convert a local/raw HTML document (body = HTML, max 10 MB). Never cached — no history entry, no share link.
POST /api/file Convert an uploaded document (raw file bytes in body; set Content-Type to the file's MIME type; filename via X-Filename header or ?filename=; max 25 MB). Returns Markdown. Requires the markitdown sidecar (MARKITDOWN_URL) for document types (PDF, Office, EPUB, ...); image and audio uploads instead use the PULLMD_VISION_* / PULLMD_STT_* tier (no markitdown container needed).
GET /s/:id Cached Markdown by share id; refreshes from source if > 1 h old.
GET /api/history Recent conversions (JSON).
GET /api/archive Paginated full archive.
GET /api/storage Cache size / hit-rate stats.
GET /api/stats Extraction telemetry (sources, quality, latency). ?window=-7 days.
GET /api/config Which optional tiers this instance has enabled (auth mode, markitdown, vision, STT, PDF OCR, YouTube). The PWA reads it to show or hide controls.
GET /api/recipes/status Which site recipes loaded, which were rejected, and any frontmatter fields dropped by the allowlist.
POST /mcp Streamable-HTTP MCP endpoint (3 tools: read_url, get_share, list_recent).
GET /pullmd.zip Claude Code skill bundle, with this instance's URL baked in (/web-reader.zip redirects here).
GET /help Bilingual user/agent setup guide.

/api parameters

Param Default Notes
url Required.
comments true Include Reddit / Hacker News comments. Ignored for other URLs.
comment_depth 3 Max nesting depth (1–10). Applies to Reddit and Hacker News.
comment_limit none Max top-level comments (uncapped by default).
frontmatter false Prepend YAML metadata.
format md text strips Markdown; json returns structured response.
nocache false Bypass the 1-hour cache.
render auto force → always render via Playwright. skip → never render. Bypasses cache.
extractor auto Force readability / trafilatura / playwright and skip the quality pick. Bypasses cache.
pdf ocr → route PDFs through the OCR tier. Bypasses cache.
yt_timecodes / yt_chunk see YouTube Transcript format overrides. Bypass cache when set.
lang de Comments-section header language (de or en).
query Set this when you need specific information from a page rather than the whole document: pass the question, get back only the matching sections (see Query-scoped extraction). Empty/whitespace-only is treated as absent — output is unchanged.
max_tokens 600 Token budget for query extraction (6420000). No effect without query; raise it when the answer likely spans several sections. Only validated when query is non-empty; an invalid value returns 400.

Both query and max_tokens are also available on the MCP read_url tool.

Response headers

  • X-Sourcereddit · cloudflare · readability · readability-fallback · trafilatura · playwright · recipe-content · coverage-guard · markitdown · youtube · image-caption · audio-transcript · pdf-ocr
  • X-Quality0.01.0 extraction confidence
  • X-Share-Id — the 8-hex permalink id
  • X-Suggested-Filename — a download filename for this conversion, e.g. YT-some-talk-dQw4w9WgXcQ.md. YouTube gets title plus video id, images/audio/documents keep the original file name from the URL, everything else uses the title; PULLMD_FILENAME_DATE_PREFIX prepends a date. Also on /s/:id, POST /api/html and POST /api/file; on /api/stream the same value arrives as suggestedFilename in the result event, since SSE has no headers to carry it.
  • X-Transcript-Status — YouTube only: ok · none · blocked · error. blocked (YouTube rate-limited the transcript fetch, HTTP 429) and error are transient and not cached — retry later; none means the video genuinely has no transcript
  • X-Extractedtrue / false. Present only when query is active.
  • X-Extract-Confidencehigh · medium · low. Present whenever X-Extracted is present, except when the page was already small enough that extraction was skipped.
  • X-Extract-Sections — number of sections/regions included in the returned markdown.
  • X-Extract-Original-Tokens / X-Extract-Returned-Tokens — estimated token counts (chars / 4) of the full page and of the returned markdown.

Query-scoped extraction

When you have a specific question about a long page, you rarely need the whole document. Pass the question as ?query=<text> on GET /api (or the query param on the MCP read_url tool) to get back only the sections relevant to it instead of the full Markdown body - typically 70-95% fewer tokens on long pages:

GET /api?url=https://example.com/long-article&query=how+does+caching+work

It's a BM25 ranking over the page's heading-based sections (falling back to paragraph-level blocks for pages with fewer than two headings) — no network calls, no LLM involved.

  • Budgetmax_tokens caps the returned size (default 600, range 6420000, estimated as chars / 4).
  • Whole-page fallback — the full page is returned unchanged when either the page is already small enough that extraction wouldn't help (estimated tokens ≤ max(800, 1.25 × max_tokens)), or nothing in the page scores against the query's terms. In both cases extracted is false; the no-match case reports confidence: low, the small-page case omits confidence.
  • Atomic blocks — code fences, tables, lists, and blockquotes are never split; if any part of one is relevant, the whole block is kept. Non-contiguous output regions are joined by an elision marker (<!-- … -->).
  • ?frontmatter=true adds extracted, extract_confidence, sections_selected, original_tokens, returned_tokens to the YAML block. format=json adds a top-level extract object with the same data (extracted, confidence, sectionsSelected, originalTokens, returnedTokens).
  • Extraction runs on the already-cached full page — the cache always stores the complete Markdown, so several different query values against the same URL cost one fetch (within the normal 1-hour cache TTL), not one per query.
  • Omitting query (or passing an empty/whitespace-only string) leaves /api behavior completely unchanged — no extra headers, no extra JSON or frontmatter fields.
  • MCP read_url has no response headers; when nothing matches, it prepends an in-band comment (<!-- query-extract: no match; returning full page -->) to the returned markdown instead.

SSRF protection

A URL-fetching service will happily be pointed at its own network if you let it. Since v3.3.0 PullMD does not: before any fetch, the target host is resolved and the resulting addresses are checked. Anything in a private, loopback, link-local, CGNAT or cloud-metadata range is refused with 403 — including 169.254.169.254 (AWS/GCP/Azure) and 100.100.100.200 (Alibaba).

  • GET /api and the MCP read_url tool check the URL up front and answer 403 before doing any work. GET /api/stream is covered by the same guard one layer down, in the Reddit / web / Playwright fetch paths, and surfaces the refusal as an SSE error event.
  • Every redirect hop is re-checked before it is followed, so a public URL that redirects inward is blocked at the hop, not after it.
  • IPv6 is covered too, including blocked v4 ranges reached through transition addressing (IPv4-mapped, NAT64, 6to4, IPv4-compatible).
  • The rejection happens before the cache is touched — a blocked URL never produces a cache row or a share id.

To reach an internal host on purpose (an intranet wiki, a document server on the same network), allowlist it:

# CIDRs, exact hostnames, or both
PULLMD_ALLOWED_HOSTS=10.0.5.0/24,wiki.internal

Empty (the default) means every internal target stays blocked.

Known residual: the guard checks resolved IPs at request time but does not pin the outbound socket to the address it checked, so a DNS-rebinding attack (where the name resolves differently between check and connect) is not fully closed. When PullMD runs behind an outbound HTTP proxy, the proxy resolves names itself — its egress filtering, not this guard, is the authoritative layer for traffic it forwards.

Reported as #41, shipped in v3.3.0.


Cache & TTLs

  • /api?url=… re-fetches from source if the cache row is older than 1 hour.
  • /s/:id does the same on-demand refresh, so share links double as live endpoints.
  • Cache rows are pruned 90 days after the last write. /s/:id hits keep the row alive (since they trigger refresh + write); read-only access does not extend the TTL.
  • If the source is unreachable on refresh, the last good snapshot is served — share links keep working even when the original URL dies.

Document conversion

When the markitdown sidecar is running (set MARKITDOWN_URL=http://markitdown:8003/convert), PullMD can convert document files to Markdown in addition to web pages.

Supported formats: PDF, DOCX, PPTX, XLSX/XLS, EPUB, ZIP (contents liste