Fork Version License GitHub all releases Android Kotlin Hybrid Engine llama.cpp stable-diffusion.cpp whisper.cpp LiteRT GGUF Import Snapdragon NPU Google Tensor G5 MediaTek Gemini Nano RAG MCP Servers Vision Document Analysis Super-Resolution Voice Mode SenseVoice Supertonic SQLCipher Biometric Offline
MusicGeneration Box Assist Image Generation Vulkan

If this project helped you, please ⭐️ star it to help others find it.

Download

Download Box v3.3.2 APK

Note: If you're using a custom ROM (LineageOS, GrapheneOS, CalyxOS), download the custom-rom-support APK from the latest release instead.

Install via Obtainium

  1. Open Obtainium on your phone
  2. Tap the + button
  3. Paste this repo URL:
    https://github.com/jegly/Box
  4. Tap Add

Recommended for most users: Main version

Which version should I install?

Version For
Main Stock Android (Pixel, Samsung, etc.)
Custom ROM GrapheneOS, LineageOS, CalyxOS — no Google services
  • The in-app updater is also available in Settings

    Setup steps

    1. Tap the badge for your version above — this opens Obtainium with the repo pre-filled
    2. Under APK filter regex, enter one of the following:
      • Main: Main
      • Custom ROM: custom-rom-support
    3. Tap Add — Obtainium will find the latest release and install it
    4. Future updates will be detected automatically

    Note: As of v2.0.0, the in-app App version matches the Box release version (2.0.0) — the earlier mismatch with the upstream Google AI Edge Gallery build number (which showed 1.0.15) is fixed (#67). Box releases are tracked via GitHub tags. Use Settings → Check for updates to see if a newer Box release is available.

Box is a security-hardened, feature rich fork of Google AI Edge Gallery — with on-device image generation (FLUX.2 klein & Z-Image Turbo diffusion), Box Assist (spoken camera assistance for blind and low-vision users), AI image upscaling, face recognition, photo erase/inpainting, music & sound generation, voice mode (speech-to-speech AI chat), voice input, multilingual text-to-speech, document analysis and Q&A, vision AI, full GPU and Snapdragon/Tensor/MediaTek NPU acceleration, a hardened security posture (biometric lock, encrypted chat history, tap-jacking protection), llama.cpp support, and GGUF model import — and more

[!IMPORTANT]

Disclaimer

Box began as a fork of Google AI Edge Gallery and is not affiliated with or endorsed by Google LLC. Google branding has been replaced throughout. Box has since diverged substantially from upstream — active merging with upstream stopped some time ago, and upstream has itself since adopted features that originated in Box. Box now carries roughly 50+ features not present in upstream Google AI Edge Gallery. Credit for the original underlying platform goes to Google and the original contributors.

Changelog v1.0.7 – v3.3.2

Version Feature Details
v3.3.2 Downloads fixed Model downloads are reliable again after 3.3.1 — no more failing mid-download or stalling at 100%. A previously stuck model downloads normally on the first try.
v3.3.2 GGUF GPU crash fix (really this time) The Snapdragon GPU crash fix from 3.3.1 now actually ships in the build.
v3.3.2 Biometric lock + database encryption The biometric app lock works alongside database encryption again — the two are independent, and the app re-locks when reopened.
v3.3.1 Live Translator (NEW, Sound tab) Two people, two languages — tap your button, speak, and the other person reads and hears it in their language. Runs on your installed Gemma audio model (E2B/E4B), each phrase translated on its own for flat latency. 24 languages, fully offline.
v3.3.1 4 new models Granite 4.0 350M (IBM's tiny fast tier, 468 MB), MiniCPM5-1B in int8 and int4 builds, and experimental Gemma 4 26B (A4B) — Google's mixture-of-experts Gemma for 16 GB+ RAM devices.
v3.3.1 Fixes GGUF models no longer crash on GPU on some Snapdragon devices (Adreno driver quirk). Rotating or folding the phone no longer unloads the model. Custom-ROM: TPU/GPU chat works again on de-Googled devices (GrapheneOS).
v3.3.0 🦯 Box Assist — a camera that talks (NEW) Spoken camera assistance for blind and low-vision users, under the Core tab. Live mode calls out people, obstacles and objects with how close they are; Reading mode reads mail, labels and menus aloud; Describe mode describes the scene, spoken as it thinks; voice questions — double-tap, ask out loud, and Box answers against what the camera sees. One download bundles everything (vision models + the Describe brain + speech recognition). Continuous autofocus with pre-capture focus sweeps, automatic flashlight in the dark, a blur check on Reading, physical volume-button controls, hold-to-repeat, TalkBack coexistence, screen never times out, and a launcher long-press shortcut straight into it. Fully offline.
v3.3.0 ⚡ GGUF engine rebuilt — real GPU acceleration The llama.cpp engine got a ground-up overhaul: full Vulkan GPU offload via the CPU/GPU chip in any GGUF chat, a massively faster CPU mode (a flaw routed CPU prompt processing through the GPU — 0.7 → 21 tok/s on a Pixel 6a), instant replies (weights read up front, reopened chats replay their history during the loading screen), a tokens/sec stat under every GGUF reply, a new Settings → GGUF Models panel (context size, CPU threads, GPU layers, mmap, mlock, Q8 KV cache), sturdier imports with byte-verification, and automatic GPU→CPU retry. llama.cpp updated to a current build.
v3.3.0 🎨 On-device image generation — FLUX.2 klein & Z-Image Turbo Two full text-to-image diffusion models running 100% on-device via LiteRT: FLUX.2 klein (4B) — photorealistic images in 4 steps (~7.4 GB download) — and Z-Image Turbo (9 steps), which shares nearly a gigabyte of files with klein so Box doesn't download them twice. Multi-gigabyte downloads now resume without refetching finished files, progress bars show honest totals, and a model only shows "downloaded" when every file is actually present.
v3.3.0 🔍 Five new vision models — bundled, work instantly Identify now hosts four model families in one picker: MobileNet V2, MobileNet V3 Large (with a Pixel Tensor G5 NPU variant), PlantNet (identify 1,081 plant species from a photo) and DM-Count crowd counting. New Erase tile — paint over anything in a photo and MI-GAN inpainting removes it (brush size, iterative erase, save to gallery). Upscale gains EDSR ×4. All bundled in the APK — no download, fully offline.
v3.3.0 📱 Android 14 support Minimum Android version lowered from 15 to Android 14 — Box now installs on a whole generation more of phones.
v3.3.0 Fixes & polish Box Assist: fixed a first-open black screen (camera and mic permission requests raced each other) and made repeat scene descriptions as fast as the first. Fixed a case where an already-loaded model would never signal "ready", leaving features waiting forever. Download cards show accurate total sizes before you tap.
v3.2.0 🎵 On-device music & sound generation Make music and sound effects from a text description — completely offline, nothing leaves your phone. Three tiers under the new Sound tab: SoundGen (quick clips & sound effects in seconds), SoundGen HD (higher-quality audio up to ~24s), and SoundGen HD Long (full pieces up to ~3 minutes). Set the length, then play, save, or share the result. The generator for each tier downloads on first use, then runs entirely on-device.
v3.2.0 Identify — on-device image recognition Point Box at a photo and it tells you what's in it — 1000+ everyday objects, animals and scenes. Pick from your gallery or take a new shot. Fully offline, hardware-accelerated on supported devices.
v3.2.0 Tabs reorganised — Sound & Core Clearer home tabs: Sound groups the audio features, Core groups chat & assistant.
v3.2.0 Chat remembers on reopen Reopening a conversation now replays recent context to the model, so it picks up where you left off — new chats still start fresh.
v3.1.0 NPU now works on Snapdragon & MediaTek — for the first time This is the first Box build where on-device NPU acceleration actually runs on Snapdragon and MediaTek phones. Previous builds shipped the NPU models but crashed on load. Box now ships the Qualcomm and MediaTek NPU dispatch libraries rebuilt to match the LiteRT runtime plus an updated Qualcomm AI stack (QNN 2.47), with per-vendor builds so each phone loads the correct driver — NPU chat and benchmarking now run on those devices. The Pixel / Tensor G5 path is unchanged. (#81, #83, #88)
v3.1.0 Smoother NPU chat on small models Long conversations on the Gemma 3 1B NPU model no longer abruptly stop or error when the context fills — Box slides the context window so the chat keeps going. Added safeguards so the small NPU model doesn't get stuck repeating itself or return empty replies. (Snapdragon / MediaTek NPU only — Tensor G5 and GPU/CPU are untouched.)
v3.1.0 Fix — NPU benchmark crash (#81) Benchmarking an NPU model no longer crashes.
v3.1.0 Polish New animated "Initializing model" loading screen; removed the "Experimental" tag from Mobile Actions; tidied up model descriptions.
v3.0.0 Major UI overhaul — Material 3 Expressive A top-to-bottom interface refresh. The app now moves with spring-physics motion: home cards bounce in and respond to taps, chat messages rise and fade in as they arrive, and screen transitions use Material 3 slide-and-fade. The jump to 3.0.0 reflects how much of the UI changed — the models and engines are unchanged.
v3.0.0 11 new themes A set of terminal-inspired palettes — Fairy Floss, Nord, Bim, Borland, C64, Cobalt Neon, Grass, Homebrew Ocean, Mono Amber, Mono Red, and Synthwave — selectable from a new dropdown in Settings, alongside the existing System, Light, Catppuccin, and Dracula themes.
v3.0.0 Custom app & chat fonts Choose from 13 bundled font families (Nunito plus Cormorant Garamond, DotGothic16, IBM Plex Mono / Serif, Instrument Serif, Playfair Display, Press Start 2P, Quicksand, Space Grotesk, Turret Road, Viaoda Libre, and more), each previewed in its own typeface — with an optional separate font just for chat messages.
v3.0.0 Text-size slider Scale text across the whole app and chat from 0.8× to 1.4×, on top of your system font size.
v3.0.0 Themed app icon With "Themed icons" enabled in your launcher, the Box icon now tints to your system Material You colours.
v3.0.0 Settings, reorganised The long settings list is now grouped into smooth, collapsible categories — Appearance, Privacy & Security, Network & Tools, Chat & Voice, and About.
v3.0.0 Theme-aware task screens Open any task (Chat, Diffusion, Voice…) and the background now follows your selected theme with the same accent tint as the home screen — no more flat black behind a colourful theme. Cleaner, icon-free task headers, a tidy box-shaped menu button, and a new Material 3 wavy download-progress indicator.
v3.0.0 Fix — NPU crash on Snapdragon & MediaTek (#82, #83) NPU models could hard-crash on load on non-Pixel devices (e.g. Galaxy S26 Ultra, Xiaomi 14T Pro) because the wrong hardware dispatch library was being loaded. Box now selects the correct Qualcomm / MediaTek runtime per device. The Pixel 10 / Tensor G5 path is unchanged and re-verified.
v3.0.0 Fix — Settings flash & jank Changing the text-size slider no longer flashes the home screen behind Settings, and opening Settings or expanding a category no longer jumps — the dialog is now fixed-size and animates its contents internally.
v2.0.2 New model tier — Gemma 3 270M Brand-new ultra-lightweight model (~460–555 MB) — fast and low-RAM, ideal for quick tasks on modest devices. Ships dedicated NPU builds for Snapdragon (SM8550 / 8650 / 8750 / 8750-AB / 8850) and MediaTek Dimensity (MT6991 / MT6993).
v2.0.2 New models — Gemma 3 1B-IT with broad NPU coverage Gemma 3 1B now ships dedicated on-device NPU builds across Snapdragon (SM8550 → SM8850, incl. the Samsung SM8750-AB) and MediaTek Dimensity (MT6989 / 6991 / 6993), plus a universal GPU/CPU build. Each device automatically downloads the build that matches its chip.
v2.0.2 New models — Gemma 3n E2B & E4B (multimodal) Text, image and audio input, up to 32K context, with Gemma 3n's selective-parameter architecture. Run on GPU/CPU on every device; NPU-accelerated on MediaTek (MT6993).
v2.0.2 Samsung Galaxy S26 Ultra (SM8850) NPU models Added SM8850 ("Snapdragon 8 Elite Gen 5") allowlist keys across the new Gemma 3 1B and 270M entries, so dedicated NPU models now appear and run on the S26 Ultra.
v2.0.1 Fix — Snapdragon 8 Elite NPU crash (SM8750 / SM8750-AB) The audio sub-graph was incorrectly routed to the NPU on all SM8750 devices, causing an instant hard crash (SIGABRT) when loading the Snapdragon NPU model — no error popup, just an immediate exit. Audio always uses CPU regardless of the primary backend, matching upstream behaviour. Fixes Red Magic NX799J, iQOO 13, and any other SM8750 or SM8750-AB device.
v2.0.1 Fix — Samsung Galaxy S25 / S26 Ultra NPU models not listed Samsung's "Snapdragon 8 Elite for Galaxy" variant reports SM8750-AB as its SoC identifier, not SM8750. The model allowlist only matched sm8750, so dedicated NPU models were invisible to all S25 and S26 Ultra users.
v2.0.1 New model — Gemma 4 E2B (Qualcomm QCS8275 / Dragonwing IQ8) Added an NPU model entry for the Qualcomm QCS8275 SoC. Appears automatically on matching hardware.
v2.0.0 Google Tensor G5 (Pixel 10) acceleration Gemma now runs on the Pixel 10's Tensor G5 TPU, not just the GPU. Supported models route to the TPU automatically and expose a dedicated TPU option in the accelerator picker.
v2.0.0 MediaTek NPU support Bundled the MediaTek dispatch runtime and added the first models that run on MediaTek Dimensity neural engines.
v2.0.0 New models Gemma 4 E2B (Tensor G5) and Gemma 4 12B (GPU); Gemma 3 1B-IT (Tensor G5), Gemma 3n E2B (MediaTek, multimodal) and Qwen3 0.6B (MediaTek)
v2.0.0 Face Recognition — on-device & encrypted New tool in the image section: detect, enroll and name people, then recognise them in photos or live from the camera, fully offline. Multi-sample enrollment with face alignment, capture-to-add, an on-screen face mesh, and a settings panel (match strictness, front camera, show %, clear all). All face data is encrypted on-device (SQLCipher) and never leaves the phone — opt-in and user-enrolled only.
v2.0.0 New Light theme + theme-aware home A crisp, wallpaper-independent Light theme, and the home background now follows your selected theme (System / Light / Catppuccin / Dracula) instead of always being black.
v2.0.0 Gemini Nano Hub on custom-ROM The full Gemini Nano hub (Summarize / Proofread / Rewrite / Describe / Chat / Speech) is now included in the custom-rom-support build too, degrading gracefully on devices without AICore (ML-Kit vision tools still work).
v2.0.0 Nano document-attach crash + leak fixes Fixed a crash when attaching a document in Summarize/Proofread/Rewrite (the file picker could be hijacked by the photo picker on Android 14+) — now uses the proper document picker with a clean fallback. Also fixed GenAI service/memory leaks when switching between Nano features.
v2.0.0 Copy button on code blocks Fenced code blocks in chat now render with a language label and a one-tap Copy code button.
v2.0.0 SenseVoice in Chat The chat mic now works with a loaded SenseVoice model (priority Whisper → SenseVoice → system) instead of dead-ending when no Whisper model is present.
v2.0.0 Speculative decoding in chat Speculative / Multi-Token-Prediction decoding is available for Gemma 4 in chat (off by default).
v2.0.0 Fix #69 — agent mode with text-only models Agent mode no longer force-loads vision on models that don't support it, which previously blocked text-only imported models entirely.
v2.0.0 Fix #67 — correct installed version Aligned versionName with the public version, so Obtainium / Android's "App version" report the right number (no more false "update available"). This is why the release jumps to 2.0.0.
v2.0.0 Smaller download Native libraries are now compressed inside the APK — the main build drops from 400 MB+ to ~278 MB (they're extracted on install).
v1.0.12 SenseVoice — multilingual speech-to-text New card in the Voice tab. Transcribes Chinese, English, Japanese, Korean and Cantonese fully offline, roughly 5× faster than Whisper on CPU. Live "listening" preview while you talk, a multi-message transcript log (copy / delete / clear), language picker, punctuation & number formatting, and optional emotion / audio-event tags. (#68)
v1.0.12 Supertonic — multilingual text-to-speech New card in the Voice tab. Lightweight (~66M param) on-device speech synthesis in English, Korean, Spanish, Portuguese and French, with multiple built-in voices and adjustable speed. Fully offline — text never leaves the device.
v1.0.12 AI Image Upscaling (super-resolution) New Upscale tool in the image tab. Enhance and enlarge any photo 4× on-device and save it to your gallery. Three models bundled in the app — XLSR (fast), Real-ESRGAN General (balanced), Real-ESRGAN x4plus (quality) — run via LiteRT, no download required. Photos are auto-rotated (EXIF-aware) before upscaling.
v1.0.12 Gemini Nano Vision — visual overlays (main) Pose detection now draws a skeleton overlay and Face Mesh a 468-point mesh directly on the camera preview and still images (previously text-only). Added copy buttons on every vision result, an adjustable live refresh rate (Fast / Balanced / Slow / Power-saver) with a Freeze/Resume toggle, front/rear camera switching on all modes, and image upload from your gallery.
v1.0.12 Models browser organised by type The model list is now grouped into Language models / Speech-to-Text / Text-to-Speech / Image generation / Other instead of one flat alphabetical list.
v1.0.12 New language models Added TinyLlama 1.1B, Phi-4-mini, TinySwallow 1.5B, VibeThinker 1.5B, and Qwen3 8B to the download list.
v1.0.12 Markdown & LaTeX rendering overhaul (#42) Headers, bullet/numbered lists and bold text now render correctly even when mixed with inline math on the same line; bold that spans a math expression no longer shows literal **; wide display equations scroll instead of being clipped.
v1.0.12 Clearer model guidance + UI cleanup Gemma 4 E2B labelled "Recommended", E4B "Best overall for flagship devices," with cleaned-up model descriptions. Removed promotional banners/links from the MCP and Agent screens (sample-prompt chips kept).
v1.0.12 Fix #59 — Snapdragon NPU crash Vision/audio sub-backends now follow the primary backend on the NPU path, fixing hard crashes on some Snapdragon devices.
v1.0.12 Fix #61 — leftover model files Orphaned model-version directories are cleaned up after app updates.
v1.0.12 Fix #65 — GrapheneOS speech hang Restored the SpeechRecognizer availability gate (custom-rom-support build).
v1.0.12 Fix — config dialog crash Opening the model settings dialog on small-context-window (<2000) models no longer crashes.
v1.0.12 Android SDK 37 + deeplink fix Updated compile/target SDK to 37 and fixed the notification tap deep link.
v1.0.11 MCP server support The Agent tab can now connect to external Model Context Protocol servers (e.g. gitmcp.io/<owner>/<repo>) and give the model access to remote tools. Off by default — enable in Settings, add a server URL, accept the disclaimer. Every tool call fires a per-call permission dialog (Allow once / Always allow / Deny). Hard Offline Mode disables MCP.
v1.0.11 "Agent Skills" renamed to "Agent" Reflects the addition of MCP tools alongside the existing 20 built-in skills. Internal IDs unchanged.
v1.0.11 Broader NPU init crash recovery (main) Snapdragon 8 Elite / Vivo OriginOS users (e.g. iQOO 13) reporting hard crashes on NPU model open now fall back silently to GPU instead. Any catchable NPU init exception is recovered, not just TF_LITE_AUX.
v1.0.11 Pixel 8/9 TPU label Tensor G3 / G4 devices now show the TPU accelerator label alongside Pixel 10 (isPixelDevice() broadened from isPixel10()).
v1.0.11 Smoother streaming render BufferedFadingMarkdownText two-layer crossfade reduces markdown re-render jank during token streaming.
v1.0.11 Chat scroll performance snapshotFlow + derivedStateOf translated to Box's LazyColumn. Significantly fewer Compose recompositions per generated token.
v1.0.11 ChatGPT-style chat layout User and assistant messages both left-aligned, restoring Box's original look.
v1.0.11 Downloaded-model tick icon Once a model is on device, the model picker chip and Model Manager show a filled-circle tick instead of the download-arrow icon.
v1.0.11 Gemma 4 model hashes refreshed Gemma 4 E2B / E4B / E2B-Snapdragon entries updated to upstream's latest commits (6e5c4f1e… / 28299f30…).
v1.0.11 R8 keep rule for tool calls Release builds preserve @Tool method names on every ToolSet subclass — MCP and Agent skills now work in release APKs (was silently broken).
v1.0.11 Upstream merged to 1.0.15 Internal versionName bumped to match upstream gallery 1.0.15 (cherry-picked over multiple sessions; chat history, model schema, and other heavily-customised Box paths preserved).
v1.0.10 Gemini Nano hub 6 on-device ML Kit features powered by Gemini Nano on Pixel 9+ (via AICore, NPU/TPU-accelerated): Summarize, Proofread, Rewrite, Chat, Describe Image, and Speech-to-Text. First use triggers an automatic background download of Gemini Nano (~1–2 GB via AICore).
v1.0.10 Nano Chat — multi-session Persistent multi-turn chat with Gemini Nano. Sessions are stored in the existing encrypted SQLCipher database, auto-titled from the first message, and fully resumable. Sessions can be renamed or deleted. Long-press any bubble to copy.
v1.0.10 Document attachment in Nano Proofread and Rewrite now accept attached documents (PDF, TXT, MD) — content is read and passed to Gemini Nano as context.
v1.0.10 Live camera in Describe Image Gallery tab + Live Camera tab. Camera tab binds an ImageCapture use case — tap Capture to send the current frame to Nano for description.
v1.0.10 Background Removal New tool powered by ML Kit Subject Segmentation (main branch). One tap removes the background from any photo with a transparency-preserving PNG output. Includes a "Trim transparent edges" toggle. Save or share the result.
v1.0.10 Catppuccin + Dracula themes Three-way theme picker in Settings: System (Material You) / Catppuccin (14 accents) / Dracula (7 accents). Accent colour persists across restarts with no first-frame flicker.
v1.0.10 Tap jacking protection toggle New toggle in Settings (on by default) — filterTouchesWhenObscured blocks touch events when an overlay is detected, preventing tap-jacking attacks.
v1.0.10 Accessibility data sensitivity toggle New Settings toggle hides app content from untrusted accessibility services. Off by default (note: incompatible with TalkBack).
v1.0.10 LaTeX in table cells Inline math inside markdown table cells no longer wraps across multiple lines. Uses Compose InlineTextContent to embed math as a single placeholder inside Text().
v1.0.10 Import button simplified Home screen import button label shortened to just "Import" (removed "GGUF · LiteRT" subtitle).
v1.0.10 NPE crash fix Fixed a null-pointer crash on startup and on Retry caused by a broken fallback comparator in groupTasksByCategory.
v1.0.9 Document Q&A New RAG pipeline: import PDFs and ask questions grounded in the document. Uses MiniLM embeddings (on-device, LiteRT) for chunk retrieval — model only sees the relevant passages. Every answer cites the source chunks it used.
v1.0.9 Model picker in Document Q&A Choose which downloaded LLM handles answering — defaults to first available, switchable mid-session.
v1.0.9 Kokoro TTS (English) Single Kokoro model (csukuangfj/kokoro-en-v0_19, ~346 MB) replaces broken individual-voice entries. Correct tensor shapes and metadata — works first time.
v1.0.9 13 Piper voices 8 new voices: LibriTTS-R, HFC Female, HFC Male, Arctic (US English); Thorsten (German); UPMC (French); MLS 10246 (Spanish); Huayan (Chinese Mandarin). 13 total across both branches.
v1.0.9 10 Whisper models Expanded from 3 hardcoded to 10: Tiny, Base, Small, Medium, Large-v3-Turbo, and Large-v3 — each in multilingual and English-only variants. Shared across Audio Scribe and Voice Input.
v1.0.9 Gemma-4-E2B-it (Snapdragon 8 Elite) NPU-optimised variant added to the model allowlist — visible only on SM8750 devices.
v1.0.9 Fix #46 — Audio Scribe OOM crash Replaced boxed List<Float> (~16 bytes/sample) with a primitive growing FloatArray (4 bytes/sample). 30-min audio at 16 kHz no longer causes ~460 MB excess allocation.
v1.0.9 Fix #47 — TTS silent with non-Amy voice Auto-init and GrapheneOS TTS fallback now filter by download status before selecting a voice model (custom-rom-support only).
v1.0.8 Saved System Prompts Save, name, and reuse system prompts from the model settings dialog. Tap to apply, swipe to delete.
v1.0.8 Restore Defaults New button in model settings resets all sliders (temperature, top-K, top-P, max tokens) back to defaults in one tap.
v1.0.8 System prompt actually applied Changing the system prompt mid-session now correctly resets the conversation with the new instruction — previously saved in UI but not passed to the model.
v1.0.8 Markdown fix in math responses Plain-text segments in chat bubbles now render through the Markdown pipeline, fixing broken formatting in responses that mix text and LaTeX math.
v1.0.8 Randomised inference seed Each conversation now uses a unique random seed for more varied outputs on CPU backend.
v1.0.8 GPU determinism root cause found LiteRT LM v0.11.0 hard-caps max_top_k: 1 on devices without a GPU sampler, forcing greedy decoding. Switch to CPU for varied outputs. Reported upstream as issue #817.
v1.0.7 Gemma 4 E2B & E4B updated Model files refreshed on HuggingFace — new commit hashes, smaller sizes, same multimodal capabilities.
v1.0.7 Speculative decoding / MTP Multi-Token Prediction reads capability from the model file itself. Gemma 4 E2B reaches 66–91 tok/s on Galaxy S26 Ultra (GPU + spec) vs 52 tok/s plain GPU.
v1.0.7 Sustained Performance Mode setSustainedPerformanceMode(true) locks clocks during inference — no mid-conversation thermal throttling on long generations.
v1.0.7 Benchmark spec decoding toggle Benchmark screen shows a speculative decoding toggle for supported models.
v1.0.7 AI Chat app shortcut Long-press the Box icon → AI Chat jumps straight into chat, even from a cold start.
v1.0.7 In-app update checker Settings → Check for updates — fetches the latest GitHub release and offers a direct download link for your variant.
v1.0.7 Model import from list Whisper and TTS models can now be imported directly from the model list.


Related

Built OfflineLLM first — a privacy-first Android chat app with a pure llama.cpp backend.


What is Box?

Box is an Android app for running AI entirely on-device — chat, voice mode, image generation, image upscaling, speech-to-text, text-to-speech, document analysis, and vision, all without a network connection. It inherits the full feature set of the upstream Google AI Edge Gallery and layers on top: encrypted conversations, biometric lock, hard offline mode, and three additional native inference engines (llama.cpp, stable-diffusion.cpp, whisper.cpp) alongside LiteRT.

Box: On-Device AI. No Cloud. No Compromise.

What makes Box unique? You can sit at your desk, tap two buttons, and have a real flowing voice conversation with an AI — no wake word, no account, no server, no subscription. It listens, thinks, and speaks back sentence by sentence before it's even finished generating. Point the camera at something and ask about it out loud. The AI sees it and answers. All of it runs on the phone in your hand, completely offline, faster than you'd expect.


Screenshots


[!NOTE]

What Box adds on top of upstream

Box is a fork of Google AI Edge Gallery. The upstream project is excellent — Box just layers on additional capabilities:

Area What Box adds
Inference engines llama.cpp (GGUF LLMs, full Vulkan GPU offload), stable-diffusion.cpp (image gen), whisper.cpp (STT) alongside LiteRT
Model import Import any local GGUF file — not limited to the curated download list
NPU / TPU All Snapdragon / Tensor / MediaTek variants bundled in one APK (upstream ships per-SoC)
Box Assist Spoken camera assistance for blind and low-vision users — Live object/proximity callouts, Reading (OCR aloud), Describe (scene answers, spoken as generated), voice questions. One bundled download, autofocus + auto-flashlight, volume-button controls, TalkBack-friendly, fully offline
Voice mode / Vision mode Free talk (continuous hands-free loop) and Vision talk (live camera + voice)
Image generation On-device Stable Diffusion via GGUF, plus FLUX.2 klein (4B) and Z-Image Turbo diffusion via LiteRT
Image recognition Identify: MobileNet V2 / V3 Large (+ Tensor G5 NPU variant), PlantNet (1,081 plant species), DM-Count crowd counting — bundled, offline
Erase (inpainting) Paint over anything in a photo and MI-GAN removes it — brush size, iterative erase, save to gallery (bundled, offline)
Music & sound generation Generate music and sound effects from a text prompt, fully offline — quick clips, higher-quality audio, or long-form pieces up to ~3 minutes (Sound tab)
Image upscaling AI super-resolution — enlarge any photo 4× on-device (XLSR / Real-ESRGAN / EDSR via LiteRT), models bundled, fully offline
Speech-to-text On-device Whisper STT, plus SenseVoice for fast multilingual transcription (Chinese / English / Japanese / Korean / Cantonese, ~5× faster than Whisper)
Text-to-speech Supertonic multilingual on-device TTS (5 languages, multiple voices) alongside Piper / Kokoro
Document analysis Attach text files (.txt, .md, .csv, .kt, etc.) directly in chat
Document Q&A RAG pipeline: import PDFs, embed with MiniLM on-device, ask questions grounded in document content — answers cite their source passages
Gemini Nano 6 on-device ML Kit features (Summarize, Proofread, Rewrite, Chat, Describe, Speech) — entirely on-device via AICore on Pixel 9+/10 and recent Samsung / Xiaomi / OnePlus / OPPO / vivo flagships (both branches as of v2.0.0). Vision modes add live camera + still-image analysis with visual overlays (pose skeleton, 468-point face mesh)
Face Recognition On-device, encrypted face recognition (both branches) — enroll and name people, then recognise them in photos or live from the camera. Multi-sample enrollment with alignment, capture-to-add, face-mesh overlay, SQLCipher-encrypted storage, fully offline and opt-in
Background Removal ML Kit Subject Segmentation — remove backgrounds from photos, output a transparency-preserving PNG (main branch)
Chat history Persisted to a SQLCipher-encrypted Room database, resumable across sessions
Security Biometric app lock, hard offline mode, prompt sanitisation, audit log, tap jacking pro