AI Skill Hub 强烈推荐:Box — AI 图像生成工具中文文档 是一款优质的AI工具。AI 综合评分 8.5 分,在同类工具中表现稳健。如果你正在寻找可靠的AI工具解决方案,这是一个值得深入了解的选择。
Android本地AI应用套件,集成llama语言模型、Whisper语音识别、Stable Diffusion图像生成等多个AI能力,支持离线运行、语音对话、视觉识别和生物识别锁定,适合隐私保护意识强的安卓用户。
Box — AI 图像生成工具中文文档 是一款基于 Kotlin 开发的开源工具,专注于 Android应用、本地AI、隐私保护 等核心功能。作为 GitHub 开源项目,它拥有活跃的社区支持和持续的版本迭代,代码完全透明可审计,支持本地部署以保护数据隐私。无论是个人使用还是集成到企业工作流,都能提供稳定可靠的解决方案。
Android本地AI应用套件,集成llama语言模型、Whisper语音识别、Stable Diffusion图像生成等多个AI能力,支持离线运行、语音对话、视觉识别和生物识别锁定,适合隐私保护意识强的安卓用户。
Box — AI 图像生成工具中文文档 是一款基于 Kotlin 开发的开源工具,专注于 Android应用、本地AI、隐私保护 等核心功能。作为 GitHub 开源项目,它拥有活跃的社区支持和持续的版本迭代,代码完全透明可审计,支持本地部署以保护数据隐私。无论是个人使用还是集成到企业工作流,都能提供稳定可靠的解决方案。
# 克隆仓库 git clone https://github.com/jegly/Box cd Box # 查看安装说明 cat README.md # 按 README 完成环境依赖安装后即可使用
# 查看帮助 box --help # 基本运行 box [options] <input> # 详细使用说明请查阅文档 # https://github.com/jegly/Box
# box 配置说明 # 查看配置选项 box --config-example > config.yml # 常见配置项 # output_dir: ./output # log_level: info # workers: 4 # 环境变量(覆盖配置文件) export BOX_CONFIG="/path/to/config.yml"
<p align="center"> <img src="https://raw.githubusercontent.com/jegly/Box/main/images/b02.svg" alt="Box Header" width="84%" /> </p>
[
]()
[
]()
-FF79C6.svg)
-FF79C6.svg)
[
]() [
]() [
]() [
]()
[
]() [
]() [
]() [
]() [
]() [
]() [
]()
⭐️ If this project helped you, please star it — it helps others find it. We've hit 28K downloads! Thank you to everyone for supporting Box.
Note: If you're using a custom ROM (LineageOS, GrapheneOS, CalyxOS), download the custom-rom-support APK from the latest release instead.
1. Open Obtainium on your phone 2. Tap the + button 3. Paste this repo URL: https://github.com/jegly/Box 4. Tap Add
Recommended for most users: Main version
### Which version should I install? | Version | For | |---|---| | Main | Stock Android (Pixel, Samsung, etc.) | | Custom ROM | GrapheneOS, LineageOS, CalyxOS — no Google services | - The in-app updater is also available in Settings ### Setup steps 1. Tap the badge for your version above — this opens Obtainium with the repo pre-filled 2. Under APK filter regex, enter one of the following: - Main: Main - Custom ROM: custom-rom-support 3. Tap Add — Obtainium will find the latest release and install it 4. Future updates will be detected automatically > Note: As of v2.0.0, the in-app App version matches the Box > release version (2.0.0) — the earlier mismatch with the upstream Google AI > Edge Gallery build number (which showed 1.0.15) is fixed (#67). Box releases > are tracked via GitHub tags. Use Settings → Check for updates to see if a > newer Box release is available. Box is a security-hardened, feature rich fork of Google AI Edge Gallery — with on-device image generation (Bonsai Image 4B, FLUX.2 klein & Z-Image Turbo diffusion), Box Assist (spoken camera assistance for blind and low-vision users), AI image upscaling, face recognition, photo erase/inpainting, music & sound generation, voice mode (speech-to-speech AI chat), voice input, multilingual text-to-speech, document analysis and Q&A, vision AI, full GPU and Snapdragon/Tensor/MediaTek NPU acceleration, a hardened security posture (biometric lock, encrypted chat history, tap-jacking protection), llama.cpp support, and GGUF model import — and more
[!IMPORTANT] ## Disclaimer
Box began as a fork of Google AI Edge Gallery and is not affiliated with or endorsed by Google LLC. Google branding has been replaced throughout. Box has since diverged substantially from upstream — active merging with upstream stopped some time ago, and upstream has itself since adopted features that originated in Box. Box now carries roughly 50+ features not present in upstream Google AI Edge Gallery. Credit for the original underlying platform goes to Google and the original contributors.
<details> <summary>
</summary>
![]() Home — Chat |
![]() Home — Diffusion |
![]() Home — Voice |
![]() AI Chat |
![]() Model Config |
![]() Model Manager |
![]() Text to Speech |
![]() Voice Input |
![]() Whisper Scribe |
![]() Image Generation |
![]() Gemini Nano Hub |
![]() MCP — Add Server |
![]() Settings — Theme & Security |
![]() Settings — Behaviour & MCP |
![]() Settings — About |
</div>
</details>
--- > [!NOTE] >## What Box adds on top of upstream
Box started off as a fork of Google AI Edge Gallery. The upstream project is excellent — Box layers on additional capabilities and features not present in upstream.
| Area | What Box adds |
|---|---|
| Inference engines | llama.cpp (GGUF LLMs, full Vulkan GPU offload), stable-diffusion.cpp (image gen), whisper.cpp (STT) alongside LiteRT |
| Model import | Import any local GGUF file — not limited to the curated download list |
| NPU / TPU | All Snapdragon / Tensor / MediaTek variants bundled in one APK (upstream ships per-SoC) |
| Box Assist | Spoken camera assistance for blind and low-vision users — Live object/proximity callouts, Reading (OCR aloud), Describe (scene answers, spoken as generated), voice questions. One bundled download, autofocus + auto-flashlight, volume-button controls, TalkBack-friendly, fully offline |
| Voice mode / Vision mode | Free talk (continuous hands-free loop) and Vision talk (live camera + voice) |
| Image generation | On-device Stable Diffusion via GGUF, plus **Bonsai Image 4B** (512×512, recommended), **FLUX.2 klein (4B)** and **Z-Image Turbo** diffusion via LiteRT |
| Image recognition | **Identify**: MobileNet V2 / V3 Large (+ Tensor G5 NPU variant), **PlantNet** (1,081 plant species), **DM-Count** crowd counting — bundled, offline |
| Erase (inpainting) | Paint over anything in a photo and MI-GAN removes it — brush size, iterative erase, save to gallery (bundled, offline) |
| Music & sound generation | Generate music and sound effects from a text prompt, fully offline — quick clips, higher-quality audio, or long-form pieces up to ~3 minutes (**Sound** tab) |
| Image upscaling | AI super-resolution — enlarge any photo 4× on-device (XLSR / Real-ESRGAN / EDSR via LiteRT), models bundled, fully offline |
| Speech-to-text | On-device Whisper STT, plus **SenseVoice** for fast multilingual transcription (Chinese / English / Japanese / Korean / Cantonese, ~5× faster than Whisper) |
| Text-to-speech | **Supertonic** multilingual on-device TTS (5 languages, multiple voices) alongside Piper / Kokoro |
| Document analysis | Attach text files (.txt, .md, .csv, .kt, etc.) directly in chat |
| Document Q&A | RAG pipeline: import PDFs, embed with MiniLM on-device, ask questions grounded in document content — answers cite their source passages |
| Gemini Nano | 6 on-device ML Kit features (Summarize, Proofread, Rewrite, Chat, Describe, Speech) — entirely on-device via AICore on Pixel 9+/10 and recent Samsung / Xiaomi / OnePlus / OPPO / vivo flagships (both branches as of v2.0.0). Vision modes add live camera + still-image analysis with visual overlays (pose skeleton, 468-point face mesh) |
| Face Recognition | On-device, encrypted face recognition (both branches) — enroll and name people, then recognise them in photos or live from the camera. Multi-sample enrollment with alignment, capture-to-add, face-mesh overlay, SQLCipher-encrypted storage, fully offline and opt-in |
| Background Removal | ML Kit Subject Segmentation — remove backgrounds from photos, output a transparency-preserving PNG (main branch) |
| Chat history | Persisted to a SQLCipher-encrypted Room database, resumable across sessions |
| Security | Biometric app lock, hard offline mode, prompt sanitisation, audit log, tap jacking protection, accessibility data sensitivity |
| Themes | Catppuccin (14 accents), Dracula (7 accents), a bright **Light** theme, and Material You — picker in Settings, with the home screen tinted to match the active theme |
| Agent (skills + MCP) | 20 built-in skills (upstream has 9) plus Model Context Protocol — connect to remote MCP servers and give the model real tools, with per-call permission prompts |
| Math rendering | LaTeX expressions rendered as Unicode in chat, including inside markdown table cells |
| App shortcuts | Long-press icon → AI Chat or Box Assist for instant cold-start navigation |
| In-app updates | Settings → Check for updates — compares against latest GitHub release, downloads correct variant |
---
优秀的隐私保护导向AI应用,整合Google开源项目和成熟推理引擎,离线运行充分保护隐私。设计实用,维护活跃,是Android平台本地AI的标杆应用。
该工具使用 NOASSERTION 协议,商用场景请仔细阅读协议条款,必要时咨询法律意见。
AI Skill Hub 为第三方内容聚合平台,本页面信息基于公开数据整理,不对工具功能和质量作任何法律背书。
建议在沙箱或测试环境中充分验证后,再部署至生产环境,并做好必要的安全评估。
📄 NOASSERTION — 请查阅原始协议条款了解具体使用限制。
总体来看,Box — AI 图像生成工具中文文档 是一款质量优秀的AI工具,在同类工具中具备一定竞争力。AI Skill Hub 将持续追踪其更新动态,建议收藏备用,结合自身场景选择合适时机引入使用。
| 原始名称 | Box |
| 原始描述 | Private on-device AI suite for Android. Fork of Google AI Edge Gallery with llama.cpp, whisper.cpp, stable-diffusion.cpp, GGUF import, voice chat, vision AI, on-device image generation, biometric lock, encrypted history, and CPU/NPU/GPU acceleration. |
| Topics | Android应用本地AI隐私保护离线运行多模态AI |
| GitHub | https://github.com/jegly/Box |
| License | NOASSERTION |
| 语言 | Kotlin |
收录时间:2026-05-22 · 更新时间:2026-05-30 · License:NOASSERTION · AI Skill Hub 不对第三方内容的准确性作法律背书。