经 AI Skill Hub 精选评估,imp AI技能包 获评「强烈推荐」。这款AI工具在功能完整性、社区活跃度和易用性方面表现出色,AI 评分 8.0 分,适合有一定技术背景的用户使用。
高性能LLM推理引擎,支持NVIDIA Blackwell GeForce
imp AI技能包 是一款基于 Cuda 开发的开源工具,专注于 cuda、cpp、cuda-graphs 等核心功能。作为 GitHub 开源项目,它拥有活跃的社区支持和持续的版本迭代,代码完全透明可审计,支持本地部署以保护数据隐私。无论是个人使用还是集成到企业工作流,都能提供稳定可靠的解决方案。
高性能LLM推理引擎,支持NVIDIA Blackwell GeForce
imp AI技能包 是一款基于 Cuda 开发的开源工具,专注于 cuda、cpp、cuda-graphs 等核心功能。作为 GitHub 开源项目,它拥有活跃的社区支持和持续的版本迭代,代码完全透明可审计,支持本地部署以保护数据隐私。无论是个人使用还是集成到企业工作流,都能提供稳定可靠的解决方案。
# 克隆仓库 git clone https://github.com/kekzl/imp cd imp # 查看安装说明 cat README.md # 按 README 完成环境依赖安装后即可使用
# 查看帮助 imp --help # 基本运行 imp [options] <input> # 详细使用说明请查阅文档 # https://github.com/kekzl/imp
# imp 配置说明 # 查看配置选项 imp --config-example > config.yml # 常见配置项 # output_dir: ./output # log_level: info # workers: 4 # 环境变量(覆盖配置文件) export IMP_CONFIG="/path/to/config.yml"
<p align="center"> <img src="docs/logo.svg" alt="imp" width="500"> </p>
<p align="center"> <a href="LICENSE"><img src="https://img.shields.io/github/license/kekzl/imp?style=flat&color=blue" alt="License"></a> <img src="https://img.shields.io/badge/CUDA-13.3-76b900?style=flat&logo=nvidia" alt="CUDA 13.3"> <img src="https://img.shields.io/badge/C++-23-00599C?style=flat&logo=cplusplus" alt="C++23"> </p>
---
imp is an LLM inference engine that targets exactly one chip: the NVIDIA RTX 5090.
What it is
sm_120a), with its own GGUF and SafeTensors loaders, tokenizer, paged KV cache and kernels.What it is not
Tracks main rather than the latest release.
git clone https://github.com/kekzl/imp.git && cd imp
docker compose build imp-server
docker run --gpus all -v ./models:/models -v imp-cache:/home/imp/.cache/imp \
-p 127.0.0.1:8080:8080 \
imp:latest --model /models/your-model.gguf
Contributor workflow, build targets and the test lanes: CONTRIBUTING.md.
Everything runs in Docker; the host needs no CUDA toolkit. The worked example is Qwen3.8-27B, a 27B multimodal model quantized to NVFP4 so it fits one 5090 with room for a real context. The 60 seconds start once the weights are on disk; the two steps before that are a build and a download.
1. Build the image (~3.5 min). scripts/stage-model.sh runs the downloader inside the imp:test image this produces and exits 1 without it (#1682). docker compose build imp-server makes imp:latest, a different tag.
make build
2. Get the weights (19.2 GiB, download only, nothing to convert). The checkpoint is kekzle/Qwen3.8-27B-NVFP4-vllm, an imp-quantize --format vllm export; the same directory also loads in vLLM (verified on 0.27.1 and 0.28.0).
scripts/stage-model.sh kekzle/Qwen3.8-27B-NVFP4-vllm ~/models/Qwen3.8-27B-NVFP4-vllm
3. Serve. The cache volume is optional and pays for itself on the second start: it holds the transformed weights (Qwen3-14B-NVFP4 init 7.9 s → 2.1 s).
docker run --gpus all -v ~/models:/models -v imp-cache:/home/imp/.cache/imp \
-p 127.0.0.1:8080:8080 ghcr.io/kekzl/imp:latest --model /models/Qwen3.8-27B-NVFP4-vllm
4. Ask. Open <http://localhost:8080>. The built-in UI streams the answer, draws one bar per token as it arrives, and shows the server's own counts for the run: prompt tokens, prefix-cache hits, reasoning tokens, context used. If the models directory holds more than one model, the header becomes a picker and choosing an unloaded entry swaps it in.
<p align="center"> <img src="docs/webui.png" width="900" alt="the built-in web UI after one answer: transcript on the left, one bar per token under it, run and usage readouts on the right"> </p>
<sub>Captured against the UI's dev mock (tools/imp-server/webui/dev/): the numbers in the picture are the mock's, not a measurement.</sub>
Or from the shell:
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"Qwen3.8-27B-NVFP4-vllm","messages":[{"role":"user","content":"Why is the sky blue?"}],"max_tokens":64}'
Expected: a JSON completion object whose choices[0].message.content explains Rayleigh scattering. The model id is the file or directory basename, GET /v1/models lists it, and the field is required. A GGUF file works the same way: swap the path.
The weights take 18.3 GiB on the card and answer at ~102 tok/s, leaving ~7 GiB for the KV cache on a 32 GB 5090.
[PROV: commit=f243179c date=2026-08-31 hw=RTX5090 model=Qwen3.8-27B-NVFP4-vllm quant=NVFP4 cuda=13.3 path=nvfp4-safetensors n=2 image=ghcr.io/kekzl/imp:latest cmd=imp-cli --model … --prompt … --max-tokens 128 --temperature 0 (102.8/101.9 tok/s)]
高性能LLM推理引擎,支持NVIDIA Blackwell GeForce
AI Skill Hub 为第三方内容聚合平台,本页面信息基于公开数据整理,不对工具功能和质量作任何法律背书。
建议在沙箱或测试环境中充分验证后,再部署至生产环境,并做好必要的安全评估。
✅ MIT 协议 — 最宽松的开源协议之一,可自由商用、修改、分发,仅需保留版权声明。
AI Skill Hub 点评:imp AI技能包 的核心功能完整,质量优秀。对于AI 技术爱好者来说,这是一个值得纳入个人工具库的选择。建议先在非生产环境试用,再逐步推广。
| 原始名称 | imp |
| 原始描述 | 开源AI工具:High-performance LLM inference engine in C++/CUDA for NVIDIA Blackwell GeForce /。⭐17 · Cuda |
| Topics | cudacppcuda-graphsgated-deltanet |
| GitHub | https://github.com/kekzl/imp |
| License | MIT |
| 语言 | Cuda |
收录时间:2026-05-16 · 更新时间:2026-05-19 · License:MIT · AI Skill Hub 不对第三方内容的准确性作法律背书。