AI Skill Hub 强烈推荐:AutoRAG Agent工作流 是一款优质的Agent工作流。已获得 4.8k 颗 GitHub Star,AI 综合评分 8.2 分,在同类工具中表现稳健。如果你正在寻找可靠的Agent工作流解决方案,这是一个值得深入了解的选择。
AutoRAG是专业的RAG评估开源框架,提供自动化的检索增强生成工作流与基准测试工具。支持文档解析、向量嵌入、工作流分析等功能,适合AI研究者、RAG应用开发者和模型评估团队使用。
AutoRAG Agent工作流 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
AutoRAG是专业的RAG评估开源框架,提供自动化的检索增强生成工作流与基准测试工具。支持文档解析、向量嵌入、工作流分析等功能,适合AI研究者、RAG应用开发者和模型评估团队使用。
AutoRAG Agent工作流 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
# 方式一:pip 安装(推荐)
pip install autorag
# 方式二:虚拟环境安装(推荐生产环境)
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install autorag
# 方式三:从源码安装(获取最新功能)
git clone https://github.com/Marker-Inc-Korea/AutoRAG
cd AutoRAG
pip install -e .
# 验证安装
python -c "import autorag; print('安装成功')"
# 命令行使用
autorag --help
# 基本用法
autorag input_file -o output_file
# Python 代码中调用
import autorag
# 示例
result = autorag.process("input")
print(result)
# autorag 配置文件示例(config.yml) app: name: "autorag" debug: false log_level: "INFO" # 运行时指定配置文件 autorag --config config.yml # 或通过环境变量配置 export AUTORAG_API_KEY="your-key" export AUTORAG_OUTPUT_DIR="./output"
A self-evolving librarian agent for document collections.
[!IMPORTANT] Looking for the original AutoRAG (RAG AutoML / pipeline optimization tool)? This repository now hosts AutoRAG 2.0, a complete reimagining of AutoRAG as a self-evolving librarian agent. The original Python-based AutoRAG — the RAG AutoML tool for automatically finding an optimal RAG pipeline for your data — now lives in thelegacy/directory of this repository. The legacy AutoRAG is NOT abandoned. It continues to be maintained (bug fixes, dependency updates, and PyPI releases viapip install AutoRAG) in maintenance mode. Existing users can keep using it exactly as before — see the legacy README for its documentation, and file issues in this repository as usual. New feature development is focused on AutoRAG 2.0.
AutoRAG searches your PDFs, wikis, notes, research papers, and knowledge bases — then curates the results into clean, numbered knowledge units. No raw grep dumps. Just answers.
AutoRAG is a customized Pi agent — the Pi agent loop configured into a librarian. The AutoRAG librarian retrieves candidates, reads source files directly, judges the evidence, and curates the structured answer. Its model and provider come from the user's authenticated runtime; AutoRAG does not ship a private provider default.
AutoRAG itself is the specialized search agent, not a coordinator for other model roles. You configure one model, and that model owns the complete retrieval, reading, judgment, and curation loop.
Published as @autorag/librarian (dist bundled with Bun, runtime Node ≥ 24 or Bun):
```bash bun add @autorag/librarian # library bun install -g @autorag/librarian # autorag CLI
import { AutoRAGAgent } from "@autorag/librarian";
const agent = new AutoRAGAgent({
searchPaths: ["/path/to/documents"],
});
const response = await agent.searchDocuments("summarize the compliance requirements");
console.log(response.answer);
for (const result of response.results) {
console.log(`[${result.number}] ${result.title} — ${result.summary}`);
}
// Mark which results were useful — AutoRAG remembers for next time
agent.recordFeedbackByNumbers(response.sessionId, [1, 3], [2]);
searchDocuments() runs the Pi agent loop — it searches, reads, consults memory, curates, and finalizes through the emit_autorag_results structured tool — then returns a typed SearchDocumentsResponse. The caller consumes the structured payload directly; no assistant text parsing.
The default home state is kept outside the workspace:
~/.autorag/
├── config.json
├── memory.json
└── logs/
└── runs.jsonl
config.json selects sources, the workspace, memory path, retrieval settings, and the agent model. Provider and model IDs must refer to a model available in the user's authenticated runtime:
{
"searchPaths": ["/path/to/documents"],
"workspacePath": "/path/to/workspace",
"memoryPath": "/Users/you/.autorag/memory.json",
"model": { "provider": "provider-name", "id": "reasoning-model" }
}
autorag init leaves model unset when no model flags are supplied. At search time AutoRAG resolves an authenticated local provider when possible; otherwise configure the model explicitly.
For fast interactive search, prefer a model with reliable tool calling, high output TPS, and low first-token latency. A query can require several short model turns while AutoRAG alternates between retrieval tools and direct source reading, so model throughput has a visible effect on end-to-end response time. It does not accelerate BM25, MinSync, Jikji, filesystem access, or indexing itself. Larger reasoning models remain useful for difficult synthesis, conflicting evidence, and specialized domain judgment, but they are not a requirement for ordinary retrieval.
Config path precedence is --config > AUTORAG_CONFIG > ~/.autorag/config.json. When the home config is absent and <cwd>/autorag.config.json exists, AutoRAG copies the legacy file to ~/.autorag/config.json without deleting or modifying the legacy file. The legacy cwd file is a migration source, not the default location.
memory.json stores retrieval memory and logs/runs.jsonl records run events. Model authentication remains with the user's configured provider or authenticated local runtime. Corpus indexes remain workspace-local: refresh keeps parsed mirrors and BM25/MinSync indexes under <workspace>/.autorag.
autorag refresh and autorag index reset|rebuild accept --method <csv> (e.g. --method bm25,minsync,parsed) to scope which indexing methods run or which index directories are removed. When omitted, all methods run. autorag init accepts --embedder-* flags to configure the MinSync embedder endpoint in the config file.
autorag health checks model/provider auth before a search — it resolves the model, verifies credential presence, and optionally probes one completion call. Use it to diagnose model, provider, auth, or timeout failures. autorag status remains the model-free index-health command (corpus freshness and BM25/MinSync readiness). When autorag search fails for a model/provider reason, the error output includes a hint pointing to autorag health.
autorag ui opens a loopback-only page (127.0.0.1) to connect local folders and datasource skills without editing JSON. It writes the same trusted datasources / datasourceAccess fields as a hand-edited config, stores env-var names rather than secrets, and refuses non-loopback binds. Use --no-open to print the URL without launching a browser.
For a deliberately deployed UI, opt in explicitly in config.json. Keep the session token in the environment, set the public URL used by the browser, and allow only the frontend origins that should make credentialed API requests:
{
"ui": {
"host": "0.0.0.0",
"port": 8787,
"allowRemote": true,
"publicOrigin": "https://autorag.example.com",
"corsOrigins": ["https://autorag.example.com"],
"tokenEnv": "AUTORAG_UI_TOKEN"
}
}
Start it with AUTORAG_UI_TOKEN set to a random value of at least 16 characters. Local use remains the safe default: omit ui (or leave allowRemote false) and AutoRAG binds to loopback, including a working localhost URL on systems that resolve it to IPv6.
node dist/cli/index.js init --search-paths /path/to/documents
Different documents need different search strategies:
| Your documents | Best method | Why |
|---|---|---|
| Plain text, config files | grep (pattern matching) | Fast, precise, literal |
| Research papers, dense prose | Vector search (semantic) | Understands meaning, not just keywords |
| Legal documents, specifications | BM25 (keyword ranking) | Handles domain terminology well |
| Mixed collections | Hybrid (vector + BM25) | Combines precision and recall |
AutoRAG supports pluggable retrieval methods. Local lexical BM25, semantic vector, and hybrid retrieval all go through MinSync over one shared CDC chunk lifecycle, wired through the RetrievalMethodRegistry. The librarian invokes retrieval tools, reads the underlying documents directly through bash, and curates one unified result set after ResultMerger score normalization and deduplication. External datasources keep their own archive/index lifecycle.
BM25, vector, and hybrid are enabled by default whenever MinSync is enabled. Disable local indexing with "minSync": false, or disable only lexical search with "bm25": false. MinSync auto-installs a verified release into <workspace>/.autorag/bin on first use (autoInstall: true); set "autoInstall": false only when managing the binary yourself. Configure minSync.embedder via autorag init --embedder-* flags for remote embedding endpoints, and set minSync.maxChunkSize (or --minsync-max-chunk-size) when a local embedder has a smaller context window. AutoRAG never forces TEI or any external embedding service.
See docs/minsync-setup.md for automatic installation, managed binary paths, and the local EmbeddingGemma QA flow.
aiskill88点评:AutoRAG是RAG领域专业工具,4.8k星证明受认可度高,完善的工作流和评估框架对RAG应用开发者很有价值,持续维护。
AI Skill Hub 为第三方内容聚合平台,本页面信息基于公开数据整理,不对工具功能和质量作任何法律背书。
建议在沙箱或测试环境中充分验证后,再部署至生产环境,并做好必要的安全评估。
✅ Apache 2.0 — 宽松开源协议,可商用,需保留版权声明和 NOTICE 文件,含专利授权条款。
总体来看,AutoRAG Agent工作流 是一款质量优秀的Agent工作流,在同类工具中具备一定竞争力。AI Skill Hub 将持续追踪其更新动态,建议收藏备用,结合自身场景选择合适时机引入使用。
| 原始名称 | AutoRAG |
| 原始描述 | 开源AI工作流:AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evalu。⭐4.8k · Python |
| Topics | RAG评估检索增强生成工作流自动化基准测试文档解析 |
| GitHub | https://github.com/Marker-Inc-Korea/AutoRAG |
| License | Apache-2.0 |
| 语言 | Python |
收录时间:2026-05-16 · 更新时间:2026-05-19 · License:Apache-2.0 · AI Skill Hub 不对第三方内容的准确性作法律背书。
选择 Agent 类型,复制安装指令后粘贴到对应客户端