经 AI Skill Hub 精选评估,Floe-Guard AI 账单卫士 获评「强烈推荐」。这款Agent工作流在功能完整性、社区活跃度和易用性方面表现出色,AI 评分 8.2 分,适合有一定技术背景的用户使用。
一个为AI Agent设计的开源统一计费护栏工具。它能为AI工作流提供硬性预算限制,在费用失控前强制停止运行,防止因Agent陷入死循环或过度调用而导致账单爆炸,非常适合开发者和企业级AI应用部署。
Floe-Guard AI 账单卫士 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
一个为AI Agent设计的开源统一计费护栏工具。它能为AI工作流提供硬性预算限制,在费用失控前强制停止运行,防止因Agent陷入死循环或过度调用而导致账单爆炸,非常适合开发者和企业级AI应用部署。
Floe-Guard AI 账单卫士 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
# 方式一:pip 安装(推荐)
pip install floe-guard
# 方式二:虚拟环境安装(推荐生产环境)
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install floe-guard
# 方式三:从源码安装(获取最新功能)
git clone https://github.com/Floe-Labs/floe-guard
cd floe-guard
pip install -e .
# 验证安装
python -c "import floe_guard; print('安装成功')"
# 命令行使用
floe-guard --help
# 基本用法
floe-guard input_file -o output_file
# Python 代码中调用
import floe_guard
# 示例
result = floe_guard.process("input")
print(result)
# floe-guard 配置文件示例(config.yml) app: name: "floe-guard" debug: false log_level: "INFO" # 运行时指定配置文件 floe-guard --config config.yml # 或通过环境变量配置 export FLOE_GUARD_API_KEY="your-key" export FLOE_GUARD_OUTPUT_DIR="./output"
Know what every AI call really costs. Your agent spends across a dozen vendors on every call — carrier, speech-to-text, model, voice, tools. floe-guard costs each call the moment it ends and keeps a live ledger of your agent's real spend. In-process, no account, no signup, no telemetry by default.
Connect free (one key) for your Coverage Score and 7-day history — the audit-grade picture of where every dollar went, up to $2K/month tracked, no card. And it hard-stops a runaway loop before it crosses your ceiling: $0.10 instead of $4,000.
Python (pip install floe-guard): plain check() / record() or adapters for OpenAI · Anthropic · Gemini · CrewAI · LiteLLM · LangChain · LangGraph; voice adapters for Pipecat · LiveKit · Vapi · Retell; pre-call admission gates.
TypeScript (npm i floe-guard): Vercel AI SDK middleware, native LiveKit · Vapi · Retell voice adapters. See the adapter matrix for what ships in Python vs TypeScript.
Reading this on PyPI? Thedocs/…andexamples/…links resolve on the GitHub README, not on the PyPI page.
The hard-stop is contract-based: adapters gate LLM calls automatically; for paid tools, reserve_tool() / settle_tool() block before the call runs (record_tool() alone meters after the fact — it can't stop a call already made).
floe-guard is a local, estimate-based guardrail. It prices tokens from a vendored cost map inside your process:
- The cost map can drift as vendors change prices — refresh it like any snapshot. - It only sees the vendors you instrument. - A determined agent or a bug could route around an in-process check. - Under heavy or cold-start concurrency it bounds steady-state spend, not the first parallel wave. Reservations default to the last call's cost (0 until the first record()) — size them to the real request with estimate_call() (the LiteLLM adapter does this for you), or use hosted Floe for a hard cap under arbitrary concurrency. - Mid-stream enforcement (guard_stream) prices chunks by a ~4 chars/token heuristic unless you supply a tokenizer, so the cut-off point is approximate; the final accrual reconciles to provider-reported usage.
It's genuinely useful on its own, and it's honest about its limits. No inflated metrics, no "zero defaults" claims — it's a free local stop, not a vault.
All runnable examples live in examples/. Use python examples/<file> from the repo root.
| Example | Description | Extra | API key / network |
|---|---|---|---|
[runaway_loop.py](examples/runaway_loop.py) | The canonical hard-stop demo — a stub loop halted before it crosses $0.10 | none | none |
[streaming_guard.py](examples/streaming_guard.py) | Pre-flight block on an oversized first call + mid-stream cut-off via guard_stream() | none | none |
[budget_aware.py](examples/budget_aware.py) | Context-aware tapering: agent downshifts to a cheap model when advisory().near_limit trips | none | none |
[budget_retry.py](examples/budget_retry.py) | Budget-aware retry / graceful degradation with with_budget_retry() | none | none |
[context_size.py](examples/context_size.py) | Context size adapts to budget: history trimmed and max_tokens capped near the ceiling | none | none |
[plan_complexity.py](examples/plan_complexity.py) | Plan complexity adapts: optional sub-tasks dropped and reasoning depth reduced near the cap | none | none |
[retrieval_depth.py](examples/retrieval_depth.py) | RAG top_k shrinks in two steps (20→12→5) as budget drains | none | none |
[step_budget.py](examples/step_budget.py) | Per-step token caps for a sequential loop: one runaway step blocked without stopping the run | none | none |
[tool_budget.py](examples/tool_budget.py) | Tool spend (Apollo lookups, Exa searches) as a first-class citizen of the same USD ceiling | none | none |
[openai_adapter.py](examples/openai_adapter.py) | guarded_completion against a duck-typed stub — exercises the real pre-flight hard-stop | none | none |
[anthropic_adapter.py](examples/anthropic_adapter.py) | Anthropic adapter with native prompt-cache pricing (cache write vs. cache read vs. uncached) | none | none |
[langgraph_budget_aware.py](examples/langgraph_budget_aware.py) | LangGraph guarded_node fan-out with an advisory-driven router that tapers before the cap | pip install floe-guard[langgraph] | none |
[langchain_groq_example.py](examples/langchain_groq_example.py) | LangChain callback handler on ChatGroq (Llama-3): call 1 succeeds, call 2 is hard-stopped | pip install floe-guard[langchain] langchain-groq | GROQ_API_KEY + network |
[voice_turn_budget.py](examples/voice_turn_budget.py) | Pipecat pipeline with FloeBudgetGuardProcessor: multi-turn voice conversation halted mid-run | pip install floe-guard[pipecat] | none |
[voice_call_cost_pipecat.py](examples/voice_call_cost_pipecat.py) | Full per-leg call cost (STT + LLM + TTS + telephony) via Pipecat, priced from the bundled map | pip install floe-guard[pipecat] | none |
[voice_call_cost_livekit.py](examples/voice_call_cost_livekit.py) | Full per-leg call cost (STT + LLM + TTS + telephony) via LiveKit, priced from the bundled map | pip install floe-guard[livekit] | none |
[voice_call_cost_vapi.py](examples/voice_call_cost_vapi.py) | Full per-leg call cost (STT + LLM + TTS + telephony) via the Vapi custom-LLM proxy, priced from the bundled map | pip install floe-guard | none |
[voice_call_cost_retell.py](examples/voice_call_cost_retell.py) | Full per-leg call cost (STT + LLM + TTS + telephony) via the Retell custom-LLM WebSocket, priced from the bundled map | pip install floe-guard | none |
Straight from the install — no repository checkout needed:
pip install floe-guard
floe-guard demo
This rigs a loop against a stub LLM — no real API key, no account, no network. It prices each fake gpt-4o call offline and the guard halts the loop after a few iterations, before it can cross the $0.10 ceiling. This is the reproducible "stop the loop" demo. Cloned the repo? The same demo is examples/runaway_loop.py (a thin wrapper around floe_guard.demo.run_demo).
pip install floe-guard # the Vapi adapter is framework-free — no extra
The custom-LLM proxy sees only the model leg, so VapiBudgetGuard guards the /chat/completions turn and admits the call via the assistant-request webhook. guard_completion (JSON) and guard_stream (SSE) reserve the estimated cost before the upstream call, settle on Vapi's real OpenAI usage afterwards, and release the hold on error/abort — so an over-budget turn gets a 402 instead of reaching your LLM. Set stream_options={"include_usage": True} on the upstream streaming request or guard_stream fails loudly (the SSE omits usage without it).
```python from floe_guard import BudgetGuard from floe_guard.errors import BudgetExceeded from floe_guard.integrations.vapi import VapiBudgetGuard
guard = BudgetGuard(limit_usd=1.00) budget = VapiBudgetGuard( guard, stt_model="deepgram-nova-3", # $/sec from the voice map tts_model="elevenlabs-flash-v2.5", # $/1k-chars from the voice map telephony="twilio-us-inbound-local", # $/min from the voice map )
completion = await budget.guard_completion( lambda: openai_client.chat.completions.create(model=model, messages=messages), model=model, )
gates.vapi(guard, assistant_id="asst_…")
from floe_guard.integrations.vapi import VapiBudgetGuard
budget = VapiBudgetGuard(guard, stt_model="deepgram-nova-3", tts_model="elevenlabs-flash-v2.5", telephony="twilio-us-inbound-local") completion = await budget.guard_completion( # reserve → run → settle on real usage lambda: openai_client.chat.completions.create(model=model, messages=body["messages"]), model=model, )
You've watched it stop a stub loop — the real payoff is protecting a real one, where the local ceiling earns its keep. Pick your stack; each is a drop-in adapter, a few lines, no rearchitecting:
guarded_completion wraps the client call.guarded_node per fan-out branch.budget_guarded_llm / guarded_completion.See the adapter matrix for what ships in Python vs TypeScript.
aiskill88点评:解决了Agent落地最痛的成本失控问题,逻辑简单且实用,是AI工程化不可或缺的安全组件。
AI Skill Hub 为第三方内容聚合平台,本页面信息基于公开数据整理,不对工具功能和质量作任何法律背书。
建议在沙箱或测试环境中充分验证后,再部署至生产环境,并做好必要的安全评估。
✅ MIT 协议 — 最宽松的开源协议之一,可自由商用、修改、分发,仅需保留版权声明。
AI Skill Hub 点评:Floe-Guard AI 账单卫士 的核心功能完整,质量优秀。对于自动化工程师和运维人员来说,这是一个值得纳入个人工具库的选择。建议先在非生产环境试用,再逐步推广。
| 原始名称 | floe-guard |
| 原始描述 | 开源AI工作流:Open-source unified billing guardrail for AI agents — hard-stop before a runaway。⭐42 · Python |
| Topics | AI安全预算控制AI Agent |
| GitHub | https://github.com/Floe-Labs/floe-guard |
| License | MIT |
| 语言 | Python |
收录时间:2026-07-10 · 更新时间:2026-07-11 · License:MIT · AI Skill Hub 不对第三方内容的准确性作法律背书。
选择 Agent 类型,复制安装指令后粘贴到对应客户端