AI Skill Hub 强烈推荐:VoiceBlender 语音控制平台 是一款优质的Agent工作流。AI 综合评分 8.2 分,在同类工具中表现稳健。如果你正在寻找可靠的Agent工作流解决方案,这是一个值得深入了解的选择。
一个可编程的开源语音平台,支持SIP和WebRTC通话控制及多方音频混音。它集成了ASR和音频处理能力,旨在为AI Agent提供实时语音交互基础设施,适合需要构建复杂语音工作流的开发者。
VoiceBlender 语音控制平台 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
一个可编程的开源语音平台,支持SIP和WebRTC通话控制及多方音频混音。它集成了ASR和音频处理能力,旨在为AI Agent提供实时语音交互基础设施,适合需要构建复杂语音工作流的开发者。
VoiceBlender 语音控制平台 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
# 方式一:go install(推荐) go install github.com/VoiceBlender/voiceblender@latest # 方式二:从源码编译 git clone https://github.com/VoiceBlender/voiceblender cd voiceblender go build -o voiceblender . # 方式三:下载预编译二进制 # 访问 Releases 页面下载对应平台二进制文件 # https://github.com/VoiceBlender/voiceblender/releases
# 查看帮助 voiceblender --help # 基本运行 voiceblender [options] <input> # 详细使用说明请查阅文档 # https://github.com/VoiceBlender/voiceblender
# voiceblender 配置说明 # 查看配置选项 voiceblender --config-example > config.yml # 常见配置项 # output_dir: ./output # log_level: info # workers: 4 # 环境变量(覆盖配置文件) export VOICEBLENDER_CONFIG="/path/to/config.yml"
A Go service that bridges SIP and WebRTC voice calls with multi-party audio mixing, a REST API, and real-time webhooks.
json_base64 framing, configurable sample rate (8/16/24/48 kHz), bidirectional text, and caller-supplied X-/P- headers — designed to also back a future generic Agent APImengelbart/moqtransport (IETF draft-11); browser interop with draft-16 clients (moqtail, moq.dev) is not expected to work out of the box. Disabled by default; enable with MOQ_ENABLED=true + MOQ_TLS_CERT_FILE / MOQ_TLS_KEY_FILErole and declare a matrix of who-hears-whom by role. Applied atomically at leg-join time so a supervisor cannot momentarily bleed into the customer's audio. See API.md.m=audio sections in one dialog (RFC 3264), each with its own RTP port, direction, language and mixer room; built for live translation, where the original audio and a translated feed are mixed separately. Follows the SIPREC (RFC 7866) wire profile for interoperability. See API.md.SDP + rs-metadata INVITEs are answered receive-only on every m=audio section, and each section is bound to the participant the metadata names. The received audio is an ordinary set of leg streams, so it can be recorded to file and attached to rooms for live STT/agents. Disabled by default; enable with SIPREC_ENABLED=true (needs SIP_TCP_ENABLED=true). The metadata is checked against the SDP it arrived with, so a client that binds a participant to the wrong a=label -- a document that is otherwise valid and would silently attribute audio to the wrong party -- is reported in warnings on GET /v1/legs/{id}/siprec instead of being recorded as if it were correct. VoiceBlender can also act as the recording client, forking a room's participants to an external recording server one stream each (POST /v1/rooms/{id}/siprec, SIPREC_SRC_ENABLED=true). See API.md.stt.turn), including eager end-of-turn signals for speculative generation; Speechmatics reports end-of-turn from server-side silence detection and supports mid-stream finalizeevent_id (also sent as X-Event-Id) for receiver-side deduplication; typed event data with CDR-style leg.disconnected (disposition, timing, quality)GET /v1/vsi streams all events and accepts commands (mute, hold, DTMF, room management) over a single persistent WebSocket; filter by app_id regex for multi-tenant isolationGET /metrics (active legs/rooms, call durations, disconnect reasons, event-egress drop/delivery counters, Go runtime). See API.md for the full metric reference. Profiling via go tool pprof is available at /debug/pprof/ when built with -tags pprof.meta.vc and the offer still carries DTLS-SRTP (a=fingerprint). A fronting SIP proxy that terminates encryption into plain RTP keeps the Meta From but takes the classic SIP path; one that re-encrypts with SDES-SRTP is rejected with 488 Not Acceptable Here rather than answered with media nothing can decrypt. On the WhatsApp path the leg comes up in ringing, fires leg.ringing (leg_type: "whatsapp_in"), and waits for POST /v1/legs/{id}/answer. The 200 OK then carries the pre-gathered ICE/DTLS-SRTP answer.POST /v1/legs {"type":"whatsapp", ...} returns 201 immediately with the leg in ringing. ICE gathering, the digest 401/407 round-trip, and the SDP-answer apply happen asynchronously; outcome is signalled via leg.connected or leg.disconnected.dtmf.received plus the standard cross-leg broadcast.leg.ringing / leg.connected / leg.disconnected / dtmf.received / speaking.started / speaking.stopped all carry leg_type set to whatsapp_in or whatsapp_out so multi-tenant filtering works as it does for SIP and WebRTC legs.go test -tags integration -v -timeout 60s ./tests/integration/
| Library | Description | Notes |
|---|---|---|
| [sipgo](https://github.com/emiago/sipgo) | SIP stack | Excellent SIP stack in go |
| [pion/webrtc](https://github.com/pion/webrtc) | WebRTC | Nothing is better than Pion |
| [go-chi](https://github.com/go-chi/chi) | HTTP router | |
| [zaf/g711](https://github.com/zaf/g711) | G.711 codec | |
| [gobwas/ws](https://github.com/gobwas/ws) | WebSocket | |
| [go-audio/wav](https://github.com/go-audio/wav) | WAV encoding | |
| [gopus](https://github.com/thesyncim/gopus) | Opus codec | Thanks Marcelo! (Claude and Codex too!) |
| [go-mp3](https://github.com/hajimehoshi/go-mp3) | MP3 decoder | Pure Go |
| [go-audio/audio](https://github.com/go-audio/audio) | Audio buffer types | |
| [google/uuid](https://github.com/google/uuid) | UUID generation | |
| [prometheus/client_golang](https://github.com/prometheus/client_golang) | Prometheus metrics | |
| [aws-sdk-go-v2](https://github.com/aws/aws-sdk-go-v2) | AWS SDK (S3, Polly) | |
| [cloud.google.com/go/texttospeech](https://cloud.google.com/go/docs/reference/cloud.google.com/go/texttospeech/latest) | Google Cloud TTS | |
| [protobuf](https://github.com/protocolbuffers/protobuf-go) | Protocol Buffers | Pipecat agent |
| [x/sync](https://pkg.go.dev/golang.org/x/sync) | Concurrency utilities |
go build -o voiceblender ./cmd/voiceblender ./voiceblender
```bash
| Example | Description |
|---|---|
[examples/call_handler.py](examples/call_handler.py) | Python webhook listener for inbound SIP calls with room conferencing |
[examples/webrtc-client/](examples/webrtc-client/) | Browser-based WebRTC voice client with room management and DTMF |
[examples/gen_test_wav.py](examples/gen_test_wav.py) | Generate test WAV files for playback testing |
All configuration is via environment variables:
| Variable | Default | Description |
|---|---|---|
INSTANCE_ID | *(auto-generated UUID)* | Instance identifier, included in API responses and webhooks |
HTTP_ADDR | :8080 | REST API listen address |
ALLOWED_IPS | _(empty = allow all)_ | Comma-separated allowlist of IPs and CIDR ranges (IPv4 and IPv6, in any mix) gating **every** HTTP endpoint, including the /v1/vsi event WebSocket, /v1/legs/websocket, the /v1/legs/moq WebTransport endpoint, /metrics, and pprof. Empty or unset disables the check. Whitespace around entries is trimmed; bare addresses are treated as host routes (/32 for v4, /128 for v6); malformed entries fail server startup. Only X-Forwarded-For is consulted as a proxy header (see TRUST_PROXY_HEADERS); X-Real-IP and RFC 7239 Forwarded are ignored. Examples: 127.0.0.1, 10.0.0.0/8,192.168.0.0/16, 2001:db8::/32,::1. |
TRUST_PROXY_HEADERS | false | When true, the client IP used for the ALLOWED_IPS check is taken from the leftmost entry in X-Forwarded-For (falling back to the socket peer when the header is absent). When false (default), X-Forwarded-For is ignored and only the socket peer (r.RemoteAddr) is consulted. Enable only when VoiceBlender sits behind a trusted reverse proxy that unconditionally overwrites X-Forwarded-For — otherwise the header is client-spoofable. |
SIP_BIND_IP | 127.0.0.1 | IPv4 address advertised in SDP/Contact/Via headers (and used as the listen address when SIP_LISTEN_IP is empty). Set to 0.0.0.0 for v4 wildcard, :: for dual-stack on Linux when bindv6only=0. |
SIP_LISTEN_IP | *(same as SIP_BIND_IP)* | UDP socket bind IP. Accepts 127.0.0.1, 0.0.0.0, ::, or any literal v4/v6 address. |
SIP_BIND_IPV6 | *(empty = v4-only)* | IPv6 address advertised in SDP/Contact/Via for IPv6 calls. Set this for IPv6-only or dual-stack deployments. |
SIP_LISTEN_IPV6 | *(same as SIP_BIND_IPV6)* | Optional separate IPv6 socket bind address (e.g. when running with both 0.0.0.0 and a specific v6 literal). |
SIP_PORT | 5060 | SIP listen port (UDP) |
SIP_TLS_PORT | *(disabled)* | SIP-over-TLS listen port (typically 5061). When set, SIP_TLS_CERT and SIP_TLS_KEY must also be provided. Required for WhatsApp Business Calling integration. |
SIP_TLS_CERT | Path to PEM-encoded TLS certificate (e.g. fullchain.pem). Meta rejects self-signed certs — use a CA-signed cert matching a public FQDN. | |
SIP_TLS_KEY | Path to PEM-encoded TLS private key (e.g. privkey.pem). | |
SIP_TLS_CA_FILE | *(system trust store only)* | Path to a PEM bundle of extra CA certificates trusted when **dialing** a remote peer over TLS (registrar, outbound proxy, carrier SBC). Added to the system roots, not a replacement. To trust a peer that presents a self-signed certificate, point this at that certificate. The name in the certificate is still checked, so a peer whose certificate has no SAN needs the per-trunk tls_insecure_skip_verify instead. Not a client certificate — VoiceBlender never presents one. |
SIP_TLS_INSECURE_SKIP_VERIFY | false | When true, any certificate a remote peer presents on an outbound TLS dial is accepted without verification — every peer, server-wide. Prefer SIP_TLS_CA_FILE, or the per-trunk sip_register.tls_insecure_skip_verify, which scopes the exemption to one peer. Never affects the inbound TLS listener. |
SIP_DEBUG | false | When true, log the full RFC 3261 wire form of every inbound and outbound SIP request and response. Very verbose — use only for troubleshooting. |
SIP_DOMAIN | *(falls back to advertised IP)* | FQDN advertised in From, Contact and Via on outbound SIP signalling (classic trunks and WhatsApp). Two exceptions apply to the From host only: a call matched to a registered SIP trunk uses that trunk's AOR realm, and a from given as a full SIP URI uses the host in that URI. Should match the SAN on SIP_TLS_CERT and any allowlist your carrier or Meta keeps. |
SIP_HOST | voiceblender | SIP User-Agent name |
ICE_SERVERS | stun:stun.l.google.com:19302 | STUN/TURN URLs (comma-separated) |
WEBRTC_EXTERNAL_IPS | *(empty)* | Comma-separated public IPs advertised as host ICE candidates (pion SetNAT1To1IPs). Set this when VoiceBlender runs behind NAT/Docker and the gathered host interface IPs aren't routable from the remote peer — otherwise WebRTC peers behind firewalls won't be able to reach VB. Supports IPv4 and IPv6 literals. The literal value auto performs STUN-based public-IP discovery at startup against the first reachable ICE_SERVERS entry; discovery failure is non-fatal and logs a warning. |
RECORDING_DIR | /tmp/recordings | Local recording output directory |
LOG_LEVEL | info | Log level (debug, info, warn, error). Verbatim transcript text, DTMF digits and full event payloads are logged only at debug. |
WEBHOOK_URL | Default webhook URL for inbound calls | |
WEBHOOK_SECRET | HMAC-SHA256 signing secret for the global webhook. Applied to events delivered to WEBHOOK_URL; per-leg/per-room webhooks set via the API can supply their own secret. | |
CUSTOM_DATA_MAX_BYTES | 1024 | Maximum size in bytes of a leg's custom_data JSON. The value is repeated on every event for that leg, so this caps webhook payload growth. 0 = unlimited. |
ELEVENLABS_API_KEY | API key for ElevenLabs TTS, STT, and Agent | |
VAPI_API_KEY | API key for VAPI Agent provider | |
DEEPGRAM_API_KEY | API key for Deepgram STT and TTS | |
AZURE_SPEECH_KEY | Subscription key for Azure Cognitive Speech Services (TTS and STT) | |
AZURE_SPEECH_REGION | eastus | Azure region for Speech Services (e.g. eastus, westeurope) |
SPEECHMATICS_API_KEY | API key for Speechmatics STT | |
SPEECHMATICS_URL | wss://eu2.rt.speechmatics.com/v2 | Speechmatics realtime endpoint. Change it for another region (eu, us, global) or a self-hosted realtime container. |
S3_BUCKET | S3 bucket for recording uploads. See [S3 bucket preflight](#s3-bucket-preflight). | |
S3_REGION | us-east-1 | AWS region |
S3_ENDPOINT | Custom S3 endpoint (MinIO, etc.). Must include an http:// or https:// scheme. | |
S3_PREFIX | Key prefix for S3 objects | |
S3_ALLOW_INSECURE_ENDPOINT | false | Allow a plaintext http:// S3_ENDPOINT on a **non-local** host. Loopback and private addresses never need this. See [S3 bucket preflight](#s3-bucket-preflight). |
S3_PREFLIGHT_TIMEOUT | 10s | Budget for the startup bucket probe. 0 disables it. |
S3_REQUEST_PREFLIGHT_TIMEOUT | 2s | Budget for the bucket probe on a per-request S3 backend. 0 disables it. |
GCS_BUCKET | Google Cloud Storage bucket for recording uploads via the native GCS API. Prefer this over S3_ENDPOINT=https://storage.googleapis.com on GKE — Workload Identity / ADC works directly, with no HMAC interop keys. | |
GCS_OBJECT_NAME_PREFIX | Object name prefix for GCS uploads (e.g. recordings or a bare workspace id like dev). A trailing slash is added automatically when missing. | |
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN, AWS_PROFILE | AWS credentials for S3 uploads and AWS Polly TTS. Resolved by the AWS SDK's default credential chain (env vars → ~/.aws/credentials → EC2/ECS/EKS instance role), not by VoiceBlender directly. AWS_REGION is honored only when S3_REGION is empty. | |
GOOGLE_APPLICATION_CREDENTIALS | Path to a Google Cloud service-account JSON file used by Google Cloud TTS and by GCS recording uploads when no other credential source is available. Resolved by Google's Application Default Credentials chain (env var → ~/.config/gcloud/application_default_credentials.json → GCE/Cloud Run/GKE metadata / Workload Identity), not by VoiceBlender directly. | |
TTS_CACHE_ENABLED | false | Enable disk-backed TTS audio cache. Cached audio persists across restarts. |
TTS_CACHE_DIR | /tmp/tts_cache | Directory for cached TTS audio files (used when TTS_CACHE_ENABLED=true) |
TTS_CACHE_INCLUDE_API_KEY | false | Include API key in TTS cache key (set true if different keys map to different voice clones) |
TTS_PREFLIGHT_TTL | 30s | How long a staged (preflight) TTS utterance is held before being discarded |
TTS_PREFLIGHT_MAX_PER_LEG | 3 | Maximum staged TTS utterances per leg; staging past the cap returns 409 |
TTS_PREFLIGHT_MAX_BYTES | 4194304 | Maximum buffered audio per staged TTS utterance, in bytes |
RTP_PORT_MIN | 10000 | Minimum UDP port for RTP/RTCP media |
RTP_PORT_MAX | 20000 | Maximum UDP port for RTP/RTCP media |
SIP_JITTER_BUFFER_MS | 0 | SIP ingress jitter buffer target delay in ms (0 = disabled passthrough). Applies to every SIP leg. |
SIP_JITTER_BUFFER_MAX_MS | 300 | Max depth of the SIP ingress jitter buffer (ms); frames beyond this are dropped oldest-first. |
WS_JITTER_BUFFER_MS | 0 | WebSocket ingress playout lead in ms (0 = disabled passthrough). Applies to every websocket leg and room WS participant. Non-zero also enables clock-drift compensation, which absorbs the producer/mixer rate difference during pauses instead of punching a 20 ms hole in the mix. Size it to the transport's worst-case stall, not to the drift — 40–80 suits a healthy link, and every millisecond is added one-way latency. |
COMFORT_NOISE_ENABLED | true | Inject low-level comfort noise (~−75 dBFS) into otherwise silent mixer frames. |
SIP_SDP_STRICT_MLINE_ANSWER | false | Emit a port-0 placeholder for every offered m= section we do not accept, so answers carry the same m-line count and order as the offer (RFC 3264 §6). Gated separately from multi-stream because it changes the SDP **single-stream** calls emit whenever a peer offers a section we don't handle, such as video. |
SIP_EXTERNAL_IP | *(empty)* | Public IPv4 address for NAT/Docker deployments. When set, used in SIP Contact headers and SDP media (c=) lines instead of the auto-detected or bind IP. IPv6 has no equivalent: set SIP_BIND_IPV6 directly to the address you want advertised. |
DEFAULT_SAMPLE_RATE | 16000 | Default mixer sample rate (Hz) for new rooms when sample_rate is not specified. Allowed: 8000, 16000, 48000. |
SIP_CODECS | PCMU,PCMA | Comma-separated, preference-ordered list of codecs the SIP engine offers on outbound INVITEs **and** accepts on inbound INVITEs (a codec absent from this list cannot be negotiated in either direction). Recognized names (case-insensitive): PCMU, PCMA, G722, opus, AMR-WB, AMR-NB (the bare token AMR also resolves to AMR-NB per RFC 4867 §8.1). Unknown names and duplicates are dropped silently; if the parsed list ends up empty the default is used. Example: SIP_CODECS=opus,G722,PCMU,PCMA,AMR-WB,AMR-NB enables every supported codec, ranked Opus-first. |
SIP_REFER_AUTO_DIAL | false | When true, the server accepts an incoming SIP REFER (202) and **dials the target itself**. When false (default), the REFER is parked and surfaced as leg.transfer_requested for the app to drive via the transfer commands (accept/progress/complete/decline); an undecided REFER auto-declines (603, **default-deny** — toll-fraud risk). Outbound transfers via the REST API are unaffected. |
SIP_REFER_CONSULT_TIMEOUT_MS | 2000 | How long an inbound REFER is parked awaiting an app accept/decline decision before it auto-declines with 603 (fail-closed). Only used when SIP_REFER_AUTO_DIAL=false. |
SIP_AUTO_RINGING | false | **Behavior change vs prior releases**: previously the server always sent 180 Ringing after 100 Trying. The new default sends only 100 Trying; the API caller drives ringing explicitly via POST /v1/legs/{id}/ring, /early-media, or /answer. Set to true to restore the legacy auto-180 behavior. |
SIP_TCP_ENABLED | false | Listen for SIP over TCP on SIP_PORT alongside the UDP listener. Recommended with SIPREC: a recording session's INVITE carries the metadata document alongside the SDP and is larger than RFC 3261 §18.1.1 allows over UDP. |
SIP_USE_SOURCE_SOCKET | false | When true, route SIP responses **and** in-dialog requests (BYE, re-INVITE, UPDATE, INFO, NOTIFY, REFER) back to the request's source UDP socket instead of the peer's Contact URI / Via sent-by. Enable when peers advertise unroutable addresses (e.g. private IPs in Contact from behind NAT, or Via sent-by hosts that don't resolve). Equivalent to sipgo's DialogUA.RewriteContact plus per-response SetDestination(req.Source()). |
SIP_REGISTRATION_DEFAULT_EXPIRES_SECONDS | 3600 | Expiry used when an inbound REGISTER carries no Expires value. |
SIP_REGISTRATION_MAX_EXPIRES_SECONDS | 7200 | Upper clamp on the granted expiry. Requests above this value are honored at this maximum. |
SIP_REGISTRATION_SWEEP_INTERVAL_MS | 1000 | Sweeper period for evicting expired AOR bindings. |
SIP_REGISTRATION_ALLOW_MULTIPLE_CONTACTS | true | When true, the same AOR may be bound from multiple Contacts simultaneously (and POST /v1/legs parallel-forks to every bound contact). When false, each REGISTER replaces any prior Contacts for the AOR. |
SIP_INBOUND_AUTH_CONSULT_TIMEOUT_MS | 2000 | How long an inbound REGISTER is parked awaiting a challenge/accept/reject decision (surfaced via the sip.registration_attempt event) before the fallback (SIP_INBOUND_REGISTER_DEFAULT) applies. Every REGISTER is surfaced for a decision — symmetric with inbound INVITE, which always surfaces leg.ringing and waits for the client. |
SIP_INBOUND_REGISTER_DEFAULT | reject | Fallback for an inbound REGISTER that no client decides within the consult window: reject (reply 403, **fail-closed default**) or accept (bind and reply 200 OK — the legacy fail-open behaviour). |
SIP_INBOUND_AUTH_NONCE_TTL_SECONDS | 60 | Lifetime of an issued inbound-auth digest challenge nonce. A credentialed retry arriving after this elapses must be re-challenged. |
SIP_OUTBOUND_REGISTRATION_DEFAULT_EXPIRES_SECONDS | 3600 | Default Expires value sent on outbound REGISTER (sip_register trunks) when the create-trunk request does not specify one. |
SIP_OUTBOUND_REGISTRATION_MIN_EXPIRES_SECONDS | 60 | Lower clamp on the requested outbound REGISTER expiry. |
SIP_OUTBOUND_REGISTRATION_MAX_EXPIRES_SECONDS | 7200 | Upper clamp on the requested outbound REGISTER expiry. |
SIP_OUTBOUND_REGISTRATION_REFRESH_RATIO | 0.5 | Fraction of the **granted** expiry at which the trunk refreshes (e.g. 0.5 of a 600 s grant → refresh every 300 s). Must be (0, 1); out-of-range values fall back to 0.5. |
SIP_OUTBOUND_REGISTRATION_FAILURE_BACKOFF_MAX_MS | 300000 | Upper cap on the exponential backoff between failed outbound REGISTER attempts. Failures emit sip.outbound_registration_failed; the trunk stays in the manager and keeps retrying. |
SIP_OUTBOUND_PROXY | _(empty = route at the registrar / dialed URI)_ | Default next hop for outbound REGISTERs and INVITEs, attached as a loose Route: <sip:proxy;lr> header with the Request-URI left unchanged. Overridden per-trunk by sip_register.outbound_proxy on POST /v1/sip/trunks and per-call by outbound_proxy on POST /v1/legs; a to that resolves to an AOR registered here outranks all three. Digest auth still targets the registrar, not the proxy. Not applied to SIPREC SRC or WhatsApp legs. A malformed value fails startup. **Caveat:** inbound legs are tagged with trunk_id by the peer socket they arrive on, so when several trunks share one proxy that tag becomes ambiguous — it is informational, never an authorization gate. |
SPEECH_DETECTION_ENABLED | false | Emit speaking.started / speaking.stopped events for every connected leg by default. Per-call speech_detection on POST /v1/legs or POST /v1/legs/{id}/answer overrides this. |
AMRWB_MODE | 2 | AMR-WB (G.722.2) encoder speech-mode **ceiling** 0..8: 0=6.60, 1=8.85, 2=12.65, 3=14.25, 4=15.85, 5=18.25, 6=19.85, 7=23.05, 8=23.85 kbit/s. The actual transmit mode is this ceiling clamped to the peer's negotiated mode-set (so e.g. 8 yields HD 23.85 only when the peer allows it, falling back automatically). Default 2 (12.65) matches the GSMA IR.92 / VoLTE common rate. Out-of-range values clamp to 0..8. |
AMRWB_OCTET_ALIGNED | true | Offer octet-aligned AMR-WB framing (RFC 4867) in outbound SDP. When false, offers bandwidth-efficient framing. On answers, VoiceBlender always echoes the framing the peer negotiated. |
AMRNB_MODE | 7 | AMR-NB (RFC 4867) encoder speech-mode **ceiling** 0..7: 0=4.75, 1=5.15, 2=5.90, 3=6.70, 4=7.40, 5=7.95, 6=10.2, 7=12.2 kbit/s. The actual transmit mode is this ceiling clamped to the peer's negotiated mode-set. Default 7 is the GSM-EFR-equivalent 12.2 kbit/s, the highest AMR-NB quality and the rate most enterprise PBXes and mobile networks default to. Out-of-range values clamp to 0..7. |
AMRNB_OCTET_ALIGNED | true | Offer octet-aligned AMR-NB framing (RFC 4867) in outbound SDP. When false, offers bandwidth-efficient framing. On answers, VoiceBlender always echoes the framing the peer negotiated. |
VSI_EVENT_BUFFER_SIZE | 256 | Per-client buffer (in events) on the /v1/vsi WebSocket. When the client consumes events slower than they're produced, the buffer fills and new events are dropped (with a warn log on the leading edge of each drop burst and at every 10× threshold; the next delivered event also includes an events_dropped notification to the client). Clamped to [16, 1_000_000]. **Tuning:** larger values absorb longer back-pressure spikes at the cost of higher peak memory per client (roughly the average JSON event size × buffer size, e.g. ~1 KB × 256 ≈ 256 KB per connection at the default) and longer end-to-end latency for buffered events when the client recovers. Increase only if you observe drops on legitimate slow-consumer scenarios you can't fix at the client. |
SIPREC_ENABLED | false | Accept inbound SIPREC recording sessions (RFC 7866), where an SBC or PBX forks a call's media to VoiceBlender. When off, an INVITE carrying Require: siprec is rejected with 420 Bad Extension and one that only hints at SIPREC with 488. Requires SIP_TCP_ENABLED=true. |
SIPREC_AUTO_ANSWER | true | Answer an inbound recording session immediately instead of parking it until POST /v1/legs/{id}/answer. A session recording client does not wait for an application decision, so leaving this on is usually correct; turn it off to gate sessions from a controller. |
SIPREC_MAX_STREAMS | 8 | Maximum number of m=audio sections accepted on one recording session. A session offering more is rejected with 486, bounding the RTP ports and goroutines a single peer can claim. |
SIPREC_METADATA_MAX_BYTES | 65536 | Maximum size of the rs-metadata XML document in a SIPREC INVITE. A larger document is rejected with 413 rather than parsed. |
SIPREC_SRC_ENABLED | false | Allow POST /v1/rooms/{id}/siprec to originate outbound recording sessions, forking a room's participants to an external session recording server. Off by default: it lets an API caller stream a room's audio to an arbitrary SIP destination. |
SIPREC_AUTO_RECORD | false | Start multi-channel recording automatically when a SIPREC session is accepted, one channel per recorded participant. When false, recording is driven through the usual /v1/legs/{id}/record endpoint. |
SIPREC_ROOM_MODE | none | Where a recording session's audio streams are mixed. none attaches nothing and leaves placement to the stream API; per_session creates a room named siprec-<legID> and attaches every stream, which is what makes live STT and agents apply to a recorded call; fixed attaches every session's streams into SIPREC_ROOM_ID. |
SIPREC_ROOM_ID | _(empty)_ | Room every recording session's streams join when SIPREC_ROOM_MODE=fixed. Ignored in the other modes. |
MOQ_ENABLED | false | Enable the experimental MoQ (Media over QUIC) inbound leg endpoint at CONNECT /v1/legs/moq over WebTransport/HTTP/3. PoC quality: tracks IETF draft-11 via mengelbart/moqtransport, single MoQ session per leg, Opus framed one frame per MoQ Object (LOC-style). When enabled, both MOQ_TLS_CERT_FILE and MOQ_TLS_KEY_FILE must be set. |
MOQ_LISTEN_ADDR | :8443 | UDP address for the HTTP/3 listener that backs the MoQ leg. Independent of HTTP_ADDR — TCP/:8080 and UDP/:8443 can run side-by-side. |
MOQ_TLS_CERT_FILE | _(none)_ | Path to the TLS certificate used by the HTTP/3 listener. Required when MOQ_ENABLED=true. |
MOQ_TLS_KEY_FILE | _(none)_ | Path to the TLS private key used by the HTTP/3 listener. Required when MOQ_ENABLED=true. |
MOQ_OPUS_BITRATE | 24000 | Target bitrate (bps) for the Opus encoder feeding the MoQ leg's mix track. Must be in 6000..510000. |
LIVEKIT_ENABLED | false | Enable the livekit_room leg type at POST /v1/legs (type=livekit_room). Lets VoiceBlender join a LiveKit room as a participant and bridge audio between SIP and LiveKit. No LiveKit SDK is used — the signaling protocol is spoken directly via github.com/livekit/protocol protobufs over the existing pion stack. |
LIVEKIT_URL | _(none)_ | Default LiveKit server endpoint (wss://...). Required when LIVEKIT_ENABLED=true unless every request supplies livekit.url. Overridable per-request. |
LIVEKIT_OPUS_BITRATE | 24000 | Target bitrate (bps) for the Opus encoder publishing audio into LiveKit. Must be in 6000..510000. Overridable per-request via livekit.opus_bitrate. |
LIVEKIT_TOKEN_SIGNING_ENABLED | false | Opt-in: when true, callers may omit livekit.token and instead pass {room,identity,permissions}; VoiceBlender mints the JWT itself. **Security caveat:** enabling this stores the LiveKit API secret (a high-privilege credential that can mint tokens for any room/identity on the LiveKit deployment) in VoiceBlender. Keep off in multi-tenant deployments. |
LIVEKIT_API_KEY | _(none)_ | LiveKit API key used to sign minted JWTs. Required only when LIVEKIT_TOKEN_SIGNING_ENABLED=true. |
LIVEKIT_API_SECRET | _(none)_ | LiveKit API secret used to sign minted JWTs. Required only when LIVEKIT_TOKEN_SIGNING_ENABLED=true. Treat as a high-value secret; redact in logs. |
LIVEKIT_DEFAULT_TOKEN_TTL | 6h | Default TTL applied to minted JWTs when the request omits livekit.token_ttl. Go duration string. LiveKit recommends ≤ 6 hours. |
Verbatim transcript text, DTMF digits and full event payloads appear only at LOG_LEVEL=debug. Debug output is therefore PII-bearing and should not be shipped to a general-purpose log sink.
Set these env vars before starting voiceblender:
| Variable | Value |
|---|---|
SIP_TLS_PORT | 5061 |
SIP_TLS_CERT | path to fullchain.pem for your FQDN |
SIP_TLS_KEY | path to privkey.pem |
SIP_DOMAIN | the FQDN you registered with Meta (must match the cert SAN) |
Make a test outbound call:
curl -X POST http://localhost:8080/v1/legs \
-H 'Content-Type: application/json' \
-d '{
"type": "whatsapp",
"to": "+447900000000",
"from": "+441300000000",
"auth": { "password": "<meta-issued-digest-password>" },
"room_id": "wa-test"
}'
The HTTP response returns immediately with the leg in ringing; subscribe to the webhook or /v1/vsi event stream to see leg.connected (or leg.disconnected with a reason if Meta rejects the INVITE).
Full reference: API.md
API authentication. The HTTP surface — REST, the/v1/vsievent WebSocket,/v1/legs/websocket, the/v1/legs/moqWebTransport endpoint,/metrics, and pprof — has no built-in credential authentication (no API key, bearer token, or session). Access is gated solely by theALLOWED_IPSallowlist and your network placement. Do not expose it to untrusted networks; front it with a reverse proxy or gateway that enforces auth if you need per-caller credentials. This is distinct from SIP-layer auth (digest challenge of inbound INVITE/REGISTER) and outbound webhook signing (WEBHOOK_SECRET), which are covered separately below.
1. Register a webhook POST /v1/webhooks
2. Receive inbound call --> webhook: leg.ringing {leg_id, from, to}
3. Answer POST /v1/legs/{id}/answer
4. Create a room POST /v1/rooms
5. Add legs to room POST /v1/rooms/{id}/legs
6. Attach AI agent POST /v1/legs/{id}/agent
7. Start recording POST /v1/legs/{id}/record
8. Hang up DELETE /v1/legs/{id}
403 SIP server X.X.X.X from INVITE does not match any SIP server configured for phone number ... — SIP_DOMAIN doesn't match what's registered with Meta. Set it to the FQDN, not the IP, and confirm via the GET /settings query above.404 Not Found on outbound — usually means the recipient phone number isn't a valid WhatsApp user, or the destination URI is malformed. Confirm the digits in to are the actual user's E.164 number.Reason: ... not receiving any media for a long time — your audio path (RTP/UDP egress) is being dropped before reaching Meta. Check firewall rules for outbound UDP from the RTP_PORT_MIN–RTP_PORT_MAX range and that ICE-srflx candidates are correct.setup:actpass + ice-lite, and they don't initiate DTLS. VoiceBlender forces setup:active automatically; if you see pcmedia: DTLS state state=connecting for >5 s, run with LOG_LEVEL=debug and inspect pion's DTLS scope for the actual error.SIP_DEBUG=true to log the full RFC 3261 wire form of every SIP message, including the auth-bearing retry after the 401/407 challenge — that's the most useful diagnostic for any signalling-layer issue.VoiceBlender 是一个基于 Go 语言开发的专业级语音服务,旨在实现 SIP 与 WebRTC 语音通话之间的无缝桥接。它支持多方音频混音(Multi-party audio mixing),并提供完善的 REST API 和实时 Webhooks 机制,能够帮助开发者构建复杂的实时语音交互应用。
本项目具备强大的语音处理能力:支持 SIP 入站与出站通话,兼容 PCMU、PCMA、G.722 及 Opus 等多种编解码器,并���持 Digest Auth 与 RFC 4028 会话定时器。此外,它支持 SIP over TLS 以满足 WhatsApp 等平台的安全性要求,并具备 Early media 功能,允许在通话接通前播放自定义彩铃或 IVR 语音。通过灵活的 Inbound/Outbound 逻辑,可轻松实现与 Meta/WhatsApp 的集成。
在进行集成测试时,系统需要部署两个 SIP 实例。项目核心依赖于高性能的 Go 语言库,包括用于处理 SIP 协议栈的 sipgo,以及业界领先的 WebRTC 实现库 pion/webrtc,确保了语音传输的高效与稳定。
您可以通过 Go 编译环境直接构建并运行本项目。在项目根目录下执行 `go build -o voiceblender ./cmd/voiceblender` 生成二进制文件,随后通过 `./voiceblender` 命令启动服务。建议在生产环境中使用预编译的二进制文件以获得最佳性能。
项目提供了丰富的示例代码以帮助快速上手。您可以参考 `examples/call_handler.py` 使用 Python 编写 Webhook 监听器来处理入站 SIP 通话及会议室逻辑;使用 `examples/webrtc-client/` 构建基于浏览器的 WebRTC 语音客户端,支持房间管理与 DTMF 按键;此外,还提供 `gen_test_wav.py` 用于生成测试用的 WAV 音频文件。
VoiceBlender 的所有配置均通过环境变量(Environment Variables)进行管理。启动前需根据需求设置 `INSTANCE_ID`、`HTTP_ADDR` 及 `SIP_BIND_IP` 等参数。若需启用 SIP TLS 功能,必须配置 `SIP_TLS_PORT`、`SIP_TLS_CERT`、`SIP_TLS_KEY` 以及与 Meta 注册信息一致的 `SIP_DOMAIN`。
本项目提供完整的 REST API 接口用于控制通话生命周期、管理会议室及挂载 AI Agent。详细的接口定义、请求参数及响应格式请参阅项目中的 `API.md` 文档,以获取完整的 API Reference。
典型的业务工作流如下:首先通过 `POST /v1/webhooks` 注册 Webhook 接收通知;当收到入站通话的 `leg.ringing` 事件后,通过 `POST /v1/legs/{id}/answer` 接听;随后创建 Room 并将通话 Leg 加入其中;最后通过 `POST /v1/legs/{id}/agent` 挂载 AI Agent 并根据需要开启录音功能。
针对常见问题,若遇到 SIP 403 错误,通常是因为 `SIP_DOMAIN` 与 Meta 注册的 FQDN 不匹配,请确保配置的是域名而非 IP;若在发起出站通话时收到 404 Not Found,请检查接收方号码的有效性及路由配置。通过 `GET /settings` 接口可以快速核对当前配置状态。
aiskill88点评:底层能力扎实,将传统通信协议与现代AI语音流结合,是构建高性能语音Agent的理想基座。
AI Skill Hub 为第三方内容聚合平台,本页面信息基于公开数据整理,不对工具功能和质量作任何法律背书。
建议在沙箱或测试环境中充分验证后,再部署至生产环境,并做好必要的安全评估。
✅ MIT 协议 — 最宽松的开源协议之一,可自由商用、修改、分发,仅需保留版权声明。
总体来看,VoiceBlender 语音控制平台 是一款质量优秀的Agent工作流,在同类工具中具备一定竞争力。AI Skill Hub 将持续追踪其更新动态,建议收藏备用,结合自身场景选择合适时机引入使用。
| 原始名称 | voiceblender |
| 原始描述 | 开源AI工作流:A programmable voice platform: SIP and WebRTC call control, multi-party mixing, 。⭐68 · Go |
| Topics | 语音AIWebRTC实时通信 |
| GitHub | https://github.com/VoiceBlender/voiceblender |
| License | MIT |
| 语言 | Go |
收录时间:2026-05-26 · 更新时间:2026-05-30 · License:MIT · AI Skill Hub 不对第三方内容的准确性作法律背书。
选择 Agent 类型,复制安装指令后粘贴到对应客户端