AI Skill Hub 强烈推荐:风漫画 是一款优质的Agent工作流。AI 综合评分 8.0 分,在同类工具中表现稳健。如果你正在寻找可靠的Agent工作流解决方案,这是一个值得深入了解的选择。
风漫画 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
风漫画 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
# 方式一:npm 全局安装 npm install -g wind-comic # 方式二:npx 直接运行(无需安装) npx wind-comic --help # 方式三:项目依赖安装 npm install wind-comic # 方式四:从源码运行 git clone https://github.com/ChrisChen667788/wind-comic cd wind-comic npm install npm start
# 命令行使用
wind-comic --help
# 基本用法
wind-comic [options] <input>
# Node.js 代码中使用
const wind_comic = require('wind-comic');
const result = await wind_comic.run(options);
console.log(result);
# wind-comic 配置说明 # 查看配置选项 wind-comic --config-example > config.yml # 常见配置项 # output_dir: ./output # log_level: info # workers: 4 # 环境变量(覆盖配置文件) export WIND_COMIC_CONFIG="/path/to/config.yml"
<p align="center"> <img src="assets/banner.jpg" alt="Wind Comic — One line of text. One finished short drama." width="100%" /> </p>
<p align="center"> <b>One sentence in. A finished short-form drama out — script, cast, storyboards, voiceover, timeline, mp4.</b><br/> Multi-agent AI studio · reusable characters · novel→season splitting · director's control room · real-time collab · bring-your-own LLM. </p> <p align="center"> <b>一句话进,整片短剧出 —— 剧本 · 角色 · 分镜 · 配音 · 时间线 · mp4 一条龙。</b><br/> 多 Agent AI 创作工作室 · 可复用角色 · 长篇小说→自动分集 · 导演级控片台 · 实时协作 · 自带 LLM。 </p>
<p align="center"> <a href="https://github.com/ChrisChen667788/wind-comic/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-MIT-blue.svg" alt="MIT License" /></a> <a href="https://github.com/ChrisChen667788/wind-comic/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/ChrisChen667788/wind-comic/ci.yml?branch=main&label=CI&logo=github" alt="CI" /></a> <a href="https://github.com/ChrisChen667788/wind-comic/stargazers"><img src="https://img.shields.io/github/stars/ChrisChen667788/wind-comic?style=social" alt="GitHub stars" /></a> <img src="https://img.shields.io/badge/Tests-5397%2F5397-2ea44f" alt="5397 tests passing" /> <img src="https://img.shields.io/badge/Node-20%2B-339933?logo=node.js&logoColor=white" alt="Node 20+" /> <img src="https://img.shields.io/badge/Next.js-16-black?logo=next.js" alt="Next.js 16" /> </p>
<p align="center"> <b>English</b> · <a href="README.zh-CN.md">简体中文</a> · <a href="docs/MARKETING-en.md">🔥 Pitch</a> · <a href="docs/llm-providers.md">🔌 BYO LLM</a> </p>
<p align="center"> <a href="https://github.com/ChrisChen667788/wind-comic/raw/main/assets/promo/wind-comic-promo-en.mp4"> <img src="assets/promo/wind-comic-promo.gif" alt="Wind Comic — the full 39-second promo, condensed into a silent 15-second loop (click to watch with voiceover & sound)" width="100%" /> </a> </p> <p align="center"> ▶ <a href="https://github.com/ChrisChen667788/wind-comic/raw/main/assets/promo/wind-comic-promo-en.mp4"><b>Watch the full 39-second promo — with voiceover & sound</b></a><br/> <sub>Real cinematic footage woven with motion-graphics · 8 distinct art styles · English narration · scored by the platform's own MiniMax music engine.</sub> </p>
---
The 创作总览 dashboard: 99 projects + 4 case studies + recent activity feed + system status (engines in use, model versions). <p align="center"><img src="assets/screenshot-dashboard-v3.1.3.jpg" width="100%" /></p>
| What it does | Where it lives | |
|---|---|---|
| **Multi-agent pipeline** | Director / Writer / Char Designer / Storyboard / Editor — 8 agents | services/hybrid-orchestrator.ts |
| **Style Bible Frame** | One canonical key-art frame locks visual identity across all shots | lib/style-bible.ts |
| **Character DNA** | 8-dim vision-extracted character signature + per-shot prompt injection | lib/character-dna.ts |
| **Style Vision Audit** | Auto-regen any shot scoring <70 on palette/lighting/colorTemp/texture | lib/style-audit.ts |
| **Cameo Vision Retry** | Auto-regen any shot scoring <75 on character resemblance | services/cameo-retry.ts |
| **Pacing Audit** | Conflict-score / reversal-detect / cliffhanger per Chinese drama tropes | lib/pacing-audit.ts |
| **Drama Tropes** | 12 vertical-drama hook templates + 9:16 default + reversal density rules | lib/drama-tropes.ts |
| **CJK Subtitle Burner** | ffmpeg libass with system CJK font discovery | lib/text-control.ts + services/video-composer.ts |
| **Multi-track Timeline** | 3 tracks, drag/resize/snap/auto-collide, BGM waveform | components/project/cinema-timeline.tsx + lib/timeline-tracks.ts |
| **Real-time collab** | Yjs CRDT + WS server + presence + cursors + segment locks | scripts/ws-server.mjs + hooks/use-yjs.ts + hooks/use-segment-locks.ts |
| **Project invites** | viewer/commenter/editor role + token expiry + revoke | lib/project-share.ts |
| **Comments + @mentions** | Threaded comments, @-autocomplete, mention notifications | lib/comments.ts + lib/notifications.ts |
| **Lipsync** | Kling / Sync.so / Hailuo auto-select, fail-safe fallback | services/lipsync.service.ts |
| **Plan-gate billing** | Per-engine plan checks (Vidu Q3 = enterprise, etc.) | lib/plan-gate.ts |
| **API quota tracker** | Per-provider failure tracking + dashboard banner | lib/api-usage-tracker.ts |
| **18 project templates** | 霸总/重生/穿越/古装/科幻/儿童/纪实/恐怖/喜剧 etc. | lib/story-templates.ts |
| **BYO LLM docs** | 12-provider config matrix, 0-code swap | docs/llm-providers.md |
| 🆕 **Character Studio** | Multi-view turnaround + DNA lock + auto-bound voice + bio | lib/character-studio.ts |
| 🆕 **Prompt Workbench** | @-mention assets, autocomplete, compile-preview, readiness score | lib/prompt-ide.ts + components/prompt-editor.tsx |
| 🆕 **Long-form Intake** | Novel→episodes + narration modes + real TTS + season parallel | lib/story-intake.ts + lib/narration-synth.ts + lib/season-orchestrator.ts |
| 🆕 **Style Gallery** | 60 presets, 5 categories, one-click apply | lib/style-presets.ts + app/dashboard/styles |
| 🆕 **Director Console** | 4-stage pipeline model + stale detection + single-stage rerun | lib/pipeline-stages.ts + components/director-console.tsx |
| 🆕 **Team Workspace** | Credit pool + per-member allocations + RBAC + real invites | lib/team-credits.ts + lib/team-invite.ts |
| 🆕 **Postgres cutover (v9)** | SQLite↔PG dual-driver; **all write paths on async repos** (project_assets/projects/users/notifications/comments cleared), tx commit+rollback verified, DB_DRIVER=pg opt-in | lib/db-driver.ts + lib/repos/* + scripts/pg-migrate.ts |
| 🆕 **API Health Board** | Live model/gateway status + balance + out-of-credits detection | lib/provider-health.ts + app/dashboard/health |
---
```bash
git clone https://github.com/ChrisChen667788/wind-comic.git cd wind-comic npm install
v3 shipped the pipeline. v6 turned it into a production studio; v7–v9 hardened it into a platform; v10 closed the lip-sync, template-market, and cost loops. Reusable characters, a prompt IDE, novel→season auto-splitting with real voiceover, a 60-style gallery, a director's control room, team credit budgets, an industry-grade script audit (Polish Pro, v7.1), a premium design pass (v8.3), a fully-migrated Postgres backend (v9), a lip-sync delivery pipeline + template market + cost observability (v10), and a live API health board — every screen below is a real capture of the running app.
<p align="center"> <img src="assets/v10/landing.jpg" alt="Wind Comic landing — 青枫漫剧 AI Animation Agent Studio, 8 agents · 7 engines · 3 consistency guards, looping cinematic hero" width="100%" /> <br/><sub>v10 landing — looping cinematic hero, 8 collaborating agents · 7 media engines · 3 consistency guards.</sub> </p>
Lock a consistent visual identity before you generate. Search, filter by category, and apply any preset straight into the creation workshop. <p align="center"><img src="assets/v6/styles.jpg" width="100%" /></p>
全部由node scripts/capture-v12-425.mjs对着本机真实数据跑出来。脚本带空壳检测与近重复检测: 页面没渲染出来、或者「点了 tab 但视图没变」的,一律删掉并如实报告 —— 一张截图冒充两张,比少一张更糟。 📚 完整 18 张 + 每张的实现逻辑与工作流架构说明:docs/SCREENSHOTS-v12.425.md (单独成文是为了不让首页替访客扛下载量 —— README 引用媒体有 12MB 预算门禁。) v12.426 补记:上面「我的项目」那张(docs/screenshots/v12/01-my-projects-manage.jpg)已重拍。 旧图里卡片大半是渐变占位 —— 不是没作品,是封面被冻结成了一次失败的快照:流水线末尾抄一份第 1 镜的imageUrl就再不重算,于是出图全挂时的 mock 兜底图被永久写死(30 个项目里 12 个如此,其中 3 个明明各有 11~12 张真分镜还活着)。现改为读时解析,并清掉了 21 个夹具/无素材项目(下架,可一键恢复)。
| 角色转身图 · v12.425 修的就是这里 | 导演台 · 全链路控片 |
|---|---|
|  |  |
立绘原生 896×1152 竖构图,修前被塞进 355×200 横框 + object-cover,**裁掉 56%**,必须点全屏才看得到完整图。现在框比例跟素材走、填充改 object-contain。 | 四个环节各自可寻址:剧本 / 角色·场景 / 分镜 / 成片,任一环单独「编辑」或「重跑」,并预告重跑的下游影响。 |
| 拉片分析 · 五栏出厂真值 | 分镜规格 · 出厂参数 + 一致性仪表 |
|---|---|
|  |  |
| 每镜给出叙事要素 / 时间 / 镜头语言 / 影像处理 / 声音五栏,且是**流水线生成时的真实摄影语言,不是 AI 事后看图反推**;可导出 CSV / 剧本册 MD / PDF,也能回灌外部片子做复刻。 | 画幅 / 色彩 / 帧率 / 安全框是**工程参数不是滤镜**(Scope 2.39:1 · ACES 1.3 · 24fps),真实进入生成 prompt 与导出参数;右侧逐镜一致性打分,低分镜一键跳转重生。 |
---
All captured by node scripts/capture-v12-416.mjs against real local data (1440×900 @2x, downscaled to 1600px wide). The script has blank-page detection and near-duplicate detection: a screen that failed to render, or one where "the tab was clicked but the view never changed", is deleted and reported honestly — one screenshot masquerading as two is worse than one missing.
Only one project-page shot: clicking the tabs and scrolling the body both leave the view unchanged.This conclusion was wrong and is retracted (v12.425.) The tab bar sits below the fold, so a coordinate click missed it;scrollIntoViewfollowed byel.click()switches tabs fine. See the v12.425 section above for 12 shots taken that way.
Measured performance (local dev server, curl time-to-first-byte): landing 46 ms; console pages (projects / assets / engine health / usage) 22–38 ms. Tests: 5294 passing · 0 failing; preflight 10/10; CI all five jobs green.
Below is the foundational v3 pipeline (the v6 studio screens are in the New in v6 section above). Every panel is a real puppeteer capture of the running app (run node scripts/capture-screenshots.mjs / node scripts/capture-v6.mjs to refresh).
Every model call is provider-pluggable (priority chain + automatic fallback). Creative and high-frequency LLM traffic are split across two model tiers, and MiniMax is always the last-resort fallback on any error / out-of-credits / timeout:
| Capability | Default model (env override) | Supplement / fallback |
|---|---|---|
| **Creative LLM** (writer / director) | deepseek-v4-pro (OPENAI_CREATIVE_MODEL) + deepseek-v4-flash fast tier for drafts/polish | MiniMax-M2.7 (LLM_FALLBACK_MODEL) |
| **General LLM** (planning / validation / Vision-Audit) | claude-sonnet-4-6 (OPENAI_MODEL) | MiniMax-M2.7 |
| **Video** | veo3.1-pro (VEO_MODEL) | veo3.1 · Kling → **MiniMax Hailuo** (Sora-2 retired — API EOL 2026-09-24) |
| **Image** | flux.1-kontext-pro (IMAGE_MODEL) | Midjourney (mj_imagine) · fal FLUX Kontext · local ComfyUI → **MiniMax image** |
| **TTS / voiceover** | gpt-4o-mini-tts (VE_TTS_MODEL) | MiniMax T2A (speech-02-hd) |
| **Music / BGM** | MiniMax music | (Suno when gateway channel available) |
-pro, a reasoning model) carries writer/director quality work; the general tier (Claude sonnet-4-6) handles high-frequency planning / validation / Vision-Audit; the -flash tier keeps draft-compare & basic polish at sub-second latency..env.local (OPENAI_* / OPENAI_CREATIVE_* / VEO_* / IMAGE_MODEL / VE_TTS_MODEL / MINIMAX_*) — zero code change. See docs/llm-providers.md.---
cp .env.example .env.local
npm run dev:ws # Yjs WebSocket server on :1234
Live status for every model and gateway: 正常 / 额度用尽 / 配置缺失 / 不可达, with real balance read-out and a "去充值 / 补配置" hint. Keys are never stored or returned. <p align="center"><img src="assets/v6/health.jpg" width="100%" /></p>
docker run -p 3100:3100 -e MOCK_ENGINES=1 -e JWT_SECRET=demo -e NEXTAUTH_SECRET=demo \
ghcr.io/chrischen667788/wind-comic:latest
Then open <http://localhost:3100>. No clone, no npm install, no local build.
Prefer compose (also seeds a demo project):
docker compose -f docker-compose.demo.yml up
⚠️ Honest note: demo mode runs the deterministic mock engines (MOCK_ENGINES=1 — SVG storyboards, solid-colour clips, sine-wave audio). It demonstrates the pipeline and the product shape, and is not representative of real generation quality. Every artifact is labelled as mock — passing a placeholder off as a finished film is exactly the thing this project keeps hunting down. Add a single API key to switch to real engines.
Features that still ship. Director console · novel→season · finished-film station · team workspace · Cinema timeline are refreshed to v10 (live demo data). The style gallery, API health board, Polish Pro audit, and character turnaround are kept as earlier (v6–v8) captures because they show fuller sample output (the full style grid / live balances / a complete Pro audit / a 3-view turnaround sheet).
Director plans the story → Writer drafts dialogue under McKee structure → Style Bible Frame locks the look → Character Designer extracts an 8-dimension DNA signature of each character → Storyboard renders with Vision Audit (auto-regen on <70 score) → Video producer races multiple engines (Minimax / Veo / Kling) → Editor cuts j/l-cut on emotional beats and burns CJK subtitles.
Lineup verified 2026-08-31 (Artificial Analysis arena; 6 parallel web-research lenses + key figures independently re-checked. Full analysis:docs/COMPETITIVE-GAP-2026-08.md): the top slot changed hands, and a Chinese model took it — T2V-with-audio: Wan 3.0 (1242) → Gemini Omni Flash (1237, down from 1244). New entrant xAI Grok Imagine Video 1.5 (Jun 16): 7 simultaneous visual reference anchors (character + scene + prop + style) plus native audio and voice cloning, $0.08 (480p) / $0.14 (720p) / $0.25/s (1080p) — more reference slots than our current cref+sref pair. Kling 3.0 remains the all-round pick for short drama (native 4K / 60fps / 15s / 6 coherent shots / built-in multilingual dialogue + lip-sync, $0.084–0.112/s); Veo 3.1 (first shipped 2025-10-15; 4K and vertical added 2026-01-13) can reach 60s+ via Scene Extension — but that is stitching; the single-generation ceiling is still 8–10s. It uses its own joint audio-visual generation (Lyria 3 is a separate Google music model, not Veo's audio component), from $0.40/s. Note: Veo 3.1's own T2V+audio Elo is ~1091 — the 1237 on the board belongs to Gemini Omni Flash, a different product. ✅ The Sora 2 API shutdown (2026-09-24) is confirmed by OpenAI — and this project has been guarded since v12.173/207: warn before the date, auto-drop from the model chain after it (falling back to veo/kling), throw only if the chain empties (tests/v12-173-sora-sunset.test.tslocks it). Not an open action item. 🇨🇳 The domestic camp broke through on both shot length and reference count — and two engines we already ship are behind. Seedance 2.5 (7-31) and Wan 3.0 (8-24 GA) both do native 30s single takes; Seedance accepts 50 reference inputs (token-billed, ¥42–70/M tokens), Wan 3.0 runs ¥0.42/s at 720P and converts documents straight to video (PPT/Word/PDF → 30s). Vidu Q3 is the only Chinese model explicitly positioned for short-drama industrialisation (7 reference images with multi-subject locking, 6 cinematic effect packs, lip-sync driving; Turbo 1080p ≈¥0.41/s). PixVerse C1 (4-08) is the most narrowly short-drama-vertical of all (storyboard-grid output + multi-speaker lip-sync). ⚠️ Two findings about our own stack, which matter more than any competitor: ①services/minimax.service.ts:240defaults toMiniMax-Hailuo-2.3(nothing overrides it in.env.local, so that is what actually runs) — but 2.3 / 2.3-Fast are marked legacy by the vendor; H3 (V2 API) is the recommended path, and this vendor has pulled an endpoint before with no notice (Music API: 410 + 2153). We list MiniMax H3 as a competitor column in the table below while still calling its previous generation. ② our two Vidu call paths disagree:qyt-vidu.service.tspinsviduq3, whilevidu.service.tssends no model field at all and rides whatever the vendor defaults to — the same disease as Midjourney「no version pinned anywhere, gateway default wins」: when the vendor changes its default our behaviour changes silently, and we cannot reconstruct afterwards which model produced a given shot. Both are queued as v12.402 / v12.403. <sub>This round re-checked every high-risk figure with an independent second search, and two claims I had already written into this README were overturned: Veo 3.1 first shipped 2025-10-15 (2026-01-13 was a feature update), its 60s+ is Scene Extension stitching — the single-generation ceiling is still 8–10s, and Lyria 3 is a separate Google music model, not Veo's audio component; and Runway Gen-4.5 has had native audio since 2025-12-11 — the earlier draft called its absence a fatal flaw, which was simply wrong. Elo figures are arena snapshots and move with voting.</sub> ⭐ BYO 架构再次接住这波:榜上模型基本都开放 API,填 key 即成为本管线可调度的引擎 —— 竞品越强,本管线越强。
🔴 We overclaimed last round — four "only we have this" claims retracted (verified 2026-09-03). This is not competitors catching up; the claims were too broad to begin with: ① "EDL/AAF export is unique to us" — Descript has had it: Timeline Export covers Premiere XML, FCPXML, Reaper (EDL) and Pro Tools/Logic (AAF). ② "the only MIT, self-hostable end-to-end AI short-drama platform" — false:EvoLinkAI/ai-short-drama(MIT + Docker + novel→script→storyboard→image→video→voice→film) qualifies too, alongside LocalMiniDrama, Toonflow, huobao-drama and others. ③ "BYO multi-provider registry is unique" — PopShort.AI already ships cross-vendor models (Veo 3.1 / Kling / Seedance / Nano Banana Pro). ④ "no competitor does pacing/reversal auditing" — Volcano Drama publicly claims a multi-agent verification mechanism with 200+ shot strategies optimising conflict, reversal and climax. What survives, stated narrowly: Volcano Drama optimises inline during generation, and Descript exports but has neither pacing audit nor an open licence — so as of 2026-09-03, no single product has all three of「pacing audit hardened into an independent blocking engineering gate + EDL/AAF export + open-source self-hostable for commercial use」. Note this is narrower than last round: "pacing audit" alone no longer holds; it must be "an independent blocking gate". Open-source scale: OpenMontage 54.7k★ (AGPL), ViMax 12.2k★ (MIT), DramaClaw 4.8k★ (Elastic 2.0, not FOSS), BigBanana 1.8k★ (non-commercial), Novella AI 89★ (MIT). Full analysis:docs/COMPETITIVE-GAP-2026-09.md. v12.214→244 双线推进。产品层:GPT Image / Nano Banana(Gemini)接入插件式图像 provider 链(issue #11,社区 @flobo3 提议,OPENAI_IMAGE_ENABLED/GEMINI_API_KEY门控、原生 i2i 接角色一致性契约);多集连续生成补上「剧情记忆」(第 N 集 Writer 注入前几集前情提要 + 承接纪律,对标红果/阅文的 60~100 集连续,此前各集独立成篇)。平台/工程层:六轮独立对抗复检把安全洞从 CRITICAL 到 LOW 全清(SSRF 逐跳重验重定向 + IPv6 全隧道变体 / serve-file 签名能力 URL / WebSocket 鉴权 / 预算护栏),并把反复踩的「改了守卫却没跟到消费方」这个病固化成 CI 入库门禁(npm run gate:consumer,零容忍,上线即抓到 2 个人肉复检漏掉的真 SSRF)。 结论不变:生成层已是红海(竞品在出片/多镜/音频都第一梯队),Wind Comic 护城河收窄到「制作/平台层」——节奏审计、智能剪辑、字幕烧入、协作、自托管、开源、BYO。
| Capability | Veo 3.1 | Kling 3.0 | Seedance 2.5 | Gemini Omni Flash | MiniMax H3 | ViMax (open-source) | **Wind Comic** |
|---|---|---|---|---|---|---|---|
| Multi-shot story from one prompt | ⚠️ | ✅ storyboard mode | ✅ multi-shot native, 30s single take | ⚠️(单段生成,多镜叙事非强项) | ⚠️ (4~15s single take) | ✅ 12-agent script→video | **✅ 8-agent script→edit pipeline** |
| Character consistency across shots | ✅ | ✅ | ✅ up to 50 reference inputs | ✅ 多模态统一 | ✅ reference-to-video | ✅ character extractor agent | **✅ cref + sref + 8-dim DNA + vision retry** |
| Style coherence locked | ✅ | ✅ | ✅ | ✅ | ⚠️ | ⚠️ | **✅ Style Bible Frame** |
| Native dialogue + SFX audio | ✅ | ✅ 多语对白+口型 | ✅ | ✅ 4 模态原生一体 | ✅ 原生立体声 | ❌ 无配音模块 | **✅ per-character TTS + lip-sync** |
| Real CJK subtitles (burned-in) | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | **✅ libass + open-license CJK font burn** |
| Vertical drama tropes | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | **✅ 12 templates + 9:16 default** |
| Real-time multiplayer timeline | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | **✅ Yjs CRDT + Y.Map locks + cursors** |
| Self-hostable | ❌ | ❌ | ❌ | ❌ | ⚠️ 权重开放(排除美/欧/英/韩) | ✅ 本地部署 | **✅ Next.js + SQLite/Postgres** |
| BYO LLM (OpenAI / Claude / DeepSeek / local) | ❌ | ❌ | ❌ | ❌(它自己就是模型) | ❌ | ⚠️ 需改代码 | **✅ 12+ providers via .env** |
| Open source | ❌ | ❌ | ❌ | ❌ | ⚠️ 权重部分开放 | ✅ MIT | **✅ MIT** |
| Per-shot regenerate with custom prompt | ⚠️ | ✅ | ⚠️ | ✅ motion brush | ✅ 对话式迭代编辑(招牌能力) | ✅ video-edit 端点 | **✅ + reference image upload** |
| Pacing / conflict audit | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | **✅ v2 (v12.275): shot score + reversal detection **plus** conflict-curve shape (escalating / flat / front-loaded / no-climax), drag-segment localisation to exact shot ranges, opening-density check, and duration-rhythm analysis — every finding names the shots to fix** |
| Smart editing (beat-snap + emotion pacing + one-instruction style) | ❌ | ❌ | ❌ | ❌ | ⚠️(对话式改片,非结构化卡点/情绪剪辑) | ❌ | **✅ beat-snap · emotion pacing · emphasis · transition aesthetics · "fast & hype/slow & lyrical" in one line (BYO LLM)** |
| First+last frame lock (image_tail cut-to-cut coherence) | ❌ | ✅ | ✅ | ⚠️ | ✅ (I2V) | ❌ | **✅ Kling FLF wired into main pipeline, per-shot tail-frame picker** |
| Multi-character face cast library (post-build editable) | ❌ | ✅ 主体库 | ✅ 角色管理 | ⚠️ | ⚠️ | ❌ | **✅ 3-slot cast + cross-shot subject_reference injection** |
| One-click localization (script + re-voice) | ❌ | ⚠️ dub only | ❌ | ❌ | ❌ | ❌ | **✅ 8-lang translate → apply → re-TTS, honest degradation** |
| Royalty-free AI BGM per story | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | **✅ MiniMax music-2.6, style-prompt → project BGM** |
| Per-shot auditable decision log (engine/cost/consistency) | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | **✅ decision log + cost drill-down + quality score** |
| Emotion-aware TTS (mapped to native enum) | ⚠️ | ⚠️ | ⚠️ | ❌ | ⚠️ | ❌ | **✅ CN emotion → MiniMax speech-2.8-hd enum, live A/B verified** |
| Lip-sync wired into pipeline (auto per dialogue shot) | ✅ | ✅ | ⚠️ | ⚠️ | ✅ 原生 | ❌ | **⚠️ Kling lip-sync — zh/en only (needs public video URL + audio ≥2s); ja/ko/ru degrade to none; honest skip on non-face** |
| Full-app i18n (zh/en/ja/ko/ru, all UI) | ⚠️ | ⚠️ | ⚠️ | ⚠️ | ❌(是模型不是 app) | ❌ | **✅ 5-language core UI, ~400 keys (component-level string cleanup ongoing)** |
Cells marked ⚠️ = the feature exists but in a limited / locked-down form (e.g. "you can only do this on a paid Pro tier through a UI panel").
---
Wind Comic 是一款革命性的 AI 短剧创作工作室。只需输入一行文字,系统即可自动生成包含剧本、角色、分镜、配音、时间轴及最终 mp4 视频在内的完整短剧。本项目采用 Multi-agent AI 架构,支持将小说自动拆解为季剧模式,并提供可复用的角色资产,让创作流程从“单点生成”进化为“工业化流水线”。
Wind Comic 拥有竞品难以企及的核心技术:内置由 Director、Writer、Editor 等 8 个 Agent 组成的 Multi-agent pipeline,实现全流程自动化;独创 Style Bible Frame 技术,通过标准视觉参考帧锁定所有镜头的视觉一致性;并结合 Character DNA 系统,通过 8 维视觉特征提取,确保角色在不同分镜中保持高度统一的形象。
请通过 Git 克隆仓库并完成基础环境安装。首先执行 `git clone https://github.com/ChrisChen667788/wind-comic.git`,进入项目目录后运行 `npm install` 即可完成依赖包的安装。建议在 Node.js 环境下运行,以确保后续开发服务器的正常启动。
项目支持快速启动模式。在完成依赖安装后,通过运行 `npm run dev` 启动本地开发服务器。此外,本项目还支持实时协作功能,你可以通过开启第二个终端运行 `npm run dev:ws` 来启动基于 Yjs 的 WebSocket 服务器(端口 :1234),实现多人在线协同创作。
项目启动前需进行环境配置。请先复制 `.env.example` 为 `.env.local`,并根据 `docs/llm-providers.md` 的说明填写必要的 API 密钥。配置过程包含 3 行强制性参数,确保你的 OpenAI API 或其他兼容接口配置正确,以便 Multi-agent 系统能够正常调用大模型能力。
Wind Comic 内置了 API Health Board(API 健康看板),旨在解决开发者在使用过程中 API Key 失效或额度不足的痛点。看板会实时监控每个 Model 和 Gateway 的状态(如:正常、额度用尽、配置缺失、不可达),并直接读取余额,提供直观的“去充值/补配置”提示,确保创作流程不因接口问题中断。
本项目拒绝“黑盒模型”模式,而是采用精细化的 Multi-agent 工作流:Director 负责故事规划 $\rightarrow$ Writer 基于 McKee 结构撰写对话 $\rightarrow$ Style Bible Frame 锁定视觉风格 $\rightarrow$ Character Designer 提取角色的 8 维 DNA 特征 $\rightarrow$ Storyboard 进行分镜渲染并配合 Vision Audit 进行自动质量审计(评分低于 70 分将自动重绘) $\rightarrow$ 最后由 Video producer 调用多个引擎完成视频合成。
AI Skill Hub 为第三方内容聚合平台,本页面信息基于公开数据整理,不对工具功能和质量作任何法律背书。
建议在沙箱或测试环境中充分验证后,再部署至生产环境,并做好必要的安全评估。
✅ MIT 协议 — 最宽松的开源协议之一,可自由商用、修改、分发,仅需保留版权声明。
总体来看,风漫画 是一款质量优秀的Agent工作流,在同类工具中具备一定竞争力。AI Skill Hub 将持续追踪其更新动态,建议收藏备用,结合自身场景选择合适时机引入使用。
| 原始名称 | wind-comic |
| 原始描述 | 开源AI工作流:Multi-agent AI pipeline that turns one line of text into a finished short-form d。⭐42 · TypeScript |
| Topics | aicomictypescript |
| GitHub | https://github.com/ChrisChen667788/wind-comic |
| License | MIT |
| 语言 | TypeScript |
收录时间:2026-05-30 · 更新时间:2026-05-30 · License:MIT · AI Skill Hub 不对第三方内容的准确性作法律背书。
选择 Agent 类型,复制安装指令后粘贴到对应客户端