changelog · semver · MIT · 2026

Every shipped version.

Format follows Keep a Changelog. Versioning follows Semantic Versioning. Each entry is dated, tagged by category, and links back to its GitHub release tag.

latest v0.8.7 · 2026-04-30
first commit v0.0.1 · 2026-04-18
versions shipped 28
license MIT
v0.13.0 pocket-tts provider · up-front credentials checks release ↗ 2026-08-26
added
  • Kyutai pocket-tts as an eighth provider. Local, $0/call, no API key, ~4× realtime on CPU. 26 built-in voices plus zero-shot cloning from a reference .wav. MIT code and CC-BY-4.0 weights, so it is clean for client work. Added because Voicebox was the only $0 provider and therefore a single point of failure.
  • Second cloning-capable backend. generate_speech --referenceAudioPath now accepts --provider pocket-tts, so local is no longer the only option.
  • check_pocket (CLI --check-pocket, /gen-check) — a full credentials and capability report: reachability, built-in voices, whether cloning actually works, and whether the provider is enabled. $0 and about half a second.
changed
  • ProviderHealth gained an optional note, rendered by health_check. A provider can be reachable and still be missing a capability; that used to read as a clean pass.
notes
  • It must refuse rather than degrade. On any cloning-weights download failure the upstream library quietly loads the non-cloning model and speaks in a stock voice — and /health still answers healthy throughout. Verified against a live server. So check_pocket clones a probe clip instead of trusting /health, generate_speech runs the same probe as a preflight, and the error becomes a refusal naming the fix (hf auth login).
  • The probe seeds itself — one token synthesized with a built-in voice, fed straight back as the cloning reference — so it needs no file from the user.
  • Reference voices are keyed by their bytes, not their path. A reference can be re-recorded in place, and a path-keyed cache would serve the previous voice forever with nothing to say so.
v0.12.0 five-tier video ladder · text-to-video release ↗ 2026-08-26
added
  • generate_video gains the same five rungs as generate_avatar, adding prunaai/p-video alongside xai/grok-imagine-video-1.5: draft $0.005 · low $0.02 · normal $0.04 · high $0.08 · ultra $0.14, all per second.
  • Text-to-video. grok is image-to-video only, so imagePath used to be mandatory. On the p-video rungs it is optional — omit it and the clip comes from the prompt alone. The sidecar omits imagePath so regenerate reproduces it as text-to-video.
  • estimate_cost --modality video ranks all five rungs with per-rung pricing.
changed
  • Default video tier is now normal (p-video 1080p) — half the price of the previous default at a larger frame. This changes which model an un-tiered call uses.
  • Video slots moved out of the registry TierTable into a dedicated ladder; three slots cannot express two models with different length caps and input rules.
fixed
  • Re-rolls booked at $0 while billing for real. regenerate always passes model from the sidecar, and the explicit-model path blanked the slot params — dropping resolution/draft from the price key, falling through to "unknown pricing", and recording the call at $0. Affected generate_avatar since v0.10.0.
  • estimate_cost no longer re-derives tier→params with its own hardcoded map, which priced the wrong rung as soon as a ladder grew.
v0.11.0 talking-avatar tier ladder · 30× price range release ↗ 2026-08-07
added
  • Five-rung avatar ladder via prunaai/p-video alongside veed/fabric-1.0: draft $0.005 · low $0.02 · normal $0.04 · high $0.08 · ultra $0.15 per second. The floor drops from $0.08/sec to $0.005/sec, so iterating on framing and timing costs cents.
  • generate_avatar --prompt — motion prompt for the p-video tiers, recorded in the sidecar so regenerate reproduces it.
  • Multi-shot guidance in the skill for scripts over the cap: split the text (not the rendered audio) at sentence boundaries, target ~17s segments, vary shot size by ~20% between adjacent cuts. Per-second billing makes splitting cost-neutral.
fixed
  • Over-length audio no longer silently truncates. p-video returns status: succeeded with only the first 20 seconds when handed longer audio — no error, nothing in the response to detect it by. Verified: 35.44s in, 20.02s out. The tool now refuses pre-call and quotes both uncapped tiers priced from the real duration. It does not auto-escalate.
v0.10.0 talking avatars (lip-sync) via VEED Fabric 1.0 release ↗ 2026-08-03
added
  • generate_avatar — an image plus speech audio produce a clip whose mouth, head and subtle body motion track the voice. Built for outreach and personalized messages: generate_image → generate_speech → generate_avatar.
  • Output length equals the audio duration, probed via ffprobe so the pre-call budget guard works. Billed per second.
v0.9.6 file metadata in the gallery lightbox release ↗ 2026-08-02
added
  • Lightbox also shows type, dimensions (sharp for images, ffprobe for video), size, and the path relative to the gallery's cwd.
v0.9.5 gallery lightbox metadata panel release ↗ 2026-08-02
added
  • Clicking an item opens a lightbox with compact facts (kind, provider, model, tier, params, cost, cached, date), then the full scrollable prompt, with open-file and copy-prompt actions.
v0.9.4 HTML gallery of everything you generated release ↗ 2026-07-31
added
  • gallery scans the output dirs, reads each sidecar, and writes a self-contained theme-aware gallery.html: thumbnail grid with prompt, model, tier, params, cost and date per card, plus client-side filter, search, sort and a lightbox.
  • Sources from the persistent output dirs, not the daily-resetting session ledger, so it gathers everything that still exists. Files without sidecars still render.
v0.9.3 gpt-image-2 background + custom sizes release ↗ 2026-07-23
added
  • background: auto | opaque | transparent. The adapter fails fast on transparent + gpt-image-2 rather than spending a doomed call.
  • size: "WIDTHxHEIGHT" exact output size, validated (÷16, ≤3840px, ratio ≤3:1, 0.65–8.3 MP) and priced at the nearest tier by megapixels.
v0.9.2 gpt-image-2 up to 4K via opt-in resolution release ↗ 2026-07-23
added
  • resolution: 1K | 2K | 4K, lifting the previous 1536px cap. Default 1K keeps output and cost unchanged. Resolution-keyed pricing keeps the ledger accurate (4K-high ≈ $0.41 vs 1K-high ≈ $0.21).
v0.9.1 refresh Google Gemini models + pricing release ↗ 2026-07-23
changed
  • Refreshed the Gemini family ahead of Imagen 4's 2026-08-17 shutdown: pro → gemini-3-pro-image (now batchable), mid → gemini-3.1-flash-image, small → gemini-3.1-flash-lite-image at ~$0.034/img.
v0.9.0 image-to-video via Replicate grok-imagine-video-1.5 release ↗ 2026-07-19
added
  • Video as a first-class third modality alongside image and TTS, with replicate as a video-only provider.
  • generate_video — image-to-video, 1–15s clips, per-second pricing, audio synthesized free in the same pass.
v0.8.7 build switched from esbuild → tsc · deterministic dist release ↗ 2026-04-30
changed
  • Build switched from esbuild bundling to plain tsc compilation. Root-cause fix for the dist-sync CI failure that has been red on every push since the workflow was added 2026-04-18. esbuild produces ~3700 bytes of platform-specific output between darwin-arm64 and linux-x64 even on identical source + lockfile + esbuild version. Switching to tsc (deterministic across OSes) makes Mac-built and Linux-built dists byte-identical. No runtime behavior changes — native deps (@imgly, sharp, onnxruntime-node) were already --external in the bundle config, so they were already runtime-resolved from node_modules.
  • dist/ is now ~50 small files instead of 4 bundled megafiles. server-main.js shrank from 2.5 MB → 39 KB; cli-main.js 1.9 MB → 22 KB. Total dist is a fraction of the bundled size because the bundle was inlining dep code that's now resolved from node_modules (which we already ship via the bootstrap install).
fixed
  • Hardcoded version strings. Server identity reported v0.8.5 (last manual bump) and CLI banner reported v0.5.0 (never updated). Both now read 0.8.7. Long-term: read from package.json at startup so this can't drift again — left for a follow-up.
internal
  • Removed esbuild from devDependencies.
  • New scripts/postbuild.mjs copies pricing.json next to the compiled pricing/load.js and prepends shebangs to entry-point files. Cross-platform (Node fs API only).
  • pricing/load.ts switched from JSON import attributes (import x from "./pricing.json" with { type: "json" }) to plain readFileSync + JSON.parse. The import-attributes syntax is Node 22+ stable but unsupported on Node 18 (our minimum); the filesystem read is universal.
v0.8.6 bg-remove streams progress · misleading download copy fixed release ↗ 2026-04-29
changed
  • bg-remove now streams phase progress to stderr. Previously the first post_process --bgRemove call would appear to hang silently for ~30s. We now pass a progress callback to @imgly/background-removal-node and emit one stderr line per phase transition (fetch, compute, etc.) so the user sees activity during ONNX warmup. Subsequent calls run in <1s with one or two phase lines.
fixed
  • Misleading "downloads ~80 MB" copy in the bg-remove error. The default medium ONNX model is bundled in the npm package (chunked into 4 MB SHA-named blobs in the @imgly dist), so there's no runtime download. The slow first call is ONNX runtime warmup + loading the chunks from disk into memory. Updated the error message to describe the actual behavior.
v0.8.5 voicebox + local providers auto-detect at startup release ↗ 2026-04-29
changed
  • Voicebox + local providers now auto-detect at startup. Previously both required an explicit VOICEBOX_ENABLED=true / LOCAL_ENABLED=true env var to be eligible for the failover chain — and an explicit --provider voicebox request was rejected when the flag was unset. That meant a perfectly running Voicebox server on localhost:17493 would still be ignored unless the user happened to know the flag.

    Behavior now: *_ENABLED env vars are tristate. If explicitly set to true or false, that value wins. If unset, the server probes the endpoint at startup ({voiceboxBaseUrl}/health, {localBaseUrl}/models, 800 ms timeout) and treats reachable endpoints as enabled. Resolution is logged at info level.

    Demo flow simplification: with this change, a default demo install needs only GEMINI_API_KEY — Voicebox is picked up automatically if it's running.
fixed
  • MCP server identity reported version: "0.0.1" — hardcoded since the project started, never updated. Now matches the released version.
  • Cost reported as $0 when --model is passed explicitly. When a user passes --model gpt-image-2 (or any model whose registered slot uses tier-suffixed pricing keys like openai/gpt-image-2:medium), the inlineSlot helper in generate-image.ts and generate-speech.ts was always setting params: {} — dropping the tier→quality mapping. The pricing-key lookup then missed (openai/gpt-image-2 instead of openai/gpt-image-2:medium), unknownCostEstimate returned $0, and the call ran "free." The bigger consequence: the pre-call budget guard never fired for explicitly-named models, because $0 < any cap. Now inlineSlot consults the registry — if the explicit model matches the registered slot for the same (provider, modality, tier), it inherits the registered slot's params (and the rest of the slot config, e.g. batchable). Arbitrary unknown models still fall through to the empty-params slot as before.
v0.8.3 native deps install on first MCP-server start release ↗ 2026-04-28
fixed
  • Native deps now install on first MCP-server start. @imgly/background-removal-node, onnxruntime-node, and sharp ship as --external and were never installed by the marketplace fetch — calls relying on them (post_process --bgRemove, most post_process resizes via sharp) failed at runtime with cryptic module-not-found errors. The plugin now bootstraps npm ci --omit=dev in mcp-server/ on first launch when the sentinel module dirs are missing, prints a one-time "First-time setup..." message, and continues into the real server. Subsequent starts are instant.
  • bg-remove error message now mentions the restart requirement. If a user hand-installs deps while the MCP process is already running, the restart hint is surfaced — the running Node process can't pick up newly-installed modules without a fresh start.
added
  • Tool description guidance for bg-remove. post_process --bgRemove description now flags that the underlying @imgly model is photo-trained and works best on photographic subjects (portraits, products), not dense illustration scenes (crowds, where's-waldo-style art) where it tends to isolate the largest object and erase everything else as "background."
internal
  • Server and CLI entry points split into a tiny bootstrap shim (src/server.ts, src/cli.ts) and the heavy main bundle (src/server-main.ts, src/cli-main.ts). The shims use only Node built-ins so they run before any third-party module resolution; they install deps if missing, then dynamic-import the main bundle.
v0.8.2 check_voicebox · engine capability matrix release ↗ 2026-04-27
added
  • check_voicebox MCP tool (CLI: --check-voicebox). Probes the running Voicebox server and reports:
    • Server health (model loaded, GPU type, backend)
    • User profiles (id, name, language, voice_type, default_engine)
    • All 7 engines with verified capabilities — voice cloning, paralinguistic tags ([laugh]/[sigh]/etc.), instruct-field natural-language delivery hints, language count, preset voice counts, tradeoffs (model size, speed, VRAM)
  • Engine capability matrix (providers/voicebox-engines.ts) hand-maintained from Voicebox's docs, timestamped via VOICEBOX_CAPABILITIES_LAST_VERIFIED. Re-verify quarterly or when Voicebox ships new engines.
  • Speech-generation skill updated to instruct Claude to call check_voicebox before picking an engine — so requests like "narrate this with a giggle and a sigh" route to chatterbox_turbo, voice cloning routes to qwen (Qwen3-TTS), and broad-language work routes to chatterbox.
notes
  • Voicebox's API doesn't expose engine capabilities, so the matrix is hand-coded. Preset voice counts ARE fetched live via GET /profiles/presets/{engine}, which validates the engine list is still current at runtime (a missing engine means our matrix is out of date).
  • The capability matrix doubles as a recommendEngine() helper — callers can ask "I need cloning + 23 languages" and get back the best-fit engine ID, falling back to null if no engine matches.
v0.8.1 voicebox maxCharsPerCall lowered · maxCharsPerChunk arg release ↗ 2026-04-27
changed
  • Voicebox default maxCharsPerCall lowered from 5000 to 300. Neural TTS engines (Qwen3-TTS, Chatterbox, Kokoro) produce noticeably better prosody and pacing on short inputs; quality drifts past ~300 chars per call. The chunker is already sentence-aware (with clause/word fallback), so long text still produces a single stitched deliverable — just with more, smaller chunks. Voicebox is $0/call so the extra round-trips have no cost. Cloud providers (Gemini, OpenAI, ElevenLabs) keep their existing larger limits since they handle long inputs cleanly.
added
  • maxCharsPerChunk arg on generate_speech (CLI: --max-chars-per-chunk <n>). Overrides the slot's maxCharsPerCall for a single call. Useful for dialing in any provider that you observe degrading on long inputs. Cache key includes the override so different chunk sizes correctly invalidate (different boundaries → different prosody at the seams).
v0.8.0 visual + voice asset suite — voicebox + bg-remove release ↗ 2026-04-27
added
  • Voicebox provider (voicebox.sh / jamiepine/voicebox). Local-first voice studio with 7 TTS engines (Qwen3-TTS, Chatterbox, Kokoro, LuxTTS, TADA, ...), zero-shot voice cloning, and 23-language support. $0/call, no API key, runs on Mac/Windows/Linux with GPU acceleration. Selected via --provider voicebox --tier small --voice <profile_id>. Profiles are listed at GET /profiles on the running Voicebox server (default port 17493).
  • Env vars: VOICEBOX_BASE_URL (default http://localhost:17493), VOICEBOX_ENABLED (opt-in to failover; off by default), and VOICEBOX_DEFAULT_VOICE (a profile_id used when --voice is omitted).
  • Pricing entry voicebox/voicebox at $0/M chars. The actual engine and model_size live in sidecar params (selected by the profile or via params.engine / params.model_size), so the cost ledger stays clean while reproducibility is preserved.
  • Background remover in post_process via bgRemove: true (CLI: --bg-remove). Uses @imgly/background-removal-node (local ONNX, ~80MB model auto-downloaded on first call). $0/call, offline after first use. When combined with presets, the cutout becomes the source for downstream resizes — --bg-remove --presets og,instagram-square produces transparent-background variants in one pass.
  • Health check now pings Voicebox's /health endpoint when VOICEBOX_ENABLED=true, alongside the other configured providers.
notes
  • This is the visual + voice asset suite release: image generation (gpt-image-2, Imagen 4, Gemini Flash Image), TTS across 4 cloud + 2 local providers (now including Voicebox), and post-processing (resize, webp, bg-remove) all share the same cost ledger, budget guard, sidecar/regenerate, and failover. One install, one budget, one history across modalities.
  • The Voicebox integration uses the server's custom REST API (not OpenAI-compatible) — POST /generate → poll GET /history/{id} → fetch GET /audio/{id}. Generation is async on the server side; the provider polls every 500ms with a 5-minute deadline.
  • @imgly/background-removal-node and onnxruntime-node are marked external in the bundle (esbuild can't bundle the native ONNX binary). Users who want bg-remove run npm install in mcp-server/ once; the structured "install required" error guides them when missing.
v0.7.11 openai gpt-image-2 across all three tiers release ↗ 2026-04-27
added
  • OpenAI gpt-image-2 registered as the OpenAI image model across all three tiers (release announcement, pricing). The openai provider now resolves small | mid | pro to gpt-image-2 with quality: low | medium | high respectively. gpt-image-1 remains callable via explicit --model gpt-image-1 and keeps its pricing entries for backward compat.
  • Per-image pricing entries openai/gpt-image-2:{low,medium,high} at $0.006 / $0.053 / $0.211 standard (50% off via Batch API). Figures are derived from OpenAI's published token rates ($5/M text in, $8/M image in, $30/M image out) — verify against the official calculator before relying on them for tight budgets. Notes on each pricing entry call this out.
  • last_updated field in pricing.json bumped to 2026-04-27.
notes
  • At the small tier, openai/gpt-image-2 ($0.006) is now ~6.5× cheaper than the default google/gemini-2.5-flash-image ($0.039). The default provider remains google (locked decision in CLAUDE.md); switch via --provider openai when minimizing image spend matters.
  • The aspect-ratio bucket map (util/aspect.ts) still routes through the same three OpenAI sizes (1024x1024 | 1024x1536 | 1536x1024) for parity, even though gpt-image-2 accepts more flexible sizes per the model card.
v0.7.10 elevenlabs eleven_v3 registered as pro tier release ↗ 2026-04-21
added
  • ElevenLabs eleven_v3 registered as the pro tier. Previously usable only via explicit --model eleven_v3 and reporting (unknown pricing) / $0 cost. Now provider: elevenlabs, tier: pro resolves to eleven_v3, with accurate per-char pricing so session ledgers and budget caps reflect real spend. Pricing set at $180 / million chars, same as eleven_multilingual_v2 — per the v3 launch blog, post-June-2025 rate is "Same as Multilingual V2." The 80% launch promo ended before this version. If pricing has changed since, run pricing:refresh.
  • Supports emotion and audio tags ([giggle], [sigh], [whisper], multi-speaker) — pass the text with tags inline.
  • last_updated field in pricing.json bumped to 2026-04-21.
notes
  • eleven_turbo_v2_5 (small) and eleven_multilingual_v2 (mid) remain unchanged; v3 slots in above them.
  • Conservative maxCharsPerCall: 5000 retained even though v3 supports up to ~10k chars per call; the plugin's auto-chunk + stitch path handles longer inputs transparently, and the INPUT_TOO_LONG reactive-retry catches provider-side limits.
v0.7.9 env-default voices honored on customVoicesAllowed slots release ↗ 2026-04-21
fixed
  • ELEVENLABS_DEFAULT_VOICE and LOCAL_DEFAULT_VOICE env vars were silently ignored for voice IDs not in the slot's known-voice list. resolveVoice only honored an env default when the name appeared in slot.voices — a deliberate check to prevent a Gemini voice name (e.g. Charon) from leaking into an ElevenLabs call. But for providers where slot.customVoicesAllowed is true (ElevenLabs with raw voice IDs from Voice Lab, local backends where voice names depend on the running server), the env default should be trusted — there's no list to validate against. resolveVoice now accepts env defaults for slots with customVoicesAllowed: true, restoring the ability to set a persistent ELEVENLABS_DEFAULT_VOICE=<id> or LOCAL_DEFAULT_VOICE=<backend-voice> without having to pass --voice on every call.
  • Gemini and OpenAI behavior unchanged — their fixed voice lists are still enforced, cross-provider leak protection preserved.
v0.7.8 manifest reverted to 0.6.x shell-env pattern release ↗ 2026-04-21
changed
  • Reverted plugin.json to 0.6.x-style shell env var pattern. After 0.7.0-0.7.7 (seven hotfix releases, six different approaches), ${user_config.foo} substitution in the mcpServers.env block was conclusively the blocker that stopped Claude Code from spawning the plugin's MCP server (silent failure, no error surfaced). 0.6.0 — still running cleanly in older sessions per ps — uses ${SHELL_VAR:-default} shell-style interpolation with no userConfig. That pattern is now restored.
  • userConfig schema removed from plugin.json. Keeping it while env refers to shell env vars would be misleading: userConfig values would collect in the install prompt but never reach the MCP server. Users configure API keys and preferences via shell env vars (see README.md → Configuration). Keychain-stored secrets and the install-time prompt UX are trade-offs accepted to get a working plugin back.
  • All of today's v0.7.0 code improvements are retained — auto-chunk on length errors (INPUT_TOO_LONG), voiceDefaulted signal, per-provider default voice env vars (GEMINI_DEFAULT_VOICE, OPENAI_DEFAULT_VOICE, ELEVENLABS_DEFAULT_VOICE, LOCAL_DEFAULT_VOICE), debug: true flag for chunk debugging. Only the plugin manifest surface reverted.
migration
  • Users installing for the first time on 0.7.8+ need to set at least GEMINI_API_KEY as a shell env var before starting Claude Code. If you had values stored via the (now-removed) userConfig prompt in 0.7.0-0.7.7, copy them out of ~/.claude/settings.json under pluginConfigs[<plugin-id>].options and from the keychain, and set them as shell env vars.
v0.7.7 mcpServers moved back inline in plugin.json release ↗ 2026-04-21
fixed
  • Plugin MCP server silently failed to spawn with external plugin-mcp.json. The mcpServers pointer in plugin.json ("mcpServers": "./plugin-mcp.json") loaded the config correctly per claude plugin list --json but the MCP server never appeared in /mcp — at session start, Claude Code didn't resolve ${CLAUDE_PLUGIN_ROOT} for externally-referenced mcp configs, leading to a silent spawn failure. Confirmed by v0.6.0 (inline mcpServers, still running in older sessions per ps) working fine with the exact same ${CLAUDE_PLUGIN_ROOT} pattern. Moved mcpServers back inline in plugin.json, matching the pattern chrome-devtools-mcp and other working plugins use. Removed plugin-mcp.json.
v0.7.6 renamed .mcp.json to plugin-mcp.json (avoids project-MCP collision) release ↗ 2026-04-21
fixed
  • MCP server "✗ Failed to connect" caused by filename collision with Claude Code's project-MCP discovery. .mcp.json at the plugin root doubles as a valid Claude Code project-level MCP config. When a user runs Claude Code inside the plugin's source repo (or any repo that has a .mcp.json), Claude Code picks it up as a project MCP config, where ${CLAUDE_PLUGIN_ROOT} doesn't resolve — causing a failed-to-connect spawn. Renamed .mcp.json → plugin-mcp.json and updated the mcpServers pointer in plugin.json accordingly. Non-discovery name avoids the collision.

    To clean up the stale project-scope entry Claude Code may have registered while this was broken, run:
    claude mcp remove claude-image-tts-gen -s project
v0.7.5 explicit mcpServers pointer for 2.1.112 auto-discovery regression release ↗ 2026-04-21
fixed
  • MCP server not auto-discovered from .mcp.json. Under Claude Code 2.1.112, the plugin installed and enabled cleanly but its MCP server never appeared in /mcp — the .mcp.json at the plugin root wasn't being picked up by auto-discovery (unlike housecallpro-mcp and other reference plugins). Added an explicit "mcpServers": "./.mcp.json" pointer in plugin.json to force the plugin loader to find it. Likely a 2.1.112 regression; the explicit path works across versions.
v0.7.4 gemini_api_key marked required in install prompt release ↗ 2026-04-21
fixed
  • Install completed silently with no userConfig prompt. 0.7.3 left every userConfig field optional (no required: true flags), which meant Claude Code skipped the configuration prompt entirely on install — the plugin landed with no API key set, silently unusable. gemini_api_key is now marked required: true so the install flow actually asks for it. Every other field stays optional and can be left blank to accept built-in defaults.
v0.7.3 moved mcpServers to .mcp.json (intermediate attempt) release ↗ 2026-04-21
fixed
  • Plugin still failed to enable after 0.7.2's fallback-syntax fix. The ${user_config.foo:-<default>} shell-style default didn't get evaluated by Claude Code's inline-manifest substitution path — the runtime kept rejecting the plugin with "Missing required user configuration value." Turns out the bare ${user_config.foo} syntax works fine when the mcpServers block lives in a separate .mcp.json file at the plugin root, but not when inlined in plugin.json. The mcpServers block has been moved to .mcp.json (the pattern used by working plugins like housecallpro-mcp), and fallback syntax dropped since it's no longer needed.
v0.7.2 shell-style fallback syntax on user_config refs release ↗ 2026-04-21
fixed
  • Plugin failed to enable when userConfig fields were left empty. 0.7.1's manifest referenced user_config values in mcpServers.env with bare ${user_config.foo} syntax. The Claude Code runtime treats every such reference as a required input and refused to enable the plugin with "Missing required user configuration value: gemini_api_key" any time a field wasn't filled in. Every env reference now uses the shell-style fallback form ${user_config.foo:-<default>}, matching how the original 0.6.x plugin.json read shell env vars:
    • String fields default to empty; config.ts applies its own defaults.
    • local_enabled / autoplay default to false; rewrite_prompts / emit_sidecar default to true; log_level defaults to info — matching the existing config.ts fallbacks.
v0.7.1 userConfig fields gain required type + title keys release ↗ 2026-04-21
fixed
  • Plugin manifest validation failure on install. userConfig fields in 0.7.0 were missing the required type and title keys, causing Claude Code to reject the manifest with "expected one of string|number|boolean|directory|file" on every field. All 18 fields now declare the correct type (string / boolean / directory) and a human-readable title. Boolean fields (local_enabled, autoplay, rewrite_prompts, emit_sidecar) render as toggles in the install prompt; directory fields (image_output_dir, audio_output_dir, state_dir) render as path pickers.
v0.7.0 reactive chunk-on-length-error · per-provider default voices · voiceDefaulted release ↗ 2026-04-21
fixed
  • Long-text TTS no longer fails on provider length errors. Single-call TTS that exceeded a provider's output duration or input token limit previously threw VALIDATION_ERROR, leaving callers to chunk the text externally. That escape hatch was the root cause of several downstream bugs: lost voice parameters across N separate calls, manual ffmpeg concat, cache misses, sidecar fragmentation. mapProviderError now detects length-related rejections and returns a new INPUT_TOO_LONG code; generate_speech catches it and auto-retries on the same provider via the built-in chunker, producing one stitched output file. Chunking triggers both pre-emptively (when text exceeds maxCharsPerCall) and reactively (when the provider rejects a shorter input for output-duration / token-limit reasons).
  • chunkFiles no longer appears in the default response. The per-chunk file paths were only meant for debugging but showed up in every chunked call's JSON, inviting callers to stitch a second time. Gated behind a new debug: true argument; files[0] is always the sole deliverable.
added
  • voiceDefaulted flag on the generate_speech response. true when neither voice nor voicePreset was passed and the slot default was used — surfaces the "you didn't spec a voice" case so callers catch voice mismatches before spending on a long run.
  • Per-provider default voice env vars: GEMINI_DEFAULT_VOICE, OPENAI_DEFAULT_VOICE, ELEVENLABS_DEFAULT_VOICE, LOCAL_DEFAULT_VOICE. Each wins over the slot default when no explicit --voice or preset is passed, but only when the value is valid for the resolved slot's voice list — a Gemini name like Charon will be silently skipped on ElevenLabs instead of producing a cryptic provider 400. Applied at every slot resolution point (initial, per-chunk, per-failover-attempt), so defaults survive provider swaps and chunked retries.
  • Plugin userConfig schema in .claude-plugin/plugin.json. Claude Code now prompts for API keys and preferences at install time; keys flagged sensitive: true are stored in the system keychain. Covers all 18 env vars the MCP server reads. Shell env vars still work for direct CLI invocation.
changed
  • generate_speech tool description updated to tell callers: pass the full text in one call; the tool chunks and stitches automatically; pre-chunking externally loses voice/cache/sidecar fidelity.
removed
  • BUDGET_USD_PER_DAY env var dropped from plugin.json. It was declared there but never read anywhere in the MCP server — budget lives in ~/.claude-image-tts-gen/budget.json and is managed via the set_budget tool.
  • LMSTUDIO_BASE_URL / LMSTUDIO_ENABLED removed from the install prompt surface. config.ts still reads them for backward compat if set via shell env, but new users configure LOCAL_* via userConfig.
v0.6.1 local provider sniffs real audio mime · WAV-at-mp3 root-fix release ↗ 2026-04-19
fixed
  • Local provider lied about output mime. providers/local.ts hard-coded mimeType: "audio/mpeg" regardless of what the server actually returned, even though local backends routinely ignore the response_format: "mp3" hint (Chatterbox-TTS, for one, always returns WAV). The saveAudioRespectingPath helper added in 0.5.2 saw the claimed mpeg mime agreed with the .mp3 output path and skipped the transcode — producing WAV bytes saved at .mp3, same bug shape as 0.5.2 but one layer deeper. The provider now sniffs the first bytes (RIFF / ID3 / MPEG sync / OggS / fLaC) and reports the real mime. The existing save path then transcodes when the caller's extension disagrees.
  • Affects both regular local TTS and voice-cloning calls.
v0.6.0 zero-shot voice cloning · referenceAudioPath + pinToPreferred release ↗ 2026-04-19
added
  • Zero-shot voice cloning via generate_speech --referenceAudioPath <path> (CLI: --reference-audio). Pass a short .wav/.mp3 sample and the local provider forwards it to a cloning-capable backend. Accepted shapes cover Chatterbox-TTS (reference_audio base64 + audio_prompt_path) and Coqui-TTS / XTTS-style servers (speaker_wav path) — backends ignore fields they don't recognize, so whichever key matches wins.
  • The reference file's sha256 fingerprint is mixed into the cache key so identical text + voice with a different reference cache separately.
  • The reference path is recorded in the sidecar input so regenerate and iterate reproduce the cloned voice without re-specifying it.
  • New pinToPreferred option on the failover helper — cloning calls skip the fallback chain so we never silently swap to a provider that would ignore the reference audio.
  • For ElevenLabs cloning, no plugin change is needed: create the voice on elevenlabs.io/voice-lab and pass its voice ID via --voice (raw IDs were already accepted by the ElevenLabs adapter).
rejected
  • referenceAudioPath on providers other than local throws a VALIDATION_ERROR pointing at ElevenLabs's voice lab for managed cloning, instead of silently ignoring the reference.
v0.5.2 absolute chunk paths for ffmpeg concat · transcode on extension mismatch release ↗ 2026-04-19
fixed
  • Chunked TTS concat (generate_speech for long text): chunks were written under a relative ./generated-audio/.chunks/ path while the ffmpeg concat listfile lived in /tmp/. ffmpeg's concat demuxer resolves relative paths against the listfile's directory, so it looked for the chunks under /tmp/ and failed with Error opening input: No such file or directory. Chunk paths and the concat listfile are now absolute.
  • WAV bytes saved at .mp3 path (generate_speech with explicit outputPath): when the provider returned audio/wav but the user asked for foo.mp3, raw WAV bytes were written to foo.mp3. The file on disk now matches its extension — if the extensions differ, the file is transcoded via ffmpeg (libmp3lame for .mp3, pcm_s16le for .wav, etc.), and the response mimeType reflects what actually landed on disk. Missing ffmpeg produces a structured CONFIG_ERROR instead of a misnamed file.
  • Applied uniformly across cached hits, chunked output, explicit-model, and failover paths.
added
  • saveAudioRespectingPath, copyAudioRespectingPath, transcodeAudio, and audioMimeForPath helpers in post/concat.ts. concatAudioFiles now picks codec from the output extension, so a mixed-format concat (e.g. wav chunks → mp3 final) works in a single ffmpeg pass.
v0.5.1 gemini image batch payload shape corrected release ↗ 2026-04-19
fixed
  • Google Gemini image batch (batch_submit for google/image): corrected the outbound payload shape against the @google/genai SDK. submit now passes src: InlinedRequest[] instead of the ignored requests key (which was causing a 400 "Must specify either an input file or a non-empty list of inlined requests" from the Gemini Batch API). poll now reads results from BatchJob.dest.inlinedResponses[] and status from the JOB_STATE_* enum (was the non-existent BATCH_STATE_*). Batches of Gemini Flash Image now actually run and save the advertised 50% vs sync.
v0.5.0 lmstudio → local · default to Kokoro-FastAPI port release ↗ 2026-04-18
changed
  • Renamed lmstudio provider to the more general local (works with any OpenAI-compatible local server: Kokoro-FastAPI, Speaches, Orpheus-FastAPI, Chatterbox-TTS, LM Studio, …). LMSTUDIO_BASE_URL / LMSTUDIO_ENABLED env vars are still honored as deprecated aliases.
  • Default LOCAL_BASE_URL is now http://localhost:8880/v1 (Kokoro-FastAPI's port). Override per-backend.
  • All provider adapters now validate response bodies before surfacing them.
v0.4.0 imagen 4 pro tier · gemini TTS small + pro sync release ↗ 2026-04-18
added
  • Google image/pro tier implemented via Imagen 4 (imagen-4.0-generate-001).
  • Google TTS small and pro tiers implemented sync (gemini-2.5-flash-preview-tts and gemini-2.5-pro-preview-tts).
v0.3.0 regenerate / iterate forward full recipe · dotfile sidecars release ↗ 2026-04-18
fixed
  • regenerate and iterate now forward the full original recipe (model, tier, params, voice, aspect ratio) so re-runs are truly reproducible.
  • Sidecars are written as dotfiles (.regenerate.json) by default; opt-out via env.
v0.2.0 user-selected provider honored · structured tier errors release ↗ 2026-04-18
fixed
  • Tools no longer silently swap providers on failure — the user-selected provider is honored, and tier errors now list concrete alternatives (availableTiers, providersForTier) in the structured error.
v0.1.0 aspectRatio param on generate_image release ↗ 2026-04-18
added
  • aspectRatio parameter on generate_image. For Imagen it routes through imageConfig.aspectRatio; for Gemini Flash Image it's injected into the prompt.
v0.0.1 initial v0+v1 release · multi-provider MCP server release ↗ 2026-04-18
added
  • Initial v0+v1 release: multi-provider MCP server (Google, OpenAI, OpenRouter, ElevenLabs, local) with tier abstraction, cost tracking, budget enforcement, batch submission (Google image + OpenAI image), cache, presets, sidecar-based regenerate, health check, and the plugin bundle (skills, slash commands, hooks).