MCP plugin · Claude Code · v0.13.0
genkettle

The ledger for image, TTS & video in Claude Code.

Eight providers behind one tier knob — images, speech, video, and talking avatars. One ledger. One budget cap. Hard-enforced before the call, sidecar-tracked after, with two independent $0/call local engines so the free tier isn't a single point of failure. Generate without the bill being a surprise.

providers 8
default cap $5 / day
video floor $0.005 / sec
local cost $0.00 / call
license MIT · semver
§ 01 tier matrix · provider × modality × slot single source · registry.ts

One knob — small | mid | pro — that spans every provider.

Code written for Gemini runs unchanged against OpenAI or a local Kokoro model — swap --provider and the call still runs. No per-vendor quirks in your prompt code, no special-casing voice IDs. Empty cells are honest: that slot just isn't shipped.

provider modality small mid pro batch
cloud · image
Google Geminidefault · cheapest image flash-lite$0.034 3.1-flash$0.067 3-pro-image$0.134 all · 50%
OpenAIgpt-image-2 · to 4K image low$0.006 medium$0.053 high$0.211 yes · 50%
OpenRouterpassthrough image passthrough passthrough passthrough —
cloud · tts
Google Geminiflash · studio tts 2.5 flash tts 3.1 flash tts 2.5 pro tts yes · 50%
OpenAItts-1 ··· hd tts tts-1 gpt-4o-mini-tts tts-1-hd —
ElevenLabs+ word timestamps tts turbo multilingual v2 — —
cloud · video & avatars — own ladder, see § 02
Replicate3 models · per sec video draft · low · normal · high · ultra$0.005 — $0.15 / sec —
local · $0 · auto-detected
Local serveropenai-compatible i · t --model <id>$0.0000 --model <id>$0.0000 --model <id>$0.0000 —
Voicebox7 engines · cloning tts kokoro · piper$0.0000 orpheus · qwen3$0.0000 chatterbox · xtts$0.0000 —
pocket-ttskyutai · 26 voices · cloning tts pocket-tts$0.0000 — — —

§ 02 the ladder · video & avatars · per second same five rungs · both tools

Five rungs, 28× apart. Iterate on the cheap one.

Motion prompts and lip-sync are hard to predict, so the bottom rung exists to be thrown away. At half a cent a second you can try twenty framings for the price of one finished take, then move up once the movement is right. Video and talking avatars use the same five words — one vocabulary, learned once.

tier generate_video $/sec generate_avatar $/sec ceiling
cheap rungs · prunaai/p-video · one engine, both tools
draftpreview · throwaway 720pdraft mode $0.005 720pdraft mode $0.005 20s
low 720p $0.02 720p $0.02 20s
normaldefault 1080p $0.04 1080p $0.04 20s
specialist rungs · purpose-built, no length cap
high grok 480pricher motion $0.08 fabric 480pbetter sync $0.08 15s / audio
ultra grok 720p $0.14 fabric 720p $0.15 15s / audio
01

The ladder is quality, not pixels.

Measured on a square input, normal renders 1408×1408 while ultra renders 960×960. The specialist rungs cost more for motion coherence and lip-sync fidelity, not frame size. If you want a bigger picture, normal is both cheaper and larger — the docs say so out loud, because the names imply the opposite.

02

Text-to-video, no input frame.

The cheap rungs generate straight from the prompt — no still to prepare, nothing to pay for twice. The specialist rungs are image-to-video only, so asking for one without a frame gets a refusal that names the tiers that would have worked rather than a stack trace.

03

A 20-second wall that refuses instead of truncating.

Hand p-video 35 seconds of audio and it returns a successful-looking 20-second clip — no error, nothing in the response to catch it by. Verified: 35.44s in, 20.02s out, status: succeeded. So the tool refuses before spending, and prices the uncapped alternatives from your actual audio length. It will not silently double your bill by escalating on your behalf.


§ 03 live ledger · pre-call enforcement 5-beat replay · auto-loops

A $5 daily cap means $5. Enforced before the API call.

Per-call, session, and per-project ledgers, all written through state/store.ts with a lockfile so parallel calls don't race. Hard caps fire pre-call, not after the charge — the budget guard returns a structured BUDGET_EXCEEDED with a suggested fix. Identical repeats hit the cache and return at $0.

estimate_cost ranks every implemented (provider, tier) combo as a dry-run before you spend. CSV / JSON receipts export by month. Pricing lives in a versioned JSON with a 30-day staleness warning surfaced by health_check.

scenario · the demo on the right A $0.05 daily cap. One image (gemini · flash · $0.0039). One long TTS via Voicebox ($0.0000). The ledger reveals the running total. The next image is blocked pre-call — no provider hit, no spend.

§ 04 spec sheet · the cross-cutting work 8 entries

Eight things most provider wrappers leave to you.

Plenty of MCP servers wrap a vendor. genkettle wraps eight and adds the cross-cutting work that, without it, lands on the caller: budgets, ledger, sidecars, failover, chunking, batch, captions, cloning.

01

Cost-aware before the call.

Per-call, session, and per-project ledgers. Hard daily / weekly / monthly caps enforced pre-call, not after the charge. Dry-run estimate_cost ranks every (provider, tier). A $0 cache returns identical repeats for free.

02

Reproducibility in the file system.

Every output writes a hidden .regenerate.json sidecar — prompt, provider, model, tier, params, lineage. regenerate re-runs the brief; iterate adds a tweak and threads parent → child.

03

Long-text TTS, chunked when it has to be.

Sentence-aware splitter, ffmpeg stitch, single deliverable. Triggers pre-emptively (text > slot max) and reactively when a provider rejects shorter input as too long — INPUT_TOO_LONG retries on the same provider so voice + cache stay stable.

04

Provider failover, with a paper trail.

Automatic on rate-limit, 5xx, or timeout. Each retry logs its cost delta. Voice cloning pins to the preferred provider so the reference isn't silently dropped when a non-cloning backend would have been chosen.

05

Free local escape hatch — now with a spare.

Same plugin, same skills, same sidecars, no API key, no network, no bill. Route to Kokoro-FastAPI, Speaches, Orpheus-FastAPI, Chatterbox, Voicebox, or pocket-tts. Two independent $0 engines matters: a single free provider is a single point of failure, and when one stops answering the other still ships. Auto-detected at startup.

06

Batch mode where the vendor supports it.

50% off, ≤ 24h. batch_submit queues, batch_status polls and fires notifications/message on completion. Image batch on Gemini Flash + OpenAI; TTS batch on Gemini.

07

Captions and a gallery of everything you made.

SRT / VTT from ElevenLabs word-level timestamps. Share-target presets (og, twitter, favicon, linkedin, instagram-square) plus a local ONNX background remover. And gallery reads every sidecar on disk into one self-contained page — thumbnail grid, prompt, model, params, and what each one cost.

08

MCP-native UX you'd expect.

Elicitation when ≥2 prompts queue (batch vs. sync). Sampling for the prompt rewriter. Resources for recent outputs in the asset panel. Structured errors with code, message, suggestedFix — never raw provider blobs.


§ 05 local · openai-compatible · $0/call auto-detected · two engines

Run it 100% local. Zero keys.

Point genkettle at any OpenAI-compatible local server and generate images or speech without a key, network round-trip, or dollar spent. Sidecars, cache, regenerate, iterate, post-processing — everything keeps working.

Auto-detected at startup. If LOCAL_BASE_URL/models responds, the local provider joins the failover chain automatically. Set LOCAL_ENABLED=false only to force-exclude a reachable server. Voicebox and pocket-tts resolve the same way via /health.

receipt tts · local · kokoro · 4 chunks · stitched $0.0000
shell · zsh · local-tts.sh
# 1. start a local TTS backend (Kokoro-FastAPI shown)
docker run -p 8880:8880 \
  ghcr.io/remsky/kokoro-fastapi-cpu:latest

# 2. confirm reachable + which models loaded
node mcp-server/dist/cli.js --check-local

# 3. generate against the local server, $0/call
node mcp-server/dist/cli.js --speech \
  -p "hello world" \
  --provider local --model kokoro

# 4. zero-shot voice cloning (Chatterbox / XTTS backend)
export LOCAL_BASE_URL=http://localhost:4123/v1
node mcp-server/dist/cli.js --speech \
  -p "read this in my voice" \
  --provider local --model chatterbox \
  --reference-audio ~/voice-samples/me.wav

§ 06 balance sheet · what's in the box no editorializing

Same modalities. Different bill of materials.

A side-by-side that doesn't editorialize. Either the row ships, or it doesn't.

genkettle raw OpenAI / Gemini ElevenLabs sub guinacio's claude-image-gen
multi-provider · one interface 8 providers — — OpenAI only
pre-call budget enforcement daily / weekly / monthly — plan limits —
per-call + per-session ledger JSON + CSV export — dashboard —
batch mode (50% off) image + TTS manual — —
sidecar / regenerate .regenerate.json — — —
local $0/call escape hatch pocket-tts + Voicebox + 4 more — — —
voice cloning 2 local engines + ElevenLabs IDs n/a subscription —
video · text-to-video · talking avatars $0.005–$0.15 / sec — — —
long-text TTS auto-chunk sentence-aware · reactive manual partial n/a
provider failover logged cost delta — — —
license MIT · semver · changelog vendor terms vendor terms MIT

§ 07 install · two commands · no clone claude code session

Two commands. No clone, no build.

Inside an existing Claude Code session. The plugin registers slash commands and MCP tools automatically on enable.

step 01 · marketplace

Add the marketplace and install.

You'll get the slash commands (/gen-image, /gen-speech, /gen-cost, /gen-budget, /gen-batch-status, /gen-presets, /gen-health) and the MCP tool surface in one go.

claude code prompt
/plugin marketplace add sherifButt/claude-image-tts-gen
/plugin install claude-image-tts-gen@claude-image-tts-gen-marketplace
step 02 · configure

Export at least one key, or run a local server.

Grab a free Gemini key at aistudio.google.com/apikey. Append to ~/.zshrc for persistence — new terminals + sessions inherit it.

~/.zshrc
# required: at least one provider key OR a local server
export GEMINI_API_KEY='your-key'
export OPENAI_API_KEY='your-key'

# optional: per-provider default voices
export GEMINI_DEFAULT_VOICE=Charon
export OPENAI_DEFAULT_VOICE=onyx
export ELEVENLABS_DEFAULT_VOICE='<voice-id>'

# optional: video + talking avatars
export REPLICATE_API_TOKEN='your-token'

# optional: local $0 engines (auto-detected at startup)
export LOCAL_BASE_URL=http://localhost:8880/v1
export POCKET_TTS_BASE_URL=http://localhost:8000
§ 08 end of manifest · ship something MIT · semver · v0.13.0

Stop paying for identical regenerations.

MIT-licensed, semver'd, and shipping. Open an issue, send a PR, or just install it and ship a feature with $5/day worth of cloud spend you can actually account for.