post-mortem · by Sherif Butt · 2026-04-21 · 12 min read · v0.7.0 → v0.7.8

Seven hotfixes to ship a Claude Code plugin.

v0.7.0 set out to add an install-time prompt for API keys. Eight releases later, we deleted the feature. This is the chain of silent failures that got us there — and why "revert" is a feature, not a defeat.

RELEASE SEQUENCE — V0.7.X v0.7.0 userConfig schema introduced — install-time prompt for keys × FAILED v0.7.1 add type + title to userConfig fields — manifest validation × FAILED v0.7.2 ${user_config.foo:-default} fallback syntax × FAILED v0.7.3 moved mcpServers into separate .mcp.json file × FAILED v0.7.4 marked gemini_api_key required:true — install prompt fires × FAILED v0.7.5 explicit "mcpServers": "./.mcp.json" pointer in plugin.json × FAILED v0.7.6 renamed to plugin-mcp.json — avoid project-MCP collision × FAILED v0.7.7 moved mcpServers back inline in plugin.json × FAILED v0.7.8 reverted to 0.6.x shell-env pattern — kept all v0.7.0 code ✓ SHIPPED 7 attempts → 1 working revert
FIG. 01 · v0.7.x release sequence — bone outlines = failed; amber = shipped. 2026-04-21

The premise was small. When a user installs a Claude Code plugin, the runtime can prompt them for configuration — API keys, defaults, paths — and stash the answers somewhere safe. Sensitive values land in the system keychain. Non-sensitive values live in ~/.claude/settings.json under pluginConfigs[<plugin-id>].options. The plugin's manifest declares the schema with a userConfig block, and references the values inside its mcpServers.env block as ${user_config.foo}.

For genkettle, this would have been a meaningful UX improvement. Before v0.7.0, users had to set GEMINI_API_KEY as a shell env var before launching Claude Code, get the timing wrong, kill the session, and retry. Plenty of plugins ship with the install-prompt pattern; chrome-devtools-mcp does. So we set out to do the same.

What followed was a chain of seven releases over a single afternoon, each fixing the previous one's failure mode and uncovering a deeper one. Then we reverted the whole thing.

Day one: the obvious failures.

The first three failures were ordinary configuration bugs. They were fast to diagnose and didn't shake confidence in the approach.

v0.7.1 manifest validation rejected the new userConfig fields — missing type and title keys.
v0.7.2 plugin refused to enable when fields were left empty — bare ${user_config.foo} made every reference required.
v0.7.3 moved mcpServers into a separate .mcp.json; bare ${user_config.foo} resolved there.
v0.7.4 install completed silently with no prompt — every field was optional, so Claude Code skipped the prompt entirely.

Each one was a real fix. v0.7.4, in particular, taught us a legibility lesson that would echo through the rest of the day: "the install completed and showed no error" is the worst possible failure mode. The plugin landed unconfigured. The user pasted a slash command. Nothing happened. No suggestion of why.

We marked gemini_api_key as required: true and pushed v0.7.4. The install prompt appeared. We tabbed through it, pasted a key, hit enter. /plugin install reported success. We typed /gen-image -p "a teal cube".

Nothing.

The auto-discovery rabbit hole.

claude mcp list showed no MCP server for the plugin. claude plugin list --json showed it loaded fine, configuration values populated. The mcpServers definition was sitting in .mcp.json at the plugin root, exactly where the reference plugins put theirs.

The install completed. The configuration was loaded. The MCP server never existed. There was no error to grep for. — field notes, 2026-04-21

v0.7.5 added an explicit "mcpServers": "./.mcp.json" pointer to plugin.json, hypothesizing that under Claude Code 2.1.112, the auto-discovery for plugin-root .mcp.json had regressed. The pointer made claude plugin list acknowledge the file. The MCP server still didn't spawn.

v0.7.6 was the moment the diagnosis got interesting. .mcp.json at a directory root is also Claude Code's project-MCP discovery filename. When you launch Claude Code inside the plugin's source repo — which we, as the plugin's authors, were doing — Claude Code finds .mcp.json first as a project config, where ${CLAUDE_PLUGIN_ROOT} doesn't resolve. The plugin registers as a project-scoped MCP, fails to spawn, never recovers.

Renaming the file to plugin-mcp.json killed the collision. We pushed v0.7.6. The MCP server still didn't spawn.

v0.7.7 — when the rename also failed

Same configuration, this time read from plugin-mcp.json via the explicit pointer. claude plugin list --json reported the config loaded correctly. claude mcp list: no entry.

v0.6.0 — still running on machines that had installed it weeks earlier — used the exact same ${CLAUDE_PLUGIN_ROOT} pattern with no issues. The difference was where the mcpServers block lived: 0.6.0 had it inline in plugin.json. We had moved it out in 0.7.3 because the bare ${user_config.foo} substitution didn't work inline.

v0.7.7 moved mcpServers back inline. The MCP server still didn't spawn. By this point the suspicion was crystallizing: it wasn't the file location. It was the substitution.

The silent spawn-time failure.

${user_config.foo} references in mcpServers.env resolved fine for the claude plugin list machinery — that machinery reads the manifest in-process and substitutes from the in-memory config. But at MCP-server spawn time, a different code path interpolates the env block before launching the child process. That second path, under the version of Claude Code we were targeting, didn't resolve ${user_config.foo} references. It got a literal string with the brace syntax in it. The spawn aborted. No error surfaced.

The spawn path did resolve ${SHELL_VAR:-default} shell-style references — which is exactly what 0.6.x had used. So 0.6.x ran cleanly; 0.7.x didn't; the only difference at the manifest layer was the substitution pattern.

The revert.

We had two options. Option one: keep the userConfig schema, ship a workaround that wrote shell-env exports into a file the plugin could source at startup, document the keychain limitation, and hope a Claude Code release fixed the spawn-path substitution. Option two: delete the feature.

v0.7.8 deleted the feature.

The userConfig schema was removed from plugin.json. The mcpServers.env block went back to the proven 0.6.x pattern: ${SHELL_VAR:-default} shell-style interpolation, no install-time prompt, configuration via ~/.zshrc exports.

Every code change from v0.7.0 was kept. The reactive INPUT_TOO_LONG chunking, the voiceDefaulted signal, the per-provider default-voice env vars, the debug: true chunk-debugging flag — all shipped, all in production. Only the manifest surface reverted. Users who had populated values via the doomed userConfig prompt were given a migration note pointing at ~/.claude/settings.json and the keychain.

Lessons.

Silent failures are worse than loud ones

The hardest bug to fix is the one that produces zero output. claude mcp list being empty looks identical whether the plugin doesn't exist, the plugin failed to register, or the plugin registered but failed to spawn. We lost three releases bisecting between those states because we couldn't tell them apart from the command line. Diagnostic output — even noisy diagnostic output — is a feature.

Runtime substitution paths are sharp edges

Anywhere a runtime reads a config and substitutes values into it, ask: does this same substitution path get re-invoked on every relevant code path? In our case, claude plugin list and claude mcp spawn were two different code paths reading the same manifest with different substitution capabilities. The surface contract — "the manifest is the manifest" — implied a uniformity that didn't exist.

Revert is a feature

It's tempting, after seven hotfixes, to keep grinding. Each one is cheap. Each one feels close. None of them work. The discipline of saying this isn't going to land in this iteration; revert and keep what works took longer than it should have, but it's what shipped a working plugin in v0.7.8. The userConfig schema can come back when the upstream substitution-path bug is fixed. Until then, shell env vars work, and working beats elegant.

Working in your own repo is a hazard

We were testing a plugin from inside the plugin's own source directory. Claude Code's project-level .mcp.json auto-discovery treated our plugin manifest as a project MCP, and the collision masked the real bug. Plugin authors should test from a neutral directory. The default development environment — your own repo — has the most configuration overlap with the runtime you're trying to integrate with. That's where false positives breed.


The full timeline lives in the changelog, one entry per release, with the diagnostic threading through. v0.7.0 introduces the feature. v0.7.1 through v0.7.7 try to make it work. v0.7.8 reverts and keeps the engineering. v0.8.0 ships Voicebox + the visual asset suite on the working manifest. The install prompt is on the deferred list and will return.

Until then: export GEMINI_API_KEY=..., append to ~/.zshrc, ship features.

← newer post The 3,700 bytes that broke our CI for two weeks. all posts back to the blog index →