Week of September 12 — sentinels, stuck menus, and a mount that lies
A week of finding the small lies our own infrastructure was telling us — sessions that looked finished, dashboards that looked healthy, and a mount that looked full.
What We Built
The most satisfying fix this week was also the pettiest bug: council-action sessions never closed themselves. The session manager was watching /tmp/ikeos-done/council-action-<slug> for a per-session sentinel file, while the skill itself was writing to the fixed path /tmp/ikeos-done/council-action — no slug, no match, ever. Thirteen finished sessions were sitting idle when we went looking, each one with its own unconsumed sentinel sitting right there on disk, just not the name anyone was checking for. launch_session now exports the correct path as IKEOS_DONE_SENTINEL, and the four affected skills write to whatever they’re handed instead of guessing their own name. blog-publish and blog-rewrite had it worse — they were watching paths nothing writes to at all.
We also went after sessions that freeze without telling anyone. Four of twenty-one live sessions were stuck on interactive TUI menus while the dashboard reported them as healthy — no chip, no warning, just a pane quietly waiting for a keystroke that was never going to arrive on its own. Part of the problem was a footer string: our menu-detection only matched “Enter to confirm / Esc to cancel,” which missed the exact dialog our own remote-control feature opens (“Enter to select — Esc to continue”). The other part was a race — polling once on a fixed delay after sending a command missed the confirmation dialog whenever the registration call behind it ran slow. Now that we can see a session is stuck, we can also do something about it: a new /sessions/<id>/keys pass-through and a Keys row in the session panel let arrows, Enter, Esc, Tab, and Space reach a blocked pane straight from the browser — something that used to require attaching to tmux by hand.
Shipping that Keys button surfaced a second, unrelated bug: the heading rendered from new HTML, but the cached agents.js had no code behind it, because static assets carry a deliberate one-hour cache and nothing was telling the browser a new build had shipped. We now stamp the build’s git SHA into every static URL, so a deploy is a new URL rather than a stale cache hit disguised as a working feature.
On the claude-config side, sync had been quietly deploying to the wrong place. apply’s target was hardcoded to ~/.claude, but CLAUDE_CONFIG_DIR points at the repo itself — so anything it wrote landed somewhere nothing loads. Thirty-four files, twenty-eight commands and six rules, had been inert, silently compensated for by a command-mirror workaround that existed specifically because of this. We fixed resolution to read ~/.claude/settings.json instead of trusting the environment, since a plain login shell never has CLAUDE_CONFIG_DIR set in the first place. In the same vein, four personal skills — close-session, promote, schema-check, triage — turned out to have only ever existed as slash commands, never as real Skills, because a bare SKILL.md isn’t a plugin without a manifest beside it.
Challenges & How We Solved Them
The theme this week was infrastructure lying to us in ways that looked like everything was fine. The best example: on 2026-09-05, a rebuild picked up VAULT_PATH from the surrounding shell instead of the Windows path in .env — because docker-compose gives the environment precedence over the file — and mounted an empty directory at /vault. The only symptom was an empty project picker. VAULT_PATH wasn’t even on the list of variables we checked for fallback mounts, because the old check only asked “is this variable set,” not “does what it points to actually contain anything.” We fixed that: each mount now declares whether emptiness is itself evidence of a problem, since /weekly-reviews being empty before the first review is legitimately fine and /vault being empty never is.
That fix earned its keep almost immediately, in a way we didn’t expect. As of this run, the /vault bind mount is empty inside the container again — not a misdirected environment variable this time, a bind mount that shows empty in-container even after a container restart. Different failure, same symptom, and it’s still open. We’re noting it rather than chasing it further today: the detector did its job and told us something’s wrong, which is a real improvement over an empty project picker and no explanation. Diagnosing why a restart doesn’t clear it is next.
The Skill Stack
close-session, promote, schema-check, triage — these four went from slash-command-only to real Claude Code Skills. The manifest each needs (.claude-plugin/plugin.json) is now generated alongside its SKILL.md, with name and description read from the skill’s own frontmatter so the two can’t drift apart. They were also being deployed to ~/.claude/skills/, the wrong directory for the same reason apply was — now fixed together.
council-action, council-discuss (and the blog commands) — updated to signal completion through IKEOS_DONE_SENTINEL instead of a fixed per-type path, with the literal fallback preserved for anything invoked outside a managed session, where the variable is unset.
What’s Next
- Purge the pre-rewrite ikeos history on GitHub. The repo stays private until GitHub Support garbage-collects the commits a force-push made unreachable but not deleted — every old SHA was broadcast publicly for the seven weeks the repo was public, so re-publishing before the purge lands would hand the pre-scrub history right back out.
- Reconcile
housekeeping.mdand the council skills into the adapter layer, once the council pipeline itself has been stable for a while — it’s real, actively-developed functionality right now, and porting it mid-flux would mean reconciling twice. - Evaluate Cactus Compute’s Needle as a document-processing service for upcoming agent work — early-stage curiosity, not yet a commitment.