Week of August 28 — the bugs we'd been mis-blaming
A shared WSL2 memory pool took the whole platform offline, a two-key-strokes-at-once bug turns out to explain two months of unrelated 'fixed' bugs, and a full-history rewrite scrubs the public repo of everything it never should have carried.
What We Built
The week started with an outage that looked like a container problem and wasn’t. Every IkeOS chat session went dark; a restart of the container didn’t help because the restart never actually happened — dockerd was alive but couldn’t answer any API call, including docker version. The real cause was one level down: the WSL2 utility VM was out of memory, swap fully consumed, load average climbing toward 256, and the kernel’s OOM killer had already picked off a live Claude session before we even started looking. The finding worth keeping: Ubuntu-22.04 and docker-desktop share one WSL2 VM and one memory pool — confirmed by running free -m in both and getting identical numbers. Process-level isolation, which we’d already built and trusted, doesn’t buy resource-level isolation. Set a real ceiling (memory=12GB, swap=8GB) where none had existed before — WSL2 had been quietly capping itself at half the host’s RAM by default — and capped all 25 running containers individually so no single one could take the rest down again. The six stopped containers were deliberately left uncapped rather than guessed at; there’s no usage data to size them from yet.
The same investigation turned up a startup task that had been silently reverting our own fixes on every reboot: the Windows Task Scheduler action still launched the session manager tmux-wrapped and bypassing start.sh, which meant every boot-started manager was quietly missing all seven of its .env keys. It only looked fine because it had never actually been restarted that way. Fixed to call restart.sh properly. In the same pass, a second scheduled task — one that opened with tmux kill-server and had been failing daily with exit 127 since a folder move months ago — got deleted rather than repaired, because its silent failure was the only thing standing between us and a nightly wipe of every live session.
Then a second root cause, unrelated to the first, explained a lot more than expected. A “stuck” publish session turned out to have its whole command sitting typed but unsubmitted — sending Enter by hand fixed it instantly. The bug: our own code was sending a command and its submitting Enter as one burst, which Claude Code’s paste-detection heuristic reads as part of a pasted block rather than a keystroke. Splitting them into two sends with a real gap between them fixed it — and made us go back and look at a string of vault bugs stretching back to June, all previously closed as one-off “rename stopped working” reports, that were almost certainly this same bug wearing different clothes. Same day, a second layer of the same problem showed up under load: five council sessions launching at once could interleave each other’s keystrokes across threads, which a single lock around just the keystroke-sending phase fixed without giving up concurrent launches elsewhere.
What We Considered (and Said No)
IkeOS has been public while being built on a private homelab, and three months of prose, fixtures, and planning docs had quietly picked up the developer’s usernames, absolute host paths, an internal hostname, and LAN addresses along the way — none of it credentials, all of it describing a private environment in a public repo. We rewrote the full history (497 commits, ~63 files) rather than scrub going forward, because going forward would have left the old story sitting there in plain sight. Two things we considered removing and didn’t: this blog’s own project name, which is a live env-var contract between two deployments and isn’t personally identifying, and commit authorship, which is attribution rather than an environment leak and doesn’t hide anything GitHub doesn’t already show. The standing rule now: anything host-, account-, path-, or network-specific lives in .env, never in a tracked file — including the worked examples and test fixtures that are actually where this crept in.
Challenges & How We Solved Them
Writing up a reliability plan for the week’s fixes turned up a checker that was lying to itself. check-drift.sh — the tool that exists to catch our deployed skill files drifting from their source — was flagging all four generated command files as drifted, including two with no real content difference at all. The generator had been changed to insert its header after the frontmatter instead of on line one, so a description wouldn’t get corrupted in /help — but the drift checker was still stripping by line position instead of by the header’s actual content, so it no longer agreed with its own generator about where the generated part started. Not a content bug; a shape mismatch between two halves of one pipeline that had drifted from each other the same way everything else drifts. Fixed by matching on the header text itself, and folded into a nine-task plan spanning both repos — six of which shipped this week, including the Status page now surfacing the actually-deployed commit and a fallback-mount warning, corrected dead placeholder paths, and a reconciled contradiction in our own permission docs around git merge/rebase.
The council review workflow got reworked this week too — from firing a request per click to queuing decisions locally and submitting them as one batch — and the reasoning ties straight back to Monday’s outage: submission is deliberately sequential, not parallel, because firing a whole queue of approvals at once on this VM is exactly the kind of thing that caused the OOM in the first place. Review cards with per-persona stance chips replaced a table that had stopped scaling once each row needed a variable-length list of opinions attached to it.
What’s Next
A quieter fix from the plugin-resolution side of things is worth a mention: the session launcher had a hardcoded path to a specific superpowers version that had since been superseded, and Claude Code accepts a nonexistent --plugin-dir without erroring — so sessions had been launching and looking healthy while silently missing the brainstorming, writing-plans, and subagent-driven-development skills. Now resolved to the highest installed version at import instead of pinned.
Two fresh grill-me items are sitting open going into next week, both aimed at the same question: whether the council is actually opining or just relaying research back at us unfiltered, including a direct comparison against how a similar weekly-review process works elsewhere. That’s a real discussion, not a quick fix, and it’s next. Smaller items in the queue: seven stale test fixtures that need updating to match an intentional status-bar contract change, a missing completion sentinel on blog-publish sessions that leaves them lingering instead of cleaning up promptly, and an idea — filed, not started — to purge the pre-rewrite commits from GitHub now that the history scrub is verified clean locally.