skip to content
Rohan
Rohan sends a voice note from his phone in a direct Telegram chat with Elaine, the chief of staff, while a watchdog keeps that connection alive. Elaine delegates to eight specialists (content, video, SEO, inbox triage, a sys-admin and three engineers) and runs scheduled loops and a shared memory graph.

My Claude Code chief of staff: one VM, eight specialists and one Telegram chat

How I run my work through one Claude Code session on Telegram that delegates to eight specialist sessions, with the watchdog, costs and failures of the first week.

Table of Contents

For the last week, most of my work outside meetings has gone through one conversation. I send a text or a voice note to a Claude Code session on Telegram, and it hands the work to one of eight other Claude sessions, each with its own job. By evening I have an experiment measured, an inbox turned into a task list and a research note, and I have not opened a terminal.

That sounds tidier than it has been. The setup broke in new ways almost every day, and most of what I know about it came from those breaks. This post walks through the agent architecture, what it costs and the failures that shaped it, at the level of detail where you could build your own.

The shape of the system

Everything runs on one Azure VM with a permanently logged-in desktop, because some work needs a real browser. On it run three kinds of process:

  • A chief of staff. One long-running Claude Code session I talk to. I named it Elaine, mostly so voice notes would sound like a person rather than a process. It never does the work itself. It routes each task to a specialist and relays the result.
  • Specialists. Background Claude Code sessions (claude --bg --name <role>), each with a written brief: content, SEO, video, inbox triage, a sysadmin for the VM itself, and engineers for the product repositories. They are separate processes rather than subagents, so they survive the chief of staff restarting.
  • Timer loops. systemd user timers that run headless claude -p jobs on a schedule, with no conversation at all.

Flowchart: Telegram to the chief of staff, which delegates to six specialists; a watchdog restarts it; timers run three loops that report to the same chat; long reads go to Notion; memory is shared by all sessions

Why the chief of staff never does the work

The first version let the chief of staff help with small tasks. Claude Code handles one turn at a time, so every minute it spent grepping a log was a minute my next request sat in a queue.

So the rule became strict. On any incoming task it makes at most one quick read to decide who should do it, sends the brief to that specialist, and ends its turn. If no specialist fits, it launches one. Verification is delegated too: a claim of "done" gets a request for evidence or a second specialist's check.

A side benefit I did not plan for: every specialist's context stays about its own domain. The content session knows the content experiment ledger in detail and nothing about database migrations, which keeps its turns short and its mistakes easy to trace.

Telegram as the only interface

I am rarely at the VM's terminal, so anything that shows up only there never reaches me. Claude Code's channels feature connects a session to a Telegram bot, and that bot became the only way the system talks to me.

  • Voice notes both ways. Incoming voice notes are transcribed with a speech-to-text deployment on Azure, and the chief of staff can reply as a native voice bubble, all for about a dollar a month. The one trap: Telegram sends .oga files, and the endpoint rejects that extension until you rename it to .ogg.
  • Tappable decisions. A decision arrives as a reply keyboard with the options as buttons in Telegram. Inline buttons looked nicer, but their taps never reached the session, because the plugin forwards messages, not callbacks.
  • One chat, long reads elsewhere. Every message names its project, and anything long, such as a draft or a plan, goes to a Notion page with a link in chat. I briefly tried one topic per project but for one person and one assistant it cost more attention than it saved.

The watchdog, and what it taught me about Telegram

Telegram allows exactly one poller per bot token, and the Claude Code plugin enforces that by killing any existing poller when a new one starts. That is reasonable for one session and dangerous for a machine running nine. In the first two days the Telegram bridge went down four ways:

  1. Another session started the plugin, killed the chief of staff's poller and took over, so my messages arrived in a specialist's session.
  2. A health check spawned a second copy. claude mcp list starts each MCP server to check it, which starts a second poller, which kills the live one and then exits. The result was no poller at all.
  3. Deleting old sessions in the agent view killed the live poller within a minute.
  4. Idle retirement stopped the session after about an hour; pinning it exempts it.

The watchdog is a systemd timer that runs a short shell script every two minutes: 30 to 40 milliseconds of CPU and about 5 MB of memory per healthy tick. It checks that the poller is alive, runs from the plugin's directory, and has the chief of staff session as its ancestor. That last check came from failure 1: for three ticks the poller was alive and healthy and belonged to the wrong session. "A poller exists" and "my session owns the poller" are different health signals, and only the second tells you whether messages will reach you.

After two failed checks it restarts the chief of staff, with a lock and cool-downs. The restart has its own trap: claude --bg --resume <session-id> wakes the session in place only with no other flags. Pass any flag, even the original ones, and you get a copy with a new id, no name and no Telegram channel.

Timer loops: move the deterministic work out of the model

Each loop runs in two stages. The first is plain bash with no model: check whether anything is due, gather the inputs, and exit if there is nothing to do. Only then does the runner start claude -p on a cheaper model, with the inputs already in the prompt.

The lesson we kept relearning: anything deterministic belongs in a script, not in the model's instructions.

  • A headless run cannot call sleep, so waiting lives in a helper.
  • A browser helper re-reads the page before any retry, so it never submits twice.
  • A skill file must never point at memory: a headless run from another directory loads none, and the reference silently does nothing.

Memory as markdown files

Every session reads a shared directory of markdown files, one fact per file with a name, a one-line description and a type, linked with [[name]] into a small graph. A one-line-per-memory index loads into every session at start.

Two rules keep it from rotting: specialists may correct facts in their own domain but must say so, and nobody deletes a memory without the chief of staff agreeing, because a stale-looking memory is often the only record of why a rule exists. The weak spot was the decision log, which grew to 188 KB in five days and was re-read at every session start. It is now a short live queue plus an archive.

What it costs

On one day early on, the Anthropic dashboard showed about $195, and I did not know why. The audit found that 88% of spend was the model re-reading its own context: 48% cache reads and 40% cache writes, with output under 12%. The chief of staff was re-reading around 440,000 tokens on every turn, because on a 1M-token window auto-compaction never triggered.

What helped was context and turns, not clever prompts. A standard window cut the chief of staff's cost per turn from about 28 cents to 13. It runs on Claude Opus 5.5, and I capped its context window by hand at 256,000 tokens, so auto-compaction kicks in early, keeping turns cheaper and the thread responsive. The specialists run on Claude Opus 5.5 too, the scheduled jobs on Claude Sonnet 5.5 at a few cents a turn, and nothing uses Fable. And one complete brief per task, because cost follows turns more closely than anything else we measured. It is still not cheap: daily spend ranges from about $30 on a quiet day to over $300 on a heavy build day.

The failures worth copying

Most bugs in this system failed quietly: clean logs, exit code 0, and nothing happened.

  • Stale checkouts. The timers run scripts from each repository's main checkout, while specialists work in separate git worktrees. One day five fixes were pushed and none of them ran for five hours, because nobody had fast-forwarded the main checkout. Every report now states that the main checkout matches the remote.
  • Rules that read correctly. Before one loop went live, simulating every timer fire of its first day against the real data found four bugs, including a lock the loop could never acquire. Each read fine on paper and would have exited 0.
  • Checks that confirmed instead of tested. Grepping a log for the lines you expect confirms your expectation; it does not test the code.
  • A test that was not a test. A dry run once sent me a real notification because only the model call was stubbed. Runners now have a switch that silences every side channel.

The habit behind all of these: we do not trust a rule until it has run against the real state, and we do not trust a report until the result is on disk.

Where it stands

The system does useful work daily, and it still needs a person watching it. What I would keep if I rebuilt it is small: one conversation that only delegates, specialists with written briefs, deterministic code for everything the model does not need to decide, and a watchdog that checks the thing I care about, which is whether my message will reach the session.

If you are building something similar, ask the watchdog's question of every component you add: what signal would tell me this part has stopped working, and would it reach me on my phone?

Stay in the loop

Get practical notes on backend systems, databases, and building with AI in your inbox.

Email subscriptions are handled by Substack. Unsubscribe anytime. Form not loading? Subscribe on Substack.