In my last post I said that running Claude Code headless on a timer is the killer feature. InboxProcessor is the clearest example of that, but it only got a couple of lines, so this post is the walkthrough: what it does, how it works, what goes wrong, and where I’ve drawn the line on what it’s allowed to do.
What it is
The idea is simple. I write a task as a note and drop it into a Tasks/Inbox folder in my Obsidian vault. Every 15 minutes the Pi checks that folder, picks up anything waiting for it, and hands each task to Claude Code to carry out. When it’s finished, the note is updated with what was done and filed away.
Because the vault syncs everywhere, I can create a task from my phone, my PC, or (most often) by asking Claude in a chat to “save this as a task”, which writes the note into the vault through Google Drive. It was built in August and has worked through a few dozen tasks since.
There’s also a Windows version. The Pi has no screen or browser, so anything that needs a real, logged-in browser (looking something up on LinkedIn or Facebook, for example) is marked for the Windows PC instead. That version runs while I’m logged in, checks the same Inbox folder every 30 seconds, and uses Claude in Chrome. The two never fight over a task because each only picks up the ones addressed to it.
What a task looks like
A task is just a Markdown note, created from a template so it always has the same shape:
- The filename is
YYYY-MM-DD-HHmm-short-slug.md, for example2026-09-30-0838-fix-regression-suite-failures.md. It sorts by date and never clashes with another task. - The frontmatter says when the task was created, where it came from (chat, me, n8n), who should do it (the Pi, the Windows PC, n8n, or either), a free-form type, a priority, and a status.
- The body has to be self-contained. It’s written assuming whoever picks it up knows nothing about the conversation that created it: exact notes to change, exact text, and what “done” looks like.
The status moves from pending to in-progress to done, or to blocked if it gets stuck. Finished tasks move to a Tasks/Archive folder rather than being deleted, so the archive is a running log of everything that’s been done and how.
How a run works
- A systemd timer starts a short Python script every 15 minutes. A lock stops a slow run overlapping with the next one.
- The script looks for notes with status
pendingthat are addressed to the Pi, oldest first, up to five per run. - For each one it first sets the status to
in-progress, which claims the task. - It then runs Claude Code non-interactively, with a short prompt that points at the task file and restates the rules: do exactly what the task asks, don’t guess, and when finished either mark it done and move it to the archive, or mark it blocked and explain why.
- Claude runs from my home directory, so it picks up the same ground rules as any other session: write a changelog note in the vault for any system change, keep the READMEs current, and update and run the regression tests.
- When Claude exits, the script doesn’t just trust the exit code. It re-reads the task file and goes by what the frontmatter now says. Every outcome goes into a history file that feeds a “recent executions” panel on my dashboard.
What goes wrong, and how it’s handled
Quite a lot can go wrong with something that runs unattended, so most of the code is about noticing when it does:
- Ambiguous or impossible tasks get marked blocked with a section saying exactly why, and stay in the Inbox for me to look at. I get an email alert.
- Crashes and timeouts. Each task gets 20 minutes. If Claude fails, times out, or exits cleanly without updating the task, the script marks it blocked itself and alerts me.
- Orphaned tasks. If a task has been
in-progressfor more than 45 minutes (three runs), the previous run must have died, perhaps because the Pi restarted. It gets marked blocked rather than sitting there forever. - Tasks that are silently ignored. This has been the most common real problem. When Claude in a chat writes a task with slightly the wrong frontmatter (
Claude Codeinstead ofclaude-code, an invented target name, or a status written as a list because it was edited in Obsidian’s properties panel), the script simply didn’t see it. Nothing failed, so nothing alerted. Each of these was fixed as it was found, with a changelog note explaining what happened. - The Windows version stepping on the Pi. When the Windows version arrived, the Pi’s orphan check would have wrongly “rescued” Windows tasks left in progress when the PC was switched off. A review of the rollout caught it, and the Pi now only judges its own tasks.
InboxProcessor has its own regression test, like everything else. It checks the timer is running, the status page responds, and the task-finding logic picks up a test task correctly. It deliberately doesn’t run Claude, so the test suite doesn’t spend money or act on real tasks every time it runs. And every change to InboxProcessor itself has a changelog note in the vault under Systems/, the same as every other project.
So far it has handled 35 tasks: 27 completed, 7 blocked and 1 failure.
Where the line is
This is the part I thought about most. Before it was built I had a choice of three options: a strict list of the commands Claude may run, limiting it to editing files in the vault, or full autonomy. I chose full autonomy. A strict list would block the infrastructure tasks I actually want it to do, and a vault-only limit would defeat the point.
That’s a real trade-off, so it’s backed up by the safety nets above: the time limit, the claim-and-orphan check, the lock, alerts on anything other than a clean finish, and an instruction in every prompt to block rather than guess. A blocked task costs me nothing more than a note to read, so I’d much rather it stopped than guessed.
Some things are still off limits whatever the task says. As I mentioned last time, my financial data is kept away from the AI except where I’ve made a deliberate exception, and a task in the Inbox doesn’t change that.
A couple of real examples
- Fixing a failing test. One morning the regression tests started emailing me every 15 minutes about Watchtower. From a chat I wrote a task asking Claude to investigate and fix it. It found that Watchtower’s nightly update of n8n had tried to send a notification to n8n while n8n was still restarting, a one-off race. It tightened the test to ignore exactly that case and nothing else, checked it against five made-up log samples, re-ran the full suite (all passing), confirmed the alerts had stopped, and wrote the cause and fix into the task note with a link to the changelog entry. I read about it afterwards rather than doing any of it myself.
- Updating the garden notes. When a bed in the garden was planted, a task asked for the bed note to be updated with what actually went in (including a substitution from the nursery) and a photo added. Claude made the text changes, but the photo was only in the chat that wrote the task, not in the vault, so it marked the task blocked and said so. Once I’d shared the photo it finished the job, then created notes for each new plant and a diary entry, keeping the garden folders consistent.
The second one is the pattern I like most: it did what it could, stopped at the bit it couldn’t do, and told me exactly why.
A note on how this post was written: fittingly, it was a task in the Inbox. Claude Code picked it up, went through the InboxProcessor code and README, its git history, its notes in my vault and the archive of past tasks, so that everything here could be checked against what was actually built. It wrote the draft directly into WordPress through the WordPress API, and I then reviewed and edited it before publishing.










