Unattended Claude Code agents for your ticket backlog.

  1. a ticket
  2. its own sandbox
  3. a second agent reviews
  4. your lint, tests, build
  5. your local branch

Queue tickets, start a run and read the report in the morning. Nothing is pushed until you push it.

$ git clone https://github.com/henkisdabro/sandcastle-kit.git ~/sandcastle-kit && cd ~/sandcastle-kit && pnpm install && ./bin/sandcastle setup

DRY_RUN=1 sandcastle run the first time: it implements, reviews and gates your tickets, and merges nothing.

macOS or Linux, Docker (or OrbStack, or Podman), Node 22+, and a Claude subscription or API key.

A run, as the status view shows it

An example night, played back as sandcastle status shows it: thirteen tickets, nine merged, two left for you, two waiting on those.

  • Ticket 41, rate limiting: implemented and reviewed, failed its tests, repaired, reviewed again, merged.
  • Tickets 42, 43, 45, 48, 49, 52 and 53: implemented, reviewed, gates green, each merged as it went green while the others worked.
  • Ticket 44, a date library upgrade: still failing three tests after two repairs; left for you.
  • Ticket 46: gates green, but it changes CI config, so it is held for you to merge by hand.
  • Ticket 47: waited for ticket 41; once that landed, it started in a free sandbox in the same run, and was implemented, reviewed and merged.
  • Tickets 50 and 51 wait for 44 and 46 to close.
  • By morning: nine branches merged into the local main branch, none of it pushed.

A real run: the kit, on its own backlog

The kit is built by its own runs. This is one of them, on 4 October 2026, from the run record. Every ticket is a public issue you can open.

tickets
18
merged
18
left open for a person
2, a criterion unmet
wall clock
1 h 40 min
tokens
28.4M in / 343k out
models
Sonnet 5.5 implements, Opus 5.5 reviews
sandboxes
5 at once
  1. #252Size: a read-only sandcastle size command that recommends the pool's limits25m1.9M / 25k
  2. #249The gates phase time includes the wait for a machine-wide gates slot16m1.4M / 14k
  3. #244A green resolution held by the stray-path check reports 0 commits and no gates22m1.9M / 14k
  4. #258backup-size test fails intermittently: 2 packs after a drop, not 131m1.0M / 20k
  5. #251An ended run's status view shows the next run's settings beside the report's recorded ones18m1.1M / 12k
  6. #250One ticket can spend an hour and 19M tokens with no ceiling short of idle15m878k / 10k
  7. #248The pool tests take 10-66 seconds each and slow the test gate 2-8x under load19m2.2M / 38k
  8. #243The <ungated>, <unmet> and <changelog> readers match a tag named in prose10m623k / 12k
  9. #242Tickets that edit protected paths always end held; say so before queueing and requeueing35m2.5M / 24k
  10. #241Report and status treat a hand-merged held branch as one the agent handed back22m3.3M / 30k
  11. #240Idle mark: the skill trigger reads the count before the skill's own labelling27m1.7M / 17k
  12. #239Doctor: the kit's own checkout is never a project, so its project checks are skipped27m660k / 7k
  13. #237pnpmStore mounts the versioned store dir, so sandboxes nest a second store inside the host's9m396k / 7k
  14. #254Size: recommend the pool's limits from measured sandbox and gate peaks19m2.5M / 41k
  15. #253Size: setup and doctor point to sandcastle size while the pool limits are untouched defaults22m1.3M / 14k
  16. #247The run estimate is about half the real cost for a run of carried or conflicting branches29m2.6M / 30k
  17. #246A requeued or resolver-held green branch is re-implemented and re-reviewed in full7m1.3M / 19k
  18. #245Needs you says a stray-held branch changes test files, hiding the real reason10m994k / 9k

bar = minutes in worktokens in / outwent back once, after a red gate or a conflict, and merged on its second try

A test that did not test: "a run below its share does hold back a younger wait" filled the pool. The younger run waited only because no slot was free, so the test also passed with the wait-order check removed.

The review agent on #248, in that run. It rewrote the test so that removing the check fails it. The gates were green either way.

Over the 26 runs on this repository so far, a ticket ended merged 110 times, in a conflict 8 times and held for a person 6 times. It is a TypeScript project whose tests take a few minutes; a slower suite or a bigger model costs more.

You already run Claude Code all day. It still needs you in the loop.

+ it fitsGitHub Issues or ticket filessmall, closed ticketslint · types · tests · buildyou would rather read a report

- not yettickets that say "make it faster"nothing tests anythingyou want to pair livetickets only in Linear or Jira

Why not a loop of claude -p, subagents or worktrees?

You can, and then you build the rest yourself.

Wired up by handsandcastle-kit
Where the agent runsyour machine, with your credentialsa Docker sandbox per ticket, holding a token that cannot push
Who checks the workthe agent that wrote ita second, stronger agent, against the ticket
Who runs the teststhe agent, when it remembersthe kit, on every branch and again on the merge
Mergingyou, one branch at a timeeach branch as it goes green; a conflict gets a second try
Changes to CI, hooks, install scriptsmerged like the restheld for you
The morning afterscrollback in six terminalsone report: what merged, what is red and why, what needs you

One night, one run

  1. sandcastle runYour gates run on main first. Red there, and nothing is spent.
  2. six sandboxesOne ticket each, on its own branch. A green one lands while the rest work.
  3. a red gate, repaired#41 fails its tests, is repaired, reviewed again and lands. #47 waited for it, and starts.
  4. nine merged#44 and #46 are left for you. Nothing is pushed.
  5. you read the reportPush once you have read it. sandcastle requeue 44 sends a ticket back.

With autonomy: "drain", the run keeps taking turns by itself until the queue is drained, and says why it stopped.

How it works

  1. 1/5

    Install

    sandcastle setup
  2. 2/5

    Point at a project

    sandcastle init
  3. 3/5

    Queue tickets

    ready-for-agent
  4. 4/5

    Run, and walk away

    sandcastle run
    1. implement
    2. review
    3. your gates
    4. land

    one sandbox per ticket, in parallel

  5. 5/5

    You push

    sandcastle report

First time? DRY_RUN=1 sandcastle run implements, reviews and gates, and merges nothing.

Each step in full
  1. Install once

    ./bin/sandcastle setup puts sandcastle on your PATH, installs the agent skill for Claude Code, Codex and OpenCode, walks you through your tokens and ends with a health check.

  2. Point it at a project

    sandcastle init reads your stack and fills in the gate commands for you to check against CI. sandcastle build makes the project's image.

  3. Queue tickets

    Label a GitHub issue ready-for-agent, or give a ticket file that status. Ask your agent for /sandcastle queue to tighten the vague ones first, or /sandcastle audit to find work in a repo with no tickets yet.

  4. Run

    sandcastle run, or DRY_RUN=1 sandcastle run the first time: implement, review and gate, but never merge or close. Watch it with sandcastle status, or from Herdr's sidebar.

  5. Land and push

    sandcastle report prints the summary again. land merges a branch you fixed, requeue sends a ticket back. Then you push.

Or ask your agent

Before the first run

  • initsets up your gates and a lean sandbox
  • auditfinds work and files it as tickets
  • queueasks one question per vague ticket

Every run after

  • runasks, then starts the run and writes its summary
  • statusreads a live run, or the last one
  • updatepulls the latest kit and says what it changes

/sandcastle Claude Code$sandcastle CodexOpenCode too

What each action does

Before the first run

  • init reads your CI, sets up the gate commands and rules for agents, agrees with you which skills and MCP servers agents see, and checks that every hook runs in the sandbox. It ends with your gates green on main.
  • audit has read-only agents check the repo for bugs, security issues, missing tests and stale docs, and files what they can prove as tickets once you have approved the list.
  • queue walks every open ticket with you. A ticket is queued only when an agent with no chat context could finish it; for the rest you get one question each, with a recommended answer.

Every run after

  • run shows the queue and models, asks before starting, and leaves the run working in the background. When it ends, you get the summary: what merged, what needs you, what to push.
  • status reads a live run, or the last one, and the log behind any red row.
  • update pulls the latest kit, rebuilds the project's image and walks you through what the new version changes for it.

Built for long runs

Fails early

Checked before a token is spent.

  • your gates, green on main
  • every model answers a one-line preflight
  • your hooks run in the sandbox

agents start

Spends less

A sandbox loads only what you keep.

Shares the machine

Live runs split the sandbox slots.

six sandbox slots, two projects running

A run outlives the session that started it: sandcastle wait · stop

The detail

Fails early

  • Your gates on main first. If one fails there, the run stops before any agent starts.
  • Preflight. One short reply from every model, so an exhausted plan or an old CLI stops the run at the start, not halfway through.
  • Hooks proven. The kit checks that your Claude Code hooks run inside the sandbox, and a test can prove a blocking hook still blocks.

Spends less

  • Lean sandboxes. Sandbox agents load only the skills, subagents, commands and MCP servers you keep, and never your plugins: each one costs context on every turn.
  • Picks up where it stopped. A ticket sent back builds on its old branch, and a branch that was already green costs only the gates.

Shares the machine

  • Several projects at once. Runs in different projects split the machine's sandbox slots between them rather than starving each other.
  • Detached. A run started by your agent outlives the session that started it; sandcastle wait and sandcastle stop find it again.

Watch it where you already work

sandcastle status works in any terminal. Two optional extras put the run in front of you.

In Claude Code

A Claude Code session with a band above the prompt: the castle, then sandcastle, my-app, implementing, and counts of tickets working, ready and merged.

An optional mod (a Claude Code plugin, in early access) shows the run above the prompt of the session that started it, and names the tickets that need you.

The detail

Above the prompt

  • A band while it runs. The run's stage and tokens, and how many tickets are working, need you, are ready to land, are blocked or have merged.
  • A line when you are needed. A conflict, a red gate, a held branch or a crash pins a line naming the tickets, until they are dealt with or the run ends.

When it ends

  • A closing prompt. When the run's process is gone, the session where you used /sandcastle gets a prompt and writes the closing summary, even if you quit Claude Code and resumed that session in between.
  • Ticket status on demand. /sandcastle-status lists every ticket and where it is, without a model turn.

Between runs

  • Ready tickets. Between runs, a quiet row of sand above the prompt shows how many tickets a run would start now: ♜ sandcastle · 4 ready - /sandcastle run. The count is read in the background, at most once every ten minutes per project.
  • Hide the reminder. /sandcastle-mark dismiss hides the count until another ticket becomes ready; hide turns the mark off for the project.

In Herdr

♜ 4/9 · 1 needs you

Herdr is a terminal multiplexer for coding agents. There the run is one line in the sidebar, red when it needs you, and a key opens the status view over any tab.

The detail

Over any tab

  • The status view on a key. Herdr's prefix key, then Shift+S, opens it over whatever you are doing; q puts you back. Prefix, then Shift+E, shows the last run's report.
  • A ticket a click away. Ctrl-click a ticket in the status view to open its card: its state, each pass with its outcome and time, and why it is held or red, with a key for each pass's log.

In the sidebar

  • The run at a glance. The run is one entry in the sidebar, and its workspace reads ♜ 4/9 · 1 needs you, red when something needs you. The tab bar lists every live run on the machine.
  • Each sandbox, if you want them. With sandbox panes on (herdr: { panes: "all" } in the project config, or SANDBOX_PANES=all for one run), each one is named after its ticket with its step and time in it: review · 12m.

Out of your way

  • One tab per run. A run started by your agent runs detached, in a tab that holds only its status view, and outlives the session that started it.
  • Set up when you say. sandcastle herdr configure shows what it adds to Herdr's config and asks first; --remove takes it out again.

Safe to leave alone

Agents run with permission prompts off, inside Docker. The container is the boundary, and only a green, reviewed branch gets through the wall.

Docker sandbox

An agent, prompts off

  • a token that cannot push
  • a ticket, which grants no permissions
  1. Reviewa second agent, against the ticket
  2. Your gateslint, tests and build, run by the kit itself
  3. Risky? It waits for youCI, hooks, install scripts, protected paths

Your machine

Merged into your local branch

  • .git is fingerprinted
  • host hooks stay off

GitHub

Nothing pushed

  • only labels, comments and closed issues
The detail

Into the sandbox

  • A token that cannot push. Only a fine-grained GitHub token for issues goes in, and sandcastle doctor --verify checks it cannot push. An empty value is refused, so a host API key cannot leak in.
  • Tickets grant no permissions. A ticket that asks an agent to touch .git/, credentials or a push is not the work, and the guards below hold if an agent tries anyway.

Before a merge

  • Review, then your gates. A second agent reviews every branch against its ticket and fixes what it finds. Then the orchestrator, never the agent, runs your lint, tests and build. Only a green branch merges, and the merged branch is gated again.
  • Risky changes wait. A green branch that changes hooks, CI, agent settings, package-manager config, install scripts or a path you protect is labelled ready-for-human and left for you. So is one that adds a file over 50 MB, a repair whose review failed, and a conflict resolution that drops lines git had merged.

On your machine

  • Nothing is pushed. Green branches merge into your local base branch. The kit never pushes or deploys; on GitHub it only labels, comments on and closes issues.
  • .git is fingerprinted. If a sandbox changes .git/config, HEAD, info/ or hooks, or moves your base branch, the run stops before your machine runs another git command there.
  • Host hooks stay off. During a run the kit's own git calls on your machine ignore hooks, so a branch that adds one cannot run it as it lands.

Install

  1. Install the kit, once per machine:

    $ git clone https://github.com/henkisdabro/sandcastle-kit.git ~/sandcastle-kit && cd ~/sandcastle-kit && pnpm install && ./bin/sandcastle setup
  2. Point it at a project:

    $ cd ~/code/your-project && sandcastle init && sandcastle build
  3. Mark a couple of tickets ready-for-agent and try a dry run. It implements, reviews and gates them, but never merges or closes a ticket:

    $ DRY_RUN=1 sandcastle run

Needs macOS or Linux, Node 22+, pnpm, git, jq, the GitHub CLI signed in (skippable with ticket files) and a container runtime. Full requirements.

Questions

What does it cost to run?

The kit is free and MIT-licensed. Agents use your Claude subscription or API key through Claude Code. On the kit's own repository, 18 tickets took 1 h 40 min and 28.4M input and 343k output tokens, with Sonnet implementing and Opus reviewing; one ticket ran from 0.4M to 3.3M input tokens. Every run on your machine shares one plan allowance; the first ticket that hits the usage limit stops that run's queue, and USAGE_CHECK=1 stops it starting tickets before then.

Which models does it use?

By default Claude Sonnet 5.5 implements and Claude Opus 5.5 reviews. A model: or effort: label sets the implementer's model or effort for one ticket, and CROSS_REVIEW=1 adds a third review by an OpenAI model through Codex.

Does it push to GitHub?

No. Green branches merge into your local base branch and the issue is closed with a comment saying so. You push when you are happy with what landed.

What about my own work in that repository?

A run refuses to start on a dirty tree or off the base branch, and it merges into that checkout while it runs. Leave the checkout alone until the run ends, and do your own work in another clone.

What if a branch is green but wrong?

Then it is in your local branch and nowhere else. A second agent reads every branch against its ticket before your gates run, and the report lists what merged, so you read before you push. Behaviour that no test covers has only that review between it and your branch.

What happens on a merge conflict?

A branch that conflicts as it lands goes back for one more try in the same run. A second conflict is left for you, and so is a resolution that drops lines git had merged.

Can the text of a ticket make an agent do something else?

It can try. The sandbox holds a token that only reaches issues, the kit runs your gates itself, and a branch that changes CI, hooks, install scripts or a path you protect waits for you whatever the ticket said.

Can a run go all night?

Yes. A run keeps the machine awake until it ends (caffeinate on macOS, systemd-inhibit on Linux). Closing a MacBook's lid still sleeps it unless it is on power with an external display. A ticket whose blocker lands starts in the same run, and with autonomy: "drain" one run keeps going until the queue is drained or something stops it, and says what.

Do I need Matt Pocock's skills?

No, but they help. His /to-tickets and /triage are a good way to write tickets an unattended agent can finish, and the kit reads the tracker and labels his setup skill records.

Linear, Jira, Windows?

Linear issues can block a ticket today, but the queues are GitHub Issues and Markdown ticket files. The kit runs on macOS and Linux.

Built on Sandcastle

Sandcastle by Matt Pocock does the sandboxing, worktrees and agent orchestration underneath. This kit is one opinionated way to put it to work. If it helps you, go and star the original.