Guardrails for a coding agent: hooks that just say no

A rule you put in a prompt is a polite request the model usually honors. A rule you put in a hook is a wall it cannot walk through. Here is where that difference earns its keep.

Tell an AI coding agent never to force-push to main and it will almost always listen. Almost. The trouble with instructions is that they live in the same soft, probabilistic layer as everything else the model does, so 'almost always' is the best you can hope for. For the handful of actions you truly never want, you need something harder than a request. That is what hooks are for.

How a blocking hook works

A hook is just a shell command the harness runs at a specific moment, for example right before the agent executes a terminal command. The event arrives as JSON on standard input, your script inspects it, and the exit code decides what happens next. Exit clean and the action proceeds. Exit with the blocking code and the action is stopped and your message is handed back to the agent, which reads it and corrects course. No model judgment is involved, and that is the entire point: the guardrail is deterministic.

Block the commands you never want

My first guardrail refuses a small set of genuinely catastrophic commands: a recursive delete aimed at the filesystem or home root, a plain force-push that can clobber shared history, and the classic fork bomb. It is deliberately narrow. Deleting a build folder or a scratch directory sails straight through, and the safer force-push variant that refuses to overwrite other people's work is allowed. The goal is to catch the disasters, not to nanny every command.

Block commits that skip your checks

The second one guards commits and pushes. It blocks the flag that skips pre-commit hooks, because that flag is exactly how unformatted or unlinted code sneaks past the checks you set up on purpose. It also blocks a direct push to a shared branch like main or develop, nudging me onto a feature branch instead. Both are things I know better than to do, and precisely the things a tired human or an eager agent does anyway.

Protect files that should never be hand-edited

The third blocks edits to files that are generated or vendored rather than authored: lockfiles, build output, dependency folders, environment files with secrets, and machine-generated code. The agent should change the source that produces those, not the artifacts themselves. Hand-editing a generated file is the kind of mistake that looks fine until the next regeneration silently wipes it out.

A rule in a prompt is a request the model can rationalize its way around. A rule in a hook is a wall. Save the walls for the things you truly never want.

One principle keeps these from becoming their own source of chaos: keep them pure and self-guarding. No network calls, no external services, and a graceful no-op when a tool they rely on is missing. A guardrail that fails loudly or blocks the wrong thing is worse than none, because you will rip it out within a day. Done right, you forget they exist until the moment one quietly saves you.