This tutorial builds a Claude Code Stop hook that runs Hypothesis property tests, so Claude Code has to verify its own work before it can finish. Here is what Claude gets back when it tries to end a turn while your chunker still loses text:
decision: blockProperty tests failed. Fix the code in chunker.py; do not edit the tests.test_no_text_is_lost AssertionError: assert '' == '0' - 0 Failing test case: test_no_text_is_lost( text='0', budget=(2, 0), )Nobody wrote that input. Hypothesis generated a few hundred texts and token budgets, found one that breaks the chunker, and shrank it to the smallest case that still fails: a one-word document with a two-token window comes back as zero chunks. In a RAG pipeline, that is the short document that never reaches your index. A Stop hook hands the shrunk case to Claude as the reason it may not stop yet.
By the end you will have a Claude Code project with three parts. A Python text chunker is guarded by four Hypothesis properties and a differential test. A Stop hook blocks the agent from finishing while any of them fails. A PreToolUse hook and two deny rules stop the agent from editing the properties to make them pass. The level is intermediate: you know Python and pytest, you have used Claude Code, and you have not used Hypothesis or written a hook before. Plan on about 45 minutes.
Verified against Python 3.13.9, hypothesis==6.168.4, pytest==9.1.1 and Claude Code 2.1.289 on 2026-10-05. The automated check did not run the live Claude Code session in Step 10, because it needs a signed-in account; that session was run once by hand, and Step 10 quotes its transcript.
The idea comes from a post by Addy Osmani, Give your agent a way to check its own work. He names four kinds of test worth investing in: end-to-end flows, property-based tests, old-versus-new comparison, and tests that are fast and deterministic. This tutorial builds on his summary line: "Example tests say what should happen; properties say what must not." The post stops at what to invest in. Putting those tests in the agent's path, and keeping the agent from weakening them, is the part you build here.
Prerequisites
You need:
- Python 3.13 (3.10 or later works; the versions below were verified on 3.13.9)
- git 2.28 or later (for
git init -b), because the Stop hook usesgit statusto detect edits to the tests - Claude Code 2.1.289, installed with the native installer. Only Step 10 needs it signed in, with a Claude subscription or an Anthropic API key; Claude Code reported a cost of USD 0.20 for the recorded run of that step. Every other step runs without an account and costs nothing
- A bash shell: Terminal on macOS or Linux, or Git Bash on Windows. Every command below is written for bash
Create the project and a virtual environment. On macOS, use python3 for this one command if python is not found; inside the activated environment, python always works:
mkdir chunker-verifier && cd chunker-verifierpython -m venv .venvActivate it. The path differs by platform, so run only the line for yours. On macOS and Linux:
source .venv/bin/activateOn Windows, in Git Bash:
source .venv/Scripts/activateThen install the two pinned packages. The && runs the install only if python now points into .venv, so a missed activation cannot install into your system Python. If this command prints nothing at all, the activation did not take; run the activation line again:
python -c "import sys; sys.exit(sys.prefix == sys.base_prefix)" \ && python -m pip install hypothesis==6.168.4 pytest==9.1.1Keep this virtual environment active for the whole tutorial, including when you start Claude Code in Step 10. The hooks call python, and only the active environment lets python find Hypothesis.
Check the install:
python -c "import hypothesis, pytest; print(hypothesis.__version__, pytest.__version__)"claude --version6.168.4 9.1.12.1.289 (Claude Code)Claude Code updates itself, so a newer version than 2.1.289 is fine; the hook fields this tutorial uses are documented for current releases. If your shell exports a CI environment variable, unset it for this tutorial (unset CI). Hypothesis loads its built-in ci settings profile whenever CI is set, and your results will not match the ones on this page.
How a Claude Code Stop hook verifier loop works
Claude edits chunker.py. When it tries to end its turn, Claude Code runs a Stop hook, which checks two things in order: did anyone change the files in tests/properties/, and do all the properties still pass? If both answers are good, the hook exits silently and Claude stops. If either fails, the hook replies decision: block with a short reason, and Claude keeps working with that reason in front of it.
A second guard sits in front of every edit. A PreToolUse hook and two permission deny rules (entries in .claude/settings.json that forbid a tool on a path, added in Step 9) refuse writes to tests/properties/ and .claude/, which keeps the spec and the hooks out of the agent's reach. A hook is enforced by Claude Code itself, not by the model, which is the difference Skills vs Hooks in Claude Code is about. An exit code from a test run is also the kind of check that the oracle problem says an agent loop needs.
direction: down
classes: {
agent: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
guard: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}}
logic: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
ok: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
bad: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
}
claude: "Claude Code session\nedits chunker.py" {class: agent}
guard: "before every edit:\nPreToolUse hook + deny rules\nrefuse tests/properties/, .claude/" {class: guard}
stop: "Stop hook: verify_properties.py" {
label.near: top-center
style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
grid-rows: 1
grid-gap: 16
tamper: "check 1: git status\ntests/properties/" {class: logic}
props: "check 2: pytest + Hypothesis\nprofile agent" {class: logic}
}
pass: "exit 0, no output\nClaude may stop" {class: ok}
fail: "decision: block\nreason = smallest failing case" {class: bad}
claude -> guard: "1. Edit / Write"
claude -> stop: "2. tries to end the turn"
stop -> pass: "3a. both checks pass"
stop -> fail: "3b. spec changed or a property fails"
fail -> claude: "4. keeps working"
In the diagram, every attempt to end a turn goes through the Stop hook, and the only path to "Claude may stop" needs both a clean spec and passing properties. The steps below build it one box at a time, starting with the code under test.
Step 1: Write a token-window text chunker in Python
Goal: create chunker.py, a function that splits text into overlapping chunks under a token budget.
Why this step: the gate needs code to guard, and a RAG chunker is a good choice because its bugs are silent. A chunker that drops text still returns a list of chunks and nothing crashes. The missing text is simply never embedded, and you find out weeks later when a search misses a document. This version has a real bug of exactly that kind. Do not look for it yet; the properties in Step 3 will find it for you.
The chunker counts tokens with a simple rule: a token is a run of non-space characters plus the whitespace after it. A real tokenizer such as tiktoken would work the same way in principle, but it needs a network download on first use, and the hook you build later should run offline in about a second. Each chunk keeps its start and end character offsets, because a RAG system needs them to cite where an answer came from. For why chunk boundaries decide retrieval quality in the first place, see Why Your RAG Chunks Are Lying to Your Retriever.
The diagram below shows how 7 tokens become 2 chunks with max_tokens=4 and overlap=1. A dot marks a trailing space.
grid-rows: 3
grid-columns: 8
grid-gap: 6
classes: {
head: {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; bold: true}}
tok: {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}
in: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
shared: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}}
empty: {style: {fill: "#FFFFFF"; stroke: "#E5E8E8"; font-color: "#2C2C2A"}}
}
h: "token" {class: head}
t0: "Agents·" {class: tok}
t1: "stop·" {class: tok}
t2: "when·" {class: tok}
t3: "the·" {class: tok}
t4: "work·" {class: tok}
t5: "looks·" {class: tok}
t6: "done." {class: tok}
c0: "chunk 0\nchars 0-21" {class: head}
a0: "in" {class: in}
a1: "in" {class: in}
a2: "in" {class: in}
a3: "in" {class: in}
a4: " " {class: empty}
a5: " " {class: empty}
a6: " " {class: empty}
c1: "chunk 1\nchars 17-37" {class: head}
b0: " " {class: empty}
b1: " " {class: empty}
b2: " " {class: empty}
b3: "overlap" {class: shared}
b4: "in" {class: in}
b5: "in" {class: in}
b6: "in" {class: in}
Chunk 1 starts 3 tokens after chunk 0 (the stride is max_tokens - overlap), so the token the· lands in both. Overlap buys you exactly this: a sentence cut at a chunk boundary still appears whole in one of the two chunks.
Create chunker.py in the project root:
# chunker.py"""Split text into overlapping chunks under a token budget, for RAG indexing."""import refrom dataclasses import dataclass# A token is a run of non-space characters plus the whitespace after it.# Leading whitespace at the very start of the text is a token of its own.TOKEN = re.compile(r"\S+\s*|\s+")@dataclass(frozen=True)class Chunk: start: int # character offset into the original text end: int text: strdef count_tokens(text: str) -> int: return len(TOKEN.findall(text))def chunk(text: str, max_tokens: int, overlap: int) -> list[Chunk]: if max_tokens < 1: raise ValueError("max_tokens must be at least 1") if not 0 <= overlap < max_tokens: raise ValueError("overlap must be in [0, max_tokens)") spans = [m.span() for m in TOKEN.finditer(text)] stride = max_tokens - overlap chunks = [] for i in range(0, len(spans) - max_tokens + 1, stride): window = spans[i : i + max_tokens] start, end = window[0][0], window[-1][1] chunks.append(Chunk(start, end, text[start:end])) return chunksRun it on the sentence from the diagram:
python -c "from chunker import chunkfor c in chunk('Agents stop when the work looks done.', max_tokens=4, overlap=1): print(c)"Expected output:
Chunk(start=0, end=21, text='Agents stop when the ')Chunk(start=17, end=37, text='the work looks done.')What just happened: two chunks, the right offsets, one shared token. The chunker looks right on the one input you tried, and that is all the evidence an agent has when it declares a task done.
Step 2: Write an example-based pytest test
Goal: pin the behaviour you just saw with an ordinary pytest test.
Why this step: most projects already have tests like this, and it is what an agent writes when you ask it to "add tests": one input you chose, one answer you expected. It is the baseline that Step 3 improves on. Writing it first also proves that pytest can import chunker.py from the project root.
Create the test folders. tests/properties/ stays empty until Step 3:
mkdir -p tests/propertiesCreate tests/test_examples.py:
# tests/test_examples.pyfrom chunker import chunkdef test_two_chunks_share_one_token(): chunks = chunk("Agents stop when the work looks done.", max_tokens=4, overlap=1) assert [c.text for c in chunks] == [ "Agents stop when the ", "the work looks done.", ]Run it from the project root. Use python -m pytest rather than plain pytest for the whole tutorial, because python -m puts the current directory on the import path, and that is how the tests find chunker.py.
python -m pytest -q tests/test_examples.pyExpected output (your timing will differ):
. [100%]1 passed in 1.13sWhat just happened: this test will keep passing for the rest of the tutorial. It also says nothing about the bug in chunker.py, because its input happens to divide evenly into windows.
Step 3: Write a Hypothesis property test that finds lost text
Goal: state one thing the chunker must never do, and let Hypothesis search for inputs that make it happen.
Why this step: an example test checks one input against one answer. A property test checks a rule against hundreds of generated inputs. For a chunker, the most important rule is that no text is lost: if you glue the chunks back together, skipping the part each chunk repeats from the one before, you must get the original text back. The rule holds for every possible input, so you never have to work out an expected answer by hand.
The diagram puts the two kinds of test side by side.
grid-rows: 1
grid-gap: 40
classes: {
you: {style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}}
gen: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
rule: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
ok: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
bad: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
}
example: "Example test (Step 2)" {
label.near: top-center
style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
in: "one input you chose\n7 tokens, max 4, overlap 1" {class: you}
out: "one answer you wrote\n2 chunks" {class: you}
verdict: "passes" {class: ok}
in -> out
out -> verdict
}
property: "Property test (Step 3)" {
label.near: top-center
style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
gen: "Hypothesis generates\ntexts up to 300 chars\nand (max_tokens, overlap)" {class: gen}
rule: "rule: reassembled\nchunks == text" {class: rule}
shrink: "fails, shrunk to\ntext='0', budget=(2, 0)" {class: bad}
gen -> rule: "hundreds of cases"
rule -> shrink: "first break, then\nsmallest break"
}
On the left you supply both the question and the answer, so the test can only find bugs you already imagined. On the right you supply only the rule. Hypothesis supplies the questions, including ones you would not think to ask, such as a document shorter than one window.
Create tests/properties/test_chunker_properties.py:
# tests/properties/test_chunker_properties.pyfrom hypothesis import given, strategies as stfrom chunker import chunk@st.compositedef budgets(draw): max_tokens = draw(st.integers(min_value=1, max_value=20)) overlap = draw(st.integers(min_value=0, max_value=max_tokens - 1)) return max_tokens, overlapdef reassemble(chunks): """Glue chunks back together, skipping the part each one repeats.""" out, pos = [], 0 for c in chunks: assert c.start <= pos, f"gap: text[{pos}:{c.start}] is in no chunk" out.append(c.text[pos - c.start :]) pos = c.end return "".join(out)@given(text=st.text(max_size=300), budget=budgets())def test_no_text_is_lost(text, budget): max_tokens, overlap = budget assert reassemble(chunk(text, max_tokens, overlap)) == textThree Hypothesis pieces do the work here:
st.text(max_size=300)generates strings of up to 300 characters from the whole Unicode range, including empty strings, tabs, carriage returns and characters from other scripts.budgets()is a composite strategy: a function decorated with@st.compositethat draws values in order. It drawsmax_tokensfirst and then anoverlapbelow it, so every generated pair is valid. Filtering invalid pairs out after the fact withassume()would also work, but it wastes generated cases.@given(...)turns the test into a property: Hypothesis calls it many times with generated arguments, 100 by default.
Run it:
python -m pytest -q tests/propertiesExpected output (on macOS and Linux the paths print with /; your timing will differ). A default Hypothesis run is random, so your shrunk case may differ slightly, for example text='a' or a different budget. Any text with fewer tokens than max_tokens is the same bug, and Step 5 makes the result repeatable:
F [100%]================================== FAILURES ===================================____________________________ test_no_text_is_lost _____________________________ @given(text=st.text(max_size=300), budget=budgets())> def test_no_text_is_lost(text, budget): ^^^tests\properties\test_chunker_properties.py:25: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _text = '0', budget = (2, 0) @given(text=st.text(max_size=300), budget=budgets()) def test_no_text_is_lost(text, budget): max_tokens, overlap = budget> assert reassemble(chunk(text, max_tokens, overlap)) == textE AssertionError: assert '' == '0'E E - 0E Failing test case: test_no_text_is_lost(E text='0',E budget=(2, 0),E )tests\properties\test_chunker_properties.py:27: AssertionError=========================== short test summary info ===========================FAILED tests/properties/test_chunker_properties.py::test_no_text_is_lost - As...1 failed in 2.53sWhat just happened: Hypothesis found an input that breaks the rule, then shrank it. Shrinking means Hypothesis kept simplifying the failing input, with shorter text, smaller numbers and plainer characters, for as long as the test still failed. It reports the smallest failure it could reach: the text '0' with max_tokens=2 and overlap=0 produces no chunks at all, so reassembly returns ''.
With that input in hand, the bug is easy to read. The loop runs range(0, len(spans) - max_tokens + 1, stride). With 1 token and a window of 2, that is range(0, 0, 2), which is empty. Any document with fewer tokens than max_tokens disappears, and longer documents lose their last partial window. Leave the bug in place. In Step 10 Claude will fix it, and the hook will decide whether the fix counts.
The output says Failing test case: because Hypothesis 6.159.0 and later print that label. Older versions, and most blog posts written before mid-2026, show Falsifying example: instead.
Step 4: Add properties for the token budget, verbatim slices, and overlap
Goal: add the three other rules a chunker must never break.
Why this step: a chunker that returned the whole document as one chunk would pass "no text is lost". That rule is necessary, but the full spec also has to say what every chunk must look like:
- Verbatim slice: each chunk's
textequalstext[start:end]of the original. A RAG citation that points at offsets the chunk text does not match is worse than no citation. - Budget: no chunk holds more than
max_tokenstokens, or it will not fit the embedding model's input. - Overlap bound: two neighbouring chunks never share more than
overlaptokens, or you pay to embed the same text twice.
Replace the whole file tests/properties/test_chunker_properties.py with this version. The import line changes too, because the new properties use count_tokens:
# tests/properties/test_chunker_properties.pyfrom hypothesis import given, strategies as stfrom chunker import chunk, count_tokens@st.compositedef budgets(draw): max_tokens = draw(st.integers(min_value=1, max_value=20)) overlap = draw(st.integers(min_value=0, max_value=max_tokens - 1)) return max_tokens, overlapdef reassemble(chunks): """Glue chunks back together, skipping the part each one repeats.""" out, pos = [], 0 for c in chunks: assert c.start <= pos, f"gap: text[{pos}:{c.start}] is in no chunk" out.append(c.text[pos - c.start :]) pos = c.end return "".join(out)@given(text=st.text(max_size=300), budget=budgets())def test_no_text_is_lost(text, budget): max_tokens, overlap = budget assert reassemble(chunk(text, max_tokens, overlap)) == text@given(text=st.text(max_size=300), budget=budgets())def test_every_chunk_is_a_verbatim_slice(text, budget): max_tokens, overlap = budget for c in chunk(text, max_tokens, overlap): assert c.text == text[c.start : c.end]@given(text=st.text(max_size=300), budget=budgets())def test_no_chunk_exceeds_the_budget(text, budget): max_tokens, overlap = budget for c in chunk(text, max_tokens, overlap): assert count_tokens(c.text) <= max_tokens@given(text=st.text(max_size=300), budget=budgets())def test_overlap_never_exceeds_the_setting(text, budget): max_tokens, overlap = budget chunks = chunk(text, max_tokens, overlap) for prev, nxt in zip(chunks, chunks[1:]): shared = text[nxt.start : prev.end] if nxt.start < prev.end else "" assert count_tokens(shared) <= overlapRun it with --tb=no, which prints the summary without the tracebacks:
python -m pytest -q tests/properties --tb=noExpected output:
F... [100%]=========================== short test summary info ===========================FAILED tests/properties/test_chunker_properties.py::test_no_text_is_lost - As...1 failed, 3 passed in 0.97sWhat just happened: the three new properties pass on the buggy chunker. Each property guards a different way to be wrong, and a given bug usually breaks only one of them. Had you written only the budget and overlap rules, this chunker would look correct. The same idea drives contract tests for a LangGraph agent: test the property, not the setting.
Step 5: Make Hypothesis deterministic with a settings profile
Goal: register a Hypothesis settings profile, agent, that gives the same verdict every time it runs on the same code.
Why this step: by default Hypothesis is random. Each run draws new inputs, keeps failing inputs in a local database at .hypothesis/examples, and fails any single test case that takes longer than 200 ms. Those defaults suit a human at a terminal. In an agent loop, a rare bug can fail one run and pass the next, and a slow machine can fail on timing alone. An agent that sees a check fail and then pass with no code change learns that re-running is a fix. Addy Osmani's post makes the same point: slow, flaky loops teach the agent to retry, not to fix.
The profile changes four settings:
grid-rows: 5
grid-columns: 3
grid-gap: 6
classes: {
head: {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; bold: true}}
name: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
def: {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}
agent: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
}
h1: "setting" {class: head}
h2: "default" {class: head}
h3: "profile agent" {class: head}
n1: "derandomize" {class: name}
d1: "False: new random inputs\nevery run" {class: def}
a1: "True: inputs derived from\nthe test itself, same every run" {class: agent}
n2: "database" {class: name}
d2: ".hypothesis/examples:\nreplays earlier failures" {class: def}
a2: "None: no memory\nbetween runs" {class: agent}
n3: "deadline" {class: name}
d3: "200 ms per test case" {class: def}
a3: "None: never fail\non timing" {class: agent}
n4: "max_examples" {class: name}
d4: "100" {class: def}
a4: "200" {class: agent}
The first three green rows each remove one way a verdict can change without a code change. The fourth, max_examples=200, spends some of the time saved on a wider search, and the suite still finishes in about a second.
Create tests/conftest.py. pytest loads it automatically for every test under tests/:
# tests/conftest.pyfrom hypothesis import settings# The profile the Stop hook runs. Same inputs every run, no timing failures,# and no memory of earlier runs, so a verdict depends only on the code.settings.register_profile( "agent", derandomize=True, database=None, deadline=None, max_examples=200,)Run it twice with the profile selected and compare the failing case. The --hypothesis-profile option comes from the pytest plugin that Hypothesis installs, so it works with no extra setup:
for run in 1 2; do python -m pytest -q tests/properties --hypothesis-profile=agent \ | grep -A3 "Failing test case"doneExpected output:
E Failing test case: test_no_text_is_lost(E text='0',E budget=(2, 0),E )E Failing test case: test_no_text_is_lost(E text='0',E budget=(2, 0),E )What just happened: both runs fail on the same case. With derandomize=True, Hypothesis derives its random seed from the test function, so the same test and the same code always explore the same inputs. A plain python -m pytest still uses the defaults; only the hook selects agent. One limit: "same inputs every run" holds only until you upgrade Hypothesis or Python, or edit the test.
Step 6: Commit the properties as the spec
Goal: put the project under git, with the properties in the first commit.
Why this step: to answer "did anyone change the spec?", the Stop hook in Step 7 compares the working tree against the last commit. Committing the properties now, before any agent touches the code, marks them as yours.
printf '.venv/\n__pycache__/\n.hypothesis/\n.pytest_cache/\n' > .gitignoregit init -q -b maingit add .git commit -q -m "chunker and its properties"git log --format=%sgit ls-filesExpected output (on Windows, this and every later git add or git commit may also print "LF will be replaced by CRLF" warnings; they are harmless and are left out of the expected output below and in later steps):
chunker and its properties.gitignorechunker.pytests/conftest.pytests/properties/test_chunker_properties.pytests/test_examples.pyWhat just happened: the first line is the commit message, which proves the commit exists; git ls-files alone would list staged files even if the commit had failed. From now on, git status --porcelain -- tests/properties prints nothing unless something under tests/properties/ changes. If git commit printed Please tell me who you are and git log then printed an error, git had no identity to commit with. Set one with git config user.name "Your Name" and git config user.email "you@example.com", then run the git commit line again.
Step 7: Write a Claude Code Stop hook that runs the Hypothesis properties
Goal: write .claude/hooks/verify_properties.py, the script Claude Code runs every time Claude tries to end its turn.
Why this step: without it, Claude stops when the work looks done, and "looks done" is all it has to go on. Anthropic's own Claude Code best practices describe a Stop hook as the deterministic version of "check your work": a script that blocks the turn from ending until a check passes.
A Stop hook receives a JSON event on stdin. If it prints nothing and exits 0, Claude stops. If it prints {"decision": "block", "reason": "..."}, Claude keeps working and reads the reason. The script asks two questions, in order:
grid-rows: 3
grid-columns: 2
grid-gap: 48
classes: {
q: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
bad: {style: {fill: "#E74C3C"; stroke: "#B03A2E"; font-color: "#FFFFFF"}}
ok: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
}
q1: "1. Did tests/properties/ change\nsince the last commit?" {class: q}
a1: "Yes: block, ask for\ngit checkout -- tests/properties" {class: bad}
q2: "2. Do all properties pass\nunder profile agent?" {class: q}
a2: "No: block with the failing\ntest and its smallest input" {class: bad}
q3: "Both pass: exit 0,\nprint nothing" {class: ok}
a3: "Claude may stop" {class: ok}
q1 -> q2: "no"
q2 -> q3: "yes"
q1 -> a1: "yes"
q2 -> a2: "no"
q3 -> a3
The spec check runs first because a passing suite proves nothing if the agent edited the suite.
Create the hook. mkdir -p makes the .claude/hooks/ folder:
mkdir -p .claude/hooks# .claude/hooks/verify_properties.py"""Stop hook: Claude may not finish while a property fails."""import jsonimport osimport subprocessimport sysevent = json.load(sys.stdin)root = os.environ.get("CLAUDE_PROJECT_DIR", event["cwd"])def block(reason: str) -> None: print(json.dumps({"decision": "block", "reason": reason})) sys.exit(0)# 1. The properties are the contract. If they changed, nothing else counts.changed = subprocess.run( ["git", "status", "--porcelain", "--", "tests/properties"], cwd=root, capture_output=True, text=True,).stdout.strip()if changed: block( "tests/properties/ changed. Those files are the spec, not yours to edit.\n" "Restore them with `git checkout -- tests/properties` and fix the code.\n" + changed )# 2. Run the properties with the deterministic profile.try: result = subprocess.run( [sys.executable, "-m", "pytest", "tests/properties", "-q", "-p", "no:cacheprovider", "--hypothesis-profile=agent", "--tb=short"], cwd=root, capture_output=True, text=True, timeout=120, )except subprocess.TimeoutExpired as exc: block(f"Property tests ran for {exc.timeout:.0f} s without finishing. " "chunker.py probably loops forever on some input.")if result.returncode == 0: sys.exit(0)# 3. Send back only what Claude needs: which property, and the smallest input.keep = []for line in result.stdout.splitlines(): if line.startswith("____"): # one header line per failing test keep.append(line.strip("_ ")) elif line.startswith("E ") and line[1:].strip(): keep.append(line[1:].rstrip()) # the assertion and the test caseblock( "Property tests failed. Fix the code in chunker.py; do not edit the tests.\n" + "\n".join(keep[:40]))A few lines are worth a second look:
CLAUDE_PROJECT_DIRis set by Claude Code to the project root. When you run the script by hand it is not set, so the script falls back to thecwdfield of the event.sys.executableruns pytest with the same Python that runs the hook, which is the virtual environment's Python.-p no:cacheproviderstops pytest from writing.pytest_cache, so the check leaves no files behind.- The
tryaround the pytest run exists because a hook that crashes does not block anything. Claude Code reads exit code 0 as success and exit code 2 as a block, and treats any other exit code as a non-blocking error and lets Claude stop. An edit that makeschunk()loop forever would hang pytest,subprocess.runwould raiseTimeoutExpired, and without thetrythe gate would open on the worst bug of all. With it, a hang becomes a block with a reason. - Part 3 of the script trims the pytest output to the failing test names and the
Elines. Claude gets the assertion and the shrunk case, not 30 lines of traceback.
Hook output is JSON, which is hard to read once long reasons are in it. Create a small helper, peek.py, in the project root. It prints a hook's reply the way Claude reads it:
# peek.py"""Print a hook's JSON reply the way Claude reads it: decision first, then reason."""import jsonimport sysreply = json.load(sys.stdin)reply = reply.get("hookSpecificOutput", reply)for key in ("decision", "permissionDecision"): if key in reply: print(f"{key}: {reply[key]}")for key in ("reason", "permissionDecisionReason"): if key in reply: print(reply[key])Run it by piping in a minimal Stop event, the same way Claude Code will:
echo '{"cwd": ".", "hook_event_name": "Stop"}' \ | python .claude/hooks/verify_properties.py | python peek.pyExpected output:
decision: blockProperty tests failed. Fix the code in chunker.py; do not edit the tests.test_no_text_is_lost AssertionError: assert '' == '0' - 0 Failing test case: test_no_text_is_lost( text='0', budget=(2, 0), )Now check the spec guard. Append a line to the property file, run the hook, and restore the file:
echo "# relax this later" >> tests/properties/test_chunker_properties.pyecho '{"cwd": ".", "hook_event_name": "Stop"}' \ | python .claude/hooks/verify_properties.py | python peek.pygit checkout -- tests/propertiesExpected output:
decision: blocktests/properties/ changed. Those files are the spec, not yours to edit.Restore them with `git checkout -- tests/properties` and fix the code.M tests/properties/test_chunker_properties.pyWhat just happened: the hook blocked both times, and each reason tells Claude what to do next. The first hands over the shrunk case from Step 3. The second fires on any change to the spec, even a comment, and never gets as far as running the tests. The final git checkout put the property file back.
Step 8: Stop Claude Code editing your tests with a PreToolUse hook
Goal: write .claude/hooks/protect_properties.py, which refuses any Edit or Write to tests/properties/ or .claude/ before it happens.
Why this step: the Stop hook catches a changed spec only at the end of a turn, after the agent has already spent effort on the wrong fix. Models do take that route: the AWS Kiro team wrote when Kiro became generally available in November 2025 that models "often 'game' the solution by modifying tests instead of fixing code". A PreToolUse hook stops the edit before it happens and tells Claude why. Claude Code runs PreToolUse hooks before every tool call, and a deny from one holds even in bypassPermissions mode, the permission mode that skips every approval prompt.
No single layer covers every way to write a file, so this tutorial uses three. The grid shows which layer stops which kind of write:
grid-rows: 5
grid-columns: 4
grid-gap: 6
classes: {
head: {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; bold: true}}
path: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
yes: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
no: {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}
}
h0: "how the agent writes" {class: head}
h1: "deny rule\n(Step 9)" {class: head}
h2: "PreToolUse hook\n(this step)" {class: head}
h3: "Stop hook git check\n(Step 7)" {class: head}
p1: "Edit tool" {class: path}
p1a: "refused" {class: yes}
p1b: "refused, with reason" {class: yes}
p1c: "caught at end of turn" {class: yes}
p2: "Write tool" {class: path}
p2a: "refused" {class: yes}
p2b: "refused, with reason" {class: yes}
p2c: "caught at end of turn" {class: yes}
p3: "Bash: sed -i, echo >" {class: path}
p3a: "refused" {class: yes}
p3b: "not seen\n(matcher is Edit|Write)" {class: no}
p3c: "caught at end of turn" {class: yes}
p4: "Bash: python -c\n\"open(...).write(...)\"" {class: path}
p4a: "not seen" {class: no}
p4b: "not seen" {class: no}
p4c: "caught at end of turn" {class: yes}
Only the right-hand column is green all the way down. The permissions documentation says plainly that deny rules do not apply to a script that opens files itself, which is why the git status check in the Stop hook stays even after you add the two guards in front of it.
Create the hook:
# .claude/hooks/protect_properties.py"""PreToolUse hook: Claude may not edit the spec or the hooks that enforce it."""import jsonimport osimport sysfrom pathlib import Pathevent = json.load(sys.stdin)root = Path(os.environ.get("CLAUDE_PROJECT_DIR", event["cwd"])).resolve()target = Path(event["tool_input"].get("file_path", "")).resolve()protected = [root / "tests" / "properties", root / ".claude"]if any(target.is_relative_to(p) for p in protected): print(json.dumps({ "hookSpecificOutput": { "hookEventName": "PreToolUse", "permissionDecision": "deny", "permissionDecisionReason": ( f"{target.relative_to(root).as_posix()} is part of the spec.\n" "Change chunker.py so the properties pass instead." ), } }))sys.exit(0)The hook compares resolved Path objects instead of strings. Claude Code always sends an absolute file_path, and on Windows it arrives with backslashes, so a check like "tests/properties/" in path would silently never match there.
Run it with one protected path and one allowed path. Real events carry absolute paths; a relative path resolves against the current directory, which is the project root here:
echo '{"cwd": ".", "tool_name": "Edit", "tool_input": {"file_path": "tests/properties/test_chunker_properties.py"}}' \ | python .claude/hooks/protect_properties.py | python peek.pyecho '{"cwd": ".", "tool_name": "Edit", "tool_input": {"file_path": "chunker.py"}}' \ | python .claude/hooks/protect_properties.pyecho "exit code: $?"Expected output:
permissionDecision: denytests/properties/test_chunker_properties.py is part of the spec.Change chunker.py so the properties pass instead.exit code: 0What just happened: the edit to the property file is denied with a reason Claude can act on. For chunker.py the hook printed nothing and exited 0, which Claude Code reads as "no objection".
Step 9: Register the hooks in .claude/settings.json
Goal: tell Claude Code when to run each script, and add the two deny rules.
Why this step: Claude Code runs only the hooks that a settings file registers. Project settings in .claude/settings.json are read by every Claude Code session started in this folder, and they are committed with the code, so anyone who clones the repository gets the same gate.
Create .claude/settings.json:
{ "permissions": { "deny": ["Edit(/tests/properties/**)", "Edit(/.claude/**)"] }, "hooks": { "PreToolUse": [ { "matcher": "Edit|Write", "hooks": [ { "type": "command", "command": "python", "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/protect_properties.py"] } ] } ], "Stop": [ { "hooks": [ { "type": "command", "command": "python", "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/verify_properties.py"], "timeout": 180 } ] } ] }}What each part does:
permissions.deny:Edit(...)rules cover every built-in tool that edits files, plus the Bash file commands Claude Code recognises. The leading/anchors the path at the project root."command": "python"with"args": this is the hooks exec form. Claude Code startspythondirectly with the script path as one argument, with no shell in between, so a project path with spaces still works. Claude Code itself replaces${CLAUDE_PROJECT_DIR}inargswith the project path before it starts the process, so the placeholder works without a shell. Usepython, notpython3. On Windows,python3can resolve to the Microsoft Store stub, which exits with code 49 and no output. Claude Code treats that as a non-blocking error, so the gate silently turns off.matcher: PreToolUse fires only for theEditandWritetools. The Stop hook takes no matcher.timeout: seconds. The default for command hooks is 600. A hook that times out does not block anything, so 180 seconds caps a stuck run without ever cutting off a normal one, which takes about a second.
Run it: check the JSON parses, then commit the hooks and the helper:
python -m json.tool .claude/settings.json > /dev/null \ && echo "settings.json is valid JSON"git add .claude peek.pygit commit -q -m "verifier hooks"git ls-files .claude peek.pyExpected output:
settings.json is valid JSON.claude/hooks/protect_properties.py.claude/hooks/verify_properties.py.claude/settings.jsonpeek.pyWhat just happened: the gate is configured and committed. The json.tool check comes first because Claude Code ignores a settings file with invalid JSON, along with every hook in it.
Step 10: Let Claude Code fix the chunker under the Stop hook
Goal: run a real Claude Code session against the buggy chunker and watch the Stop hook hold it until the properties pass.
Why this step: so far you have run the hooks by hand; here Claude Code runs them. The prompt deliberately asks for something unrelated to the bug, because the gate does not depend on the task: whatever you asked for, Claude cannot end the turn while the spec fails.
No Claude Code account? Read the recorded run below, skip the claude command, and paste the fixed chunker.py given under "whichever path you took".
sequenceDiagram
participant You
participant Claude
participant Stop as Stop hook
participant Props as pytest + Hypothesis
You->>Claude: Add a docstring to chunk()
Claude->>Claude: Edit chunker.py (docstring)
Claude->>Stop: tries to end the turn
Stop->>Props: profile agent
Props-->>Stop: test_no_text_is_lost fails, text='0'
Stop-->>Claude: decision: block + smallest failing case
Claude->>Claude: Edit chunker.py (fix the loop)
Claude->>Stop: tries to end the turn
Stop->>Props: profile agent
Props-->>Stop: all pass
Stop-->>You: exit 0, turn ends
The diagram traces a real run, recorded for this tutorial on 2026-10-05 with Claude Code 2.1.289.
Start Claude Code from the project root, with the virtual environment still active. The recorded run used headless mode (-p), which prints Claude's final message and exits:
claude -p "Add a one-line docstring to chunk() in chunker.py." \ --permission-mode acceptEdits --setting-sources project--permission-mode acceptEdits lets Claude edit files without asking. --setting-sources project loads only this project's .claude/settings.json, so hooks from your personal settings do not join the run. A -p session treats the folder as trusted, so the project hooks run.
Plain -p output shows only Claude's final message, not the hook's feedback. To see the hook fire, add --output-format stream-json --verbose to the command and pipe it through grep -A10 "Stop hook feedback". If you prefer the interactive UI, start claude --permission-mode acceptEdits --setting-sources project with no prompt, accept the workspace trust prompt (Claude Code holds back every hook until you do), and type the same request. A block then shows up as a Stop hook error occurred notice, and ctrl+o shows the reason. Either way, if you see no block at all, Claude fixed the bug on its first pass and the hook simply let it stop. That is a valid outcome too.
The recorded run used --output-format stream-json --verbose, and the quotes below come from its transcript. Claude read chunker.py, added a docstring, and noticed the bug on its own. Its first reply tried to end the turn with this (excerpt):
I wrote "exactly" on purpose, because of what the code does now: the loop only produces full-size windows. [...] I didn't change the code because you only asked for a docstring. If you want it fixed, I can add a final shorter chunk to cover the leftover tokens and update the docstring to match.
That is a fair answer to the task it was given. Without the hook, the session would have ended there, with the bug described in a docstring instead of fixed. The Stop hook ran, and the transcript records Claude Code sending this back to Claude:
Stop hook feedback:Property tests failed. Fix the code in chunker.py; do not edit the tests.test_no_text_is_lost AssertionError: assert '' == '0' - 0 Failing test case: test_no_text_is_lost( text='0', budget=(2, 0), )Claude then read both test files, changed the loop in chunk() to emit a final, shorter window, and corrected its docstring to "at most max_tokens tokens". It tried twice to run pytest itself, and both commands were denied, because acceptEdits approves edits, not shell commands. That did not matter. When Claude tried to stop again, the hook ran the properties, they passed, and the turn ended. The whole run took 11 turns, 43 seconds and USD 0.20.
The Stop hook error occurred label in the interactive UI is how Claude Code shows a decision: block reply. It does not mean your script crashed.
Claude's replies vary from run to run, and so can its fix. The four properties accept more than one correct chunker. For example, changing the loop to for i in range(0, len(spans), stride): passes all four, but on the Step 1 sentence it emits a third chunk holding only done., so the example test from Step 2 fails while the Stop hook, which runs only tests/properties, lets Claude stop. Properties say what must never happen. They do not pin down exact chunk boundaries, and that gap is what Step 11 closes.
So, whichever path you took, replace the whole file with this version now. It has the loop change Claude made in the recorded run, without its docstring. Steps 11 and 12 compare exact chunk boundaries, so they need this exact chunker. If you have no Claude Code account, this paste is your fix:
# chunker.py"""Split text into overlapping chunks under a token budget, for RAG indexing."""import refrom dataclasses import dataclass# A token is a run of non-space characters plus the whitespace after it.# Leading whitespace at the very start of the text is a token of its own.TOKEN = re.compile(r"\S+\s*|\s+")@dataclass(frozen=True)class Chunk: start: int # character offset into the original text end: int text: strdef count_tokens(text: str) -> int: return len(TOKEN.findall(text))def chunk(text: str, max_tokens: int, overlap: int) -> list[Chunk]: if max_tokens < 1: raise ValueError("max_tokens must be at least 1") if not 0 <= overlap < max_tokens: raise ValueError("overlap must be in [0, max_tokens)") spans = [m.span() for m in TOKEN.finditer(text)] stride = max_tokens - overlap chunks = [] i = 0 while i < len(spans): window = spans[i : i + max_tokens] start, end = window[0][0], window[-1][1] chunks.append(Chunk(start, end, text[start:end])) if i + max_tokens >= len(spans): break # this window reached the last token i += stride return chunksRun it: with the listed version in place, run the hook by hand and then the whole suite:
echo '{"cwd": ".", "hook_event_name": "Stop"}' | python .claude/hooks/verify_properties.pyecho "exit code: $?"python -m pytest -qgit status --shortExpected output:
exit code: 0..... [100%]5 passed in 0.66s M chunker.pyThen commit the fix:
git commit -qam "fix: keep short documents"What just happened: a silent exit 0 from the hook means it now lets Claude stop. The five passing tests are the example test and four properties. chunker.py is the only changed file in git status, so the fix went into the code, not the spec. If Claude also edited another file, such as tests/test_examples.py, restore it with git checkout -- <file> before committing.
Two built-in limits keep a Stop hook from trapping a session. The event carries a stop_hook_active field, which is true when Claude is already continuing because of a Stop hook. On top of that, Claude Code ends the turn anyway after 8 blocks in a row with no tool call in between. The count resets each time Claude calls a tool, so an agent that keeps editing keeps getting feedback, and an agent that only replies "done" gets cut off. This hook ignores stop_hook_active on purpose: a spec failure should keep blocking for as long as the agent is still making edits. For a loop that needs a hard stop of its own, see How to Build a Claude Code Agent Loop That Cannot Run Away.
Step 11: Catch a behaviour change with a Hypothesis differential test
Goal: freeze the working chunker as a reference, and add a test that the current chunker.py returns exactly the same chunks for every generated input.
Why this step: the four properties say what must never happen. They do not say what must stay the same. Suppose you ask Claude to simplify or speed up the tokenizer. A rewrite can pass every property and still put chunk boundaries in different places. In RAG, that means every stored embedding now points at offsets the new chunker would never produce, and your re-index quietly differs from your index. This is Addy Osmani's third test: when you replace a system, feed the old and new versions the same random inputs and compare. The usual name for it is differential testing.
direction: down
classes: {
gen: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
ref: {style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}}
new: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
cmp: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}}
}
gen: "Hypothesis: the same text and budget\nfor both sides" {class: gen}
pair: "two implementations, one input" {
label.near: top-center
style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
grid-rows: 1
grid-gap: 30
ref: "reference_chunker.chunk\nfrozen copy in tests/properties/\n(the agent cannot edit it)" {class: ref}
new: "chunker.chunk\nthe code the agent changes" {class: new}
}
cmp: "assert: the same list of\n(start, end) offsets" {class: cmp}
gen -> pair: "1. same input"
pair -> cmp: "2. compare"
The reference lives inside tests/properties/, so the deny rules, the PreToolUse hook and the git status check that guard the spec also guard the oracle. The agent can change the new side as much as it likes, but it cannot move the target.
Copy the fixed chunker into the spec folder as the reference:
cp chunker.py tests/properties/reference_chunker.pyCreate tests/properties/test_matches_reference.py. It reuses the budgets() strategy from the property file, which pytest can import because it puts each test file's folder on the import path:
# tests/properties/test_matches_reference.pyfrom hypothesis import given, strategies as stimport chunkerimport reference_chunkerfrom test_chunker_properties import budgets@given(text=st.text(max_size=300), budget=budgets())def test_new_chunker_matches_reference(text, budget): max_tokens, overlap = budget expected = reference_chunker.chunk(text, max_tokens, overlap) actual = chunker.chunk(text, max_tokens, overlap) assert [(c.start, c.end) for c in actual] == [(c.start, c.end) for c in expected]Run it: commit the two new spec files first, because the Stop hook treats any uncommitted change under tests/properties/, including a new file, as tampering. Then run the property folder:
git add tests/propertiesgit commit -q -m "freeze reference chunker"python -m pytest -q tests/properties --hypothesis-profile=agent --tb=noExpected output:
..... [100%]5 passed in 1.26sWhat just happened: the reference is now part of the committed spec. It cannot fail yet, because chunker.py and the reference are the same file. Step 12 changes that.
Step 12: Review an agent's refactor against the reference
Goal: apply the kind of rewrite an agent produces when asked to "simplify the tokenizer", and watch the differential test catch what the properties miss.
Why this step: the properties alone would let this rewrite through. It drops the regular-expression alternation and builds tokens from word start positions instead. It is shorter, it reads well, and it passes every property. It also changes behaviour on text that starts with whitespace.
Replace chunker.py with this version, which is what an agent might return:
# chunker.py"""Split text into overlapping chunks under a token budget, for RAG indexing."""import refrom dataclasses import dataclassWORD = re.compile(r"\S+")@dataclass(frozen=True)class Chunk: start: int # character offset into the original text end: int text: strdef token_spans(text: str) -> list[tuple[int, int]]: """Each token runs from the start of one word to the start of the next.""" starts = [m.start() for m in WORD.finditer(text)] if not starts: return [(0, len(text))] if text else [] starts[0] = 0 # fold leading whitespace into the first token return list(zip(starts, starts[1:] + [len(text)]))def count_tokens(text: str) -> int: return len(token_spans(text))def chunk(text: str, max_tokens: int, overlap: int) -> list[Chunk]: if max_tokens < 1: raise ValueError("max_tokens must be at least 1") if not 0 <= overlap < max_tokens: raise ValueError("overlap must be in [0, max_tokens)") spans = token_spans(text) stride = max_tokens - overlap chunks = [] i = 0 while i < len(spans): window = spans[i : i + max_tokens] start, end = window[0][0], window[-1][1] chunks.append(Chunk(start, end, text[start:end])) if i + max_tokens >= len(spans): break # this window reached the last token i += stride return chunksRun it: the property folder first, then the hook:
python -m pytest -q tests/properties --hypothesis-profile=agent --tb=noecho '{"cwd": ".", "hook_event_name": "Stop"}' \ | python .claude/hooks/verify_properties.py | python peek.pyExpected output:
....F [100%]=========================== short test summary info ===========================FAILED tests/properties/test_matches_reference.py::test_new_chunker_matches_reference1 failed, 4 passed in 1.38sdecision: blockProperty tests failed. Fix the code in chunker.py; do not edit the tests.test_new_chunker_matches_reference assert [(0, 2)] == [(0, 1), (1, 2)] At index 0 diff: (0, 2) != (0, 1) Right contains one more item: (1, 2) Use -v to get more diff Failing test case: test_new_chunker_matches_reference( text='\r0', budget=(1, 0), )All four properties pass, because the rewrite is internally consistent: it never loses text, never breaks the budget by its own count, and never overlaps too much. Only the differential test fails. Its shrunk case is a carriage return followed by one character. The reference treats the leading \r as a token of its own and makes two one-token chunks. The rewrite folds the \r into the first word and makes one.
Fix it the way the reason asks, in chunker.py. Replace the token_spans function with this version, which keeps leading whitespace as a separate token:
def token_spans(text: str) -> list[tuple[int, int]]: """Each token runs from the start of one word to the start of the next.""" starts = [m.start() for m in WORD.finditer(text)] if not starts: return [(0, len(text))] if text else [] if starts[0] > 0: starts.insert(0, 0) # leading whitespace is a token of its own return list(zip(starts, starts[1:] + [len(text)]))Run the hook and the full suite again:
echo '{"cwd": ".", "hook_event_name": "Stop"}' | python .claude/hooks/verify_properties.pyecho "exit code: $?"python -m pytest -qExpected output:
exit code: 0...... [100%]6 passed in 0.71sWhat just happened: the rewrite now produces the same chunks as the reference for every generated input, so the hook lets Claude stop again. The six passing tests are one example, four properties and one differential test. The differential test never measures speed. It checks that the replacement behaves exactly like the system it replaces.
Fix common Hypothesis and Claude Code hook errors
| Error you see | Root cause | Fix |
|---|---|---|
E ModuleNotFoundError: No module named 'chunker' | You ran pytest instead of python -m pytest. Plain pytest does not put the project root on the import path. | Run python -m pytest, as every command here does. The hook already uses sys.executable -m pytest. |
hypothesis.errors.DeadlineExceeded: Test took 300.80ms, which exceeds the deadline of 200.00ms. | You ran without --hypothesis-profile=agent, so the default 200 ms deadline applies, and one test case was slow. | Use the agent profile, which sets deadline=None. In the default profile, add @settings(deadline=None) to a test that is slow by design. |
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0) from peek.py | The hook printed nothing, which means it allowed the action. peek.py has no JSON to read. | Not a bug in the hook. Check echo "exit code: $?" instead, as Steps 10 and 12 do. |
decision: block with ?? tests/properties/reference_chunker.py | You added a file to the spec folder without committing it. The git status check treats any uncommitted change there, including a new file, as tampering. | Commit your own spec changes: git add tests/properties && git commit. |
Stop hook error occurred · ctrl+o to see | Not an error in your script. Claude Code shows this notice whenever a Stop hook replies decision: block. | Press ctrl+o to read the reason Claude received. If you want the feedback without the notice, a Stop hook can reply with hookSpecificOutput.additionalContext instead of decision: block. |
| Hooks never run in an interactive session | The folder is not trusted yet. Claude Code holds back every settings-file hook until you accept the workspace trust dialog. | Accept the trust prompt when Claude Code starts, or start it from a folder you have already trusted. |
Failed with non-blocking status code: ... in the transcript | The hook command failed to start: a wrong path in settings.json, or python3 resolving to the Windows Store stub (exit code 49). Exit codes other than 0 and 2 never block. | Use "command": "python" with the virtual environment active, keep the script path in args, and run the hook by hand as in Step 7 to see its real error. |
Results differ from this page, or derandomize seems ignored | Your shell exports CI, so Hypothesis loaded its built-in ci profile. | unset CI before running, or always pass --hypothesis-profile=agent. |
Fix "DeadlineExceeded: Test took 300.80ms, which exceeds the deadline of 200.00ms"
Hypothesis fails any single test case that runs longer than 200 ms by default, and a slow or loaded machine can trip it with no bug in your code. Run with --hypothesis-profile=agent, which sets deadline=None, or add @settings(deadline=None) to a test that is slow by design.
Fix "Stop hook error occurred" in Claude Code
This notice means your Stop hook replied decision: block; the hook worked. Press ctrl+o to read the reason Claude received. If the hook never seems to run at all, check for Failed with non-blocking status code in the transcript and run the hook by hand as in Step 7.
Fix "ModuleNotFoundError: No module named 'chunker'" in pytest
Plain pytest does not put the project root on the import path, so the tests cannot import chunker.py. Run python -m pytest from the project root instead. The Stop hook already does this through sys.executable -m pytest.
How the Stop hook, properties and chunker fit together
The loop diagram at the start showed two checks. This one shows the files you actually built: which file runs which, and which file guards which.
direction: down
classes: {
cfg: {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}}
hook: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
spec: {style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}}
code: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
tool: {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}}
}
settings: ".claude/settings.json, read when claude starts" {
label.near: top-center
style: {fill: "#FFFFFF"; stroke: "#C9A800"; font-color: "#2C2C2A"}
grid-rows: 1
grid-gap: 14
deny: "permissions.deny\nEdit(/tests/properties/**)\nEdit(/.claude/**)" {class: cfg}
pre: "PreToolUse, matcher Edit|Write\npython protect_properties.py" {class: hook}
stop: "Stop\npython verify_properties.py" {class: hook}
}
checks: "verify_properties.py runs" {
label.near: top-center
style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
grid-rows: 1
grid-gap: 14
git: "git status --porcelain\n-- tests/properties" {class: tool}
pytest: "python -m pytest tests/properties\n--hypothesis-profile=agent" {class: tool}
}
spec: "tests/ (committed spec)" {
label.near: top-center
style: {fill: "#FFFFFF"; stroke: "#5FA898"; font-color: "#2C2C2A"}
grid-rows: 1
grid-gap: 14
conf: "conftest.py\nprofile agent" {class: spec}
props: "properties/\ntest_chunker_properties.py\ntest_matches_reference.py\nreference_chunker.py" {class: spec}
}
code: "chunker.py\nthe only file the agent may change" {class: code}
settings -> checks: "on every end of turn"
checks -> spec: "runs"
spec -> code: "imports and checks"
Read it from the top. The blue box at the bottom is the only thing Claude is free to edit, and everything between Claude and that box is either a rule in settings.json or a committed file that the rules protect.
Complete code: Claude Code Stop hook and Hypothesis tests
The final file tree, not counting .venv/, .git/ and the caches listed in .gitignore:
chunker-verifier/├── .claude/│ ├── hooks/│ │ ├── protect_properties.py│ │ └── verify_properties.py│ └── settings.json├── .gitignore├── chunker.py├── peek.py└── tests/ ├── conftest.py ├── properties/ │ ├── reference_chunker.py │ ├── test_chunker_properties.py │ └── test_matches_reference.py └── test_examples.pyEvery file below is the final version, exactly as it stands at the end of Step 12. If your project drifted at any point, copy these over it.
.gitignore:
.venv/__pycache__/.hypothesis/.pytest_cache/The code under test, chunker.py, as rewritten and fixed in Step 12:
# chunker.py"""Split text into overlapping chunks under a token budget, for RAG indexing."""import refrom dataclasses import dataclassWORD = re.compile(r"\S+")@dataclass(frozen=True)class Chunk: start: int # character offset into the original text end: int text: strdef token_spans(text: str) -> list[tuple[int, int]]: """Each token runs from the start of one word to the start of the next.""" starts = [m.start() for m in WORD.finditer(text)] if not starts: return [(0, len(text))] if text else [] if starts[0] > 0: starts.insert(0, 0) # leading whitespace is a token of its own return list(zip(starts, starts[1:] + [len(text)]))def count_tokens(text: str) -> int: return len(token_spans(text))def chunk(text: str, max_tokens: int, overlap: int) -> list[Chunk]: if max_tokens < 1: raise ValueError("max_tokens must be at least 1") if not 0 <= overlap < max_tokens: raise ValueError("overlap must be in [0, max_tokens)") spans = token_spans(text) stride = max_tokens - overlap chunks = [] i = 0 while i < len(spans): window = spans[i : i + max_tokens] start, end = window[0][0], window[-1][1] chunks.append(Chunk(start, end, text[start:end])) if i + max_tokens >= len(spans): break # this window reached the last token i += stride return chunksThe example test from Step 2:
# tests/test_examples.pyfrom chunker import chunkdef test_two_chunks_share_one_token(): chunks = chunk("Agents stop when the work looks done.", max_tokens=4, overlap=1) assert [c.text for c in chunks] == [ "Agents stop when the ", "the work looks done.", ]The settings profile from Step 5:
# tests/conftest.pyfrom hypothesis import settings# The profile the Stop hook runs. Same inputs every run, no timing failures,# and no memory of earlier runs, so a verdict depends only on the code.settings.register_profile( "agent", derandomize=True, database=None, deadline=None, max_examples=200,)The four properties from Step 4:
# tests/properties/test_chunker_properties.pyfrom hypothesis import given, strategies as stfrom chunker import chunk, count_tokens@st.compositedef budgets(draw): max_tokens = draw(st.integers(min_value=1, max_value=20)) overlap = draw(st.integers(min_value=0, max_value=max_tokens - 1)) return max_tokens, overlapdef reassemble(chunks): """Glue chunks back together, skipping the part each one repeats.""" out, pos = [], 0 for c in chunks: assert c.start <= pos, f"gap: text[{pos}:{c.start}] is in no chunk" out.append(c.text[pos - c.start :]) pos = c.end return "".join(out)@given(text=st.text(max_size=300), budget=budgets())def test_no_text_is_lost(text, budget): max_tokens, overlap = budget assert reassemble(chunk(text, max_tokens, overlap)) == text@given(text=st.text(max_size=300), budget=budgets())def test_every_chunk_is_a_verbatim_slice(text, budget): max_tokens, overlap = budget for c in chunk(text, max_tokens, overlap): assert c.text == text[c.start : c.end]@given(text=st.text(max_size=300), budget=budgets())def test_no_chunk_exceeds_the_budget(text, budget): max_tokens, overlap = budget for c in chunk(text, max_tokens, overlap): assert count_tokens(c.text) <= max_tokens@given(text=st.text(max_size=300), budget=budgets())def test_overlap_never_exceeds_the_setting(text, budget): max_tokens, overlap = budget chunks = chunk(text, max_tokens, overlap) for prev, nxt in zip(chunks, chunks[1:]): shared = text[nxt.start : prev.end] if nxt.start < prev.end else "" assert count_tokens(shared) <= overlapThe differential test from Step 11:
# tests/properties/test_matches_reference.pyfrom hypothesis import given, strategies as stimport chunkerimport reference_chunkerfrom test_chunker_properties import budgets@given(text=st.text(max_size=300), budget=budgets())def test_new_chunker_matches_reference(text, budget): max_tokens, overlap = budget expected = reference_chunker.chunk(text, max_tokens, overlap) actual = chunker.chunk(text, max_tokens, overlap) assert [(c.start, c.end) for c in actual] == [(c.start, c.end) for c in expected]tests/properties/reference_chunker.py is the Step 10 chunker.py, copied unchanged in Step 11, so its first line still reads # chunker.py:
# chunker.py"""Split text into overlapping chunks under a token budget, for RAG indexing."""import refrom dataclasses import dataclass# A token is a run of non-space characters plus the whitespace after it.# Leading whitespace at the very start of the text is a token of its own.TOKEN = re.compile(r"\S+\s*|\s+")@dataclass(frozen=True)class Chunk: start: int # character offset into the original text end: int text: strdef count_tokens(text: str) -> int: return len(TOKEN.findall(text))def chunk(text: str, max_tokens: int, overlap: int) -> list[Chunk]: if max_tokens < 1: raise ValueError("max_tokens must be at least 1") if not 0 <= overlap < max_tokens: raise ValueError("overlap must be in [0, max_tokens)") spans = [m.span() for m in TOKEN.finditer(text)] stride = max_tokens - overlap chunks = [] i = 0 while i < len(spans): window = spans[i : i + max_tokens] start, end = window[0][0], window[-1][1] chunks.append(Chunk(start, end, text[start:end])) if i + max_tokens >= len(spans): break # this window reached the last token i += stride return chunksThe Stop hook from Step 7:
# .claude/hooks/verify_properties.py"""Stop hook: Claude may not finish while a property fails."""import jsonimport osimport subprocessimport sysevent = json.load(sys.stdin)root = os.environ.get("CLAUDE_PROJECT_DIR", event["cwd"])def block(reason: str) -> None: print(json.dumps({"decision": "block", "reason": reason})) sys.exit(0)# 1. The properties are the contract. If they changed, nothing else counts.changed = subprocess.run( ["git", "status", "--porcelain", "--", "tests/properties"], cwd=root, capture_output=True, text=True,).stdout.strip()if changed: block( "tests/properties/ changed. Those files are the spec, not yours to edit.\n" "Restore them with `git checkout -- tests/properties` and fix the code.\n" + changed )# 2. Run the properties with the deterministic profile.try: result = subprocess.run( [sys.executable, "-m", "pytest", "tests/properties", "-q", "-p", "no:cacheprovider", "--hypothesis-profile=agent", "--tb=short"], cwd=root, capture_output=True, text=True, timeout=120, )except subprocess.TimeoutExpired as exc: block(f"Property tests ran for {exc.timeout:.0f} s without finishing. " "chunker.py probably loops forever on some input.")if result.returncode == 0: sys.exit(0)# 3. Send back only what Claude needs: which property, and the smallest input.keep = []for line in result.stdout.splitlines(): if line.startswith("____"): # one header line per failing test keep.append(line.strip("_ ")) elif line.startswith("E ") and line[1:].strip(): keep.append(line[1:].rstrip()) # the assertion and the test caseblock( "Property tests failed. Fix the code in chunker.py; do not edit the tests.\n" + "\n".join(keep[:40]))The PreToolUse hook from Step 8:
# .claude/hooks/protect_properties.py"""PreToolUse hook: Claude may not edit the spec or the hooks that enforce it."""import jsonimport osimport sysfrom pathlib import Pathevent = json.load(sys.stdin)root = Path(os.environ.get("CLAUDE_PROJECT_DIR", event["cwd"])).resolve()target = Path(event["tool_input"].get("file_path", "")).resolve()protected = [root / "tests" / "properties", root / ".claude"]if any(target.is_relative_to(p) for p in protected): print(json.dumps({ "hookSpecificOutput": { "hookEventName": "PreToolUse", "permissionDecision": "deny", "permissionDecisionReason": ( f"{target.relative_to(root).as_posix()} is part of the spec.\n" "Change chunker.py so the properties pass instead." ), } }))sys.exit(0).claude/settings.json from Step 9:
{ "permissions": { "deny": ["Edit(/tests/properties/**)", "Edit(/.claude/**)"] }, "hooks": { "PreToolUse": [ { "matcher": "Edit|Write", "hooks": [ { "type": "command", "command": "python", "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/protect_properties.py"] } ] } ], "Stop": [ { "hooks": [ { "type": "command", "command": "python", "args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/verify_properties.py"], "timeout": 180 } ] } ] }}The peek.py helper from Step 7:
# peek.py"""Print a hook's JSON reply the way Claude reads it: decision first, then reason."""import jsonimport sysreply = json.load(sys.stdin)reply = reply.get("hookSpecificOutput", reply)for key in ("decision", "permissionDecision"): if key in reply: print(f"{key}: {reply[key]}")for key in ("reason", "permissionDecisionReason"): if key in reply: print(reply[key])Where to go next
- Swap in a real tokenizer. Replace the regular-expression token counter with
tiktoken0.14.0 and keep all six tests. The properties will quickly find two traps.Encoding.decode()replaces invalid UTF-8 by default, so decoding a token slice can split a multi-byte character and break "no text is lost". Andencode()raises on text that contains a special token such as<|endoftext|>unless you useencode_ordinary(). SetTIKTOKEN_CACHE_DIRand download the encoding once, so the hook still runs offline. - Add an end-to-end property. Addy Osmani's first test is an end-to-end flow, the path this tutorial stops short of. Index a generated document set into an in-memory vector store and add a property that every document is retrievable by an exact quote of its shortest chunk.
- Turn more team rules into hooks. The same pattern, a script that blocks with a reason, enforces other policy too; Hooks: The Enforcement Layer collects the patterns.
- Let Hypothesis write the differential test.
hypothesis write --equivalent chunker.chunk reference_chunker.chunkgenerates a starting version of Step 11's test, which is useful when the function under test has many arguments.
References
- Osmani, A. (2026). Give your agent a way to check its own work. X. https://x.com/addyosmani/status/2106995301802541481
- Anthropic. Best practices for Claude Code, section "Give Claude a way to verify its work". https://docs.claude.com/en/docs/claude-code/best-practices
- Anthropic. Hooks reference (Claude Code 2.1.289). https://docs.claude.com/en/docs/claude-code/hooks
- Anthropic. Automate actions with hooks. https://docs.claude.com/en/docs/claude-code/hooks-guide
- Anthropic. Configure permissions. https://docs.claude.com/en/docs/claude-code/permissions
- Anthropic. Set up Claude Code. https://docs.claude.com/en/docs/claude-code/setup
- Hypothesis 6.168.4. Quickstart. https://hypothesis.readthedocs.io/en/latest/quickstart.html
- Hypothesis 6.168.4. API reference (
@given,@settings, profiles). https://hypothesis.readthedocs.io/en/latest/reference/api.html - Hypothesis 6.168.4. Strategies reference (
text,integers,composite). https://hypothesis.readthedocs.io/en/latest/reference/strategies.html - Hypothesis 6.168.4. Integrations reference (pytest plugin options, Ghostwriter). https://hypothesis.readthedocs.io/en/latest/reference/integrations.html
- Hypothesis. Changelog (6.159.0: "example" becomes "test case" in messages). https://hypothesis.readthedocs.io/en/latest/changelog.html
- MacIver, D. R. (2016). Testing performance optimizations. https://hypothesis.works/articles/testing-performance-optimizations/
- MacIver, D. R., Hatfield-Dodds, Z., et al. (2019). Hypothesis: A new approach to property-based testing. Journal of Open Source Software. https://joss.theoj.org/papers/10.21105/joss.01891
- Maaz, M., DeVoe, L., Hatfield-Dodds, Z., Carlini, N. (2025). Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem. arXiv:2510.09907. https://arxiv.org/abs/2510.09907
- pytest 9.1.1. Changelog. https://docs.pytest.org/en/stable/changelog.html
- Kiro. (2025). Kiro is generally available. https://kiro.dev/blog/general-availability/
- anthropics/claude-code issue #57946:
python3hook silently bypassed on Windows. https://github.com/anthropics/claude-code/issues/57946 - OpenAI. tiktoken 0.14.0. https://github.com/openai/tiktoken
Related Articles
- How to Build a Claude Code Agent Loop That Cannot Run Away
- How to Reduce Claude Code Token Usage: A Measured Setup
- Contract Tests for a LangGraph Agent: Test the Property, Not the Setting
- Superpowers Plugin for Claude Code: Install and Verify




