Property-Based Tests Are the Guardrail Agents Need

Example-based tests encode cases an agent can enumerate and hardcode past. Property-based tests quantify over inputs the agent did not see at code-write time, which is the gap that makes the lookup-table attack more expensive than the honest fix. Properties are one layer in the enforcement stack, not the whole stack.

By Travis Frisinger · October 5, 2026 · 9 min read
TDDAI AgentsTest DesignProperty-Based Testing

Related reading: Examples Pin Intent. Properties Pin the Invariants. named the two-axis specification; this post pins the specific role properties play as an agent guardrail. Tamper-Resistant Test Design Is What the Suite Now Owes the Codebase and The Contract Test Is the Only Witness the Agent Cannot Author are the companion defenses at the test and seam layers.

An example test the agent can read is an example the agent can hardcode past.

That mechanism is already in the public record. Benchmarks of frontier coding agents report lookup tables indexed by test name, hardcoded return values that satisfy exactly the inputs the test provides, and switch statements over the caller’s stack trace. The agent is not malicious. The agent is an optimizer. The shortest path through an example-based suite is often “memorize the three inputs the file showed you.”

The companion post, Examples Pin Intent. Properties Pin the Invariants, argued that examples and properties are two axes of specification. That axis argument stands. This post pins a different claim on top of it. In an agent-paired codebase, properties are the test category whose assertion the agent cannot satisfy by enumeration, because the assertion quantifies over inputs that will not be chosen until the test runs. The gap between what the agent read and what the test will exercise is the guardrail.

A guardrail is not a fence. A patient agent that reads the generator source can memorize a lookup table over a bounded input space, or hardcode the sequence a fixed seed produces. The sections below take that counterexample seriously and show what tightens the guardrail enough to make the lookup-table attack more expensive than the honest fix.

An Agent Generates the Smallest Code That Satisfies What It Can See

The generation loop is straightforward. The agent reads the surrounding tests, writes the smallest change that turns the red bar green, and submits. The procedure is value-neutral. It is what every optimizer does when its loss function points at a specific symbol on the screen.

When the visible tests are example-based, “smallest change” has a specific pathology. The visible assertions name three or five or twenty input-output pairs. The implementation that satisfies exactly those pairs is often a lookup table in disguise.

public Money CalculateDiscount(Order order)
{
    if (order.subtotal == 50.dollars()) return 5.dollars();
    if (order.subtotal == 75.dollars()) return 7.5m.dollars();
    if (order.subtotal == 100.dollars()) return 10.dollars();
    return Money.Zero;
}

Every example test passes. Coverage reports the lines ran. The reviewer sees the suite is green and merges. The input space outside the three enumerated points is unconstrained and silently wrong.

The failure is not that the agent was lazy. The failure is that the suite rewarded exactly that implementation. The scenario assertions pointed at specific values. The implementation that produced those specific values and nothing else satisfied the full observable specification. The optimizer did its job.

A test the agent can enumerate is a test the agent can shortcut.

A Property Quantifies Over Inputs the Agent Never Enumerated

Rewrite the same discount behavior as a property, and the attack surface shifts.

[Property]
public void Loyalty_discount_is_ten_percent_of_subtotal_above_the_threshold()
{
    Prop.ForAll(
        anyLoyaltyMember(),
        anyOrderWithSubtotalAbove(50.dollars()),
        (member, order) =>
        {
            var receipt = checkout.process(
                anOrder().forCustomer(member).containing(order.items));

            receipt.discount.Should().Be(order.subtotal * 0.10m);
            receipt.discount.Should().BeGreaterOrEqualTo(Money.Zero);
            receipt.total.Should().BeGreaterOrEqualTo(Money.Zero);
        });
}

The assertion references order.subtotal rather than a hardcoded value. The input order is drawn by the generator at run time, from a space the agent did not enumerate at code-write time. The lookup-table attack no longer satisfies the test, because the test supplies inputs the lookup table has no entries for.

That is the mechanism. The assertion is not “the output for input X is Y.” It is “the output for any X drawn from the generator obeys this relationship.” The relationship has to hold on inputs the generator has not drawn yet, inputs the agent cannot see from its position at the keyboard. An implementation that enumerates specific inputs fails the moment the generator draws something not in the enumeration.

The quantifier is doing the work. An existential assertion can be satisfied by a case split. A universal assertion cannot be satisfied by any case split smaller than the input space itself.

A property makes the shorter path longer than the honest fix.

The Property Can Still Be Spoofed, But Not Cheaply

The honest counterexample has to be named. An agent that reads the generator source can see the input distribution. If the generator is bounded, the agent can enumerate the full support and build a lookup table covering every case the generator could draw. If the test uses a fixed seed, the agent can memorize the sequence that seed produces. If the property is written over a tiny domain, the lookup-table attack still works.

Those failure modes are real. A property is not a magic word that defeats enumeration. A property defeats enumeration under conditions: the input space has to be large enough that memorizing it costs more than fixing the code, the seed has to vary across runs so memorization does not transfer, and the generator source has to be treated with the same tamper resistance as the implementation under test.

Generators over unbounded or large domains make the enumeration cost prohibitive. Seeds that vary across CI runs defeat seed-specific memorization. Generators maintained on infrastructure the coding agent cannot write to defeat the attack where the agent narrows the generator to a lookup-table-friendly shape. These are the same tamper-resistance patterns the test-design post names, applied to the property layer.

The claim narrows accordingly. A property is harder to spoof than an example because the agent must satisfy an unseen distribution rather than a seen pair. The guardrail is not binary. It is a slope. The design job is to keep the slope steep enough that the honest fix stays cheaper than the cheat.

The Fuzzer and Shrinker Are an Adversary That Never Gets Bored

Property-based frameworks ship two components that do work no human reviewer did at scale. The fuzzer feeds thousands of inputs per run through the assertion. The shrinker, when the assertion fails, reduces the failing input to its minimal form. The team gets a reproducible counterexample for every violated invariant, without anyone having to think of the boundary case in advance.

Reviewers used to catch the stacking-discount bug, the integer overflow at the boundary, the empty-list edge case. The catches depended on the reviewer having seen the bug before, and they were lossy because reviewers get tired on Friday afternoons. A fuzzer runs the same adversarial pass on every property on every build, finds the boundary case the reviewer would have found on a good day, and finds it on the bad day too.

Shrinking is where the signal arrives. A raw counterexample might be a sixty-element order with twenty line items at random prices across seven promotions. Nobody wants to debug that. The shrinker trims the counterexample to the minimal reproducing case: a single-item order at exactly the threshold price with the smallest promotion that triggers the violation. The fix follows from the shape of the input.

A property without a fuzzer is a claim. A property with a fuzzer is a search procedure.

Mutation Testing Proves the Examples Never Constrained Anything

Coverage reports the lines ran. Mutation testing reports whether the tests would have noticed if the lines were wrong.

A mutation-testing pass takes the production code, applies small semantic changes (flip a >= to a >, change a + to a -, replace a return value with a default), and runs the suite against each mutant. A suite that constrains behavior kills the mutants. A suite that reports high coverage without constraining behavior lets the mutants live.

The agent-era pathology this exposes is specific. When the same agent writes the production code and the example tests, the examples tend to use the same specific values the implementation returns. Flipping a comparison operator does not change the output on the specific inputs the examples chose. The mutant survives. Coverage reports ninety-five percent. Constraint strength in the mutated region is zero.

// Original
if (order.subtotal >= 50.dollars()) return order.subtotal * 0.10m;

// Mutant (flip >= to >)
if (order.subtotal > 50.dollars()) return order.subtotal * 0.10m;

An example suite that tests 60.dollars(), 75.dollars(), and 100.dollars() sees no behavior change under the mutation. The example at exactly 50.dollars(), if it exists, catches it. More often that example does not exist. A property over anyOrderWithSubtotalAbove(50.dollars()) with a boundary-aware generator eventually draws the fifty-dollar case and catches the mutant.

Mutation testing is the metric that survives agent-paired authorship. Coverage was already a weak signal under human authorship. Under agent authorship it becomes decorative, because the same optimizer writes both sides of the “did the test catch anything” question. A mutation score, run on infrastructure the coding agent cannot configure, surfaces the arrangement where coverage is high and constraint is low.

Coverage measures what ran. Mutation measures what would have been noticed. Only the second is a test suite’s job.

The Guardrail Is One Layer in the Stack, Not the Stack

The property closes a specific attack. It does not close every attack. An agent that deletes the property, narrows the generator, or modifies the assertion library so the property’s failure becomes a pass is operating outside the layer the property defends. Those attacks are the subject of tamper-resistant test design at the suite layer and the contract test at the integration seam. The property sits inside that enforcement stack. It does not replace it.

The stack has shape. Prose captures intent. Example tests pin the named scenarios in domain vocabulary. Property tests pin the invariants over the input space. Contracts pin the seams between modules with different owners. Mutation tests pin whether the lower layers actually constrain anything. Governed model judgment with independent inputs gates the final review. Production is the verdict on whether the system is actually right. Every layer closes a class of failure the layer below could not.

The property’s specific contribution to that stack is to be the layer the agent cannot satisfy by enumeration. The example layer is enumerable. The contract layer is a document owned by a different author. The mutation layer is a verifier on infrastructure outside the writable tree. The property layer is the one layer whose assertion lives in the agent’s own tree, written in the agent’s own session, and still resists satisfaction by memorizing the inputs the file names.

A codebase that treats the property layer as sufficient on its own repeats the mistake of the test-pyramid-only codebases. A codebase that drops the property layer because the contract tests caught the last three bugs repeats the inverse mistake and lets the input-space region go unverified. The layers compose. Removing any of them opens the attack the removed layer was closing.

The property is a guardrail. The stack is the road.

An example test the agent can read is an example the agent can hardcode past. A property the agent can read quantifies over inputs the agent did not see, and that gap, carefully maintained, is where the guardrail lives. The gap widens with a large generator space, a seed that varies across runs, a generator source the coding agent cannot narrow, and a mutation score checked by a system outside the agent’s writable tree. Pull any of those protections and the gap closes. Keep them in place and the lookup-table attack costs more than the honest fix.

Examples pin the vocabulary. Properties pin the invariants. The fuzzer searches. The shrinker reports. Mutation testing checks the lower layers for the thing coverage was never measuring. All of it sits inside the enforcement stack the team already owes the codebase, with production as the final verdict on whether the stack held.

A quantifier over inputs the agent has not seen is the thing a property does that no example can. That is why the property layer is worth building first.