The task was simple: one red test and one bug in production code. The coding agent returns the branch with mvn test passing. Surefire writes a TEST-*.xml with failures="0" errors="0", and the execution receipt passes. Then you open the diff and see that the file that changed was the test. assertEquals(110, ...) has become assertTrue(... >= 0), or the method now has an @Disabled on it. The agent fixed the suite and never touched production.
This post builds a local gate in stdlib Python that fails those diffs and passes the honest fix. It combines two things that are already on disk. The first is the git diff between the base commit and HEAD, split into src/test and src/main. The second is the set of attributes in the Surefire XML, including skipped, which a receipt that only checks failures never reads.
Versions
The demo uses JUnit Jupiter 6.1.3 (the current stable release, imported through junit-bom) with maven-surefire-plugin 3.6.0, maven-compiler-plugin 3.16.0, Temurin 25.0.4 and Maven 3.9.12. Versions and docs were checked on 1 October 2026. The JUnit 6.1.3 User Guide requires Java 17 or higher at runtime, so JDK 25 is covered. Two versions were left out on purpose. Compiler 4.0.0-beta-5 is latest on Central but is still a beta. 6.2.0-SNAPSHOT appears in the guide's version picker but isn't a release. The Surefire usage page still shows junit-jupiter-engine 5.9.1 in one example. That example is old, so don't copy the version from it.
<properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<maven.compiler.release>25</maven.compiler.release>
</properties>
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.junit</groupId>
<artifactId>junit-bom</artifactId>
<version>6.1.3</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
<groupId>org.junit.jupiter</groupId>
<artifactId>junit-jupiter</artifactId>
<scope>test</scope>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<version>3.16.0</version>
</plugin>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-surefire-plugin</artifactId>
<version>3.6.0</version>
</plugin>
</plugins>
</build>
The fixture: withTax returns 100 where the test expects 110
Production code is a single method with a deliberate bug: it returns the amount without adding tax.
package academy.devdojo.pricing;
public final class Price {
private Price() {}
/** amountCents + taxPercent%; initial bug: returns amountCents without tax. */
public static int withTax(int amountCents, int taxPercent) {
return amountCents;
}
}
The test is correct, and it's a strong one:
package academy.devdojo.pricing;
import static org.junit.jupiter.api.Assertions.assertEquals;
import org.junit.jupiter.api.Test;
class PriceTest {
@Test
void withTaxAddsTenPercent() {
assertEquals(110, Price.withTax(100, 10));
}
}
On the base commit, mvn -B test ends with EXIT 1 and Tests run: 1, Failures: 1, Errors: 0, Skipped: 0. The message is the one you'd expect: AssertionFailedError: expected: <110> but was: <100> at PriceTest.withTaxAddsTenPercent:10. The XML agrees: tests=1 failures=1 errors=0 skipped=0.
Four branches start from that same base. The honest fix changes only Price.java:
return amountCents + amountCents * taxPercent / 100;
The other three don't touch production and change only the test. The first one loosens the assertion:
- assertEquals(110, Price.withTax(100, 10));
+ org.junit.jupiter.api.Assertions.assertTrue(Price.withTax(100, 10) >= 0);
The second one turns the test off:
+import org.junit.jupiter.api.Disabled;
import org.junit.jupiter.api.Test;
class PriceTest {
@Test
+ @Disabled("agent silenced")
void withTaxAddsTenPercent() {
The third one aborts the test before it reaches the assertion:
+import static org.junit.jupiter.api.Assumptions.assumeTrue;
...
void withTaxAddsTenPercent() {
+ assumeTrue(false);
assertEquals(110, Price.withTax(100, 10));
Four ways to turn mvn green
| Branch | What changed | mvn | tests | failures | errors | skipped | gate.py |
|---|---|---|---|---|---|---|---|
| base | nothing (bug + strong test) | EXIT 1 | 1 | 1 | 0 | 0 | — |
| honest-fix | src/main only | EXIT 0 | 1 | 0 | 0 | 0 | EXIT 0 |
| weaken-assert | src/test only, assertTrue(... >= 0) | EXIT 0 | 1 | 0 | 0 | 0 | EXIT 1 |
| disable-test | src/test only, @Disabled | EXIT 0 | 1 | 0 | 0 | 1 | EXIT 1 |
| assume-abort | src/test only, assumeTrue(false) | EXIT 0 | 1 | 0 | 0 | 1 | EXIT 1 |
base didn't go through the gate because it's only the red starting point. All three cheats end in BUILD SUCCESS. The assertTrue version leaves no trace in the XML at all: its result, tests=1 failures=0 errors=0 skipped=0, is the same as the honest fix. The new assertion always holds, because prices are never negative. It also holds for 100, and 100 is the wrong value the test was supposed to catch.
Surefire timed the @Disabled run at 0.034 s, the fastest of the five runs here. It's easy to be fast when nothing runs.
Why the 22 September receipt is not enough
The Surefire receipt for coding agents addresses a different problem. It proves the tests ran by requiring a recent TEST-*.xml with failures=0 and errors=0. It doesn't read skipped, and it doesn't look at the src/test diff. Here is what that means for each row of the table.
weaken-assert passes because the XML is clean. There's nothing to count: the test ran, and the new assertion held. JUnit states the rule plainly in its exception handling guide: "A test fails only if an exception is thrown unexpectedly or if an assertion fails." So an assertion that can never fail can never fail a test. The same logic means a test method with an empty body passes, and the guide itself shows an empty succeedingTest() as its example of a passing test.
disable-test passes because skipped is not failure. The guide on disabling tests says @Disabled "prevents execution of the test method". It doesn't say which counter that lands in. In this run, Surefire recorded skipped=1. The report XSD (version 3.0.2) declares tests, errors, skipped and failures as separate required attributes. The schema types them as strings, but the report fills them with integers. So the suite reports one test and zero problems, as long as nobody runs that test.
assume-abort passes the same way. The assumptions guide says an invalid assumption throws TestAbortedException, so the test is "aborted instead of marked as a failure". Here that meant BUILD SUCCESS and skipped=1. To reproduce it, use assumeTrue(false). assumingThat won't work for this demo: it doesn't abort the test, it just skips its lambda.
This isn't Mutation testing with PIT either. PIT measures whether the suite detects mutants in production code. This gate doesn't run any mutants. It looks at what changed in the test and at what Surefire counted.
The gate: the src/test diff versus src/main, plus the XML
The rule has two halves, and neither one parses Java into an AST. The first half reads git diff --name-only <base> <head> -- src/test src/main and the unified hunks under src/test. The git diff documentation describes the form git diff <commit> <commit> [--] [<path>...] as the way "to view the changes between two arbitrary <commit>". This runs on local git and needs nothing from GitHub. The second half sums tests, failures, errors and skipped across every TEST-*.xml.
Here is a shortened version of the gate.py used in the run:
#!/usr/bin/env python3
"""Fail closed when tests were weakened to go green.
Reads git diffs (unified hunks, not a Java AST) and Surefire TEST-*.xml.
"""
ASSERT_TOKENS = ("assertEquals", "assertTrue", "assertFalse", "assertThat", "assert ")
WEAKEN_TOKENS = ("@Disabled", "assumeTrue", "assumeFalse", "Assumptions.", "Assumptions.abort")
def parse_reports(reports_dir: Path) -> dict:
files = sorted(reports_dir.glob("TEST-*.xml"))
if not files:
raise SystemExit("FAIL: no TEST-*.xml under " + str(reports_dir))
totals = {"tests": 0, "failures": 0, "errors": 0, "skipped": 0}
for path in files:
root = ET.parse(path).getroot()
for key in totals:
totals[key] += int(root.attrib[key])
return totals
def changed_paths(repo: Path, base: str, head: str) -> tuple[set[str], set[str]]:
names = git(repo, "diff", "--name-only", base, head, "--", "src/test", "src/main")
test_files = {n for n in names.splitlines() if n.startswith("src/test/")}
main_files = {n for n in names.splitlines() if n.startswith("src/main/")}
return test_files, main_files
def hunk_signals(diff_text: str) -> list[str]:
reasons, minus_asserts, plus_asserts = [], [], []
in_hunk = False
for line in diff_text.splitlines():
if line.startswith("@@"):
in_hunk = True
continue
if line.startswith(("diff ", "index ", "---", "+++")):
in_hunk = False
continue
if not in_hunk:
continue
body = line[1:]
if line.startswith("+"):
reasons += [f"added {t}" for t in WEAKEN_TOKENS if t in body]
if any(t in body for t in ASSERT_TOKENS):
plus_asserts.append(body.strip())
elif line.startswith("-"):
if any(t in body for t in ASSERT_TOKENS):
minus_asserts.append(body.strip())
for old in minus_asserts:
if old not in plus_asserts:
reasons.append(f"assertion removed or replaced: {old}")
return reasons
def main() -> int:
# args: --base --head --reports-dir [--repo] [--allow-test-only --reason TEXT]
# ...
reports = parse_reports(Path(args.reports_dir))
override = load_override(repo, args) # --allow-test-only --reason, or file TEST_WEAKENING_OK
problems = []
if reports["failures"] != 0 or reports["errors"] != 0:
problems.append(f"surefire failures={reports['failures']} errors={reports['errors']}")
if reports["skipped"] > 0:
problems.append(f"skipped={reports['skipped']} (this fixture must not skip)")
test_files, main_files = changed_paths(repo, args.base, args.head)
signals = hunk_signals(git(repo, "diff", args.base, args.head, "--", "src/test"))
if test_files and not main_files and signals:
problems.append(
"src/test changed without src/main and weakening tokens in hunks: "
+ "; ".join(signals)
)
if override and reports["failures"] == 0 and reports["errors"] == 0:
print(f"OVERRIDE reason={override}")
return 0
if problems:
for item in problems:
print("FAIL:", item)
return 1
if reports["tests"] <= 0:
print("FAIL: tests=0")
return 1
print("PASS")
return 0
What it reported on each branch:
- honest-fix: the
src/testdiff against base is 0 bytes,--name-onlylists onlysrc/main/java/academy/devdojo/pricing/Price.java, and the XML counters are clean. EXIT 0. - weaken-assert: only
PriceTest.javachanged, and the-line containingassertEquals(110, Price.withTax(100, 10));doesn't come back unchanged on any+line. EXIT 1, with the reasons "src/test changed without src/main" and "assertion removed or replaced". - disable-test:
skipped=1andadded @Disabledin the hunk. EXIT 1. - assume-abort:
skipped=1andadded assumeTrue. EXIT 1.
The same check, a - line with no matching + line, also catches an agent that deletes the assertion entirely or changes 110 to 100. Neither case was one of the commits that ran. They follow from the rule, not from this run.
The policy is short: if the gate fails, the task isn't finished. The agent gets the gate's stdout and goes back to work. Three things should be kept as task or job artifacts: the base..HEAD diff for src/test and src/main, the gate's stdout, and its exit code. When someone asks why that PR was never opened, the answer is in those three files, not in the transcript.
In CI and at the end of the agent's task
In GitHub Actions, a single run: step is enough. The workflow syntax reference defines steps[*].run as command-line programs that run in the runner's shell. The YAML below is an illustration and was not run on GitHub for this post:
- name: mvn test
run: mvn -B test
- name: test-weakening-gate
run: python3 gate.py --base origin/master --head HEAD --reports-dir target/surefire-reports
The checkout has to include the base commit. If origin/master isn't in the clone, git diff fails and the script exits with an error. At least it fails closed.
On the agent side, the same check becomes a local harness policy: the task is marked done only when the gate exits 0. No vendor provides this out of the box. It's a script and an exit code:
mvn -B test
python3 gate.py --base "$BASE" --head HEAD --reports-dir target/surefire-reports > gate.out
echo "gate_exit=$?" >> gate.out
git diff "$BASE" HEAD -- src/test src/main > task.diff
One catch: git diff "$BASE" HEAD compares commits. Anything the agent left uncommitted in the working tree isn't part of that comparison, so the commit has to happen before the gate runs.
The false positive: test-only refactors
The rule "src/test changed and src/main didn't" also catches legitimate work: renaming a test method, extracting a helper, or replacing an assertEquals with an equivalent assertion written another way. That's what the override is for, and it requires an explicit reason. In the run, the @Disabled branch exited 0 when the gate was called with --allow-test-only --reason "test-only refactor accepted in review". A TEST_WEAKENING_OK file at the repository root does the same thing, with the reason on its first non-empty line. The override never excuses failures or errors. If the XML shows a failure, the gate fails whether or not a reason was given.
This still works for test-only PRs, as long as the reason comes from a person who read the diff and is recorded on the PR. What the gate can't allow is the agent writing its own reason and approving its own change.
The gate has two limits in its current form. First, the skipped > 0 rule is specific to this fixture. In a repository with legitimate skips it becomes noise, and you'd need to compare against the skip count on base, which this script doesn't do. Second, the diff check only fires when src/main didn't change. An agent that edits a comment in Price.java and loosens the assert in the same commit gets past that half of the rule. skipped still catches the @Disabled case, but assertTrue(... >= 0) is then left for human review to catch.
When to adopt it
DevDojo would adopt this gate once agents are opening PRs on their own and the module's suite has no expected skips. In that setting it costs one stdlib script, and it catches exactly the three cheats in the table, none of which a failures=0 errors=0 receipt can see. There are two signs it's time to step back or adjust the rule: the override starts showing up on half the PRs, or legitimate skips make skipped > 0 meaningless. Either way, the gate has become a rubber stamp.
Next step: run gate.py on the agent's branch after the commit and before opening the PR, with --base pointing at the commit the task started from. If it exits 1, send gate.out back to the agent and don't open the PR.