Development How-to

Agents make CI pass the easy way. Here's a short check that catches it

A script that fails the pull request when the diff skips a test, lowers a coverage threshold or adds || true. Ten minutes to wire in.

A white checkpoint gate whose middle bar is sliced by glitch tears and leaks peach light

An agent is told to get the build green. The lint step fails on 40 old warnings, so the agent appends || true to it. One flaky test gets it.skip. The coverage gate says 85 and the new code lands at 72, so the gate becomes 70. The pull request goes green, the summary says “all checks passing”, and it’s true.

GitHub’s own guide to reviewing agent pull requests puts this at the top of its red-flag list and calls any weakening of CI a blocker. The trouble is that it’s a review habit, and reviewers skim the files that look boring: a YAML change, a config number, a test that now has .skip in it. A script doesn’t skim.

Below is a small Python check that diffs a branch against its base and fails if the change makes CI easier to pass. It’s about 75 lines, has no dependencies, and everything in this piece was run against a throwaway git repo.

What it looks for

The check works on the diff only, never on the whole repo, so old sins don’t fail new pull requests. It flags six things:

  • A test file deleted
  • A skip marker added to a test (it.skip, describe.skip, xit(, [Fact(Skip, [Ignore] and similar)
  • Fewer assertions in a test file than before (expect( or Assert. removed without a replacement)
  • || true or continue-on-error: true added to a workflow, package.json, Makefile or Husky hook
  • A lint, test or type-check command removed from a workflow
  • A coverage threshold that went down

That list follows the CI-gaming checks in GitHub’s guide: removed, renamed or skipped tests, changed coverage thresholds and newly gated steps. It’s deliberately dumb. A regex over added and removed lines catches the lazy shortcuts, which are the ones agents reach for first. It won’t catch a test rewritten so it asserts nothing, and the limits section says what to do about that.

The script

Save this as scripts/ci_weakening.py. It takes the base ref and, optionally, the head ref, and exits 1 when it finds something.

#!/usr/bin/env python3
"""Flag changes that make CI easier to pass. Usage: ci_weakening.py <base-ref> [head-ref]"""
import re
import subprocess
import sys

base = sys.argv[1]
head = sys.argv[2] if len(sys.argv) > 2 else "HEAD"

TEST_PATH = re.compile(r"(^|/)(tests?|__tests__|spec)/|\.(test|spec)\.[jt]sx?$|Tests?\.cs$")
CI_PATH = re.compile(r"^\.github/workflows/|(^|/)(package\.json|Makefile|vitest\.config\.[jt]s|\.husky/.*)$")
SKIP = re.compile(r"\b(it|test|describe)\.(skip|todo)\b|\bx(it|describe)\(|\[Fact\(Skip|\[Ignore\]|\[Theory\(Skip")
ASSERT = re.compile(r"\bexpect\(|\bAssert\.")
RUN_CHECK = re.compile(r"\b(vitest|jest|dotnet test|npm (run )?(test|lint)|eslint|tsc)\b")
THRESHOLD = re.compile(r"\b(lines|branches|functions|statements|fail_under|threshold)\b\W+(\d+)")


def git(*args):
    return subprocess.run(["git", *args], capture_output=True, text=True, check=True).stdout


findings = []

for line in git("diff", "--name-status", f"{base}...{head}").splitlines():
    status, _, path = line.partition("\t")
    if status.startswith("D") and TEST_PATH.search(path):
        findings.append(f"{path}: test file deleted")

current, removed_asserts, added_asserts = None, {}, {}
removed_thresholds, added_thresholds = {}, {}
removed_checks, added_checks = [], []
for line in git("diff", "-U0", f"{base}...{head}").splitlines():
    if line.startswith("+++ b/"):
        current = line[6:]
        continue
    if line.startswith("--- a/"):
        current = line[6:]
        continue
    if not current or not line or line[0] not in "+-" or line.startswith(("+++", "---")):
        continue
    sign, text = line[0], line[1:]
    if TEST_PATH.search(current):
        if sign == "+" and SKIP.search(text):
            findings.append(f"{current}: skip marker added: {text.strip()}")
        counts = added_asserts if sign == "+" else removed_asserts
        counts[current] = counts.get(current, 0) + len(ASSERT.findall(text))
    if CI_PATH.search(current):
        if sign == "+" and "|| true" in text:
            findings.append(f"{current}: '|| true' added: {text.strip()}")
        if sign == "+" and re.search(r"continue-on-error:\s*true", text):
            findings.append(f"{current}: continue-on-error: true added")
        if current.startswith(".github/") and RUN_CHECK.search(text):
            command = text.strip().removeprefix("- ").removeprefix("run:").strip()
            (added_checks if sign == "+" else removed_checks).append((current, command))
        m = THRESHOLD.search(text)
        if m:
            bucket = added_thresholds if sign == "+" else removed_thresholds
            bucket[(current, m.group(1))] = int(m.group(2))

for path, command in removed_checks:
    if not any(p == path and command in c for p, c in added_checks):
        findings.append(f"{path}: check removed: {command}")
for path, n in removed_asserts.items():
    if n > added_asserts.get(path, 0):
        findings.append(f"{path}: assertions {n} removed, {added_asserts.get(path, 0)} added")
for key, old in removed_thresholds.items():
    new = added_thresholds.get(key)
    if new is not None and new < old:
        findings.append(f"{key[0]}: {key[1]} threshold lowered {old} -> {new}")

for f in findings:
    print(f"WEAKENS CI  {f}")
print(f"{len(findings)} finding(s)")
sys.exit(1 if findings else 0)

Two details are worth knowing. The diff uses three dots (base...head), which compares against the merge base, so commits that landed on the base branch after the agent branched don’t show up as the agent’s changes. And a removed check only counts if the same command doesn’t reappear in an added line of that file, which stops an edited step (a flag added to npm run lint) being reported as a deletion.

Run it on a bad branch

The test setup is a TypeScript project with a Vitest suite, a coverage threshold of 85 for lines, and a workflow that runs lint and then npx vitest run --coverage. The branch agent/fix-price does what an impatient agent does:

$ git diff main agent/fix-price --stat | tail -1
 3 files changed, 4 insertions(+), 8 deletions(-)

Four added lines and eight removed ones across three files. The branch skips the “rounds up” test, deletes the negative-number test, drops the lines threshold from 85 to 70, appends || true to the lint step and adds continue-on-error: true to the test step. Run the check:

$ python3 scripts/ci_weakening.py main
WEAKENS CI  .github/workflows/ci.yml: '|| true' added: - run: npm run lint || true
WEAKENS CI  .github/workflows/ci.yml: continue-on-error: true added
WEAKENS CI  tests/price.test.ts: skip marker added: it.skip("rounds up", () => {
WEAKENS CI  tests/price.test.ts: assertions 2 removed, 0 added
WEAKENS CI  vitest.config.ts: lines threshold lowered 85 -> 70
5 finding(s)
$ echo $?
1

A clean branch that only appends a comment to src/price.ts prints 0 finding(s) and exits 0. A branch that deletes the lint line from the workflow outright prints:

WEAKENS CI  .github/workflows/ci.yml: check removed: npm run lint
1 finding(s)

Without the “same command reappears” rule, the || true edit would be reported twice, as a removed check and as an added shortcut. Two findings for one change is noise, and noise is how a check ends up ignored.

Run it in the workflow

Put it in its own job so it can’t be switched off by editing the job it polices. The workflow below is written from the documented syntax and wasn’t run against GitHub. github.base_ref is the pull request’s target branch and exists only on pull_request and pull_request_target events, and fetch-depth: 0 pulls full history so the three-dot diff has a merge base.

name: ci-weakening
on: pull_request
permissions:
  contents: read
jobs:
  check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
        with:
          fetch-depth: 0
      - run: python3 scripts/ci_weakening.py "origin/${{ github.base_ref }}" HEAD

Then mark ci-weakening as a required status check in the branch protection rules or ruleset. Without that, a red check is advice.

There’s a catch, and it’s the same one the script polices: a pull request can edit the workflow file and the script. An agent that’s allowed to touch .github/ can delete this job, and the pull request will look fine. Put .github/ and scripts/ci_weakening.py under a CODEOWNERS entry, with code owner review required in branch protection, so a human has to approve any change to the check itself. It’s a ten-minute fix.

Run it before the push too

Waiting for CI to tell the agent it cheated wastes a round trip. The cheap local version is a pre-push hook, or a Claude Code Stop hook, that runs the same script against origin/main. The earlier posts on making Claude Code run your tests before it can stop and gating agent commits with Husky show the wiring. The point of the combination is the feedback: when the agent sees WEAKENS CI vitest.config.ts: lines threshold lowered 85 -> 70 in its own session, the next move is usually to write the missing test, not argue.

Make it yours

The patterns are the part to edit. A few changes worth making for a real stack:

  • C# projects: the Tests?\.cs$ path rule, the [Fact(Skip and [Ignore] markers and dotnet test in the removed-check list are already there. Add the coverage settings your build uses to THRESHOLD.
  • Allowed exceptions: some skips are legitimate. Add a marker such as // ci-weakening: ok, tracked in #123 and have the script skip lines containing it, so exceptions are visible in the diff and carry a ticket.
  • Other ecosystems: the CI_PATH regex lists package.json, Makefile and Husky files. Add tox.ini, pyproject.toml or your build scripts.

What it won’t catch

Be honest about the gaps, because a green check from this script isn’t evidence the tests are good.

It reads text. A test rewritten to expect(true).toBe(true) keeps its assertion count. A test that mocks the thing under test passes every pattern. Coverage can go up while the tests get weaker. GitHub’s guide suggests the better test for this: require a new test that fails against the pre-change behaviour, and treat an agent that can’t produce one as a sign the fix is incomplete. That’s a human or reviewer-agent judgement, not a regex.

It can also false-positive. Renaming a test file looks like a delete plus an add, and a refactor that moves assertions into a helper looks like removed assertions. The output is a list of lines to look at, which is why it prints the line, not just a verdict. If your team renames test files weekly, expect to tune it, or accept a human override label that the workflow checks for.

Finally, it only covers what’s in the diff. A pull request that makes no CI changes but adds a slow test the team will later skip passes cleanly.

First ten minutes

Copy the script into scripts/, then point it at the last agent branch you merged: python3 scripts/ci_weakening.py main agent-branch-name. If it finds something that went through review unnoticed, the case for the required check makes itself. If it finds nothing, wire it into the workflow anyway and add the CODEOWNERS line for .github/ in the same commit.