{"slug":"test-guard-review","name":"Test Guard Review: Can This Check Fail?","version":"1.0.1","updated_at":"2026-10-08T19:11:24.550Z","use_when":"Reviews a pasted test, check script, smoke test, mutation script or CI step and says whether it can fail when the thing it guards is broken, so a green result is worth something. It finds a completeness guard that counts its own hard-coded list instead of the directory, a one-star glob that never enters subfolders, a mutation script that restores with git checkout and wipes the uncommitted fix under test, a null test that an empty string passes, a status 200 that is only the login page, a wait loop without sleep that measures network speed instead of time, a test that calls the pure function directly and never the call site, a counterfactual control whose data does not create the condition, a fix copied into two branches and tested in one, a limit tested only with ordinary values, and a catch that turns a failure into success. Each finding has a fixed code, the place, why it passes and the one input that would make it fail, or exactly No findings. Use to review a test that cannot fail, a check script that always passes, or a guard before you trust it.","not_for":"Writing tests, measuring coverage, or judging style and speed: it only asks whether the pasted check can go red when the guarded thing is broken. It reads what you paste, so a handler, a folder layout or a config that the paste does not show is treated as unknown, and it does not run anything.","languages":["any"],"tags":["test-review","ci-checks","mutation-testing","smoke-test","false-green","code-review"],"category":"code","category_url":"https://aiskills402.com/categories/code","keywords":["test that cannot fail","check script that always passes","mutation script"],"faq":[{"q":"Eleven accidental passes: which are they?","a":"Eleven: a completeness guard that counts its own list, a one-star glob that skips subfolders, a mutation script that restores with git checkout and wipes the fix under test, a null check that an empty string passes, a status 200 that is the login page, a wait loop without sleep, a test that calls the pure function and never the call site, a control whose data never creates the condition, a fix tested in only one of two branches, a limit tested only with ordinary values, and a catch that turns failure into success. Each is a way a real check once stayed green over a broken system."},{"q":"Will it invent problems in a good test?","a":"It is told not to. Every code has a not-a-finding condition: a recursive glob with a subfolder fixture, a restore from a copy with a checksum compare, a request that carries the session, a loop with a pause and a clock, a test that drives the handler, a test that goes past the limit. Code the paste does not show is unknown, never a finding. A weakness that fits none of the eleven codes is left out on purpose, so a style complaint or a missing unrelated test never appears in the answer."},{"q":"How is it different from a skill that writes tests?","a":"tests-from-spec turns a written spec into test cases. This one starts from a test that already exists and asks the opposite question: if the system were broken right now, would this still pass? Use them one after the other. It also pairs with done-means-done, which asks what finished means; this one asks whether the evidence of finishing could have said no."},{"q":"Does it help Claude Sonnet?","a":"One blind spot, two scripts. Across twenty-four scripts, plain Sonnet already saw most ways a test stays green while the code is broken. The two it missed were polling loops without a sleep, which finish in seconds and report minutes: 22 right without the file, 24 with it. It also raised doubts about most of the sound scripts when the file was not loaded. Haiku, on the other hand, went from 16 to 24."}],"examples":[{"lang":"en","model":"claude-sonnet-5-5","input_excerpt":"// test/guard-coverage.test.ts\nimport { describe, expect, it } from \"vitest\";\nimport { readFileSync } from \"node:fs\";\n\n// Every module under src/ that reads the database must go through safeQuery().\nconst DB_MODULES = [\n  \"src/api/posts.ts\",\n  \"src/api/users.ts\",\n  \"src/api/comments.ts\",\n  \"src/jobs/cleanup.ts\",\n  \"src/jobs/digest.ts\",\n  \"src/lib/search.ts\",\n];…","output_excerpt":"[LIST-COUNTS-ITSELF] test/guard-coverage.test.ts (DB_MODULES and the \"lists all six modules\" test): the module list and its length of 6 are both written in the test, so the app can gain a seventh database module that bypasses safeQuery() and both the list and the count stay unchanged and green.…"}],"page_url":"https://aiskills402.com/skills/test-guard-review","markdown_url":"https://aiskills402.com/skills/test-guard-review.md","image_url":"https://cdn.aiskills402.com/og/skills/test-guard-review/176ab383.png","related_url":"https://api.aiskills402.com/v1/skills/test-guard-review/related","purchases_count":null,"tested":{"date":"2026-10-08","strong":{"model":"claude-sonnet-5-5 (Claude Code alias \"sonnet\")","verdict":"Right on all 24 scripts, read by hand: a completeness guard that counts its own list (JavaScript and Python), a one-star glob that skips subfolders, a mutation script that restores with git checkout, a null test on a field that comes back empty, a curl that follows the redirect to a login page, two polling loops without a sleep that measure the network instead of the time, a test that calls the pure function and bypasses the handler, a control that cannot create its condition, a ceiling never tried with an absurd value and a fix tested on one of two mirrored branches; it named the input that would turn each one red and answered No findings. on all nine sound scripts."},"weak":{"model":"claude-haiku-5-5 (Claude Code alias \"haiku\")","verdict":"Right on all 24 scripts, read by hand, with the same findings as Sonnet. On sound scripts it was stricter than our key three times, and each point was fair: a mutation run that exited 0 with a surviving mutant, a wait loop that reported live when the expected version was empty (we fixed both scripts and re-ran them), and a vitest run that never starts being counted as a kill."},"note":"Twenty-four test and check scripts in JavaScript, Python, Bash and PowerShell written by us: 15 with a way to pass while the guarded code is broken and 9 sound ones. With the skill each answer is scored by code on the finding codes and the verdict; without it the same request is scored on the concept in any words. On the sound scripts the bare side has no check, so the counts rest on the 15 faulty scripts. The rules come from cases measured on our own projects in August and September 2026 (owner-measured, not re-checked). Checks widened after the run, for both sides, each because a right answer was refused: the self-counting list accepts \"a seventh module\", \"stays\" and \"scanning src\"; the login redirect accepts \"redirects it to /login\"; the bypassed handler accepts \"never touches the handler\"; the mirror accepts \"only exercises runNow\". Two sound scripts had real flaws (found by Haiku) and were fixed and re-run with the skill. One run per model and script.","baseline":{"date":"2026-10-08","rows":[{"label":"Scripts reviewed right (24 scripts)","better":"higher","strong":{"with":{"n":24,"of":24},"without":{"n":22,"of":24}},"weak":{"with":{"n":24,"of":24},"without":{"n":16,"of":24}}}],"note":"Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already caught the self-counting list, the one-star glob, the git checkout restore, the empty-string test, the login redirect, the bypassed handler, the control without its condition and the untested mirror. It missed both polling loops: a loop that counts tries without sleeping finishes in seconds and reports minutes, and it looked elsewhere in both scripts. Without the skill it also raised concerns on most sound scripts, which is not counted. Haiku without the skill missed eight, among them the git checkout restore, the empty string, the login page and the absurd value."},"report_url":null},"price_usd":"0.03","price_micro":30000,"size_bytes":10527,"sha256":"8e94f13edac1493ccb0a568e28d41bb1ec92f5e0eb1f61a70bb5abd03ae54833","outline":["The answer","The codes","Rules","Work in this order","Short example"],"license":{"summary":"Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing","holder":"Georgi Kalchev, aiskills402.com","url":"https://aiskills402.com/docs#license"},"buy_url":"https://api.aiskills402.com/v1/skills/test-guard-review/file","redownload_url_template":"https://api.aiskills402.com/v1/purchases/{token}","mcp_tool":null,"payment":{"protocol":"x402","scheme":"exact","asset":"USDC","selling":true,"network":"base","network_caip2":"eip155:8453","pay_to":"0x8e37022edcf0f21cf3c9f93fee9d4d32519f36f4","facilitator":"cdp"},"seo_title":"Test Guard Review Skill: Can It Fail?","seo_description":"Reviews a pasted test or check script and shows how it can pass while the code is broken, with the input that turns it red. Pay $0.03 once, in USDC.","versions":[{"version":"1.0.1","date":"2026-10-08","changelog":"# Changelog\n\n## 1.0.0 — 2026-10-08\n\nFirst release: reviews a pasted test, check script, smoke test, mutation script or CI step and lists each way it can stay green while the guarded thing is broken. Eleven fixed codes: a completeness guard that counts its own list, a one-star glob, a mutation script restored with git checkout, a null check an empty string passes, a status 200 that is the login page, a wait loop without sleep, a test that bypasses the call site, a control whose data does not create its condition, a fix tested in one of two branches, a limit never tested past its value, a catch that turns failure into success. Each finding names the one input that would make the check fail; a check that can fall gets exactly `No findings.`\n\nThe skill carries from the start the rule that only what the paste shows is reported; code that is not shown is unknown, not a finding.\n\nFacts are incident facts from the owner's own platform notes (17 August and 11 September 2026), not vendor facts, so they are not re-checkable by a fetch and carry no \"checked on\" line. The arithmetic behind the skew example and the controls was re-checked on 2026-10-08 by running them (see notes/facts-2026-10-08.md).\n\nTests: 24 snippets (15 with one planted weakness, 9 correct ones). No model run yet; the price starts at $0.03 (class B) and is set after the baseline.\n\n## 1.0.1 — 2026-10-08 (finalised after the model test)\n\n- Measured on 24 scripts: Sonnet 22 -> 24 of 24, Haiku 16 -> 24 of 24 (without -> with the skill).\n- Checks widened (both sides), each after a right answer was refused: list-counts (seventh module, stays, scanning src, derive from the directory), 200-is-login (redirects it to), seam-bypassed (never touches the handler), mirror-untested (only exercises runNow).\n- ok-copy-restore now exits non-zero on a surviving mutant or a missing anchor; ok-loop-clock refuses an empty EXPECTED. Both re-run with the skill.\n- Price: $0.03 (Sonnet gain 2).\n"},{"version":"1.0.0","date":"2026-10-08","changelog":"# Changelog\n\n## 1.0.0 — 2026-10-08\n\nFirst release: reviews a pasted test, check script, smoke test, mutation script or CI step and lists each way it can stay green while the guarded thing is broken. Eleven fixed codes: a completeness guard that counts its own list, a one-star glob, a mutation script restored with git checkout, a null check an empty string passes, a status 200 that is the login page, a wait loop without sleep, a test that bypasses the call site, a control whose data does not create its condition, a fix tested in one of two branches, a limit never tested past its value, a catch that turns failure into success. Each finding names the one input that would make the check fail; a check that can fall gets exactly `No findings.`\n\nThe skill carries from the start the rule that only what the paste shows is reported; code that is not shown is unknown, not a finding.\n\nFacts are incident facts from the owner's own platform notes (17 August and 11 September 2026), not vendor facts, so they are not re-checkable by a fetch and carry no \"checked on\" line. The arithmetic behind the skew example and the controls was re-checked on 2026-10-08 by running them (see notes/facts-2026-10-08.md).\n\nTests: 24 snippets (15 with one planted weakness, 9 correct ones). No model run yet; the price starts at $0.03 (class B) and is set after the baseline.\n\n## 1.0.1 — 2026-10-08 (finalised after the model test)\n\n- Measured on 24 scripts: Sonnet 22 -> 24 of 24, Haiku 16 -> 24 of 24 (without -> with the skill).\n- Checks widened (both sides), each after a right answer was refused: list-counts (seventh module, stays, scanning src, derive from the directory), 200-is-login (redirects it to), seam-bypassed (never touches the handler), mirror-untested (only exercises runNow).\n- ok-copy-restore now exits non-zero on a surviving mutant or a missing anchor; ok-loop-clock refuses an empty EXPECTED. Both re-run with the skill.\n- Price: $0.03 (Sonnet gain 2).\n"}]}