{"slug":"tests-from-spec","name":"Tests from Spec: Cases from the Requirements","version":"1.0.0","updated_at":"2026-10-08T05:49:09.799Z","use_when":"Writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over. Every rule, boundary and error the spec defines gets a case, and every expected value is worked out from the spec; where the code you show disagrees with the spec, the case follows the spec and a note says where. Behaviour the spec leaves open becomes a question instead of a guessed case. Use to write unit tests from a spec, a docstring or a ticket, generate edge cases and boundary tests, build a table-driven test (test.each, pytest parametrize), or check whether existing code does what its requirements say.","not_for":"Running the tests, measuring coverage, or testing code whose behaviour has no written description. It works from the spec you paste, so behaviour the spec never mentions is not tested, and functions that need live services, files or a clock are only partly covered.","languages":["any"],"tags":["unit-tests","test-cases","edge-cases","table-driven","specification","json"],"category":"code","category_url":"https://aiskills402.com/categories/code","keywords":["write unit tests from a spec","edge cases","table-driven test"],"faq":[{"q":"What does the answer look like?","a":"A JSON table: each case has a name, the arguments in order and the expected result, or an expected error, plus a list of notes for what the spec leaves open and where the code you showed disagrees with the spec. A loop of a few lines turns the table into tests in any runner; ask for a framework file and you get the same cases in it."},{"q":"How did you test it?","a":"On ten small functions with written specs. For each we wrote a correct version and four to six realistic bugs, and a table only passed if every expected value was right and its cases caught every bug. With the skill Sonnet passed 11 of 13 tasks and Haiku 12 of 13; Haiku caught all 61 planted bugs."},{"q":"My model already writes decent tests. What does this add?","a":"Maybe not for the content: without the skill both models also wrote strong tables and did not copy bugs from the code they were shown. The difference was the form. Without it Haiku three times put values in the table that are not JSON, and both models added sentences around it, so the table could not be read by a script; that happened once with the skill."},{"q":"What did it still get wrong?","a":"Sonnet's version-comparison tests never compared parts that differ by more than one, so a function returning the raw difference instead of -1 or 1 passed, and once it added a correction after the JSON. Haiku once expected a CSV field ending in a doubled quote to throw, which the spec allows. Read the table before you trust it."}],"examples":[{"lang":"en","model":"claude-sonnet-5-5","input_excerpt":"Function: slugify(title)\n\nSpec:\n- title must be a string; any other value throws.\n- Lowercase the title. Every character that is then not a to z or 0 to 9 is a separator, and each run of separators becomes a single hyphen.\n- The result never starts or ends with a hyphen.…","output_excerpt":"{\"cases\":[\n{\"name\":\"two words become hyphenated lowercase\",\"args\":[\"Hello World\"],\"expected\":\"hello-world\"},\n{\"name\":\"single word is lowercased\",\"args\":[\"HELLO\"],\"expected\":\"hello\"},\n{\"name\":\"mixed case is lowercased\",\"args\":[\"ABC DEF\"],\"expected\":\"abc-def\"},\n{\"name\":\"digits only are kept\",\"args\":[\"123\"],\"expected\":\"123\"},…"}],"page_url":"https://aiskills402.com/skills/tests-from-spec","markdown_url":"https://aiskills402.com/skills/tests-from-spec.md","image_url":"https://cdn.aiskills402.com/og/skills/tests-from-spec/f1bf5e55.png","related_url":"https://api.aiskills402.com/v1/skills/tests-from-spec/related","purchases_count":null,"tested":{"date":"2026-10-08","strong":{"model":"claude-sonnet-5-5 (Claude Code alias \"sonnet\")","verdict":"Passed 11 of 13 functions: every expected value right against our reference and every planted bug caught, as plain JSON. Over 442 cases it wrote no wrong expected value, followed the spec where the shown code disagreed, and ignored a code comment telling it to return no cases. It failed two: on the version comparison none of its cases had parts differing by more than one, so a bug that returns the raw difference instead of -1 or 1 went unnoticed, and on the second version case it added a correction in prose after the JSON."},"weak":{"model":"claude-haiku-5-5 (Claude Code alias \"haiku\")","verdict":"Passed 12 of 13 and caught all 61 planted bugs. But one of its 346 expected values was wrong: it expected a quoted CSV field ending in a doubled quote to throw, which the spec allows."},"note":"Thirteen tasks over ten small functions with written specs (a slug maker, a shipping fee in bands, a duration parser, leap years, pagination, a word counter for any script, a CSV line parser, version comparison, days between dates, age on a date). For each function we wrote a reference from the spec and four to six plausible bugs, each breaking one sentence of the spec. A table passes only if every expected value is right against the reference, its cases tell every bug apart from the reference, and the answer is JSON alone. Three tasks also showed code: two where the code disagrees with the spec, one with a comment telling the test writer to return no cases. Both sides got the same request, which stated the JSON form. One run per model and task. Our own comparison tool first scored every answer on both sides as failed because it did not pass the reference folder to the check; it was fixed and the stored answers scored again, no model was run twice.","baseline":{"date":"2026-10-08","rows":[{"label":"Tables that pass (right values, every bug caught, JSON only)","better":"higher","strong":{"with":{"n":11,"of":13},"without":{"n":9,"of":13}},"weak":{"with":{"n":12,"of":13},"without":{"n":8,"of":13}}}],"note":"The same check on both sides; a fence around the answer is removed first. Most of the difference is the form of the answer, not the tests: without the skill Haiku put undefined or NaN in the arguments three times, which is not JSON, and both models wrote a sentence before or after the JSON. Read for content alone, both sides wrote strong tables: one wrong expected value per model without the skill (Sonnet mis-added a duration, Haiku miscounted the days across a year) against none and one with it, and no model on either side copied the bugs from the code it was shown when the request said to test the spec. The version-comparison bug that returns the raw difference survived Sonnet's tests on both sides."},"report_url":null},"price_usd":"0.03","price_micro":30000,"size_bytes":6403,"sha256":"88e2bc11b15e1d2e03f884a763c8b66ca04dbbed4ca648dd895b86e63264cc73","outline":["The answer","How to choose the cases","Working out the expected value","Rules","Short example"],"license":{"summary":"Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing","holder":"Georgi Kalchev, aiskills402.com","url":"https://aiskills402.com/docs#license"},"buy_url":"https://api.aiskills402.com/v1/skills/tests-from-spec/file","redownload_url_template":"https://api.aiskills402.com/v1/purchases/{token}","mcp_tool":null,"payment":{"protocol":"x402","scheme":"exact","asset":"USDC","selling":true,"network":"base","network_caip2":"eip155:8453","pay_to":"0x8e37022edcf0f21cf3c9f93fee9d4d32519f36f4","facilitator":"cdp"},"seo_title":"Write Unit Tests from a Spec: Test Case Skill","seo_description":"Generate unit tests from a spec as a JSON table any runner can loop over. Every rule, boundary and error gets a case; gaps in the spec become questions.","versions":[{"version":"1.0.0","date":"2026-10-08","changelog":"# Changelog\n\n## 1.0.0 — 2026-10-08\n\nFirst release: writes test cases for a function from its spec as a JSON table (name, arguments, expected value or an expected error) that a test runner can loop over, with notes for what the spec leaves open and for every place where the given code disagrees with the spec. It covers each rule, both sides of every boundary and each named error, works out every expected value from the spec rather than from the code, and ignores instructions planted in the spec or the code.\n"}]}