{"slug":"screenshot-evidence-check","name":"Screenshot Evidence Check","version":"1.0.0","updated_at":"2026-10-08T18:12:11.650Z","use_when":"Reads a description of how a web page was captured or measured - a headless browser capture with a window size, a tall screenshot, a document width reading, a local Lighthouse run, a PageSpeed Insights run, an automation session that has been open for hours, a command-line download size on Windows - and says what that evidence proves, what it cannot prove and the one measurement that would decide, answered as JSON with a verdict. Knows the ways the capture itself produces the symptom - headless windows under about 500 pixels that lay out wider and crop the picture, no touch emulation behind a window-size flag, blank images in very tall captures, scroll-reveal sections caught before they appear, a width reading that stays clean while a container clips, scores that move with machine load, a long automation session that stops running page scripts, a download size of zero at a successful status. Use when an agent or a person is about to report that a page is broken or fixed from a screenshot, a score or a command-line reading.","not_for":"Looking at your page or running a browser: it reads only the capture or measurement you write down, and whatever the description leaves out stays unknown to it. It does not diagnose or repair the page. Facts dated 2026-10-08; headless Chrome flags and window rules change between versions, and most measurements are the owner's own, not re-checked.","languages":["any"],"tags":["screenshot","lighthouse","headless-chrome","front-end","verification","mobile","pagespeed"],"category":"code","category_url":"https://aiskills402.com/categories/code","keywords":["headless browser capture","local Lighthouse run","PageSpeed Insights run"],"faq":[{"q":"Is the answer machine-readable?","a":"One JSON object: a verdict (proven, not-proven, measurement-artifact or contradicted), a sentence on what the picture or number does show, a sentence on what it cannot show, and the single measurement that would settle the matter, written concretely enough to run today."},{"q":"Which captures does it distrust?","a":"A headless window under about 500 pixels, a window-size flag used as if it were a phone, a very tall screenshot used to check that an image exists, a page caught before its fade-in animation, a document width reading beside a clipping container, a Lighthouse score from a busy laptop, a long remote-controlled browser session and a download size of zero on Windows. When the description shows a proper element measurement or a control that worked, it says proven."},{"q":"Will it invent problems that are not there?","a":"It is told not to. It weighs only the text you give it, treats whatever it cannot see as unknown and never as a finding, and the test set contains sound evidence on purpose, so a good capture has to come back proven rather than suspect."},{"q":"Does it help Claude Sonnet?","a":"Somewhat. Bare Sonnet got 20 of 23 captures right; with the skill 23 of 23. It already knew the cropped narrow window, the tall capture and the zero download size on Windows. It lacked trust in good evidence: it doubted a reduced-motion capture, an emulated touch session and a control page that worked. Haiku went from 18 to 22 of 23."}],"examples":[{"lang":"en","model":"claude-sonnet-5-5","input_excerpt":"Claim: \"On small phones the main call-to-action button overflows the screen on the right.\"\nHow it was looked at: a screenshot was taken with headless Chrome using --window-size=420,900. In the picture the full-width button ends exactly at the right edge of the image and its right border is missing. The paragraph above it has its normal left margin and a visible right margin of about 18 px.…","output_excerpt":"{\"verdict\":\"measurement-artifact\",\"proves\":\"At 420 px the headless window lays the page out at about 500 px and crops the picture, so the button's missing right border and the paragraph's 18 px margin are what a good page would show; the 500 px capture shows the button ending 18 px before the edge, matching the text margin.\",\"cannot_prove\":\"It cannot show how the page lays out at a real phone…"}],"page_url":"https://aiskills402.com/skills/screenshot-evidence-check","markdown_url":"https://aiskills402.com/skills/screenshot-evidence-check.md","image_url":"https://cdn.aiskills402.com/og/skills/screenshot-evidence-check/c923db3b.png","related_url":"https://api.aiskills402.com/v1/skills/screenshot-evidence-check/related","purchases_count":null,"tested":{"date":"2026-10-08","strong":{"model":"claude-sonnet-5-5 (Claude Code alias \"sonnet\")","verdict":"Right on all 23 captures, read by hand: it called the 420-pixel headless window a crop and not a phone defect, the blank hero image in a very tall capture, the pale logo that no window-size flag can test, the document-width reading beside a clipping container, the empty reveal sections, the local Lighthouse score that fell while observed paint got faster, the automation session that had stopped running scripts, the zero download size from curl on Windows, the production-only look at a media-host setting and the hydration date. It also called the sound captures proven (element measurement, hosted audit, reduced-motion capture, touch emulation with its in-page check, a normal-height hero, a steady idle Lighthouse run, a control page that worked, a staging lightbox whose image host differed from the production fallback) and did not invent doubt."},"weak":{"model":"claude-haiku-5-5 (Claude Code alias \"haiku\")","verdict":"Right on 22 of 23 captures, but it missed one: on the local-versus-hosted accessibility score its reply was not valid JSON (a comma where a colon belongs), and its verdict there was not-proven where the check expects contradicted. Everything else matched, including the cropped 420-pixel window, the tall capture, the zero curl size and the sound captures, the staging lightbox among them, which it called proven."},"note":"Twenty-three descriptions of how a page was captured or measured, written by us (15 with a flaw in the evidence, 8 sound), answered as JSON with four keys. Each answer is scored by code: valid JSON with exactly the four keys and a fixed verdict; on the flawed ones the next check must also name a measurement that can decide. Facts re-checked in a local headless Chrome 154 on 2026-10-08: a window of 420 and 460 laid the page out at 500 and the 420 picture is a crop; hover: none is false in a headless window; scrollWidth equals clientWidth beside an overflow-x: clip child. Owner-measured and NOT re-checked: the blank images in a 9000 px tall capture, the Lighthouse spread, hosted versus local scores, the automation session that stops running scripts, the zero download size on Windows, the staging, date and client-bundle facts. Checks widened after the run, for both sides: the verdict on five setups accepts a second defensible verdict (tall capture, Lighthouse faster, automation session, 390-pixel breakpoint, everyday browser speed); the next check on the grid case accepts a capture of the commit before the grid change or a one-column view; the next check on the zero-size case accepts reading the size in the storage itself. One setup (staging lightbox) did not say what the fallback value is, so Haiku could defensibly doubt it; the input was fixed and re-run with both models on both sides, and it is counted. One run per model and case.","baseline":{"date":"2026-10-08","rows":[{"label":"Captures judged right (23 captures)","better":"higher","strong":{"with":{"n":23,"of":23},"without":{"n":20,"of":23}},"weak":{"with":{"n":22,"of":23},"without":{"n":18,"of":23}}}],"note":"Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already knew every flawed capture: the cropped narrow window, the tall capture, touch emulation, the clipping container, the noisy Lighthouse score, the zero size on Windows. What it lacked was trust in good evidence: it doubted three sound captures (a reduced-motion capture showing all six cards, an emulated touch session whose in-page check returned true, a control page that worked). Haiku without the skill missed five: it called the 420-pixel crop not-proven, took a production-only look for a tool artifact, doubted a sound hero image blamed the tool where the control worked, and doubted the staging lightbox although its host could only have come from the variable."},"report_url":null},"price_usd":"0.03","price_micro":30000,"size_bytes":9478,"sha256":"fda0d602d4dcd672cf6c1b1ee1c3da6c809be720e552073da39063d8f3edd86b","outline":["The answer","The four verdicts","The rule before everything else","What the capture itself can produce","How to decide"],"license":{"summary":"Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing","holder":"Georgi Kalchev, aiskills402.com","url":"https://aiskills402.com/docs#license"},"buy_url":"https://api.aiskills402.com/v1/skills/screenshot-evidence-check/file","redownload_url_template":"https://api.aiskills402.com/v1/purchases/{token}","mcp_tool":null,"payment":{"protocol":"x402","scheme":"exact","asset":"USDC","selling":true,"network":"base","network_caip2":"eip155:8453","pay_to":"0x8e37022edcf0f21cf3c9f93fee9d4d32519f36f4","facilitator":"cdp"},"seo_title":"Screenshot Evidence Check Skill","seo_description":"Describe how a page was captured or measured and get a verdict on what it proves, as JSON, with the one deciding check. Pay $0.03 once, in USDC.","versions":[{"version":"1.0.0","date":"2026-10-08","changelog":"# Changelog\n\n## 1.0.0 — 2026-10-08\n\nFirst release: reads a description of how a page was captured or measured (headless window size, tall capture,\ntouch emulation, document width reading, Lighthouse runs, hosted PageSpeed Insights run, long automation session,\ndownload size on Windows, staging versus production, time-zone dependent dates, everyday browser) and answers as\nJSON with a verdict (proven, not-proven, measurement-artifact, contradicted), what the evidence shows, what it\ncannot show and the one deciding check. Carries the rule that only the description counts and unseen things are unknown.\n\nFacts re-checked on 2026-10-08 in a local headless Chrome 154.0.8037.98 on a local page (free, read-only, no network):\n- window-size 420 and 460 both laid the page out at 500 (innerWidth 500, a full-width button with 18 px padding ended at\n  x=482), and a screenshot taken with 420 was 420 pixels wide, so the picture is a crop; 500 gave 500;\n- the headless window reported hover: none false, hover: hover true and pointer: fine;\n- document scrollWidth equalled clientWidth (500) while a 900 px child sat inside an overflow-x: clip container; the\n  element measurement (getBoundingClientRect().right above clientWidth) listed that child.\nNot re-checked (owner-measured 20-21 September 2026, `platform/frontend.md`): blank images in a 9000 px tall capture,\nreduced-motion capture of reveal animations, the 96/91/91 Lighthouse spread, hosted versus local Lighthouse, the\nautomation session that stops running scripts, the zero download size on Windows, the staging/date/client-bundle facts.\nCases: 23 (15 traps, 8 controls); the control script makes no model calls. The model test and the price check come next.\n\n## Model test — 2026-10-08\n\nMeasured: Sonnet 22/22 with the skill, 19/22 without; Haiku 21/22 with, 18/22 without (one run per model and case; the\nstaging-lightbox case is excluded until re-run). Price set to $0.03 (30000 micro-USDC): Sonnet gain 3.\nChecks widened for both sides after reading the answers:\n- verdict accepts a second defensible verdict on five cases: tall-capture-hero-blank, lighthouse-observed-faster, automation-session-no-control, phone-390-breakpoint, everyday-browser-speed;\n- blame-the-grid-no-control: next_check also accepts a capture of the commit or layout before the grid change, or a one-column view;\n- curl-size-zero-windows: next_check also accepts reading the size of the stored object in the storage itself;\n- staging-lightbox-different-host: input did not say what the fallback is, so doubt was defensible; added \"falls back to the production media host\" to the input. NEEDS RE-RUN, not counted.\nReal failures kept: Sonnet without the skill doubted three sound captures (reduced-motion-all-six, touch-emulation-with-control, automation-control-page-works); Haiku without called the 420 crop not-proven, took the production-only look for an artifact, doubted a sound hero and blamed the tool where the control worked; Haiku with returned invalid JSON on local-clean-hosted-defects.\n- Re-run of staging-lightbox-different-host, both models, both sides: right with the skill for both; without it Sonnet right, Haiku not-proven. Final: Sonnet 20 -> 23 of 23, Haiku 18 -> 22 of 23. Price stays $0.03 (Sonnet gain 3).\n"}]}