{"slug":"regex-from-examples","name":"Regex From Examples for Any Script","version":"1.0.0","updated_at":"2026-10-08T10:45:26.186Z","use_when":"Writes a JavaScript regular expression from a description plus lists of strings that must match and must not match, answered as JSON with the pattern and flags. Knows the traps measured in Node 24 - \\b and \\w are ASCII-only even with the u flag, so \\bдума\\b never matches Cyrillic text and \\w+ cuts \"café\" to \"caf\"; a word boundary for any script is a pair of lookarounds with \\p{L} and the u flag. Anchors a validator so extra text around the value is rejected, adds the s or m flag when the input spans lines, and avoids nested quantifiers such as (\\w+\\s?)* that hang on a near-miss string. Use when asked to write, fix or tighten a regex, a validation pattern or a text-search pattern, especially for non-English text.","not_for":"Other regex flavours such as PCRE, Python or RE2 (their escapes and flags differ), parsing nested formats like HTML, and checking that an email or phone number really exists. It writes a JavaScript pattern for the strings you list, so cases you do not list are only as good as its reading of your description.","languages":["any"],"tags":["regex","javascript","unicode","validation","redos","text-search"],"category":"code","category_url":"https://aiskills402.com/categories/code","keywords":["regular expression","validation pattern","word boundary"],"faq":[{"q":"What does the answer look like in practice?","a":"JSON with a pattern and flags, ready for new RegExp(pattern, flags). Whole-value checks are anchored, text in Cyrillic, Greek or accented Latin is matched with Unicode properties and the u flag, and multi-line input gets the s or m flag. Nothing else is printed, so a script can parse the reply directly and hand it to the compiler without stripping a fence or a greeting."},{"q":"Why would a model get this wrong?","a":"In JavaScript the \\b and \\w shortcuts only know ASCII letters, even with the u flag. A pattern like \\bдума\\b never matches Bulgarian text and \\w+ cuts café to caf, so a check can return zero findings for months. We measured this in Node 24; the skill uses lookarounds with \\p{L} instead. The same trap hides in word counters, linters and search boxes built for one alphabet."},{"q":"Does it protect against patterns that hang?","a":"Yes. Nested quantifiers such as (\\d+,?)+ took one second on 24 digits followed by a letter and double with every character. The skill writes the separator as required, which finishes a 100 000 character miss in milliseconds. Our test runs each pattern on long near-miss strings with a 200 ms limit. Greedy lazy tricks are not needed: the fix is structural, so the engine has only one way to split the input and cannot retry endlessly."},{"q":"Does it help Claude Sonnet?","a":"Not on our 22 tasks. Sonnet and Haiku both passed all 22 with and without the skill, because the tasks list the strings that must match and both models already used lookarounds with \\p{L} instead of \\b. The skill gives you the measured JavaScript notes and a fixed JSON answer; do not expect a gain on tasks like ours."}],"examples":[{"lang":"en","model":"claude-sonnet-5-5","input_excerpt":"Task: Find the Bulgarian word \"дума\" as a whole word, in any letter case, anywhere in a sentence. It must not match inside longer words.\n\nMust match (each string is one test input, shown as JSON):\n  \"една дума тук\"\n  \"Дума е\"\n  \"дума\"\n  \"(дума)\"\n  \"Каква дума.\"\n  \"ДУМА!\"\n\nMust NOT match:\n  \"думата\"\n  \"задума\"\n  \"думи\"\n  \"продумам\"\n  \"дума1\"\n  \"нищо\"…","output_excerpt":"{\"pattern\": \"(?<![\\\\p{L}\\\\p{N}_])дума(?![\\\\p{L}\\\\p{N}_])\", \"flags\": \"iu\"}"}],"page_url":"https://aiskills402.com/skills/regex-from-examples","markdown_url":"https://aiskills402.com/skills/regex-from-examples.md","image_url":"https://cdn.aiskills402.com/og/skills/regex-from-examples/6827c411.png","related_url":"https://api.aiskills402.com/v1/skills/regex-from-examples/related","purchases_count":null,"tested":{"date":"2026-10-08","strong":{"model":"claude-sonnet-5-5 (Claude Code alias \"sonnet\")","verdict":"Wrote a regex that passes all 22 tasks: every must-match and must-not-match string, and long near-miss inputs inside 200 ms. Word boundaries in Cyrillic, Greek and accented text use lookarounds with Unicode properties and the u flag, validators are anchored, multi-line input gets the right flag, and repeated groups avoid nested quantifiers."},"weak":{"model":"claude-haiku-5-5 (Claude Code alias \"haiku\")","verdict":"Also 22 of 22, with the same kinds of pattern as Sonnet."},"note":"Twenty-two tasks written by us, each with strings that must match, strings that must not, and long near-miss strings that must finish within 200 ms. About a third are plain tasks where the obvious regex is right; the rest set traps: word boundaries in Bulgarian, Greek and accented text, whole-value validators, multi-line input, names with letters outside A-Z, and patterns that hang on nested quantifiers. Each returned pattern was compiled and run against all strings. The JavaScript facts the skill states were measured in Node 24 on 8 October 2026. One run per model and task.","baseline":{"date":"2026-10-08","rows":[{"label":"Regexes that pass every string (22 tasks)","better":"higher","strong":{"with":{"n":22,"of":22},"without":{"n":22,"of":22}},"weak":{"with":{"n":22,"of":22},"without":{"n":22,"of":22}}}],"note":"No measurable gain on these tasks, for either model. The same request, asking for JSON only, went to both sides. When the task lists the strings that must match, both models already wrote lookarounds with p{L} and the u flag for Cyrillic and Greek instead of \b, anchored the validators and avoided nested quantifiers. What the skill adds is the measured notes on why those choices matter and a fixed JSON answer; it may help more when no examples are given or on weaker models than the two tested."},"report_url":null},"price_usd":"0.01","price_micro":10000,"size_bytes":6342,"sha256":"77f1ec3198ae87e95cca5b699f1c107716ec388aa1c75f90320118602597174e","outline":["The answer","Where JavaScript surprises (measured in Node 24, 8 October 2026)","Work in this order","Short examples"],"license":{"summary":"Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing","holder":"Georgi Kalchev, aiskills402.com","url":"https://aiskills402.com/docs#license"},"buy_url":"https://api.aiskills402.com/v1/skills/regex-from-examples/file","redownload_url_template":"https://api.aiskills402.com/v1/purchases/{token}","mcp_tool":null,"payment":{"protocol":"x402","scheme":"exact","asset":"USDC","selling":true,"network":"base","network_caip2":"eip155:8453","pay_to":"0x8e37022edcf0f21cf3c9f93fee9d4d32519f36f4","facilitator":"cdp"},"seo_title":"Regex From Examples: Unicode-Safe Patterns","seo_description":"Write JavaScript regex from match and no-match lists. Handles Cyrillic and accented text, anchors and patterns that hang. $0.03, one payment.","versions":[{"version":"1.0.0","date":"2026-10-08","changelog":"# Changelog\n\n## 1.0.0 — 2026-10-08\n\nFirst release: writes a JavaScript regular expression from a description and lists of strings that must and must not match, answered as JSON (pattern and flags). Covers what was measured in Node 24: \\b and \\w are ASCII-only even with the u flag, so word boundaries for Cyrillic, Greek and accented text are written with lookarounds and \\p{L}; \\p needs the u flag; validators are anchored; the s and m flags for multi-line input; and nested quantifiers that hang on near-miss strings are rewritten with a mandatory separator.\n"}]}