How-To

How to test regular expressions — a practical workflow

Updated: September 2026

Most regex bugs are not syntax errors — they are matches that succeed when they should fail. A pattern that works on your one test string can silently accept the wrong input in production. This guide is a workflow: how to build a pattern against a checklist of positive and negative cases, and the handful of pitfalls that cause most real-world failures.

Test cases first, pattern second

Before writing a single character of pattern, write down what the regex must match and — more importantly — what it must reject. For a date pattern like \d{4}-\d{2}-\d{2}:

MUST match:      2026-09-18   1999-12-31
MUST NOT match:  26-09-18   2026-9-18   2026-13-45   hello

Then build the pattern while watching the negative column, not just the positive one. A pattern that only passes the "must match" list is usually halfway done.

Anchor everything

The single most common regex bug: forgetting that a search finds matches anywhere in the input. \d{4}-\d{2}-\d{2} happily matches the date inside id9999-00-00x. Anchors fix it:

AnchorMeaningExample
^…$Start / end of line^\d{4}-\d{2}-\d{2}$ — the whole line must be a date
\bWord boundary\bcat\b — "cat" but not "category"
\A \zStart / end of string (some flavors)Safer than ^$ for multi-line input

When validating input (forms, tokens, IDs), always anchor with ^ and $. When extracting, anchor with \b where possible to avoid substring surprises.

Refine against live matches

Type your pattern and sample text into a tester, look at what matched and where, then tighten. Typical iteration for validating a username (letters, digits, underscore, 3–16 chars):

v1:  \w{3,16}        → fails: matches inside "  bad--input  ", accepts unicode \w
v2:  ^\w{3,16}$       → better: anchored; still allow-list by default
v3:  ^[a-zA-Z0-9_]{3,16}$  → explicit allow-list, no \w surprises

Each pass closes one loophole the previous version accepted. This is also exactly how the AI Directories tester works: paste the pattern, paste the sample, and inspect the highlighted matches — including the ones you did not intend.

The four patterns that break in production

1. Catastrophic backtracking. Nested quantifiers like (a+)+$ explode exponentially on non-matching input — a 30-character string can hang a server. Keep quantifiers atomic and possessive where the flavor allows (++, *+), and never nest +/* inside another quantifier without need.

2. The dot. .* is greedy and crosses lines more often than expected. Use [^\n]* for "anything on this line" or make the dot lazy with .*? when you need the shortest match.

3. Unescaped metacharacters. A period, plus, or question mark in the text you are matching (URLs are full of them) must be escaped: example\.com, not example.com — which also matches examplexcom.

4. Trusting regex for validation. Regex can check shape, not truth: an email that matches a pattern can still not exist. Use regex to reject obviously wrong input; confirm with a real check (send the email, call the API).

Patterns worth keeping

TaskPatternNotes
Date (YYYY-MM-DD)^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$Rejects month 13 — shape and range
Hex color^#(?:[0-9a-fA-F]{3}|[0-9a-fA-F]{6})$3- or 6-digit forms
Trailing whitespace[ \t]+$For lint rules; no anchors needed
URL slug^[a-z0-9]+(?:-[a-z0-9]+)*$No leading/trailing/multiple hyphens

Regex flavors differ. JavaScript has no lookbehind in older browsers and no \A/\z; PCRE has both. Test in the engine you will actually ship.

Frequently asked questions

How do I test a regex before shipping it?

Write a fixed table of must-match and must-reject strings, run both columns against the pattern in a tester, and keep the table as a unit test. A pattern is done when the whole rejection column fails to match.

Why does my regex match only the first occurrence?

Most APIs require a global flag to find all matches: /g in JavaScript, re.findall in Python, preg_match_all in PHP. Without it the engine stops after the first match by design.

How can I tell if my pattern is vulnerable to backtracking?

Look for a quantifier inside a group that itself has a quantifier, like (\d+)* or (a|a)*$. Feed it a long non-matching string of the repeated character — if it takes noticeably longer as the string grows, restructure it.

Related guides