← Back to Blog

Regex Cheat Sheet The Patterns You'll Actually Copy-Paste

By
Regex Cheat Sheet The Patterns You'll Actually Copy-Paste

Regex syntax is dense enough that even people who use it constantly look up the same handful of patterns every time. Here's the reference, plus the two mistakes that cause the most debugging time.

The building blocks

Symbol Meaning
. any character except newline
\d any digit (0–9)
\w any word character (letters, digits, underscore)
\s any whitespace
^ start of string
$ end of string
* zero or more of the previous token
+ one or more of the previous token
? zero or one (also makes a quantifier "lazy" — see below)
{n,m} between n and m repetitions
[abc] any one of a, b, or c
(...) capture group
(?:...) non-capturing group

Patterns people actually reuse

What you need Pattern
Digits only ^\d+$
Basic email check ^[\w.+-]+@[\w-]+\.[a-zA-Z]{2,}$
URL (http/https) ^https?:\/\/[\w.-]+\.[a-zA-Z]{2,}(\/\S*)?$
Hex color ^#([0-9a-fA-F]{3}){1,2}$
Trim leading/trailing whitespace ^\s+\|\s+$
US phone number (loose) ^\(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}$
IPv4 address ^(\d{1,3}\.){3}\d{1,3}$

A note on the email one, since it's the pattern people trust the most and should trust the least: there is no regex that fully validates an email address against the actual RFC 5322 spec without becoming an unreadable multi-hundred-character monster, and even a perfect one can't tell you whether the address actually exists. The pattern above catches obvious typos and malformed input; the real validation is sending a confirmation email and seeing if anyone clicks the link.

Greedy vs. lazy: the mistake that eats the most debugging time

By default, * and + are greedy — they match as much as possible. Given the string <b>bold</b> and <i>italic</i>, the pattern <.+> doesn't stop at the first > — it grabs everything up to the last one, matching the whole string as one chunk instead of the four separate tags you probably meant.

Adding a ? after the quantifier makes it lazy — matches as little as possible instead: <.+?> stops at the first > it finds, correctly matching each tag separately.

This single character is the difference between a regex that works and one that silently grabs way more than intended, and it's the most common reason a "working" regex breaks the moment it hits real-world input slightly messier than the test case it was written against.

The other classic mistake: forgetting to escape

. matches any character in regex, not a literal period. If you're matching a filename like report.pdf and write report.pdf as your pattern, it'll also match reportXpdf — technically correct by regex rules, not what you meant. Escape it: report\.pdf. The same applies to *, +, ?, (, ), [, ], and a handful of other special characters — anywhere you mean the literal character rather than its regex meaning, it needs a backslash in front of it.

One thing regex should never be used for

Don't use regex to parse structured formats like HTML, JSON, or XML. They have nesting rules regex fundamentally can't track reliably, and a pattern that appears to work on your test case will eventually break on real-world input with unexpected structure. Use an actual parser for those — it's built for exactly this problem.

Testing without the guesswork

The Regex Tester & Matcher shows live match highlighting against your sample text as you type, so you catch a greedy-match mistake immediately instead of shipping it and finding out in production.

Try the Regex Tester →