Regular Expressions for Beginners: A Practical Guide With Examples
A regular expression (regex) is a short pattern that describes the text you want to find. One line such as \d{4}-\d{2}-\d{2} can pick out every date in a document, check a form field or rewrite thousands of lines at once. The syntax looks cryptic at first, but it's built from a small set of pieces. This guide uses JavaScript syntax, the flavour used by web browsers and our tools, and each example shows the exact matches JavaScript returns.
How a pattern matches
Most characters in a pattern simply match themselves. The pattern cat finds the letters c, a and t in that order, wherever they appear. In "The cat sat on the concatenated mat" it matches twice: the word "cat" and the middle of "concatenated". Matching is case-sensitive, so "Cat" is missed unless you turn on the i flag.
The power comes from characters with special meanings, such as "any digit", "one or more" or "start of the line". These are . * + ? ^ $ ( ) [ ] { } | and the backslash \. To match one of them literally, put a backslash in front. 3\.14 matches "3.14" only, while 3.14 also matches "3514", because an unescaped dot means any character. In JavaScript code written between slashes, such as /a\/b/, a forward slash needs escaping too.
Character classes: one character from a set
| Token | Matches | Example |
|---|---|---|
. | Any character except a line break | h.t matches "hat", "hot" and "h t" |
\d | A digit from 0 to 9 | \d+ finds "66" and "3" in "Order 66 shipped 3 boxes" |
\w | A basic Latin letter, a digit or an underscore | \w+ matches "user_01" whole |
\s | Whitespace: space, tab, line break and other Unicode spaces | \s{2,} finds runs of two or more |
[ae] | Any one of the listed characters | m[ae]n matches "man" and "men", not "mon" |
[^0-9] | Any character not listed | Anything except a digit |
\D \W \S | The opposite of \d, \w and \s | \D is any non-digit |
One trap for non-English text: in JavaScript, \w only covers A to Z, a to z, 0 to 9 and the underscore. Run \w+ over "café naïve" and you get "caf", "na" and "ve". To match letters in any script, turn on the u flag and use \p{L}, the Unicode letter class: \p{L}+ matches "café" and "naïve" whole. In the same way, \d matches only 0 to 9, not digits from other scripts such as Arabic-Indic numerals.
Quantifiers: how many times
| Quantifier | Meaning | Example |
|---|---|---|
? | Optional (0 or 1) | Mrs? matches "Mr" and "Mrs"; https?:// matches both link types |
* | 0 or more | ab*c matches "ac", "abc", "abbc" |
+ | 1 or more | \d+ matches a whole run of digits |
{3} | Exactly 3 | \d{3} finds "123" and "456" in "123456" |
{2,5} | Between 2 and 5 | \w{2,5} splits "abcdefg" into "abcde" and "fg" |
{2,} | 2 or more | \s{2,} |
A quantifier applies to the single item just before it, so ha+ matches "haaa". To repeat several characters, group them: (ha)+ matches "hahaha".
Quantifiers are greedy: they take as much text as they can. On <b>bold</b> text, the pattern <.*> matches all of <b>bold</b>. A question mark after the quantifier makes it lazy, so <.*?> matches <b> and </b> separately. Even so, a regex is no substitute for a real HTML or JSON parser (see our JSON formatting guide).
Anchors and word boundaries
^ matches the start of the text and $ the end. They match positions, not characters, so ^\d{3}$ matches "123" but not "123456". That's how you check that a whole field is exactly three digits. With the m (multiline) flag, ^ and $ match at the start and end of every line instead.
\b is a word boundary, the point between a word character and anything else. With the i flag, \bcat\b finds "cat" and "Cat" in "The cat sat on the concatenated mat. Cat." and skips "concatenated".
Groups, alternatives and lookarounds
(…)groups items and captures what they matched. In a replacement you refer to the groups as$1,$2and so on. Inside the pattern,\1means "the same text again", so\b(\w+)\s+\1\bwith theiflag finds "the the" and "Is is" in "This is the the end. Is is it?"(?:…)groups without capturing, which keeps your group numbers tidy.(?<year>\d{4})is a named group. In a replacement, write$<year>.|means "or":cat|dogmatches either word. It has the lowest priority, so^cat|dog$means "cat at the start, or dog at the end" and matches "dog" in "hotdog". Write^(?:cat|dog)$when you mean a field that is exactly one of the two.(?=…)(lookahead) and(?<=…)(lookbehind) check what comes after or before without including it.\d+(?= kg)finds "70" and "12" in "70 kg, 5 lb, 12 kg", and(?<=#)\w+returns "regex" from "#regex" without the hash.
Flags
| Flag | Effect |
|---|---|
g | Global: find every match, not just the first |
i | Ignore case |
m | Multiline: ^ and $ work per line |
s | Dot-all: . also matches line breaks |
u | Unicode: enables \p{…} classes and reads characters outside the basic range, such as most emoji, as one character instead of two |
JavaScript also has y, d and v, which you'll rarely need at first. MDN's regular expressions guide covers every flag and token in detail.
Worked examples
| Task | Pattern and flags | Result |
|---|---|---|
| Find ISO dates | \b\d{4}-\d{2}-\d{2}\b, g | In "Due 2026-10-01, not 2026-13-45 or 20261001" it finds "2026-10-01" and "2026-13-45". It checks the shape, not whether the date exists. |
| Reorder dates | (\d{4})-(\d{2})-(\d{2}), g, replace with $3/$2/$1 | "2026-10-01" becomes "01/10/2026" |
| Same, with names | (?<y>\d{4})-(?<m>\d{2})-(?<d>\d{2}), replace with $<d>.$<m>.$<y> | "2026-10-01" becomes "01.10.2026" |
| Remove trailing spaces | [ \t]+$, gm, replace with nothing | Strips spaces and tabs from the end of every line |
| Pull out numbers | -?\d+(?:\.\d+)?, g | "Temps: -3.5, 0, 12 and 7.25" gives -3.5, 0, 12 and 7.25 |
| Rough email finder | [\w.+-]+@[\w-]+(?:\.[\w-]+)+, g | Finds "[email protected]" and "[email protected]" |
Email addresses are the classic trap: the formal rules allow far more than any readable pattern covers. Use a simple pattern to find likely addresses, and confirm one works by sending it a message.
Testing a pattern with our tools
All three tools below run in your browser, so the text you paste isn't uploaded.
- Open the Regex Tester and type your pattern in the Pattern box without the surrounding slashes. If you paste a JavaScript literal such as
/\d+/gi, the tool strips the slashes and ticks the matching flags for you. - Tick the flags you need: g (all matches, on by default), i (ignore case), m (multiline ^ $), s (. matches newline) and u (unicode).
- Paste sample text into Test text. Matches are highlighted as you type, with a count, and a table lists each match, its position and every capture group. The table shows the first 200 matches.
- To preview a substitution, fill in Replace with (use $1, $2, $<name>). The Result box shows the new text and Copy result copies it. With g unticked, only the first match is replaced.
- Stuck on the syntax? Open the Cheat sheet under the tool and click a token to add it to your pattern.
To run a pattern over a whole document, paste it into Find & Replace and tick Regular expression ($1 for groups). It always replaces every match, it ignores case unless you tick Match case, and it runs with the m and u flags on. Unicode mode is strict about unnecessary backslashes, so \- or \_ outside square brackets, which the tester accepts with u unticked, is reported as a mistake. Remove the extra backslash and it works. Download .txt saves the result.
Often you don't need to write a pattern at all. Text Extractor has ready-made ones: choose Email addresses, Links (URLs), Phone numbers, Numbers, Hashtags, @mentions, IPv4 addresses, IPv6 addresses or Dates from the Find list and it lists every match, with duplicates removed and an optional A–Z sort.
Common mistakes
- An unescaped dot.
example.comalso matches "exampleXcom". Writeexample\.com. - No anchors.
\d{3}matches inside "123456". Use^…$for whole fields and\bfor whole words. - A greedy wildcard.
".*"runs from the first quote to the last one on the line. Use".*?"or"[^"]*". - Nested quantifiers. Patterns like
(a+)+$can take an extremely long time on text that almost matches, which can freeze a browser tab. Keep repeated groups simple. - Assuming every flavour is the same. Python writes named groups as
(?P<name>…)and refers to groups in replacements as\1. Command-line tools such as grep and sed have their own rules, so check the documentation of the software that will run your pattern. - Treating a match as proof. A pattern checks shape only. "2026-13-45" passes a date pattern, so parse dates and numbers properly when correctness matters.
Sources
Spotted a mistake or something out of date? Tell us and we'll fix it.