Regular Expressions for Beginners: A Practical Guide With Examples

A regular expression (regex) is a short pattern that describes the text you want to find. One line such as \d{4}-\d{2}-\d{2} can pick out every date in a document, check a form field or rewrite thousands of lines at once. The syntax looks cryptic at first, but it's built from a small set of pieces. This guide uses JavaScript syntax, the flavour used by web browsers and our tools, and each example shows the exact matches JavaScript returns.

How a pattern matches

Most characters in a pattern simply match themselves. The pattern cat finds the letters c, a and t in that order, wherever they appear. In "The cat sat on the concatenated mat" it matches twice: the word "cat" and the middle of "concatenated". Matching is case-sensitive, so "Cat" is missed unless you turn on the i flag.

The power comes from characters with special meanings, such as "any digit", "one or more" or "start of the line". These are . * + ? ^ $ ( ) [ ] { } | and the backslash \. To match one of them literally, put a backslash in front. 3\.14 matches "3.14" only, while 3.14 also matches "3514", because an unescaped dot means any character. In JavaScript code written between slashes, such as /a\/b/, a forward slash needs escaping too.

Character classes: one character from a set

TokenMatchesExample
.Any character except a line breakh.t matches "hat", "hot" and "h t"
\dA digit from 0 to 9\d+ finds "66" and "3" in "Order 66 shipped 3 boxes"
\wA basic Latin letter, a digit or an underscore\w+ matches "user_01" whole
\sWhitespace: space, tab, line break and other Unicode spaces\s{2,} finds runs of two or more
[ae]Any one of the listed charactersm[ae]n matches "man" and "men", not "mon"
[^0-9]Any character not listedAnything except a digit
\D \W \SThe opposite of \d, \w and \s\D is any non-digit

One trap for non-English text: in JavaScript, \w only covers A to Z, a to z, 0 to 9 and the underscore. Run \w+ over "café naïve" and you get "caf", "na" and "ve". To match letters in any script, turn on the u flag and use \p{L}, the Unicode letter class: \p{L}+ matches "café" and "naïve" whole. In the same way, \d matches only 0 to 9, not digits from other scripts such as Arabic-Indic numerals.

Quantifiers: how many times

QuantifierMeaningExample
?Optional (0 or 1)Mrs? matches "Mr" and "Mrs"; https?:// matches both link types
*0 or moreab*c matches "ac", "abc", "abbc"
+1 or more\d+ matches a whole run of digits
{3}Exactly 3\d{3} finds "123" and "456" in "123456"
{2,5}Between 2 and 5\w{2,5} splits "abcdefg" into "abcde" and "fg"
{2,}2 or more\s{2,}

A quantifier applies to the single item just before it, so ha+ matches "haaa". To repeat several characters, group them: (ha)+ matches "hahaha".

Quantifiers are greedy: they take as much text as they can. On <b>bold</b> text, the pattern <.*> matches all of <b>bold</b>. A question mark after the quantifier makes it lazy, so <.*?> matches <b> and </b> separately. Even so, a regex is no substitute for a real HTML or JSON parser (see our JSON formatting guide).

Anchors and word boundaries

^ matches the start of the text and $ the end. They match positions, not characters, so ^\d{3}$ matches "123" but not "123456". That's how you check that a whole field is exactly three digits. With the m (multiline) flag, ^ and $ match at the start and end of every line instead.

\b is a word boundary, the point between a word character and anything else. With the i flag, \bcat\b finds "cat" and "Cat" in "The cat sat on the concatenated mat. Cat." and skips "concatenated".

Groups, alternatives and lookarounds

  • (…) groups items and captures what they matched. In a replacement you refer to the groups as $1, $2 and so on. Inside the pattern, \1 means "the same text again", so \b(\w+)\s+\1\b with the i flag finds "the the" and "Is is" in "This is the the end. Is is it?"
  • (?:…) groups without capturing, which keeps your group numbers tidy.
  • (?<year>\d{4}) is a named group. In a replacement, write $<year>.
  • | means "or": cat|dog matches either word. It has the lowest priority, so ^cat|dog$ means "cat at the start, or dog at the end" and matches "dog" in "hotdog". Write ^(?:cat|dog)$ when you mean a field that is exactly one of the two.
  • (?=…) (lookahead) and (?<=…) (lookbehind) check what comes after or before without including it. \d+(?= kg) finds "70" and "12" in "70 kg, 5 lb, 12 kg", and (?<=#)\w+ returns "regex" from "#regex" without the hash.

Flags

FlagEffect
gGlobal: find every match, not just the first
iIgnore case
mMultiline: ^ and $ work per line
sDot-all: . also matches line breaks
uUnicode: enables \p{…} classes and reads characters outside the basic range, such as most emoji, as one character instead of two

JavaScript also has y, d and v, which you'll rarely need at first. MDN's regular expressions guide covers every flag and token in detail.

Worked examples

TaskPattern and flagsResult
Find ISO dates\b\d{4}-\d{2}-\d{2}\b, gIn "Due 2026-10-01, not 2026-13-45 or 20261001" it finds "2026-10-01" and "2026-13-45". It checks the shape, not whether the date exists.
Reorder dates(\d{4})-(\d{2})-(\d{2}), g, replace with $3/$2/$1"2026-10-01" becomes "01/10/2026"
Same, with names(?<y>\d{4})-(?<m>\d{2})-(?<d>\d{2}), replace with $<d>.$<m>.$<y>"2026-10-01" becomes "01.10.2026"
Remove trailing spaces[ \t]+$, gm, replace with nothingStrips spaces and tabs from the end of every line
Pull out numbers-?\d+(?:\.\d+)?, g"Temps: -3.5, 0, 12 and 7.25" gives -3.5, 0, 12 and 7.25
Rough email finder[\w.+-]+@[\w-]+(?:\.[\w-]+)+, gFinds "[email protected]" and "[email protected]"

Email addresses are the classic trap: the formal rules allow far more than any readable pattern covers. Use a simple pattern to find likely addresses, and confirm one works by sending it a message.

Testing a pattern with our tools

All three tools below run in your browser, so the text you paste isn't uploaded.

  1. Open the Regex Tester and type your pattern in the Pattern box without the surrounding slashes. If you paste a JavaScript literal such as /\d+/gi, the tool strips the slashes and ticks the matching flags for you.
  2. Tick the flags you need: g (all matches, on by default), i (ignore case), m (multiline ^ $), s (. matches newline) and u (unicode).
  3. Paste sample text into Test text. Matches are highlighted as you type, with a count, and a table lists each match, its position and every capture group. The table shows the first 200 matches.
  4. To preview a substitution, fill in Replace with (use $1, $2, $<name>). The Result box shows the new text and Copy result copies it. With g unticked, only the first match is replaced.
  5. Stuck on the syntax? Open the Cheat sheet under the tool and click a token to add it to your pattern.

To run a pattern over a whole document, paste it into Find & Replace and tick Regular expression ($1 for groups). It always replaces every match, it ignores case unless you tick Match case, and it runs with the m and u flags on. Unicode mode is strict about unnecessary backslashes, so \- or \_ outside square brackets, which the tester accepts with u unticked, is reported as a mistake. Remove the extra backslash and it works. Download .txt saves the result.

Often you don't need to write a pattern at all. Text Extractor has ready-made ones: choose Email addresses, Links (URLs), Phone numbers, Numbers, Hashtags, @mentions, IPv4 addresses, IPv6 addresses or Dates from the Find list and it lists every match, with duplicates removed and an optional A–Z sort.

Common mistakes

  • An unescaped dot. example.com also matches "exampleXcom". Write example\.com.
  • No anchors. \d{3} matches inside "123456". Use ^…$ for whole fields and \b for whole words.
  • A greedy wildcard. ".*" runs from the first quote to the last one on the line. Use ".*?" or "[^"]*".
  • Nested quantifiers. Patterns like (a+)+$ can take an extremely long time on text that almost matches, which can freeze a browser tab. Keep repeated groups simple.
  • Assuming every flavour is the same. Python writes named groups as (?P<name>…) and refers to groups in replacements as \1. Command-line tools such as grep and sed have their own rules, so check the documentation of the software that will run your pattern.
  • Treating a match as proof. A pattern checks shape only. "2026-13-45" passes a date pattern, so parse dates and numbers properly when correctness matters.

Sources

  1. MDN Web Docs: Regular expressions
  2. MDN Web Docs: Regular expression syntax cheat sheet
  3. MDN Web Docs: String.prototype.replace()

Spotted a mistake or something out of date? Tell us and we'll fix it.

More guides