Regex patterns every developer should know

7 min read
regex

Most regular expressions people actually need in day-to-day work come from a small, reusable set. Memorising the syntax matters less than knowing which pattern to reach for.

Whitespace and trimming

^\s+|\s+$ matches leading and trailing whitespace, which is what String.prototype.trim() does internally. \s+ collapses runs of spaces, tabs and newlines into one match, useful when normalising pasted text before storage.

Emails: good enough, not perfect

A fully RFC 5322-compliant email regex is enormous and still wrong at the edges. For form validation, a pragmatic pattern works better paired with a confirmation email:

const emailLike = /^[^\s@]+@[^\s@]+\.[^\s@]+$/;

This rejects obviously malformed input without pretending to be a full validator. Never use regex alone to decide whether an email address can actually receive mail.

URLs and paths

^https?:\/\/[\w.-]+(?:\/[\w./?%&=-]*)?$ catches the common case of an http(s) URL with an optional path and query string. For anything more precise, parse with the URL API instead — a URL parser tool exists exactly because regex is the wrong tool for full RFC 3986 compliance.

Capturing groups vs non-capturing groups

Use (?:...) when you need to group alternatives but do not need the matched text back — it is faster and keeps your capture group indices predictable when the pattern grows.

const version = /^v(?:ersion)?[\s-]?(\d+)\.(\d+)\.(\d+)$/;

Here only the three numeric groups are captured; the optional "ersion" text is grouped but ignored.

Lookaheads for validation without consuming

Password strength rules are the classic use case:

  • (?=.*[a-z]) — at least one lowercase letter
  • (?=.*[A-Z]) — at least one uppercase letter
  • (?=.*\d) — at least one digit
  • (?=.*[^\w\s]) — at least one symbol

Chain lookaheads at the start of the pattern, then follow with the length requirement: ^(?=.[a-z])(?=.[A-Z])(?=.*\d).{8,}$.

Greedy vs lazy quantifiers

. is greedy and will match as much as possible before backtracking. .? is lazy and matches as little as possible. Extracting content between HTML-like tags is the classic trap:

"<b>one</b><b>two</b>".match(/<b>(.*)<\/b>/)[1];  // "one</b><b>two"
"<b>one</b><b>two</b>".match(/<b>(.*?)<\/b>/)[1]; // "one"

Escaping special characters

If you are building a pattern from user input, escape regex metacharacters first, or the input can change what the pattern matches entirely:

function escapeRegex(s) {
  return s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
}

Catastrophic backtracking

Nested quantifiers like (a+)+ can take exponential time on pathological input — a classic ReDoS. Avoid nesting repetition, prefer atomic patterns, and test suspicious patterns against a long string of near-matches before shipping them anywhere user input reaches the pattern.

Test before you commit

Write the pattern, then run it against a batch of both matching and non-matching examples in a regex tester with live highlighting. Seeing which group matched what removes almost all of the guesswork, especially with lookaheads and nested groups.

Tools from this article

Comments

No login needed. Comments appear after a quick review.

Protected by an on-site captcha — no third-party trackers.

Optional: get an alert when your comment is published or replied to, plus new tool announcements. No sign-up.

Loading comments…

← All articles