Skip to main content

Regex tester: matches, groups, replace and split

Type a pattern, paste the text, and the matches light up while you type. Numbered and named groups go into a table. Replace and Split run the same engine your JavaScript will.

runs in your browser · the pattern and the text are never uploaded

flags
presets:

    

How to use it

  1. Type the expression without the surrounding slashes into Pattern and tick the flags you need: g for every match, i to ignore case, m for multiline, plus s, u and y.
  2. Paste the text you are testing against. Matches are highlighted as you type, and the table below gives the position and every capture group, numbered and named.
  3. Switch to Replace to see what replace returns — $1, $<name> and $& all work in the replacement field — or to Split for the split array.
  4. The preset buttons load a working pattern together with sample text: email, URL, IPv4, ISO date, UTM parameters, domain, number.
  5. A syntax error appears in red with the engine message behind it. The stats row counts the matches and the milliseconds the run took.

Which engine, and where it parts ways with PCRE

The tester uses the JavaScript regular expression engine built into your browser (ECMAScript 2018 and later) — the same one Node.js runs. What you see here is exactly what JavaScript code will do with the pattern. PHP with PCRE, Python with re and Go with RE2 all differ in places.

Supported: named groups (?<name>…) and \k<name> backreferences, lookahead (?=…) / (?!…), lookbehind (?<=…) / (?<!…), Unicode property escapes \p{L} and \p{Script=Greek} with the u flag, lazy quantifiers *? and +?, the s flag that lets the dot match a newline, and the sticky y flag.

Not supported: atomic groups (?>…) and possessive quantifiers a++, recursion (?R), conditionals (?(1)…), the x flag for comments, and inline modifiers such as (?i). In JavaScript \w means [A-Za-z0-9_] and nothing else, so café and naïve break a \w+ match halfway: use [\p{L}\p{M}]+ with the u flag for text in any script. Word boundaries \b are ASCII-only for the same reason.

What media buyers and marketers use it for

  • Filters in trackers and ad accounts: rules on the URL, the referer or a subid — check the expression before it quietly drops half the traffic.
  • Cleaning up exports: pull emails, phone numbers, click ids or UTM tags out of a log. Replace mode turns the same pattern into a one-line cleanup script.
  • Form validation on a landing page: a promo code or a postcode, and the pattern that works here goes straight into pattern="" or a JavaScript validator.

Syntax, briefly

ConstructMeaningExample
.any character except a newline; with the s flag, that one tooa.c matches abc, a-c
\d \w \sdigit, word character (ASCII), whitespace; the capitals negate\d{3}-\d{2}
[abc] [^0-9] [a-z]character class, negation, range[A-Za-z0-9_-]+
* + ? {n,m}0 or more, 1 or more, 0 or 1, n to m repeats; a trailing ? makes it lazy\d+?
^ $ \bstart and end of the text (of every line with m), word boundary^https?://
(…) (?:…) (?<n>…)capture group, non-capturing, named(?<host>[\w.-]+)
a|balternationjpe?g|png|webp
(?=…) (?!…) (?<=…) (?<!…)lookahead and lookbehind\d+(?=\s?USD)
\1 $1 $<n>backreference inside the pattern; in the replacement $1, $<n>, $& for the whole match(\w+)@(\w+) → $2:$1
\p{L} \p{Script=Greek}Unicode categories and scripts, only with the u flag; the short form \p{Greek} is a syntax error in JavaScript[\p{Script=Greek}]+

Freezes and long texts

Nested quantifiers such as (a+)+$ or (\w*\s*)* cause catastrophic backtracking: the time grows exponentially with the input and the tab stops responding. The tester warns about patterns shaped like that, caps the text at 200,000 characters and runs the match on a timeout so the status line can paint first. If a pattern does hang, closing the tab is the only way out — a JavaScript regex cannot be interrupted from inside the page.

Questions

Which engine does the regex tester run?

The JavaScript engine in your own browser. The result matches what new RegExp() gives in JavaScript and in Node.js. PCRE in PHP and re in Python differ in the ways listed above.

Why does \w miss accented letters?

Because in JavaScript \w is defined as [A-Za-z0-9_]. Use [\p{L}\p{M}]+ with the u flag, or spell out the characters in a class, and the same applies to \b.

How do I get every match instead of the first one?

Tick the g flag, which is on by default. Without it the engine stops at the first match and replace rewrites only that one.

What do $1 and $&lt;name&gt; mean in the replacement field?

$1 and $2 are numbered capture groups, $<name> is a named one, $& is the whole match and $$ is a literal dollar sign. This is the standard String.prototype.replace syntax.

Is my text sent anywhere?

No. The pattern and the text are handled by JavaScript in the page and never leave it. The last input is kept in your browser local storage so a reload does not lose it.

The pattern hangs — what now?

A pattern with nested quantifiers has met text where backtracking explodes. Simplify it: (a+)+ becomes a+, add ^ or $ anchors, and remove alternatives that can match the same thing. The warning above the result flags the risky shapes before you hit that.

Project sponsors

Companies that keep this analytics open