Menu

PHP preg_match(): Regular Expressions with Examples

preg_match($pattern, $subject, $matches) tests a string against a regular expression and returns 1 for a match, 0 for none. Learn capture groups, named groups, preg_match_all, anchors, validation patterns and the u modifier for UTF-8 and Japanese text.

This page includes runnable editors - edit, run, and see output instantly.

preg_match($pattern, $subject, $matches) checks whether a string matches a regular expression. It returns 1 on a match and 0 otherwise, and fills $matches with what it found: preg_match('/\d+/', 'Order 4521 shipped', $m) returns 1 and sets $m[0] to "4521".

The pattern is a string with delimiters around it, usually /.../, followed by modifiers such as i. Write patterns in single quotes: inside double quotes PHP first expands $name and escapes such as \x41 or \101, so the regex engine can receive a different pattern from the one you typed.

Capture groups

Parentheses in a pattern capture the part of the string they match. $m[0] is the whole match, $m[1] the first group, $m[2] the second, counted by opening parenthesis from the left.

Named groups

(?<name>...) gives a group a name. The value is stored under the name and under its number, so code that reads the result does not break when someone adds a group earlier in the pattern.

Find every match with preg_match_all

preg_match_all() keeps searching after the first match and returns the number of matches. By default $m[0] holds all full matches and $m[1] all values of the first group. With PREG_SET_ORDER you get one array per match instead, which is easier to loop over.

Validate input: anchors, phone numbers, postal codes

For validation the pattern must match the whole string, so start it with ^ and end it with $. One trap: $ also matches before a final newline, so '/^\d+$/' accepts "123\n". Add the D modifier, or use \z, which only matches at the very end.

For email addresses, a regex either rejects valid addresses or accepts invalid ones. Use filter_var($email, FILTER_VALIDATE_EMAIL) instead (see filter_var()).

The u modifier for UTF-8 and Japanese text

Without the u modifier, a pattern works on bytes. A Japanese character is three bytes in UTF-8, so . matches a third of a character and lengths come out wrong. With u, the pattern works on characters and you can use Unicode property classes.

u also validates the subject: if the string is not valid UTF-8, preg_match() returns false instead of matching.

Escape user input with preg_quote

Characters like ., +, ?, ( and / have special meanings in a pattern. If part of the pattern comes from a variable, such as a search term, pass it through preg_quote() with the delimiter as the second argument, or 1+1 will match 11 and an unescaped / breaks the pattern.

An invalid pattern does not throw. PHP prints a warning and preg_match() returns false:

Warning: preg_match(): No ending delimiter '/' found in /path/to/index.php on line 3

So check === false (or preg_last_error_msg()) when the pattern is built at runtime. For a plain substring check with no pattern at all, str_contains() is faster and clearer. To change the text you matched, use preg_replace().

Frequently Asked Questions

What does preg_match return in PHP?

1 if the pattern matches, 0 if it does not, and false if the pattern itself is invalid. Because 0 and false are both falsy, compare with === 1 when you need to tell a failed pattern from no match.

How do I get the matched text from preg_match?

Pass a third argument: preg_match('/(\d+)/', 'Order 4521', $m). Then $m[0] is the whole match and $m[1] the first capture group, both "4521" here. With a named group (?<id>\d+) the value is also in $m['id'].

What is the difference between preg_match and preg_match_all?

preg_match() stops at the first match. preg_match_all() finds every match, returns how many it found, and fills $matches with all of them: preg_match_all('/\d+/', 'a1 b22 c333', $m) returns 3 and $m[0] is ["1", "22", "333"].

How do I use preg_match with Japanese or other UTF-8 text?

Add the u modifier: preg_match('/^.{3}$/u', '日本語') returns 1. Without u, PHP treats the string as bytes, so . matches one byte of a 3-byte character. With u you can also use Unicode classes such as \p{Han}, \p{Hiragana} and \p{Katakana}.

How do I check if a string contains only numbers in PHP?

preg_match('/^\d+$/D', $s) or, without a regex, ctype_digit($s). The D modifier (or \z instead of $) matters: plain $ also accepts a trailing newline, so "123\n" would pass.

Coddy programming languages illustration

Learn to code with Coddy

GET STARTED