HardТеория7 min

Регулярные выражения

PCRE, preg_match, preg_replace, named groups, lookahead

Регулярные выражения (PCRE)

PHP использует библиотеку PCRE (Perl Compatible Regular Expressions) для работы с регулярными выражениями.

Основные функции

preg_match — первое совпадение

<?php
declare(strict_types=1);

$text = 'My phone is 123-456-7890';

// Basic match (returns 0 or 1)
$result = preg_match('/\d{3}-\d{3}-\d{4}/', $text, $matches);

echo $result;           // 1 (found)
echo $matches[0];       // 123-456-7890 (full match)

// With capturing groups
$result = preg_match('/(\d{3})-(\d{3})-(\d{4})/', $text, $matches);
echo $matches[0];  // 123-456-7890 (full match)
echo $matches[1];  // 123 (first group)
echo $matches[2];  // 456 (second group)
echo $matches[3];  // 7890 (third group)

// No match
$result = preg_match('/\d{5}/', 'hello', $matches);
echo $result;  // 0

preg_match_all — все совпадения

<?php
declare(strict_types=1);

$text = 'Emails: [email protected] and [email protected]';

$count = preg_match_all('/[\w.]+@[\w.]+/', $text, $matches);
echo $count;  // 2

// $matches[0] contains all full matches
print_r($matches[0]);
// ['[email protected]', '[email protected]']

// With groups
preg_match_all('/([\w.]+)@([\w.]+)/', $text, $matches);
print_r($matches[0]); // Full matches: ['[email protected]', '[email protected]']
print_r($matches[1]); // First group: ['alice', 'bob']
print_r($matches[2]); // Second group: ['test.com', 'test.com']

// PREG_SET_ORDER — group by match instead of by group
preg_match_all('/([\w.]+)@([\w.]+)/', $text, $matches, PREG_SET_ORDER);
// $matches[0] = ['[email protected]', 'alice', 'test.com']
// $matches[1] = ['[email protected]', 'bob', 'test.com']

preg_replace — замена

<?php
declare(strict_types=1);

// Simple replacement
$result = preg_replace('/\d+/', '#', 'Order 123 and 456');
echo $result;  // Order # and #

// With backreferences
$result = preg_replace('/(\w+)\s(\w+)/', '$2 $1', 'Hello World');
echo $result;  // World Hello

// Array of patterns and replacements
$text = 'Hello World 123';
$patterns = ['/[a-z]+/i', '/\d+/'];
$replacements = ['WORD', 'NUM'];
echo preg_replace($patterns, $replacements, $text);
// WORD WORD NUM

// With limit
$result = preg_replace('/\d/', 'X', '123456', 3);
echo $result;  // XXX456 (only first 3 replaced)

// preg_replace with count
$count = 0;
$result = preg_replace('/\d/', 'X', '123456', -1, $count);
echo $count;  // 6

preg_replace_callback

<?php
declare(strict_types=1);

$text = 'prices: $10, $25, $100';

$result = preg_replace_callback(
    '/\$(\d+)/',
    function (array $matches): string {
        $price = (int)$matches[1];
        return '$' . ($price * 2);  // Double all prices
    },
    $text
);
echo $result;  // prices: $20, $50, $200

// PHP 7.4+: With arrow function
$result = preg_replace_callback(
    '/\$(\d+)/',
    fn(array $m): string => '$' . ((int)$m[1] * 2),
    $text
);

// preg_replace_callback_array (PHP 7.0+)
$result = preg_replace_callback_array(
    [
        '/\d+/' => fn(array $m): string => (string)((int)$m[0] * 2),
        '/[a-z]+/i' => fn(array $m): string => strtoupper($m[0]),
    ],
    'hello 42 world 100'
);
echo $result;  // HELLO 84 WORLD 200

preg_split

<?php
declare(strict_types=1);

// Split by pattern
$result = preg_split('/[\s,;]+/', 'one, two; three  four');
// ['one', 'two', 'three', 'four']

// With limit
$result = preg_split('/\s+/', 'one two three four', 3);
// ['one', 'two', 'three four']

// Keep delimiters
$result = preg_split('/([,;])/', 'a,b;c', -1, PREG_SPLIT_DELIM_CAPTURE);
// ['a', ',', 'b', ';', 'c']

// Remove empty strings
$result = preg_split('/,/', ',a,,b,', -1, PREG_SPLIT_NO_EMPTY);
// ['a', 'b']

Синтаксис паттернов

Метасимволы

Символ Значение Пример
. Любой символ (кроме \n) a.c = "abc", "a1c"
^ Начало строки ^Hello
$ Конец строки world$
* 0 или более ab*c = "ac", "abc", "abbc"
+ 1 или более ab+c = "abc", "abbc"
? 0 или 1 colou?r = "color", "colour"
| Или cat|dog
() Группа (abc)+
[] Символьный класс [aeiou]
{} Квантификатор \d{3,5}
\ Экранирование \., \$

Символьные классы

<?php
declare(strict_types=1);

// Predefined character classes
// \d — digit [0-9]
// \D — non-digit [^0-9]
// \w — word character [a-zA-Z0-9_]
// \W — non-word character [^a-zA-Z0-9_]
// \s — whitespace [\t\n\r\f\v ]
// \S — non-whitespace [^\t\n\r\f\v ]

preg_match('/\d{3}/', '123abc', $m);   // $m[0] = '123'
preg_match('/\w+/', 'hello!', $m);     // $m[0] = 'hello'

// Custom character classes
preg_match('/[aeiou]+/', 'hello', $m);    // $m[0] = 'e'
preg_match('/[^aeiou]+/', 'hello', $m);   // $m[0] = 'h' (negated)
preg_match('/[a-zA-Z]+/', 'Hello!', $m);  // $m[0] = 'Hello'

// POSIX classes (inside brackets)
preg_match('/[[:alpha:]]+/', 'Hello123', $m);  // $m[0] = 'Hello'
preg_match('/[[:digit:]]+/', 'Hello123', $m);  // $m[0] = '123'

Квантификаторы

<?php
declare(strict_types=1);

// Greedy (default) — match as MUCH as possible
preg_match('/<.+>/', '<b>hello</b>', $m);
echo $m[0];  // <b>hello</b> (greedy: takes everything)

// Lazy (add ?) — match as LITTLE as possible
preg_match('/<.+?>/', '<b>hello</b>', $m);
echo $m[0];  // <b> (lazy: takes minimum)

// Quantifier table
// {n}    — exactly n
// {n,}   — n or more
// {n,m}  — between n and m
// *      — same as {0,}
// +      — same as {1,}
// ?      — same as {0,1}

// Possessive (add +) — greedy without backtracking
// preg_match('/\d++/', '123abc', $m);
// More efficient but less flexible

Ловушка экзамена: По умолчанию квантификаторы ЖАДНЫЕ (greedy). /<.+>/ захватит от первого < до ПОСЛЕДНЕГО >. Для минимального совпадения добавьте ? после квантификатора: /<.+?>/. Это один из самых частых вопросов!

Именованные группы

<?php
declare(strict_types=1);

$date = '2024-12-25';

// Named groups with (?P<name>...) or (?<name>...)
preg_match('/(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})/', $date, $matches);

echo $matches['year'];   // 2024
echo $matches['month'];  // 12
echo $matches['day'];    // 25

// Also accessible by number
echo $matches[1];  // 2024
echo $matches[2];  // 12

// Named backreference in replacement
$result = preg_replace(
    '/(?P<first>\w+)\s(?P<last>\w+)/',
    '${last}, ${first}',
    'John Smith'
);
echo $result;  // Smith, John

// Non-capturing group (?:...)
preg_match('/(?:https?|ftp):\/\/(\S+)/', 'https://example.com', $m);
echo $m[1];  // example.com (no group for protocol)

Assertions (утверждения)

Lookahead и Lookbehind

<?php
declare(strict_types=1);

// Positive lookahead (?=...) — followed by
preg_match_all('/\w+(?=\s*=)/', 'name = Alice, age = 25', $m);
// ['name', 'age'] — words followed by =

// Negative lookahead (?!...) — NOT followed by
preg_match_all('/\d+(?!\.)/', 'Price: 10.5 and 20', $m);
// ['0', '5', '20'] — digits NOT followed by dot
// Note: '10' becomes '1' (stopped before '0.') and '0' after dot

// Positive lookbehind (?<=...) — preceded by
preg_match_all('/(?<=\$)\d+/', 'Prices: $100 and $200', $m);
// ['100', '200'] — digits preceded by $

// Negative lookbehind (?<!...) — NOT preceded by
preg_match_all('/(?<!\$)\d+/', 'Item 42: $100', $m);
// ['42', '00'] — digits NOT preceded by $

Запомни: Lookahead и lookbehind НЕ включаются в результат совпадения — это "утверждения нулевой ширины" (zero-width assertions). Lookbehind в PCRE должен иметь ФИКСИРОВАННУЮ длину (переменная длина не поддерживается).

Anchors

<?php
declare(strict_types=1);

// ^ — start of string (or line with /m)
// $ — end of string (or line with /m)
// \b — word boundary
// \B — non-word boundary
// \A — absolute start of string
// \Z — end of string (before final \n if any)
// \z — absolute end of string

preg_match('/^\d+$/', '12345', $m);  // Full string is digits
preg_match('/\bword\b/', 'a word here', $m);  // 'word' as whole word

// With multiline flag
$text = "line1\nline2\nline3";
preg_match_all('/^\w+/m', $text, $m);
// ['line1', 'line2', 'line3'] — ^ matches each line start

Модификаторы (флаги)

<?php
declare(strict_types=1);

// i — case-insensitive
preg_match('/hello/i', 'Hello World', $m);  // Matches

// m — multiline (^ and $ match line boundaries)
preg_match_all('/^\w+/m', "foo\nbar", $m);  // ['foo', 'bar']

// s — dotall (. matches \n)
preg_match('/hello.world/s', "hello\nworld", $m);  // Matches

// x — extended (whitespace ignored, comments with #)
preg_match('/
    (\d{4})    # year
    -
    (\d{2})    # month
    -
    (\d{2})    # day
/x', '2024-12-25', $m);

// u — UTF-8 mode
preg_match('/\w+/u', 'Привет', $m);  // Works with Unicode

// Multiple flags
preg_match('/pattern/imsu', $text, $m);

// D — dollar matches only at end (not before \n)
// U — ungreedy mode (makes quantifiers lazy by default)
// A — anchored (as if every pattern starts with ^)
// J — allow duplicate named groups
Флаг Описание
i Регистронезависимый
m Многострочный (^/$ для каждой строки)
s Однострочный (. включает \n)
x Расширенный (игнорировать пробелы, комментарии)
u Unicode (UTF-8)
U Инвертировать жадность
D $ только конец строки
A Привязка к началу

Обработка ошибок

<?php
declare(strict_types=1);

// preg_last_error() — check for errors
$result = @preg_match('/(?:\D+|<\d+>)*[!?]/', 'too long string...');

if (preg_last_error() !== PREG_NO_ERROR) {
    echo match(preg_last_error()) {
        PREG_INTERNAL_ERROR => 'Internal error',
        PREG_BACKTRACK_LIMIT_ERROR => 'Backtrack limit reached',
        PREG_RECURSION_LIMIT_ERROR => 'Recursion limit reached',
        PREG_BAD_UTF8_ERROR => 'Invalid UTF-8',
        PREG_BAD_UTF8_OFFSET_ERROR => 'Bad UTF-8 offset',
        PREG_JIT_STACKLIMIT_ERROR => 'JIT stack limit',
        default => 'Unknown error',
    };
}

// preg_last_error_msg() (PHP 8.0+)
echo preg_last_error_msg();

Практические паттерны

<?php
declare(strict_types=1);

// Email validation (simplified)
$pattern = '/^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$/';
var_dump(preg_match($pattern, '[email protected]'));  // 1

// URL matching
$pattern = '#^https?://[^\s/$.?#].[^\s]*$#i';

// IPv4 address
$pattern = '/^(\d{1,3}\.){3}\d{1,3}$/';

// Password strength (min 8 chars, upper, lower, digit, special)
$password = 'MyP@ss1234';
$strong = preg_match('/^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[\W_]).{8,}$/', $password);

// Extract all HTML tags
preg_match_all('/<([a-z][a-z0-9]*)\b[^>]*>/i', $html, $tags);

// preg_quote — escape regex special characters
$userInput = 'price: $10.00 (USD)';
$escaped = preg_quote($userInput, '/');
// Result: 'price\: \$10\.00 \(USD\)'
preg_match('/' . $escaped . '/', $text);

Запомни: preg_quote() экранирует все специальные символы regex. Второй параметр — дополнительный символ для экранирования (обычно разделитель паттерна /). Используйте при вставке пользовательского ввода в regex.


Вопросы с экзамена ZCE

Проверь себя

5 из 15

Какие компоненты включает паттерн MVC в веб-разработке?

Выберите все правильные варианты

Какое PCRE-выражение соответствует любому пробельному символу?

Какой тип функций регулярных выражений в PHP предпочтителен по производительности?

Какой строке соответствует следующее PCRE регулярное выражение? ```php $regex = "/^([a-z]{5})[1-5]+([a-z]+)/"; ```

Выберите все правильные варианты

Какая функция экранирует все метасимволы оболочки и управляющие операторы в строке?