Resolve #3628857 "AliasCleaner regexes split UTF-8 characters when the locale treats the 0xA0 byte as whitespace (e.g. on macOS)"

Closes #3628857

What

Two preg_replace() calls in AliasCleaner::cleanString() work on bytes (no u modifier) while the string is UTF-8. In byte mode, \s follows the current locale's character tables. On macOS, under the C.UTF-8 locale set by DrupalKernel, the byte 0xA0 counts as whitespace, and that byte is the second half of à (C3 A0). So \s cuts à in two:

  1. Ignore words. Splitting the "Strings to remove" list leaves a lone 0xC3 byte in the pattern. mb_eregi_replace() warns Pattern is not valid under UTF-8 encoding and returns FALSE, so no ignore word is removed, ASCII ones included.
  2. Whitespace to separator. With transliteration off, every à in the text is cut too: Voilà le résumé becomes voil?-le-résumé.

On glibc (Debian) and musl (Alpine), 0xA0 is not whitespace under C.UTF-8, so Linux sites don't show the problem. The byte-mode matching is still wrong on every platform.

Changes

  • src/AliasCleaner.php
    • add u to the two patterns that split the ignore words list, and to the preg_replace() fallback used when mbstring is missing;
    • add u to the whitespace replacement. With u, preg_replace() returns NULL on a string that is not valid UTF-8, so that case falls back to the old byte-mode call and its output stays the same.
  • PathautoKernelTest
    • testCleanStringIgnoreWordsWithMultibyteCharacters(): transliteration off, ignore words à, de, une, Une série de fiches à lire gives série-fiches-lire;
    • testCleanStringWhitespaceWithMultibyteCharacters(): transliteration off, Voilà le résumé gives voilà-le-résumé;
    • testCleanStringWhitespaceWithInvalidUtf8(): "broken \xC3 string" still gives broken-?-string. This one doesn't depend on the locale and runs on CI.
  • .cspell-project-words.txt: résumé, serie, série.

About the tests

The first two only fail where the locale makes \s match 0xA0, so they are skipped on Linux, including Drupal CI. Checked on macOS (PHP 8.4, Drupal 11.4) by removing each part of the change:

  • without the ignore words fix: the mb_eregi_replace() warning, and une-série-de-fiches-?-lire;
  • without the whitespace fix: voil?-le-résumé;
  • with u but no byte-mode fallback: '' and a trim(): Passing null deprecation for the invalid UTF-8 string;
  • with the whole change: the 4 testCleanString* tests pass.

Ignore words are still removed after transliteration, so with transliteration on, an accented ignore word like à still can't match, because the text already reads a. That ordering is what !156 (#3311669) changes; this MR doesn't touch it.

!127 (#3059837) changes the ignore words split line, so whichever merges second needs a small rebase that keeps the u modifiers.

AI-Generated: Yes (Claude Code was used to help write the fix, its test cases and this description. I reviewed them, and on macOS each new test was confirmed to fail without its part of the fix and to pass with it.)

Edited by Frank Mably

Merge request reports

Loading
Loading