Resolve #3628857 "AliasCleaner regexes split UTF-8 characters when the locale treats the 0xA0 byte as whitespace (e.g. on macOS)"
Closes #3628857
What
Two preg_replace() calls in AliasCleaner::cleanString() work on bytes (no u modifier) while the string is UTF-8. In byte mode, \s follows the current locale's character tables. On macOS, under the C.UTF-8 locale set by DrupalKernel, the byte 0xA0 counts as whitespace, and that byte is the second half of à (C3 A0). So \s cuts à in two:
- Ignore words. Splitting the "Strings to remove" list leaves a lone
0xC3byte in the pattern.mb_eregi_replace()warnsPattern is not valid under UTF-8 encodingand returnsFALSE, so no ignore word is removed, ASCII ones included. - Whitespace to separator. With transliteration off, every
àin the text is cut too:Voilà le résumébecomesvoil?-le-résumé.
On glibc (Debian) and musl (Alpine), 0xA0 is not whitespace under C.UTF-8, so Linux sites don't show the problem. The byte-mode matching is still wrong on every platform.
Changes
src/AliasCleaner.php- add
uto the two patterns that split the ignore words list, and to thepreg_replace()fallback used when mbstring is missing; - add
uto the whitespace replacement. Withu,preg_replace()returnsNULLon a string that is not valid UTF-8, so that case falls back to the old byte-mode call and its output stays the same.
- add
PathautoKernelTesttestCleanStringIgnoreWordsWithMultibyteCharacters(): transliteration off, ignore wordsà, de, une,Une série de fiches à liregivessérie-fiches-lire;testCleanStringWhitespaceWithMultibyteCharacters(): transliteration off,Voilà le résumégivesvoilà-le-résumé;testCleanStringWhitespaceWithInvalidUtf8():"broken \xC3 string"still givesbroken-?-string. This one doesn't depend on the locale and runs on CI.
.cspell-project-words.txt:résumé,serie,série.
About the tests
The first two only fail where the locale makes \s match 0xA0, so they are skipped on Linux, including Drupal CI. Checked on macOS (PHP 8.4, Drupal 11.4) by removing each part of the change:
- without the ignore words fix: the
mb_eregi_replace()warning, andune-série-de-fiches-?-lire; - without the whitespace fix:
voil?-le-résumé; - with
ubut no byte-mode fallback:''and atrim(): Passing nulldeprecation for the invalid UTF-8 string; - with the whole change: the 4
testCleanString*tests pass.
Ignore words are still removed after transliteration, so with transliteration on, an accented ignore word like à still can't match, because the text already reads a. That ordering is what !156 (#3311669) changes; this MR doesn't touch it.
Related
!127 (#3059837) changes the ignore words split line, so whichever merges second needs a small rebase that keeps the u modifiers.
AI-Generated: Yes (Claude Code was used to help write the fix, its test cases and this description. I reviewed them, and on macOS each new test was confirmed to fail without its part of the fix and to pass with it.)