Repository navigation
ext/standard: Check for UTF-8 first when resolving the charset of html functions - #23961
nicolas-grekas wants to merge 1 commit into
Conversation
79f566a to
3d55fb6
Compare
|
Nice speedup for the UTF-8 case! But worth noting this trades in a small regression for everything else: lookup_charset() is zend_never_inline, and now every call runs the "utf-8" probe first before falling through to it, so non-UTF-8 charsets pay an extra failed comparison plus a real (non-inlinable) call they didn't before. Personally I don't think it's a real problem — UTF-8 is the default basically everywhere nowadays. |
3d55fb6 to
4517a35
Compare
|
PR updated: the walk no longer calls |
4517a35 to
8d67b20
Compare
8d67b20 to
2440019
Compare
|
Both done: the probe is now |
2440019 to
64f7542
Compare
|
LGTM |
…l functions determine_charset() ran strlen() and walked the charset map with case-insensitive compares on every call, while "UTF-8" (explicit or from default_charset) is what nearly every call asks for. It now checks for "utf-8" in any case first and moves the walk to a separate function, which skips strlen() and compares only the entries that start with the same letter as the hint, so other charsets resolve faster too. The "utf-8" entry of the charset map, which the walk can no longer reach, is removed.
64f7542 to
be56f28
Compare
Follow-up of #23957: once htmlspecialchars() returns its input when nothing needs encoding, resolving the charset is about half of what's left for short strings. determine_charset() runs strlen() and walks the charset map with case-insensitive compares on every call, while UTF-8 is what nearly every call asks for, either explicitly as Twig or Laravel's e() do, or through default_charset.
This checks for "utf-8" in any case before walking the map, and moves the walk to a separate function so that the common path doesn't pay for its stack frame. On top of #23957, htmlspecialchars() on a short string becomes about 40% cheaper. htmlentities(), html_entity_decode() and get_html_translation_table() benefit the same way.
The walk itself no longer calls strlen() and only compares the entries that start with the same letter as the hint, so other charsets resolve a bit faster than before too. The "utf-8" entry of the map, which the walk can no longer reach, is removed.