Skip to content

ext/standard: Check for UTF-8 first when resolving the charset of html functions - #23961

Open
nicolas-grekas wants to merge 1 commit into
php:masterfrom
nicolas-grekas:html-determine-charset
Open

nicolas-grekas wants to merge 1 commit into
php:masterfrom
nicolas-grekas:html-determine-charset

Conversation

@nicolas-grekas

@nicolas-grekas nicolas-grekas commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Follow-up of #23957: once htmlspecialchars() returns its input when nothing needs encoding, resolving the charset is about half of what's left for short strings. determine_charset() runs strlen() and walks the charset map with case-insensitive compares on every call, while UTF-8 is what nearly every call asks for, either explicitly as Twig or Laravel's e() do, or through default_charset.

This checks for "utf-8" in any case before walking the map, and moves the walk to a separate function so that the common path doesn't pay for its stack frame. On top of #23957, htmlspecialchars() on a short string becomes about 40% cheaper. htmlentities(), html_entity_decode() and get_html_translation_table() benefit the same way.

The walk itself no longer calls strlen() and only compares the entries that start with the same letter as the hint, so other charsets resolve a bit faster than before too. The "utf-8" entry of the map, which the walk can no longer reach, is removed.

@adapik

adapik commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

@nicolas-grekas

Nice speedup for the UTF-8 case! But worth noting this trades in a small regression for everything else: lookup_charset() is zend_never_inline, and now every call runs the "utf-8" probe first before falling through to it, so non-UTF-8 charsets pay an extra failed comparison plus a real (non-inlinable) call they didn't before.

Personally I don't think it's a real problem — UTF-8 is the default basically everywhere nowadays.

@nicolas-grekas

Copy link
Copy Markdown
Contributor Author

PR updated: the walk no longer calls strlen() and only compares entries that start with the same letter as the hint, so other charsets are now faster than before too (ISO-8859-1 resolves about 25% faster).

Comment thread ext/standard/html.c Outdated
Comment thread ext/standard/html.c
@nicolas-grekas

Copy link
Copy Markdown
Contributor Author

Both done: the probe is now charset_is_utf8(), and the "utf-8" entry is gone from charset_map and html_table_gen.php.

@nicolas-grekas
nicolas-grekas force-pushed the html-determine-charset branch from 2440019 to 64f7542 Compare October 2, 2026 10:36
@adapik

adapik commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

LGTM

…l functions

determine_charset() ran strlen() and walked the charset map with case-insensitive
compares on every call, while "UTF-8" (explicit or from default_charset) is what
nearly every call asks for. It now checks for "utf-8" in any case first and moves
the walk to a separate function, which skips strlen() and compares only the entries
that start with the same letter as the hint, so other charsets resolve faster too.
The "utf-8" entry of the charset map, which the walk can no longer reach, is removed.
@nicolas-grekas
nicolas-grekas force-pushed the html-determine-charset branch from 64f7542 to be56f28 Compare October 6, 2026 13:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants