Skip to content

ext/standard: Return the input of htmlspecialchars() when nothing needs encoding - #23957

Open
nicolas-grekas wants to merge 1 commit into
php:masterfrom
nicolas-grekas:htmlspecialchars-fast-path
Open

nicolas-grekas wants to merge 1 commit into
php:masterfrom
nicolas-grekas:htmlspecialchars-fast-path

Conversation

@nicolas-grekas

@nicolas-grekas nicolas-grekas commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

htmlspecialchars() allocates a buffer of twice the input and decodes it character by character, even when the result is the input itself. That's the common case for template engines that escape every value they print.

This scans the input first, with SSE2 (or NEON via zend_simd.h) on 16-byte blocks, and returns the input string when nothing needs encoding, as html_entity_decode() and htmlspecialchars_decode() already do since 68dc754. Otherwise, encoding starts after the scanned part, so strings that need escaping don't pay twice. On the Symfony Demo blog page, which makes 711 calls per request, htmlspecialchars() becomes about 3.5 times cheaper.

The fast path doesn't apply with ENT_DISALLOWED, nor to htmlentities() unless it falls back to the basic entities. Bytes above 0x7F go through the same UTF-8 decoder as before. The IS_STR_VALID_UTF8 flag isn't used because mbstring can set it on strings that hold encoded surrogates (e.g. mb_convert_encoding() from UCS-4), which htmlspecialchars() rejects; #23959 fixes that.

The @refcount 1 annotations are removed since both functions can now return their argument.

@github-actions

Copy link
Copy Markdown

AWS x86_64 (c6id.metal)

Attribute Value
Environment aws
Instance type c6id.metal
Architecture x86_64
CPU Intel(R) Xeon(R) Platinum 8375C CPU @ 2.90GHz, 64 cores @ 2900 MHz
CPU settings disabled deeper C-states, disabled turbo boost, disabled hyper-threading
RAM 251 GB
Kernel 6.18.38-76.139.amzn2023.x86_64
OS Amazon Linux 2023.12.20260727
GCC 14.2.1
Binary layout strategy bolt-align
Time 2026-09-28 07:33:06 UTC
Job details https://github.com/php/php-src/actions/runs/36392217564 (Artifacts)
Changeset https://github.com/php/php-src/compare/e256041a92..b9531eca92

Laravel 12.11.0 demo app - 50 iterations, 50 warmups, 100 requests (sec)

PHP Min Max Std dev Rel std dev % Mean Mean diff % Median Median diff % Skewness Z-stat P-value Memory
PHP - baseline@e256041 0.38430 0.38582 0.00030 0.08% 0.38535 0.00% 0.38540 0.00% -1.072 0.000 1.000 31.36 MB
PHP - htmlspecialchars-fast-path 0.38482 0.38595 0.00029 0.08% 0.38543 0.02% 0.38546 0.02% -0.184 -1.113 0.266 31.31 MB

Symfony 2.8.0 demo app - 50 iterations, 50 warmups, 100 requests (sec)

PHP Min Max Std dev Rel std dev % Mean Mean diff % Median Median diff % Skewness Z-stat P-value Memory
PHP - baseline@e256041 0.69701 0.70878 0.00179 0.26% 0.69790 0.00% 0.69749 0.00% 5.289 0.000 1.000 31.37 MB
PHP - htmlspecialchars-fast-path 0.69531 0.69699 0.00033 0.05% 0.69581 -0.30% 0.69575 -0.25% 1.520 8.614 0.000 31.32 MB

Wordpress 6.9 main page - 50 iterations, 20 warmups, 20 requests (sec)

PHP Min Max Std dev Rel std dev % Mean Mean diff % Median Median diff % Skewness Z-stat P-value Memory
PHP - baseline@e256041 0.59811 0.60016 0.00037 0.06% 0.59867 0.00% 0.59858 0.00% 1.553 0.000 1.000 31.37 MB
PHP - htmlspecialchars-fast-path 0.59850 0.60015 0.00033 0.05% 0.59906 0.07% 0.59900 0.07% 0.959 -5.422 0.000 31.32 MB

bench.php - 50 iterations, 20 warmups, 2 requests (sec)

PHP Min Max Std dev Rel std dev % Mean Mean diff % Median Median diff % Skewness Z-stat P-value Memory
PHP - baseline@e256041 0.45441 0.45653 0.00046 0.10% 0.45586 0.00% 0.45593 0.00% -1.139 0.000 1.000 31.37 MB
PHP - htmlspecialchars-fast-path 0.45528 0.45881 0.00067 0.15% 0.45622 0.08% 0.45606 0.03% 1.589 -2.216 0.027 31.32 MB

@kocsismate

Copy link
Copy Markdown
Member

@nicolas-grekas If you need more benchmark runs, please ping me and I can start another one.

Comment thread ext/standard/html.c Outdated
@Sjord

Sjord commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Looks good to me.

@nicolas-grekas
nicolas-grekas force-pushed the htmlspecialchars-fast-path branch 2 times, most recently from 1171518 to 45e19e7 Compare September 28, 2026 15:21
/* {{{ html.c */

/** @refcount 1 */
function htmlspecialchars(string $string, int $flags = ENT_QUOTES | ENT_SUBSTITUTE | ENT_HTML401, ?string $encoding = null, bool $double_encode = true): string {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in a typical web-application the htmlspecialchars function usually is invoked a lot. maybe its worth to make it frameless?

@nicolas-grekas nicolas-grekas Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could be worth it after this one is merged - the call overhead will then be significant enough.
As a separate PR.
FTR, frameless variants exist only for 1 to 3 arguments. Twig passes 3, Laravel's e() 4, so it wouldn't benefit there.

@ArtUkrainskiy

Copy link
Copy Markdown
Contributor

I've been working on the same loop from the other end in #24145: instead of skipping the clean prefix, it stays in the loop but copies runs of plain ASCII instead of going byte by byte, and inlines the decoder for the non-ASCII part. Since both touch the same function I built yours, mine and both together, and ran them through the same suite. They don't compete: this PR wins on strings with nothing to encode, mine on everything after the first character that needs it.

i7-13700H, GCC 13, perf stat instructions per call:

input this PR vs master this PR + #24145 vs this PR
40-byte title, clean 7.3× faster same
256 B English, clean 34× faster same
1 KB cp1251 32× faster same
40-byte title with & and quotes same 1.8× faster
1 KB HTML 1.05× slower 2.0× faster
64 KB Wikipedia HTML 1.03× slower 1.9× faster
4 KB Markdown / JS same 2.8× faster
htmlentities() 1.0–1.2× faster another 1.1–1.2×

Once the scan stops, the rest of the string still goes byte by byte through the old loop; #24145 copies plain ASCII in runs inside that loop, so it adds on top. Rebased on this branch it's +32/−2 with two trivial conflicts: https://github.com/ArtUkrainskiy/php-src/commits/hsc-runs-on-ng

Symfony Demo (callgrind, instructions per request): master 39.42M, this PR 39.05M (−0.93%), both 39.00M (−1.06%); htmlspecialchars() goes 526K → 178K → 81K per request. Output identical to master on 20,000 random inputs across 13 flag sets and 14 charsets.

Full tables and the raw runs: https://github.com/ArtUkrainskiy/php-src-bench/tree/main/reports/htmlspecialchars-vs-23957

So I'd say land this one first; I'll rebase #24145 onto it as the follow-up.

…ds encoding

htmlspecialchars() allocated a buffer of twice the input and decoded it
character by character, even when the result was the input itself.

The input is now scanned first for the characters that need encoding and
for invalid multi-byte sequences, with SSE2 on 16-byte blocks. When there
is none, the input string is returned, as html_entity_decode() and
htmlspecialchars_decode() already do since 68dc754. Otherwise, encoding
starts after the part that was scanned.

The @refcount 1 annotations of htmlspecialchars() and htmlentities() are
removed since they can now return their argument.
@nicolas-grekas
nicolas-grekas force-pushed the htmlspecialchars-fast-path branch from 384fbd6 to c1db7a4 Compare October 6, 2026 13:51
@nicolas-grekas

Copy link
Copy Markdown
Contributor Author

Thanks for checking @ArtUkrainskiy
That plan suits me well!
Anyone available to merge this one? 🙏

Kocal added a commit to Kocal/Twig that referenced this pull request Oct 7, 2026
With the UTF-8 charset, the `html` strategy calls `htmlspecialchars()`, which decodes the string one character at a time: about 5 ns per byte, so 1.6 µs for the 280-character class list of a styled button, even when nothing needs escaping. On valid UTF-8, all it does is replace `&`, `"`, `'`, `<` and `>`: `ENT_SUBSTITUTE` only matters for invalid sequences.

Strings longer than 32 bytes are now checked for UTF-8 validity with `preg_match('//u')` and escaped with `strtr()` on those five characters. Invalid UTF-8 and shorter strings still go through `htmlspecialchars()`, whose fixed cost is lower below that length. The output is identical: compared against `htmlspecialchars()` on every Unicode code point (in three positions) and on invalid sequences (lone continuation bytes, overlong forms, surrogates, code points above U+10FFFF, truncated sequences), 3.3 million strings, no difference on PHP 8.4 and 8.5. The shortcut only applies on PHP < 8.7 (`PHP_VERSION_ID < 80700`), since php/php-src#23957 and php/php-src#24145 (both still open, targeting PHP 8.7) make `htmlspecialchars()` itself skip the work on input that needs no encoding.

Benchmark below, PHP 8.4, 3 interleaved runs each side, identical output, ns per escape:

| String | Before | After |
| --- | --- | --- |
| short id (`user-42`) | **155-165 ns** | **168-175 ns** (below the threshold; the length check costs ~10 ns) |
| 40-character title | **327-334 ns** | **269-285 ns** |
| 280-character CSS class list | **1590-1630 ns** | **498-511 ns** (3.2x) |
| paragraph with quotes and tags | **1634-1663 ns** | **758-766 ns** (2.2x) |
| accented paragraph | **1310-1321 ns** | **373-380 ns** (3.5x) |

End to end, on an admin dashboard built with Symfony UX components that escapes about 2,500 attribute values per request (mostly long Tailwind class lists), CPU time per request goes from **36.7-38.7 ms** to **35.7-36.4 ms** (minimum of 6 runs of 60 requests each side), about 3%.

```php
<?php

// Run from the repository root: php bench.php
require getcwd().'/vendor/autoload.php';

$escaper = new Twig\Runtime\EscaperRuntime();
$classes = 'inline-flex shrink-0 items-center justify-center rounded-lg border border-transparent bg-clip-padding text-sm font-medium whitespace-nowrap transition-all outline-none select-none focus-visible:border-ring focus-visible:ring-3 focus-visible:ring-ring/50 disabled:pointer-events-none disabled:opacity-50';
$cases = [
    'short (id)' => 'user-42',
    'title (40 chars)' => 'How we made our dashboard ten times faster',
    'CSS classes (280 chars)' => $classes,
    'paragraph with quotes' => str_repeat('Twig escapes "quotes" & <tags> in a sentence like this one. ', 5),
    'accented paragraph' => str_repeat('Une phrase accentuée, déjà très répétée. ', 6),
];
foreach ($cases as $label => $string) {
    $times = [];
    for ($run = 0; $run < 7; ++$run) {
        $start = hrtime(true);
        for ($i = 0; $i < 200000; ++$i) {
            $escaper->escape($string);
        }
        $times[] = (hrtime(true) - $start) / 200000;
    }
    sort($times);
    printf("%-24s %5.0f ns/escape | sha1 %s\n", $label, $times[3], sha1($escaper->escape($string)));
}
```

Blackfire profiles (Symfony UX dashboard):
- Before: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/b344f602-a21f-41ef-a585-46fc865e0880/graph
- After: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/e5e16b8b-8895-40e7-9952-43282e3bd0df/graph
- Comparison: https://app.blackfire.io/envs/5f4f9a62-eaa0-45ee-b7b3-a1b879f550e9/profiles/compare/b344f602-a21f-41ef-a585-46fc865e0880...e5e16b8b-8895-40e7-9952-43282e3bd0df/graph

Blackfire barely moves (**2.29 s** -> **2.30 s** overall) because its per-call instrumentation dominates calls this short, even though `htmlspecialchars()` drops out of the profile entirely (5,398 calls per request before).

This came out of profiling Symfony UX components on https://github.com/Kocal/sf-ux-perfs-twig-components.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants