Repository navigation
User opcode handlers corrupt VM state under PHP 8.6's tail-call VM (macOS arm64) #280
Description
Activity
Root cause found — this is a php-src bug in the 8.6 tail-call VM, not a z-engine bug. z-engine now fails fast on affected builds (PR #281); below is a ready-to-file upstream report for
php/php-src(I can't open issues there from this session — please forward).
Draft php-src report
Title: User opcode handlers resume execution against a stale frame under ZEND_VM_KIND_TAILCALL
Version: PHP 8.6.0beta2 / 8.6.0-dev (any build selecting
ZEND_VM_KIND_TAILCALL, i.e. clang withmusttail/preserve_nonebut withoutHAVE_GCC_GLOBAL_REGS— notably all Homebrew/macOS arm64 builds). Hybrid and call VM builds of the same source are unaffected.Summary: A user opcode handler (
zend_set_user_opcode_handler) that returnsZEND_USER_OPCODE_DISPATCHcorrupts execution when the hooked opcode executes in any frame other than the oneexecute_ex()was entered with (any function call, any include). Symptoms range from calls dispatching to the wrong function (stale run-time cache slots of the wrong frame) to SIGSEGV inZEND_DO_UCALL_SPEC_RETVAL_USED_TAILCALL_HANDLER(KERN_INVALID_ADDRESS at 0x10, i.e. a NULLzend_functionderef).Minimal repro (only ext-ffi,
php -d ffi.enable=1 -d opcache.jit=off repro.php):<?php $engine = FFI::cdef(' typedef int (*user_opcode_handler_t)(void *execute_data); int zend_vm_kind(void); int zend_set_user_opcode_handler(unsigned char opcode, user_opcode_handler_t handler); '); echo 'vm_kind=', $engine->zend_vm_kind(), "\n"; // 5 = ZEND_VM_KIND_TAILCALL // ZEND_ADD = 1; ZEND_USER_OPCODE_DISPATCH = 2 $handler = static function ($executeData): int { return 2; }; $engine->zend_set_user_opcode_handler(1, $handler); $payload = tempnam(sys_get_temp_dir(), 'p') . '.php'; file_put_contents($payload, '<?php function f(int $a, int $b): int { return $a + $b; } var_dump(f(1, 2));'); require $payload; // compiled after install => its ADD dispatches through the handler unlink($payload); echo "OK\n";
On hybrid/call builds:
int(3)+OK. On a TAILCALL build (macOS arm64):vm_kind=5then SIGSEGV at the first ADD insidef(). Verified on GitHub'smacos-latest(arm64, crash) vsmacos-15-intel(x64, hybrid, passes): https://github.com/lisachenko/z-engine/actions/runs/33257997984Analysis (
Zend/zend_vm_execute.has generated for 8.6.0beta2):ZEND_USER_OPCODE_SPEC_TAILCALL_HANDLER's continuation cases use the genericZEND_VM_DISPATCHmacro, which is never redefined for the tail-call section:#define ZEND_VM_DISPATCH(opcode, opline) return zend_vm_get_opcode_handler_func(opcode, opline)(ZEND_OPCODE_HANDLER_ARGS_PASSTHRU); ... case ZEND_USER_OPCODE_DISPATCH: ZEND_VM_DISPATCH(opline->opcode, opline); default: ZEND_VM_DISPATCH((uint8_t)(ret & 0xff), opline);
zend_vm_get_opcode_handler_func()resolves the single-step fastcall variant (execute one op, return the next opline /ENTER_BITstate). The tail-call handler then plain-returns that value — which, because every frame transition inside the chain happened viamusttail, unwinds straight toexecute_ex(). Its loop:opline = (opline->handler)(ZEND_OPCODE_HANDLER_ARGS_PASSTHRU); if (UNEXPECTED(((uintptr_t)opline & ZEND_VM_ENTER_BIT))) { opline = (const zend_op*)((uintptr_t)opline & ~ZEND_VM_ENTER_BIT); if (EXPECTED(opline != NULL)) { execute_data = EG(current_execute_data); /* refreshed ONLY here */ ...
only refreshes its
execute_datalocal in theENTER_BITbranch. A user opcode returning DISPATCH for a normal advancing opcode yields a plain next-opline, so the loop's next iteration executes the current frame's oplines againstexecute_ex's entry frame. With mismatchedEX(run_time_cache)the very nextINIT_FCALL/DO_*CALLresolves a cached callee of the wrong frame — in our traces afwrite(STDERR, "...")in the included file executed as the outer script's cachedunlink()(unlink(): Argument #2 ($context) must be of type resource or null, string givenwith the outer file's name and the inner file's line), and object construction dispatched to an unrelated class's constructor before the SIGSEGV.ZEND_USER_OPCODE_CONTINUE/ENTER/LEAVEare unaffected (they stay on the musttail chain / go throughZEND_VM_ENTER()); DISPATCH of frame-changing opcodes survives by luck via theENTER_BITpath.Suggested fix: in the tail-call variant of the
ZEND_USER_OPCODEhandler (viazend_vm_gen.php), don't return the single-step result up the chain — resume it:case ZEND_USER_OPCODE_DISPATCH: opline = zend_vm_get_opcode_handler_func(opline->opcode, opline)(ZEND_OPCODE_HANDLER_ARGS_PASSTHRU); goto user_opcode_resume; default: opline = zend_vm_get_opcode_handler_func((uint8_t)(ret & 0xff), opline)(ZEND_OPCODE_HANDLER_ARGS_PASSTHRU); user_opcode_resume: if (UNEXPECTED(((uintptr_t)opline & ZEND_VM_ENTER_BIT))) { opline = (const zend_op*)((uintptr_t)opline & ~ZEND_VM_ENTER_BIT); if (UNEXPECTED(opline == NULL)) { ZEND_VM_RETURN(); } execute_data = EG(current_execute_data); } ZEND_VM_CONTINUE();
(
zend_vm_call_opcode_handler()'s TAILCALL branch performs exactly this unwrap, so the same pattern applied inside the handler restores the invariant.)
Supporting evidence from the diagnostics on this branch (
tools/diagnostics/issue-280/, runs on PR #281):install-onlypasses; every mode that lets the handler fire in a nested frame corrupts (probe run: https://github.com/lisachenko/z-engine/actions/runs/33256998594)- crash report:
EXC_BAD_ACCESS KERN_INVALID_ADDRESS at 0x10inZEND_DO_UCALL_SPEC_RETVAL_USED_TAILCALL_HANDLER - identical probes green on PHP 8.5 arm64 and PHP 8.6 x64/linux
Generated by Claude Code
- added a commit that references this issue
on Aug 29, 2026 - added a commit that references this issue
on Sep 30, 2026 Filed upstream: php/php-src#24081 — user opcode handlers returning
ZEND_USER_OPCODE_DISPATCHmis-resume execution against a stale frame under the tail-call VM.Reference points for anyone following along:
- Pure ext-ffi repro (no z-engine):
tools/diagnostics/issue-280/pure-ffi-repro.php; latest reproduction on PHP 8.6.0-dev @dc1145e(macOS arm64,vm_kind=5, SIGSEGV) in run 36771134832; original discovery on 8.6.0-dev @ec93b0b. - The affected code path (handler,
ZEND_VM_DISPATCHmacro,execute_ex()resume loop) is byte-identical on the PHP-8.5 branch and was already in the 8.5.0 release — 8.5 builds in the wild just don't typically selectZEND_VM_KIND_TAILCALL, so the bug stays latent there. - Mitigation on z-engine's side: the
OpCodeHook::install()guard in feat: PHP 8.6 support for user opcode hooks (tail-call VM bug diagnosed, fixed upstream) #281; once #24081 is resolved, re-dispatch theDiagnose issue 280workflow against the fixed build and the guard can be retired.
Generated by Claude Code
- Pure ext-ffi repro (no z-engine):
The bug affects the tail-call VM (ZEND_VM_KIND_TAILCALL) as such and has been present since the VM landed in 8.5.0 — the generated handler and execute_ex() loop are identical in PHP-8.5 and PHP-8.6. Reproduced on 8.6.0-dev builds that select the tail-call VM (vm_kind=5); 8.5 binaries in the wild are typically unaffected only because common builds don't satisfy the HAVE_MUSTTAIL/HAVE_PRESERVE_NONE selection and fall back to the CALL VM. Any 8.5 build that does select it should reproduce with the same ext-ffi script.
Will be fixed once php/php-src#24109 merged. @claude We can try to use this branch to see if it fixes the bug or not. Report here the results.
Summary
On macOS arm64 with PHP 8.6 (Homebrew/setup-php build), installing a user opcode handler via
OpCode::setHandler()onEXT_STMT(withCompiler::COMPILE_EXTENDED_STMTenabled) corrupts VM state as soon as the handler's PHP closure runs: closure bound variables read back as garbage (e.g. anintcounter reads as aZEngine\Coreobject), and full consumers (lisachenko/zdebug) segfault (EXC_BAD_ACCESS at 0x10insideZEND_DO_UCALL_SPEC_RETVAL_USED_TAILCALL_HANDLER).The same code is green on:
z-engine's own CI is green on macOS arm64 + 8.6 because
OpCodeHookTesthooksADDinside a small probe function withoutCOMPILE_EXTENDED_STMT; the corruption shows once EXT_STMT-instrumented code (compiled after the handler is installed) starts dispatching through the hook.Why only arm64 + 8.6
PHP 8.6 introduced a new VM dispatch kind,
ZEND_VM_KIND_TAILCALL(Zend/zend_vm_opcodes.hin php-8.6.0beta2):HAVE_GCC_GLOBAL_REGS, so the build selects TAILCALL: opcode handlers becomeconst zend_op *(*)(zend_execute_data *, const zend_op *)with thepreserve_nonecalling convention, chained viamusttail. The crash frames (*_TAILCALL_HANDLER) confirm the failing build runs this VM.zend_user_opcode_handlers[opline->opcode](execute_data)is still called with the classicint (*)(zend_execute_data *)signature fromZEND_USER_OPCODE_SPEC_TAILCALL_HANDLER, so the trampoline installation itself works (a trivial handler onADDfires fine). The breakage appears when the handler re-enters PHP execution (the FFI callback runs a userland closure) while the interrupted frame's opcode came from EXT_STMT-instrumented code — pointing at an interaction between libffi callback re-entry, nestedexecute_ex, and thepreserve_nonehandler chain. It may ultimately be a php-src beta bug rather than z-engine's, but z-engine is where consumers hit it, and where a repro/workaround can live.Minimal repro (pure z-engine, no consumer code)
probe.php:Run with
php -d ffi.enable=1 -d opcache.jit=off probe.phpon macOS arm64 + PHP 8.6.0-dev:The by-ref
use (&$fires)int reads back as aZEngine\Coreobject on the very first handler invocation. AddingRETURN/THROWhandlers fails the same way (Cannot use object of type ZEngine\Core as arrayon an array counter). On PHP 8.5 (same machine, same script) everything passes.Crash signature from a full consumer (zdebug)
(reading offset 0x10 of a NULL
zend_function, i.e. a call frame whosefuncwas clobbered). Downstream CI evidence: lisachenko/zdebug#24 — the diagnostic jobDiagnose arm64 (PHP 8.5/8.6)in run https://github.com/lisachenko/zdebug/actions/runs/33254290174 shows the layered probes (L1Core::init✅, L2COMPILE_EXTENDED_STMT✅, L3 EXT_STMT handler ❌ on 8.6 / ✅ on 8.5, L5 module registration ✅, L6 full boot ❌ SIGSEGV) plus the lldb backtrace and the macOS.ipscrash report.Suggested directions
ZEND_VM_KIND_TAILCALLatCore::init()(e.g. viaphp -i/Reflectionon the build metadata or a runtime probe) and fail fast with a clear message instead of corrupting the debuggee.