Summary
With two TCG translation threads (-accel tcg,thread=multi, the default when
-icount is not used), qemu-system-xtensa -machine esp32s3 dies with SIGSEGV
a few seconds into a guest boot, roughly one run in fifteen, but only when the
host is busy. The fault is inside generated code and the address it dies on is
a guest address, which points at a softmmu TLB entry being used while its
addend is zero.
With -accel tcg,thread=single it does not happen at all. That is what we run
with now, so this is not blocking us; it is filed because the underlying defect
is still there for anyone who does not.
Environment
- Tag
esp-develop-9.2.2-20260417 (QEMU 9.2.2), -machine esp32s3.
- Host: x86-64 Linux (Ubuntu 24.04 container), 14 cores.
- Guest: dual-core ESP32-S3 running an unmodified Arduino/ESP-IDF firmware
(Meshtastic 2.7.10 for the LilyGO T-Deck), 8 MB octal PSRAM.
What the fault looks like
Four captures, from separate processes:
=== signal 11 at address 0x3fcb45eb, faulting pc 0x740e800763bd ===
=== signal 11 at address 0x3fcb45eb, faulting pc 0x70050807693d ===
=== signal 11 at address 0x3fcb45eb, faulting pc 0x733e7007643d ===
=== signal 11 at address 0x3fcf14db, faulting pc 0x61a102c3de3b ===
Two things stand out.
si_addr is a guest address — 0x3fcb45eb and 0x3fcf14db are both in the
S3's internal SRAM (0x3FC88000+). It repeats across processes, so it is not a
wild host pointer; ASLR would have moved that. The softmmu fast path reaches
host memory as guest_addr + entry->addend, so faulting on the guest address
itself means addend was 0 at the moment the access was made — the entry was
read while it held nothing.
The faulting PC is outside the binary's executable segment
(0x374000–0x9f612d), i.e. in the generated-code buffer, which is consistent
with the access coming from a TLB fast path rather than a helper.
Reproduction
It needs load. Fourteen concurrent boots, each with its own build directory:
for i in $(seq 1 14); do
qemu-system-xtensa -machine esp32s3 -accel tcg,thread=multi \
-drive file=flash$i.bin,if=mtd,format=raw \
-m 8M -display none -no-reboot ... &
done
wait
- 14 parallel runs, thread=multi: ~1 fault per 15 runs (2/28, then 3/56).
- The same runs sequentially on an idle host: 0 faults in 28.
- Under
gdb: 0 faults in 12+ runs. The timing shift closes the window, which
is why this went unexamined for a long time here.
Since the crash does not survive being watched, the captures above come from a
SIGSEGV handler installed inside the process that prints si_addr and the
ucontext RIP, rather than from a debugger or a core file.
Caveat on the reproduction
I should be straight about this: our boot uses several out-of-tree devices of
our own (a bus bridge that forwards unmodelled MMIO to a host process over a
socket, plus GPIO/SPI/I2C/I2S models). They are not in your tree, so this is not
a reproduction you can run as-is, and I cannot rule out that they are involved.
Two things argue they are not the mechanism:
- They only ever call
memory_region_add_subregion_overlap at device realize.
Nothing of ours changes a mapping, enables/disables a region, or flushes a
TLB while the machine is running.
- The fault is on plain internal SRAM, not on any address our devices claim.
What I could not do is a clean A/B: with our devices removed the guest only
reaches about half as far into its boot before the harness stops it, so "no
crash without them" would not have meant anything.
What I ruled out
hw/misc/esp32s3_cache.c writes guest RAM directly through
memory_region_get_ram_ptr() in two places — filling a page from flash via
blk_pread(), and stamping a discarded page with 0xdeadbeef — without
telling QEMU the memory changed. That is a real hazard on its own (any code
already translated from those pages stays live over contents that no longer
exist), and it was my best candidate.
It is not this defect. Declaring both writes with memory_region_set_dirty()
changed the crash rate not at all: 3 faults in 56 runs, against the same rate
before. I reverted it rather than carry a patch that fixes nothing, but the
missing invalidation is probably still worth a separate look.
Workaround
-accel tcg,thread=single. On this workload it costs nothing measurable — the
guest's peripherals are serialised through a single socket anyway, so a second
translation thread mostly contends — and in one paired measurement it was
several times faster.
Summary
With two TCG translation threads (
-accel tcg,thread=multi, the default when-icountis not used),qemu-system-xtensa -machine esp32s3dies with SIGSEGVa few seconds into a guest boot, roughly one run in fifteen, but only when the
host is busy. The fault is inside generated code and the address it dies on is
a guest address, which points at a softmmu TLB entry being used while its
addendis zero.With
-accel tcg,thread=singleit does not happen at all. That is what we runwith now, so this is not blocking us; it is filed because the underlying defect
is still there for anyone who does not.
Environment
esp-develop-9.2.2-20260417(QEMU 9.2.2),-machine esp32s3.(Meshtastic 2.7.10 for the LilyGO T-Deck), 8 MB octal PSRAM.
What the fault looks like
Four captures, from separate processes:
Two things stand out.
si_addris a guest address —0x3fcb45eband0x3fcf14dbare both in theS3's internal SRAM (
0x3FC88000+). It repeats across processes, so it is not awild host pointer; ASLR would have moved that. The softmmu fast path reaches
host memory as
guest_addr + entry->addend, so faulting on the guest addressitself means
addendwas0at the moment the access was made — the entry wasread while it held nothing.
The faulting PC is outside the binary's executable segment
(
0x374000–0x9f612d), i.e. in the generated-code buffer, which is consistentwith the access coming from a TLB fast path rather than a helper.
Reproduction
It needs load. Fourteen concurrent boots, each with its own build directory:
gdb: 0 faults in 12+ runs. The timing shift closes the window, whichis why this went unexamined for a long time here.
Since the crash does not survive being watched, the captures above come from a
SIGSEGVhandler installed inside the process that printssi_addrand theucontextRIP, rather than from a debugger or a core file.Caveat on the reproduction
I should be straight about this: our boot uses several out-of-tree devices of
our own (a bus bridge that forwards unmodelled MMIO to a host process over a
socket, plus GPIO/SPI/I2C/I2S models). They are not in your tree, so this is not
a reproduction you can run as-is, and I cannot rule out that they are involved.
Two things argue they are not the mechanism:
memory_region_add_subregion_overlapat device realize.Nothing of ours changes a mapping, enables/disables a region, or flushes a
TLB while the machine is running.
What I could not do is a clean A/B: with our devices removed the guest only
reaches about half as far into its boot before the harness stops it, so "no
crash without them" would not have meant anything.
What I ruled out
hw/misc/esp32s3_cache.cwrites guest RAM directly throughmemory_region_get_ram_ptr()in two places — filling a page from flash viablk_pread(), and stamping a discarded page with0xdeadbeef— withouttelling QEMU the memory changed. That is a real hazard on its own (any code
already translated from those pages stays live over contents that no longer
exist), and it was my best candidate.
It is not this defect. Declaring both writes with
memory_region_set_dirty()changed the crash rate not at all: 3 faults in 56 runs, against the same rate
before. I reverted it rather than carry a patch that fixes nothing, but the
missing invalidation is probably still worth a separate look.
Workaround
-accel tcg,thread=single. On this workload it costs nothing measurable — theguest's peripherals are serialised through a single socket anyway, so a second
translation thread mostly contends — and in one paired measurement it was
several times faster.