Summary
The fix for #22 returns AVERROR(EAGAIN) from the BLOCK_PENDING wait when AVIO_FLAG_NONBLOCK is set. However, pending_since is a local in shared_read(), so it resets on every call. The cache_timeout expiry check can only fire within a single call, and a nonblocking call never gets past its first iteration: it sets pending_since and returns EAGAIN.
If the owner died while holding BLOCK_PENDING, which is persistent state in the spacemap, a nonblocking reader gets EAGAIN forever. A blocking reader recovers after cache_timeout.
Location
pending_since local:
|
int64_t pending_since = 0; |
- Wait loop:
|
case BLOCK_PENDING: |
|
/* Another thread is busy fetching this block, wait for it to finish */ |
|
if (!s->timeout) { |
|
break; /* no timeout requested, immediately race to fetch block */ |
|
} else if (pending_since) { |
|
int64_t new = av_gettime_relative(); |
|
if (new - pending_since >= s->timeout) |
|
break; /* timeout expired, try to fetch the block ourselves */ |
|
} else { |
|
pending_since = av_gettime_relative(); |
|
} |
|
|
|
if (h->flags & AVIO_FLAG_NONBLOCK) |
|
return AVERROR(EAGAIN); |
|
|
|
/* Make sure we try a few times before giving up */ |
|
av_usleep(FFMIN(s->timeout >> 4, 10000)); |
|
if (ff_check_interrupt(&h->interrupt_callback)) |
|
return AVERROR_EXIT; |
|
|
|
state = atomic_load_explicit(&block->state, memory_order_acquire); |
|
goto retry; |
|
} |
Reproduction
Cache block 0, then simulate a crashed owner by writing BLOCK_PENDING (1) into block 1's slot in the .spacemap (byte offset 128 + 4). Then seek to 32768 and read with cache_timeout=100000 (harness from #71, with AVIO_FLAG_NONBLOCK added to the open flags):
blocking: [0 EAGAINs, 0.108s] read: 4 (ok)
AVIO_FLAG_NONBLOCK: [2375 EAGAINs, 3.000s] read: -35 (Resource temporarily unavailable) # never recovers
Suggested fix
Keep the wait start in SharedContext, together with the block id it applies to. Reset it when the reader moves to a different block or the block leaves BLOCK_PENDING.
Summary
The fix for #22 returns
AVERROR(EAGAIN)from theBLOCK_PENDINGwait whenAVIO_FLAG_NONBLOCKis set. However,pending_sinceis a local inshared_read(), so it resets on every call. Thecache_timeoutexpiry check can only fire within a single call, and a nonblocking call never gets past its first iteration: it setspending_sinceand returnsEAGAIN.If the owner died while holding
BLOCK_PENDING, which is persistent state in the spacemap, a nonblocking reader getsEAGAINforever. A blocking reader recovers aftercache_timeout.Location
pending_sincelocal:FFmpeg/libavformat/shared.c
Line 660 in fe63f8f
FFmpeg/libavformat/shared.c
Lines 764 to 786 in fe63f8f
Reproduction
Cache block 0, then simulate a crashed owner by writing
BLOCK_PENDING(1) into block 1's slot in the.spacemap(byte offset 128 + 4). Then seek to 32768 and read withcache_timeout=100000(harness from #71, withAVIO_FLAG_NONBLOCKadded to the open flags):Suggested fix
Keep the wait start in
SharedContext, together with the block id it applies to. Reset it when the reader moves to a different block or the block leavesBLOCK_PENDING.