Reproduction
ReturnSubject chooses its accumulator dtype from the first non-None reward and then adds every later reward in place.
- Integer zero rewards followed by fractional rewards raise a NumPy casting error.
- Two
np.int8([120]) rewards produce an episode return of -16 instead of 240.
- Two float16 rewards of
65504 produce infinity even though the sum is representable in float64.
- Repeated small float32 rewards can disappear after a large accumulated reward.
These cases reproduce through the actual ReturnSubject and reactive subscriptions. Standard reward-free FIRST timesteps are already handled; this concerns subsequent numeric accumulation.
Expected behavior
Accumulate with at least float64 precision and allow wider later floating-point values to promote the accumulator. Preserve per-player shape, original reward arrays, episode boundaries and terminal-only emission. Integer outputs would consequently use floating-point representation, consistent with the default Melting Pot reward specification; arbitrary-size exact integer arithmetic is not proposed.
Reproduction
ReturnSubjectchooses its accumulator dtype from the first non-None reward and then adds every later reward in place.np.int8([120])rewards produce an episode return of-16instead of240.65504produce infinity even though the sum is representable in float64.These cases reproduce through the actual ReturnSubject and reactive subscriptions. Standard reward-free FIRST timesteps are already handled; this concerns subsequent numeric accumulation.
Expected behavior
Accumulate with at least float64 precision and allow wider later floating-point values to promote the accumulator. Preserve per-player shape, original reward arrays, episode boundaries and terminal-only emission. Integer outputs would consequently use floating-point representation, consistent with the default Melting Pot reward specification; arbitrary-size exact integer arithmetic is not proposed.