Population.send_timestep() unconditionally indexes timestep.reward, so a standard dm_env.restart(...) fails before any policy receives the initial timestep.
Reproduced on main at 1b4239b54303ed593e2335c219f453300b5885dd with the actual population executor:
import dm_env
from meltingpot.utils.policies import fixed_action_policy
from meltingpot.utils.scenarios import population
agents = population.Population(
policies={'bot': fixed_action_policy.FixedActionPolicy(0)},
names_by_role={'default': ['bot']}, roles=['default'])
try:
agents.reset()
agents.send_timestep(dm_env.restart([{}]))
print(agents.await_action())
finally:
agents.close()
The send raises TypeError: 'NoneType' object is not subscriptable. Preserve reward=None in each player's timestep, while continuing to select per-player entries for an explicit reward sequence.
A focused correction passes eleven regression/control cases, including one/two/four players, shared policies with separate recurrent states, repeated episodes, observable forwarding and the actual evaluation.run_episode loop over a deterministic local environment. Original code fails seven and passes four controls.
This concerns population timestep distribution, separately from the return accumulator in #360. No change to sampling, locking, shutdown, numeric rewards or scenario allocation is proposed. A focused PR is being prepared.
Population.send_timestep()unconditionally indexestimestep.reward, so a standarddm_env.restart(...)fails before any policy receives the initial timestep.Reproduced on main at
1b4239b54303ed593e2335c219f453300b5885ddwith the actual population executor:The send raises
TypeError: 'NoneType' object is not subscriptable. Preservereward=Nonein each player's timestep, while continuing to select per-player entries for an explicit reward sequence.A focused correction passes eleven regression/control cases, including one/two/four players, shared policies with separate recurrent states, repeated episodes, observable forwarding and the actual
evaluation.run_episodeloop over a deterministic local environment. Original code fails seven and passes four controls.This concerns population timestep distribution, separately from the return accumulator in #360. No change to sampling, locking, shutdown, numeric rewards or scenario allocation is proposed. A focused PR is being prepared.