Skip to content

Population cannot forward a standard reward-free restart timestep #417

Description

@sylvesterkaczmarek

Population.send_timestep() unconditionally indexes timestep.reward, so a standard dm_env.restart(...) fails before any policy receives the initial timestep.

Reproduced on main at 1b4239b54303ed593e2335c219f453300b5885dd with the actual population executor:

import dm_env
from meltingpot.utils.policies import fixed_action_policy
from meltingpot.utils.scenarios import population

agents = population.Population(
    policies={'bot': fixed_action_policy.FixedActionPolicy(0)},
    names_by_role={'default': ['bot']}, roles=['default'])
try:
    agents.reset()
    agents.send_timestep(dm_env.restart([{}]))
    print(agents.await_action())
finally:
    agents.close()

The send raises TypeError: 'NoneType' object is not subscriptable. Preserve reward=None in each player's timestep, while continuing to select per-player entries for an explicit reward sequence.

A focused correction passes eleven regression/control cases, including one/two/four players, shared policies with separate recurrent states, repeated episodes, observable forwarding and the actual evaluation.run_episode loop over a deterministic local environment. Original code fails seven and passes four controls.

This concerns population timestep distribution, separately from the return accumulator in #360. No change to sampling, locking, shutdown, numeric rewards or scenario allocation is proposed. A focused PR is being prepared.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions