Mastering Atari with Discrete World Models

Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, Jimmy Ba

2021 · ICLR

Mastering Atari with Discrete World Models

Problem

Framing

Atari world models had been sample-efficient but not competitive at 200M-frame scale. DreamerV2 closes that gap with a discrete latent RSSM and latent-space actor-critic training, reaching human-level Atari performance and beating single-GPU Rainbow and IQN on aggregate scores.

Currently Used Methods

Foundational

Proposed Method

Architecture

DreamerV2 keeps the Dreamer RSSM factorization but replaces Gaussian stochastic state with 32 categorical variables of 32 classes each. The world model predicts image, reward, and discount from latent state; actor and critic are MLPs trained on imagined latent trajectories.

World-model learning diagram: CNN encoder, recurrent state h_t, categorical latent z_t, prior \hat{z}_t, and reconstruction/reward heads.

Evaluation

Datasets

Metrics

Headline results

Ablations

Method Strengths and Weaknesses

Strengths

Weaknesses

Suggestions from the authors

Links

Prior Papers

Further Papers