Abstract
In this paper, we study distributional reinforcement learning from the perspective of statistical efficiency. We investigate distributional policy evaluation, aiming to estimate the complete return distribution (denoted ) attained by a given policy . We use the certainty-equivalence method to construct our estimator , given a generative model is available. In this circumstance we need a dataset of size to guarantee the -Wasserstein metric between and less than with high probability. This implies the distributional policy evaluation problem can be solved with sample efficiency. Also, we show that under different mild assumptions a dataset of size suffices to ensure the Kolmogorov metric and total variation metric between and is below with high probability. Furthermore, we investigate the asymptotic behavior of . We demonstrate that the ``empirical process'' converges weakly to a Gaussian process in the space of bounded functionals on Lipschitz function class , also in the space of bounded functionals on indicator function class and bounded measurable function class when some mild conditions hold. Our findings give rise to a unified approach to statistical inference of a wide class of statistical functionals of .