Policy Gradient Variance Reduction Techniques for Sparse Reward Environments
Practical techniques that tame gradient noise when rewards are scarce and rollouts are few.
Omar Delgado
Section
1 story in RL Algorithms and Methods.
Practical techniques that tame gradient noise when rewards are scarce and rollouts are few.