Policy Gradient Variance Reduction Techniques for Sparse Reward Environments
Practical techniques that tame gradient noise when rewards are scarce and rollouts are few.
Omar Delgado
Contributing Editor
Omar Delgado is a contributing editor at The Calibration Review covering rl algorithms and methods. Based in Melbourne, Omar has written for The Calibration Review since 2015.
1 story · Melbourne
Practical techniques that tame gradient noise when rewards are scarce and rollouts are few.