Making Deep Q-learning methods robust to time discretization
Citations Over Time
Abstract
Despite remarkable successes, Deep Reinforcement Learning (DRL) is not robust to hyperparameterization, implementation details, or small environment changes (Henderson et al. 2017, Zhang et al. 2018). Overcoming such sensitivity is key to making DRL applicable to real world problems. In this paper, we identify sensitivity to time discretization in near continuous-time environments as a critical factor; this covers, e.g., changing the number of frames per second, or the action frequency of the controller. Empirically, we find that Q-learning-based approaches such as Deep Q- learning (Mnih et al., 2015) and Deep Deterministic Policy Gradient (Lillicrap et al., 2015) collapse with small time steps. Formally, we prove that Q-learning does not exist in continuous time. We detail a principled way to build an off-policy RL algorithm that yields similar performances over a wide range of time discretizations, and confirm this robustness empirically.
Related Papers
- → Overview of deep learning in medical imaging(2017)1,052 cited
- → Deep learning ensemble 2D CNN approach towards the detection of lung cancer(2023)173 cited
- → Exploring Deep Learning for View-Based 3D Model Retrieval(2020)109 cited
- → A Comprehensive Analysis of Machine Learning Techniques in Biomedical Image Processing Using Convolutional Neural Network(2022)18 cited
- → Survey of Machine Learning Applications of Convolutional Neural Networks to Medical Image Analysis(2021)2 cited