Batch Constrained Q-learning (BCQ)
Batch Constrained Q-learning (BCQ)
Offline reinforcement learning algorithm that constrains policies to remain close to actions observed in the training dataset to avoid extrapolation errors. BCQ uses an action generator model to produce actions similar to those in the batch while exploring slight variations.
← Geri