Reinforcement Learning without TD Learning
BAIR Blog1881 字 (约 8 分钟)
85
The article introduces a new reinforcement learning algorithm based on divide and conquer strategy to solve the error accumulation problem in long tasks.
入选理由:分治策略可将Bellman递归次数减少到对数级别
FeaturedArticle#Reinforcement Learning#Divide and Conquer中文
