ML
Hard
+20 XP
ReinforcementLearning: on vs off policy
Off-policy methods can learn from data generated by another:
Sign in to solve this challenge
Use your college email to start solving, earn XP, build a streak, and climb your college leaderboard. It is free.
Sign in with Google