Streaming Reinforcement Learning2
- Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
Action Value Gradient enables deep policy-gradient reinforcement learning with fully incremental updates, without replay buffers, target networks, or batch updates.
NeurIPS, 2024readstreaming-rlincremental-learningpolicy-gradientactor-criticavgreinforcement-learning - Streaming Deep Reinforcement Learning Finally Works
Stream-X introduces Stream TD, Stream Q, and Stream AC, enabling deep reinforcement learning directly from a stream of experience without replay buffers or batch updates.
arXiv preprint, 2024readstreaming-rlincremental-learningstream-xtemporal-difference-learningq-learningactor-critic
Change Point Deteciton1
- Minimum-Delay Adaptation in Non-Stationary Reinforcement Learning via Online High-Confidence Change-Point Detection
Detecting recurring and novel MDP contexts online with MCUSUM, while maintaining context-specific dynamics models and policies.
AAMAS, 2021readnon-stationary-rlchange-point-detectioncusummodel-based-rlcontinual-learning