Pearl
Free while signed in. Answers cite the passages they came from.

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.
Production-oriented design: Built for real-world environments with limited observability, sparse feedback, and high stochasticity - conditions that usually break research-oriented RL libraries.
Modular components: Offers modular policy networks, exploration strategies, offline RL, and safety constraints that can be composed for specific applications.
Research + practice: Targets both researchers building new RL agents and practitioners deploying RL in production recommender systems, ranking, and control.
Meta internal use: Reflects learnings from Meta's internal deployments, making it a rare RL library that starts from production pain rather than benchmark scores.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack