BanditRLlib
Lean gate passed before this site build; local proof declarations are shown as compiled.Lean-verified build · exact declarations linked.

Planned reading map

Online Learning Book

Convex optimization and regret minimization. Existing EXP3/FTRL routes are shared references; adjacent online-learning frontier questions are indexed separately from core Bandit/RL open problems.

Online Learning: A Modern Introduction Using Convex Optimization

Francesco Orabona

arXiv:1912.13213v10, June 21, 2026; final preprint

Read the source ↗ · Official source page ↗

Bibliographic metadata checked 2026-09-09. Page and theorem mappings are separately audited.

Existing shared reading

These links reuse established pages with their original sources and exact Lean boundaries. They do not certify a chapter of the new book.

  1. 2. Probability, kernels, filtrations, and concentration
  2. 7. EXP3 and adversarial concentration
  3. 8. Tsallis-FTRL, corruption, and nonstationarity

Planned source mapping

Next: freeze source versions, chapter contracts, assumptions and theorem locators; retrieve existing declarations and prove only the missing interfaces. Chapter numbers, page coverage and completion totals will appear after that audit.

One underlying Lean graph

All references resolve to canonical declarations in the global index. Reading views do not create additional Lean modules.

Explore the graph · Download the shared reference registry