BanditRLlib
Lean gate passed before this site build; local proof declarations are shown as compiled.Lean-verified build · exact declarations linked.

Primary-source theorem index

BanditRLwiki Papers

19 primary papers expose 27 indexed theorem surfaces across 13 assumption-compatible comparison cases.

Audit boundary

Links point to primary papers, proceedings pages, DOIs, or the author-hosted textbook. Rate summaries are independently written comparison statements; source prose, templates, and data are not copied.

Sources and indexed cases

Primary paper · 1 theorem surface

Asymptotically Efficient Adaptive Allocation Rules

Tze Leung Lai and Herbert Robbins · 1985

Theorem surfaces

  • LowerLai–Robbins asymptotic information lower boundAsymptotic information lower bound · stochastic-instance-dependent-klucb

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

The Nonstochastic Multiarmed Bandit Problem

Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert Schapire · 2002

Theorem surfaces

  • UpperEXP3 expected-regret upper boundEXP3 expected-regret theorem · adversarial-exp3
  • LowerAdversarial minimax lower boundSection 5 lower bound · adversarial-exp3

Indexed cases

Primary source ↗

Primary paper · 1 theorem surface

Contextual Bandit Algorithms with Supervised Learning Guarantees

Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2010

Theorem surfaces

  • UpperExp4.P high-probability policy-regret upper boundTheorem 2 · contextual-finite-policy-exp4p

Indexed cases

Primary source ↗

Primary paper · 1 theorem surface

The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond

Aurélien Garivier and Olivier Cappé · 2011

Theorem surfaces

  • UpperKL-UCB asymptotic pull-count upper boundTheorems 1–2 and Corollary 3 · stochastic-instance-dependent-klucb

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

Stochastic Multi-Armed-Bandit Problem with Non-stationary Rewards

Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014

Theorem surfaces

  • UpperRexp3 variation-budget upper boundTheorem 2 · nonstationary-variation-budget
  • LowerVariation-budget dynamic-regret lower boundTheorem 1 · nonstationary-variation-budget

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

Anytime optimal algorithms in stochastic multi-armed bandits

Rémy Degenne and Vianney Perchet · 2016

Theorem surfaces

  • UpperAnytime MOSS regretTheorem 3 with Lemma 3 · stochastic-finite-arm-minimax
  • LowerMinimax stochastic-bandit lower boundLemma 3 · stochastic-finite-arm-minimax

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

Optimal Best Arm Identification with Fixed Confidence

Aurélien Garivier and Emilie Kaufmann · 2016

Theorem surfaces

  • UpperTrack-and-Stop asymptotic upper boundTheorem 14 · fixed-confidence-best-arm-identification
  • LowerCharacteristic-time lower boundTheorem 1 · fixed-confidence-best-arm-identification

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

Minimax Regret Bounds for Reinforcement Learning

Mohammad Azar, Ian Osband, and Rémi Munos · 2017

Theorem surfaces

  • UpperUCBVI high-probability regret upper boundsTheorems 1–2 · tabular-finite-horizon-rl-ucbvi
  • LowerTabular episodic minimax lower-bound comparisonMinimax lower-bound comparison · tabular-finite-horizon-rl-ucbvi

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

Delay and Cooperation in Nonstochastic Bandits

Nicolò Cesa-Bianchi, Claudio Gentile, and Yishay Mansour · 2019

Theorem surfaces

  • UpperFixed-delay regret upper boundCorollary 15 · delayed-adversarial-bandit
  • LowerFixed-delay minimax comparisonFixed-delay minimax comparison · delayed-adversarial-bandit

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

Nearly Minimax-Optimal Regret for Linearly Parameterized Bandits

Lihong Li, Wei Wang, and Zhihua Zhou · 2019

Theorem surfaces

  • UpperFinite-action linear contextual upper boundTheorem 1 · finite-action-linear-contextual
  • LowerFinite-action linear contextual lower boundTheorem 2 · finite-action-linear-contextual

Indexed cases

Primary source ↗

Primary paper · 1 theorem surface

Nonstochastic Multiarmed Bandits with Unrestricted Delays

Tobias Thune, Nicolò Cesa-Bianchi, and Yevgeny Seldin · 2019

Theorem surfaces

  • UpperKnown-delay DEXP3/DEW upper boundTheorem 1 and Corollary 4 · delayed-adversarial-bandit

Indexed cases

Primary source ↗

Primary paper · 1 theorem surface

A Near-Optimal Change-Detection Based Algorithm for Piecewise-Stationary Combinatorial Semi-Bandits

Zhou, Wang, Varshney, and Lim · 2020

Theorem surfaces

  • LowerPiecewise-stationary minimax lower boundTheorem 5.1 · nonstationary-best-arm-switch-budget

Indexed cases

Primary source ↗

Primary paper · 2 theorem surfaces

Tsallis-INF: An Optimal Algorithm for Stochastic and Adversarial Bandits

Julian Zimmert and Yevgeny Seldin · 2021

Theorem surfaces

  • UpperTsallis-INF best-of-both-worlds upper boundsTheorem 1 · adversarial-stochastic-best-of-both-worlds
  • LowerRate-optimality comparison used by the Tsallis-INF analysisLower-bound comparisons summarized with Theorem 1 · adversarial-stochastic-best-of-both-worlds

Indexed cases

Primary source ↗

Primary paper · 1 theorem surface

A New Look at Dynamic Regret for Non-Stationary Stochastic Bandits

Yasin Abbasi-Yadkori, András György, and Nevena Lazić · 2023

Theorem surfaces

  • UpperArmSwitch best-arm-switch upper boundTheorem 1 · nonstationary-best-arm-switch-budget

Indexed cases

Primary source ↗