Primary paper · 1 theorem surface
Asymptotically Efficient Adaptive Allocation Rules
Tze Leung Lai and Herbert Robbins · 1985
Theorem surfaces
- LowerLai–Robbins asymptotic information lower boundAsymptotic information lower bound ·
stochastic-instance-dependent-klucb
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
The Nonstochastic Multiarmed Bandit Problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert Schapire · 2002
Theorem surfaces
- UpperEXP3 expected-regret upper boundEXP3 expected-regret theorem ·
adversarial-exp3 - LowerAdversarial minimax lower boundSection 5 lower bound ·
adversarial-exp3
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Theorem surfaces
- LowerLinear-bandit minimax expected-regret lower boundTheorem 3 ·
stochastic-linear-oful
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
Contextual Bandit Algorithms with Supervised Learning Guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2010
Theorem surfaces
- UpperExp4.P high-probability policy-regret upper boundTheorem 2 ·
contextual-finite-policy-exp4p
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
Improved Algorithms for Linear Stochastic Bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Theorem surfaces
- UpperOFUL high-probability regret upper boundTheorem 13 ·
stochastic-linear-oful
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond
Aurélien Garivier and Olivier Cappé · 2011
Theorem surfaces
- UpperKL-UCB asymptotic pull-count upper boundTheorems 1–2 and Corollary 3 ·
stochastic-instance-dependent-klucb
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
Stochastic Multi-Armed-Bandit Problem with Non-stationary Rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Theorem surfaces
- UpperRexp3 variation-budget upper boundTheorem 2 ·
nonstationary-variation-budget - LowerVariation-budget dynamic-regret lower boundTheorem 1 ·
nonstationary-variation-budget
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
Anytime optimal algorithms in stochastic multi-armed bandits
Rémy Degenne and Vianney Perchet · 2016
Theorem surfaces
- UpperAnytime MOSS regretTheorem 3 with Lemma 3 ·
stochastic-finite-arm-minimax - LowerMinimax stochastic-bandit lower boundLemma 3 ·
stochastic-finite-arm-minimax
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
Optimal Best Arm Identification with Fixed Confidence
Aurélien Garivier and Emilie Kaufmann · 2016
Theorem surfaces
- UpperTrack-and-Stop asymptotic upper boundTheorem 14 ·
fixed-confidence-best-arm-identification - LowerCharacteristic-time lower boundTheorem 1 ·
fixed-confidence-best-arm-identification
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
Minimax Regret Bounds for Reinforcement Learning
Mohammad Azar, Ian Osband, and Rémi Munos · 2017
Theorem surfaces
- UpperUCBVI high-probability regret upper boundsTheorems 1–2 ·
tabular-finite-horizon-rl-ucbvi - LowerTabular episodic minimax lower-bound comparisonMinimax lower-bound comparison ·
tabular-finite-horizon-rl-ucbvi
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
Delay and Cooperation in Nonstochastic Bandits
Nicolò Cesa-Bianchi, Claudio Gentile, and Yishay Mansour · 2019
Theorem surfaces
- UpperFixed-delay regret upper boundCorollary 15 ·
delayed-adversarial-bandit - LowerFixed-delay minimax comparisonFixed-delay minimax comparison ·
delayed-adversarial-bandit
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
Nearly Minimax-Optimal Regret for Linearly Parameterized Bandits
Lihong Li, Wei Wang, and Zhihua Zhou · 2019
Theorem surfaces
- UpperFinite-action linear contextual upper boundTheorem 1 ·
finite-action-linear-contextual - LowerFinite-action linear contextual lower boundTheorem 2 ·
finite-action-linear-contextual
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
Nonstochastic Multiarmed Bandits with Unrestricted Delays
Tobias Thune, Nicolò Cesa-Bianchi, and Yevgeny Seldin · 2019
Theorem surfaces
- UpperKnown-delay DEXP3/DEW upper boundTheorem 1 and Corollary 4 ·
delayed-adversarial-bandit
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
A Near-Optimal Change-Detection Based Algorithm for Piecewise-Stationary Combinatorial Semi-Bandits
Zhou, Wang, Varshney, and Lim · 2020
Theorem surfaces
- LowerPiecewise-stationary minimax lower boundTheorem 5.1 ·
nonstationary-best-arm-switch-budget
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Theorem surfaces
- LowerChapter 17 distributional-regret lower boundTheorem 17.1 ·
distributional-high-probability-regret
Indexed cases
Primary source ↗
Primary paper · 2 theorem surfaces
Tsallis-INF: An Optimal Algorithm for Stochastic and Adversarial Bandits
Julian Zimmert and Yevgeny Seldin · 2021
Theorem surfaces
- UpperTsallis-INF best-of-both-worlds upper boundsTheorem 1 ·
adversarial-stochastic-best-of-both-worlds - LowerRate-optimality comparison used by the Tsallis-INF analysisLower-bound comparisons summarized with Theorem 1 ·
adversarial-stochastic-best-of-both-worlds
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
A New Look at Dynamic Regret for Non-Stationary Stochastic Bandits
Yasin Abbasi-Yadkori, András György, and Nevena Lazić · 2023
Theorem surfaces
- UpperArmSwitch best-arm-switch upper boundTheorem 1 ·
nonstationary-best-arm-switch-budget
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
Settling the sample complexity of online reinforcement learning
Zihan Zhang and collaborators · 2024
Theorem surfaces
- UpperModified MVP full-range minimax regret boundModified MVP full-range result ·
tabular-finite-horizon-rl-ucbvi
Indexed cases
Primary source ↗
Primary paper · 1 theorem surface
Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
Lee and Oh · 2026
Theorem surfaces
- UpperEQO+ distributional-regret upper boundTheorem 4 ·
distributional-high-probability-regret
Indexed cases
Primary source ↗