Two independent ledgers
BanditRLwiki Progress
Source coverage and Lean completion are counted separately; neither percentage estimates how much of all bandit or reinforcement-learning theory is complete.
Literature comparison ledger
Local Lean evidence ledger
By assumption family
Finite stochastic bandits2 cases
stochastic-finite-arm-minimax Minimax matched Partial local route
stochastic-instance-dependent-klucb Asymptotically matched Partial local route
Adversarial and best-of-both-worlds bandits2 cases
adversarial-exp3 Near minimax Partial local route
adversarial-stochastic-best-of-both-worlds Near minimax Partial local route
Linear and contextual bandits3 cases
contextual-finite-policy-exp4p Source audit pending Planned
stochastic-linear-oful Near minimax Partial local route
finite-action-linear-contextual Near minimax Planned
Finite-horizon tabular reinforcement learning1 cases
tabular-finite-horizon-rl-ucbvi Near minimax Partial local route
Delayed and nonstationary bandits3 cases
delayed-adversarial-bandit Near minimax Partial local route
nonstationary-variation-budget Near minimax Partial local route
nonstationary-best-arm-switch-budget Near minimax Partial local route
Pure exploration and best-arm identification1 cases
fixed-confidence-best-arm-identification Asymptotically matched Planned
Distributional and high-probability lower bounds1 cases
distributional-high-probability-regret Minimax matched Partial local route