BanditRLlib
Lean gate passed before this site build; local proof declarations are shown as compiled.Lean-verified build · exact declarations linked.

Setting → new technique → proof mechanism

Bandit Technique Map

Conceptual map from bandit/RL settings and objectives to the proof/algorithmic techniques that materially change the analysis. These relations are teaching/research overlays, not Lean theorem dependencies.

How to read this page. Taxonomy says what changes in the problem. Technique Map says what new mathematical move handles that change. Functor Hypergraph asks whether the move is structurally reusable across domains. Lean Graph records what is actually formalized.

25 technique families

confidence / optimism · mapped

Confidence sets → optimism → regret

Construct a high-probability confidence object, choose an optimistic action/value, and charge instantaneous regret to confidence width before summing widths.

Used by settings. Multi-Armed Bandits (MAB) · Stochastic bandits · Linear bandits · Generalized linear bandits (GLB) · Kernel / RKHS bandits · Gaussian-process UCB (GP-UCB) · Variance-aware bandits · MDP reinforcement learning · Linear / kernel MDPs

Functor family. family:optimism-confidence-regret

Current Lean routes. teaching:ucb · teaching:oful · teaching:finite-horizon-rl

pure exploration · mapped

Elimination / sequential testing

Maintain confidence sets, eliminate implausible arms/hypotheses, and convert fixed-confidence testing error into sample-complexity bounds.

Used by settings. Best-arm identification (BAI) · Threshold / thresholding bandits · Top-m / combinatorial pure exploration · Elimination / successive rejects / racing

Functor family. family:information-change-of-measure

Current Lean routes. teaching:etc · spine:chapter-14-information-theory · spine:chapter-16-instance-dependent

adversarial online learning · mapped

Exponential weights / FTRL / mirror descent

Telescope a regularized potential while controlling bandit estimators and stability, producing comparator regret under adversarial feedback.

Used by settings. Adversarial bandits · Semi-adversarial / best-of-both-worlds bandits · EXP3 / FTRL / mirror-descent methods · Parameter-free / adaptive bandits · Model selection / corralling

Functor family. family:regularized-potential-stability

Current Lean routes. teaching:exp3 · teaching:tsallis

metric structure · source-map

Zooming / metric localization

Localize exploration to near-optimal regions of a metric action space; covering/packing or near-optimality dimension replaces ambient cardinality.

Used by settings. Lipschitz / continuum-armed bandits

Functor family. family:metric-localization

Current Lean routes. No compiled route mapped yet.

bandit convex optimization · source-map

Bandit convex smoothing / gradient estimation

Smooth a convex loss and build a one- or multi-point gradient estimator from bandit feedback, then apply convex optimization regret machinery.

Used by settings. Bandit convex optimization

Functor family. family:regularized-potential-stability

Current Lean routes. No compiled route mapped yet.

constraints / resources · source-map

Primal–dual / Lagrangian resource control

Introduce dual prices or feasibility certificates so reward optimization and cumulative resource/safety constraints are controlled together.

Used by settings. Constrained bandits · Bandits with knapsacks / budgeted bandits · Safe bandits · Constrained / safe RL · Risk-sensitive / CVaR / survival bandits · Multi-fidelity bandits

Functor family. family:oracle-relaxation

Current Lean routes. No compiled route mapped yet.

causal structure · source-map

Causal intervention / observational transfer

Use a structural causal model to transfer information across interventions or combine observational and interventional data.

Used by settings. Causal bandits

Functor family. family:structured-information-transfer

Current Lean routes. No compiled route mapped yet.

preference feedback · source-map

Pairwise preference / dueling reductions

Replace scalar reward observations by comparisons and analyze preference matrices, Condorcet/Borda structure, or reductions to ordinary bandits.

Used by settings. Dueling / preference bandits

Functor family. family:structured-information-transfer

Current Lean routes. No compiled route mapped yet.

structural dimension · mapped

Low-dimensional / sparse / factorized structure

Exploit linear, sparse, low-rank, tensor, factorized, unimodal, or parametric structure so confidence and exploration depend on intrinsic rather than ambient complexity.

Used by settings. Linear bandits · Generalized linear bandits (GLB) · Matrix / low-rank / tensor bandits · Sparse / high-dimensional bandits · Factored bandits · Unimodal / structured-order bandits · Multinomial-logit (MNL) / assortment bandits

Functor family. family:structured-dimension

Current Lean routes. teaching:oful

time variation · source-map

Change detection / windows / variation-budget adaptation

Forget, restart, window, discount, or meta-combine learners so regret adapts to switches, drift, variation, delayed feedback, or changing arm availability.

Used by settings. Dynamic / nonstationary bandits · Delayed-feedback bandits · Sleeping / availability bandits · Ballooning / growing-arm bandits · Rotting / rising / restless bandits · Batched / limited-adaptivity bandits · Parameter-free / adaptive bandits

Functor family. family:adaptation-meta

Current Lean routes. No compiled route mapped yet.

privacy · source-map

Privacy accounting / randomized perturbation

Inject and account for privacy-preserving noise while maintaining confidence or regret guarantees under central/local/joint privacy semantics.

Used by settings. Private / JDP bandits

Functor family. family:robust-confidence

Current Lean routes. No compiled route mapped yet.

multi-objective · source-map

Scalarization / Pareto / multi-objective selection

Replace one scalar reward with a vector criterion and use scalarization, Pareto identification, or preference-dependent selection without collapsing objectives silently.

Used by settings. Multi-objective / Pareto bandits · Single-objective bandits

Functor family. family:objective-transport

Current Lean routes. No compiled route mapped yet.

offline RL · source-map

Coverage / concentrability / pessimism

Replace online exploration by assumptions on dataset coverage or concentrability and use pessimistic value estimation to avoid unsupported actions.

Used by settings. Offline / batch RL

Functor family. family:coverage-vs-exploration

Current Lean routes. No compiled route mapped yet.

partial observability · source-map

Belief-state / memory representation

Replace fully observed states by information states, beliefs, finite memory, or predictive representations before planning/learning.

Used by settings. Partially observable MDP (POMDP)

Functor family. family:representation-lift

Current Lean routes. No compiled route mapped yet.

application bridge · application-bridge

Preference / routing / allocation reductions for LLM systems

Map an LLM-system decision problem—preference selection, model routing, evaluation, or inference-time allocation—to an explicit bandit/RL feedback and objective contract before applying theory.

Used by settings. Bandits for LLM / language-model systems

Functor family. family:representation-lift

Current Lean routes. No compiled route mapped yet.

quantum access · cross-library

Quantum mean estimation / quantum testing / quantum design

Exploit coherent quantum reward-oracle access for faster mean estimation, and use quantum hypothesis-testing/polynomial-method lower bounds to characterize the new regret scale.

Used by settings. Quantum bandits

Functor family. family:quantum-estimation-testing

Current Lean routes. teaching:probability · teaching:oful · spine:chapter-14-information-theory · spine:chapter-15-minimax-lower-bounds

Primary/current source anchors
Cross-library substrates

Truth boundary

A setting→technique link is a reviewed teaching/research relation. It becomes a formal Lean dependency only after a concrete declaration calls a compiled interface. Cross-library candidates remain dashed until an adapter/import is verified.