confidence / optimism · mapped
Confidence sets → optimism → regret
Construct a high-probability confidence object, choose an optimistic action/value, and charge instantaneous regret to confidence width before summing widths.
Used by settings. Multi-Armed Bandits (MAB) · Stochastic bandits · Linear bandits · Generalized linear bandits (GLB) · Kernel / RKHS bandits · Gaussian-process UCB (GP-UCB) · Variance-aware bandits · MDP reinforcement learning · Linear / kernel MDPs
Functor family. family:optimism-confidence-regret
Current Lean routes. teaching:ucb · teaching:oful · teaching:finite-horizon-rl
pure exploration · mapped
Elimination / sequential testing
Maintain confidence sets, eliminate implausible arms/hypotheses, and convert fixed-confidence testing error into sample-complexity bounds.
Used by settings. Best-arm identification (BAI) · Threshold / thresholding bandits · Top-m / combinatorial pure exploration · Elimination / successive rejects / racing
Functor family. family:information-change-of-measure
Current Lean routes. teaching:etc · spine:chapter-14-information-theory · spine:chapter-16-instance-dependent
Bayesian · mapped
Posterior sampling / probability matching
Sample from a posterior or posterior-induced optimal-action law instead of constructing an explicit optimistic index.
Used by settings. Thompson sampling / posterior sampling · Multi-Armed Bandits (MAB) · Contextual bandits · Linear bandits · Kernel / RKHS bandits
Functor family. family:posterior-randomization
Current Lean routes. teaching:thompson
robust statistics · mapped
Robust estimation / truncation / robust confidence
Replace sub-Gaussian sample means by estimators or confidence sequences that survive heavy tails, contamination, censoring, or heterogeneous variance.
Used by settings. Heavy-tailed bandits · Corruption-tolerant bandits · Missing / censored outcome bandits · Variance-aware bandits
Functor family. family:robust-confidence
Current Lean routes. teaching:probability · teaching:ucb
variance adaptation · mapped
Second-order / empirical-Bernstein control
Use observed or intrinsic variance/second-moment information so regret scales with variance-like quantities rather than only ranges or worst-case proxies.
Used by settings. Variance-aware bandits · Semi-adversarial / best-of-both-worlds bandits · Adversarial bandits
Functor family. family:robust-confidence
Current Lean routes. teaching:probability · teaching:tsallis
adversarial online learning · mapped
Exponential weights / FTRL / mirror descent
Telescope a regularized potential while controlling bandit estimators and stability, producing comparator regret under adversarial feedback.
Used by settings. Adversarial bandits · Semi-adversarial / best-of-both-worlds bandits · EXP3 / FTRL / mirror-descent methods · Parameter-free / adaptive bandits · Model selection / corralling
Functor family. family:regularized-potential-stability
Current Lean routes. teaching:exp3 · teaching:tsallis
metric structure · source-map
Zooming / metric localization
Localize exploration to near-optimal regions of a metric action space; covering/packing or near-optimality dimension replaces ambient cardinality.
Used by settings. Lipschitz / continuum-armed bandits
Functor family. family:metric-localization
Current Lean routes. No compiled route mapped yet.
kernel methods · source-map
RKHS confidence / information gain
Control nonparametric uncertainty through RKHS geometry, posterior/kernel variance, and information-gain or effective-dimension quantities.
Used by settings. Kernel / RKHS bandits · Gaussian-process UCB (GP-UCB) · Linear / kernel MDPs
Functor family. family:metric-localization
Current Lean routes. No compiled route mapped yet.
bandit convex optimization · source-map
Bandit convex smoothing / gradient estimation
Smooth a convex loss and build a one- or multi-point gradient estimator from bandit feedback, then apply convex optimization regret machinery.
Used by settings. Bandit convex optimization
Functor family. family:regularized-potential-stability
Current Lean routes. No compiled route mapped yet.
structured actions · source-map
Combinatorial oracle / relaxation / semi-bandit estimation
Separate statistical learning from a combinatorial optimization oracle or relaxation; exploit coordinate/semi-bandit observations when available.
Used by settings. Combinatorial bandits · Matching bandits · Bandits with knapsacks / budgeted bandits · Constrained bandits
Functor family. family:oracle-relaxation
Current Lean routes. No compiled route mapped yet.
constraints / resources · source-map
Primal–dual / Lagrangian resource control
Introduce dual prices or feasibility certificates so reward optimization and cumulative resource/safety constraints are controlled together.
Used by settings. Constrained bandits · Bandits with knapsacks / budgeted bandits · Safe bandits · Constrained / safe RL · Risk-sensitive / CVaR / survival bandits · Multi-fidelity bandits
Functor family. family:oracle-relaxation
Current Lean routes. No compiled route mapped yet.
causal structure · source-map
Causal intervention / observational transfer
Use a structural causal model to transfer information across interventions or combine observational and interventional data.
Used by settings. Causal bandits
Functor family. family:structured-information-transfer
Current Lean routes. No compiled route mapped yet.
feedback structure · source-map
Feedback-graph / partial-monitoring information structure
Exploit which actions reveal information about which other actions; graph observability or feedback partitions determine the regret regime.
Used by settings. Graph-feedback / graphical bandits · Partial monitoring
Functor family. family:structured-information-transfer
Current Lean routes. No compiled route mapped yet.
preference feedback · source-map
Pairwise preference / dueling reductions
Replace scalar reward observations by comparisons and analyze preference matrices, Condorcet/Borda structure, or reductions to ordinary bandits.
Used by settings. Dueling / preference bandits
Functor family. family:structured-information-transfer
Current Lean routes. No compiled route mapped yet.
structural dimension · mapped
Low-dimensional / sparse / factorized structure
Exploit linear, sparse, low-rank, tensor, factorized, unimodal, or parametric structure so confidence and exploration depend on intrinsic rather than ambient complexity.
Used by settings. Linear bandits · Generalized linear bandits (GLB) · Matrix / low-rank / tensor bandits · Sparse / high-dimensional bandits · Factored bandits · Unimodal / structured-order bandits · Multinomial-logit (MNL) / assortment bandits
Functor family. family:structured-dimension
Current Lean routes. teaching:oful
time variation · source-map
Change detection / windows / variation-budget adaptation
Forget, restart, window, discount, or meta-combine learners so regret adapts to switches, drift, variation, delayed feedback, or changing arm availability.
Used by settings. Dynamic / nonstationary bandits · Delayed-feedback bandits · Sleeping / availability bandits · Ballooning / growing-arm bandits · Rotting / rising / restless bandits · Batched / limited-adaptivity bandits · Parameter-free / adaptive bandits
Functor family. family:adaptation-meta
Current Lean routes. No compiled route mapped yet.
adaptation · source-map
Corralling / model-selection meta-learning
Run several base learners or model classes under a master algorithm and pay only the unavoidable adaptation overhead.
Used by settings. Model selection / corralling · Parameter-free / adaptive bandits · Semi-adversarial / best-of-both-worlds bandits · Multi-bandit / multi-task bandits
Functor family. family:adaptation-meta
Current Lean routes. No compiled route mapped yet.
multi-agent / federated · source-map
Distributed coordination / communication / collision control
Coordinate exploration across agents while accounting for communication, collisions, heterogeneity, decentralization, or strategic interaction.
Used by settings. Multi-agent bandits · Federated bandits · Multi-bandit / multi-task bandits · Multi-agent RL
Functor family. family:distributed-information
Current Lean routes. No compiled route mapped yet.
privacy · source-map
Privacy accounting / randomized perturbation
Inject and account for privacy-preserving noise while maintaining confidence or regret guarantees under central/local/joint privacy semantics.
Used by settings. Private / JDP bandits
Functor family. family:robust-confidence
Current Lean routes. No compiled route mapped yet.
multi-objective · source-map
Scalarization / Pareto / multi-objective selection
Replace one scalar reward with a vector criterion and use scalarization, Pareto identification, or preference-dependent selection without collapsing objectives silently.
Used by settings. Multi-objective / Pareto bandits · Single-objective bandits
Functor family. family:objective-transport
Current Lean routes. No compiled route mapped yet.
offline RL · source-map
Coverage / concentrability / pessimism
Replace online exploration by assumptions on dataset coverage or concentrability and use pessimistic value estimation to avoid unsupported actions.
Used by settings. Offline / batch RL
Functor family. family:coverage-vs-exploration
Current Lean routes. No compiled route mapped yet.
partial observability · source-map
Belief-state / memory representation
Replace fully observed states by information states, beliefs, finite memory, or predictive representations before planning/learning.
Used by settings. Partially observable MDP (POMDP)
Functor family. family:representation-lift
Current Lean routes. No compiled route mapped yet.
reinforcement learning · mapped
Bellman recursion + optimistic bonuses
Propagate statistical uncertainty through Bellman recursion and control cumulative episodic regret via occupancy/bonus decompositions.
Used by settings. MDP reinforcement learning · Online / adversarial MDP · Linear / kernel MDPs · Constrained / safe RL
Functor family. family:optimism-confidence-regret
Current Lean routes. teaching:finite-horizon-rl
application bridge · application-bridge
Preference / routing / allocation reductions for LLM systems
Map an LLM-system decision problem—preference selection, model routing, evaluation, or inference-time allocation—to an explicit bandit/RL feedback and objective contract before applying theory.
Used by settings. Bandits for LLM / language-model systems
Functor family. family:representation-lift
Current Lean routes. No compiled route mapped yet.
quantum access · cross-library
Quantum mean estimation / quantum testing / quantum design
Exploit coherent quantum reward-oracle access for faster mean estimation, and use quantum hypothesis-testing/polynomial-method lower bounds to characterize the new regret scale.
Used by settings. Quantum bandits
Functor family. family:quantum-estimation-testing
Current Lean routes. teaching:probability · teaching:oful · spine:chapter-14-information-theory · spine:chapter-15-minimax-lower-bounds
Primary/current source anchors
Cross-library substrates